Set up Vertex AI for the Gemini API (keyless and service account)
Set up Google Cloud Vertex AI (now Gemini Enterprise Agent Platform) to call Gemini models: enable the API, create a service account, authenticate keyless with ADC or with a JSON key, install the google-genai SDK, and fix org policy, 429 and 404 errors.
On this page
- When to use Vertex AI instead of AI Studio
- What the $300 credit actually covers
- Quick reference
- Step 1: Enable the Vertex AI API
- Step 2: Create a Service Account and grant roles
- Step 3: Choose an authentication method
- Option A - User ADC (fast for dev)
- Option B - Attached Service Account (production)
- Option C - Impersonation (VPS/CI, no key)
- Option D - JSON key (when you have no choice)
- Step 4: Configure environment variables
- Step 5: Install the SDK
- Step 6: Sample code
- Python - Keyless (recommended, no key file)
- Python - Using a JSON key (option D)
- REST API (cURL) - Global endpoint
- Troubleshooting
- Org policy blocks key creation: "Service account key creation is disabled"
- 429 RESOURCE_EXHAUSTED and Dynamic Shared Quota (DSQ)
- 404 "Publisher Model ... was not found" with Gemini 3.x
- Regional vs Global endpoint
- Gemini models on Vertex AI (as of around July 2026)
- Script to test which models work
- AI Studio vs Vertex AI
- Common errors
- References
This guide sets up Google Cloud Vertex AI so you can call the Gemini models (2.5, 3.x...) over its API. Use it when your environment cannot reach the Google AI Studio API (VPS, CI/CD, production), or when you want to spend the $300 Google Cloud Free Trial Credit. It covers enabling the API, creating a service account, picking an auth method (keyless first), sample code, and fixes for the common errors.
Throughout, replace <project-id> with your Project ID, and <username> and
<server-ip> with your server user and IP. my-ai-worker is an example service account
name and service-account-key.json is an example key file name - rename them freely.
Rename (2026): Google renamed "Vertex AI" to "Gemini Enterprise Agent Platform"
(announced at Google Cloud Next 2026, rolled out by around late May 2026). The API
endpoint did NOT change - it is still aiplatform.googleapis.com, and existing code
keeps working. Docs under cloud.google.com/vertex-ai/... now 301-redirect to
.../gemini-enterprise-agent-platform/...; this guide keeps saying "Vertex AI" because
that is the familiar name.
When to use Vertex AI instead of AI Studio
- Your environment blocks the AI Studio API endpoint (
generativelanguage.googleapis.com). - You want to use the $300 Free Trial Credit (see the limits right below).
- You need proper IAM, logging, monitoring and billing.
What the $300 credit actually covers
This matters because many people get it wrong:
| How you call the model | Covered by the $300 credit? | Notes |
|---|---|---|
| Gemini through Vertex AI | Yes | Google's own models, billed via aiplatform.googleapis.com |
Gemini API through AI Studio (generativelanguage.googleapis.com) | No | Excluded by Google since around March 2026 |
| Partner models / model-as-a-service in Model Garden (Claude, Llama, Mistral...) | No | The credit does not apply to third-party models |
Quick reference
There are two ways to authenticate. Prefer Keyless (ADC) - it is safer and sidesteps the org policy that blocks key creation (see Troubleshooting).
# === 0. Enable the API (the service name did NOT change with the rebrand) ===
gcloud services enable aiplatform.googleapis.com
gcloud services enable storage.googleapis.com # if you need to upload large files via GCS
# === 1. Create a Service Account + grant roles ===
gcloud iam service-accounts create my-ai-worker --display-name="AI Worker"
gcloud projects add-iam-policy-binding <project-id> \
--member="serviceAccount:my-ai-worker@<project-id>.iam.gserviceaccount.com" \
--role="roles/aiplatform.user"
# If you use GCS: add the storage.objectAdmin role the same way
# === 2A. KEYLESS AUTH (recommended) - no key file needed ===
gcloud auth application-default login # login = your own identity (dev machine)
gcloud auth application-default set-quota-project <project-id>
# On GCE / Cloud Run: attach the SA to the machine/service -> ADC just works, no command needed
# === 2B. Or use a JSON key (when forced to, e.g. a VPS that cannot log in) ===
gcloud iam service-accounts keys create service-account-key.json \
--iam-account=my-ai-worker@<project-id>.iam.gserviceaccount.com
# NOTE: may be blocked by the org policy iam.disableServiceAccountKeyCreation -> see Troubleshooting
# === 3. Environment variables ===
export GOOGLE_CLOUD_PROJECT=<project-id>
export GOOGLE_CLOUD_LOCATION=global # "global" is required for Gemini 3.x preview models
export GOOGLE_APPLICATION_CREDENTIALS=/path/to/service-account-key.json # ONLY needed for option 2B
# === 4. Install the SDK (use google-genai, NOT the old vertexai SDK - removed on June 24, 2026) ===
pip install "google-genai>=1.51.0" google-auth
# === 5. Quick test (keyless - no credentials= passed -> the SDK uses ADC) ===
python3 -c "
from google import genai
client = genai.Client(vertexai=True, project='<project-id>', location='global')
r = client.models.generate_content(model='gemini-2.5-flash', contents='Hello')
print(r.text)
"Step 1: Enable the Vertex AI API
The most reliable way (no hunting for the name in the UI):
gcloud services enable aiplatform.googleapis.comOr open the Library page of the exact service directly (replace <project-id>):
https://console.cloud.google.com/apis/library/aiplatform.googleapis.com?project=<project-id>After the rebrand, if you go to APIs & Services -> Library, search for "Vertex AI
API" and find nothing (only "Vertex AI Search for commerce API"...), don't panic.
Usually the API is already enabled so it is hidden from the Browse area, or the UI changed
with the rename. Search for aiplatform (the service ID always matches) or use the
URL/gcloud command above.
If you need to upload large files through Cloud Storage:
gcloud services enable storage.googleapis.comStep 2: Create a Service Account and grant roles
gcloud iam service-accounts create my-ai-worker --display-name="AI Worker"Grant the minimum roles:
| Role | Purpose |
|---|---|
roles/aiplatform.user (Vertex AI User) | Call the Gemini API through Vertex AI |
roles/storage.objectAdmin (Storage Object Admin) | Upload/delete files on GCS (only if needed) |
gcloud projects add-iam-policy-binding <project-id> \
--member="serviceAccount:my-ai-worker@<project-id>.iam.gserviceaccount.com" \
--role="roles/aiplatform.user"For a quick test you can grant Owner, but production should always follow
least-privilege (only the 2 roles above).
With Keyless using your personal identity (Step 3, option A), you don't have to create an SA at all - use your own account if you are already Owner/Editor of the project. An SA is only needed for attached-SA, impersonation, or a key file.
Step 3: Choose an authentication method
| Method | When to use | Key file needed? | Notes |
|---|---|---|---|
| A. User ADC | Developing on your own machine | No | gcloud auth application-default login |
| B. Attached SA | Running on GCE / Cloud Run / GKE | No | Attach the SA to the machine -> ADC just works. Best for production |
| C. Impersonation | VPS/CI that needs the SA identity without a key | No | Log in, then impersonate the SA |
| D. JSON key | Environments where an ADC login is not possible | Yes | May be blocked by org policy (see Troubleshooting) |
Google recommends Keyless (A/B/C) over a key file (D) - long-lived JSON keys are the most common source of credential leaks.
Option A - User ADC (fast for dev)
gcloud auth application-default login
gcloud auth application-default set-quota-project <project-id>ADC is stored at ~/.config/gcloud/application_default_credentials.json. The SDK picks
it up on its own, no need to pass credentials=.
If your .env has GOOGLE_API_KEY / GEMINI_API_KEY, delete it or leave it empty -
otherwise the SDK may prefer that key over ADC.
Option B - Attached Service Account (production)
Attach the SA my-ai-worker@... to the GCE VM / Cloud Run service when you create it (or
edit it in the settings). No command, no file - ADC fetches credentials from the metadata
server.
Option C - Impersonation (VPS/CI, no key)
gcloud auth application-default login \
--impersonate-service-account=my-ai-worker@<project-id>.iam.gserviceaccount.comYour account needs the roles/iam.serviceAccountTokenCreator role on that SA.
Option D - JSON key (when you have no choice)
gcloud iam service-accounts keys create service-account-key.json \
--iam-account=my-ai-worker@<project-id>.iam.gserviceaccount.comIf you get "Service account key creation is disabled"
(iam.disableServiceAccountKeyCreation), an org policy is blocking it. See the
Troubleshooting section below, and prefer switching to Keyless (A/B/C).
Push the key to the server (if needed) and make it readable by its owner only:
scp service-account-key.json <username>@<server-ip>:/path/to/project/service-account-key.json
ssh <username>@<server-ip> "chmod 600 /path/to/project/service-account-key.json"NEVER commit the key file to Git. Add it to .gitignore as soon as you create it.
echo "service-account-key.json" >> .gitignore
# Check that it is ignored:
git check-ignore service-account-key.jsonStep 4: Configure environment variables
# === VERTEX AI CONFIG ===
GOOGLE_CLOUD_PROJECT=<project-id>
GOOGLE_CLOUD_LOCATION=global
# Only needed with a JSON key (option D); REMOVE this line when going Keyless:
GOOGLE_APPLICATION_CREDENTIALS=/path/to/service-account-key.json
# Model
GEMINI_MODEL=gemini-2.5-flash| Variable | Description | Example |
|---|---|---|
GOOGLE_CLOUD_PROJECT | Project ID (must match the project of the SA/key) | <project-id> |
GOOGLE_CLOUD_LOCATION | Endpoint location | global or us-central1 |
GOOGLE_APPLICATION_CREDENTIALS | Path to the key file - option D only | /app/service-account-key.json |
GEMINI_MODEL | Model name | gemini-2.5-flash |
Common trap: GOOGLE_CLOUD_PROJECT must be the project the SA/credential belongs to. If
the projects differ, the SA has no permission -> a 403 error.
Step 5: Install the SDK
pip install "google-genai>=1.51.0" google-auth
# google-cloud-storage if you upload files to GCS| Package (Python) | Purpose |
|---|---|
google-genai | The main SDK for calling Gemini through Vertex AI (the replacement) |
google-auth | Service Account / ADC authentication |
google-cloud-storage | Upload files to GCS (if needed) |
The old SDK is gone: the generative-AI modules of vertexai / google-cloud-aiplatform
(vertexai.generative_models, .language_models...) were deprecated on June 24, 2025 and
removed on June 24, 2026. Use google-genai; Gemini 3.x needs google-genai >= 1.51.0.
Node.js: npm install @google/genai (the old @google-cloud/vertexai package is also
being replaced by @google/genai).
Step 6: Sample code
Python - Keyless (recommended, no key file)
import os
from dotenv import load_dotenv
from google import genai
from google.genai import types
load_dotenv()
# No credentials= passed -> the SDK uses Application Default Credentials (ADC)
# (from 'gcloud auth application-default login', an attached SA, or impersonation)
client = genai.Client(
vertexai=True,
project=os.getenv('GOOGLE_CLOUD_PROJECT'),
location=os.getenv('GOOGLE_CLOUD_LOCATION', 'global'),
)
response = client.models.generate_content(
model=os.getenv('GEMINI_MODEL', 'gemini-2.5-flash'),
contents='Hello, please introduce yourself.',
config=types.GenerateContentConfig(max_output_tokens=1024, temperature=0.7),
)
print(response.text)Python - Using a JSON key (option D)
import os
from dotenv import load_dotenv
from google import genai
from google.oauth2 import service_account
load_dotenv()
creds = service_account.Credentials.from_service_account_file(
os.getenv('GOOGLE_APPLICATION_CREDENTIALS'),
scopes=['https://www.googleapis.com/auth/cloud-platform'],
)
client = genai.Client(
vertexai=True,
project=os.getenv('GOOGLE_CLOUD_PROJECT'),
location=os.getenv('GOOGLE_CLOUD_LOCATION', 'global'),
credentials=creds,
)
# ... call generate_content as aboveREST API (cURL) - Global endpoint
ACCESS_TOKEN=$(gcloud auth print-access-token)
curl -X POST \
"https://aiplatform.googleapis.com/v1/projects/<project-id>/locations/global/publishers/google/models/gemini-2.5-flash:generateContent" \
-H "Authorization: Bearer $ACCESS_TOKEN" \
-H "Content-Type: application/json" \
-d '{
"contents": [{"role": "user", "parts": [{"text": "Hello"}]}]
}'Troubleshooting
Org policy blocks key creation: "Service account key creation is disabled"
An Organization Policy that blocks service accounts key creation has been enforced
on your organization. (iam.disableServiceAccountKeyCreation)Cause: Google enables the iam.disableServiceAccountKeyCreation constraint by
default, under "Secure by Default", for every organization created on or after May 3,
2024 (including the org auto-created for a personal account). It blocks JSON key
creation.
Fixes (in recommended order):
- Switch to Keyless (Step 3, option A/B/C) - avoids the problem entirely, no policy change needed, and it is safer.
- Turn the policy off (only if you truly need a key). You need the
roles/orgpolicy.policyAdminrole (Organization Policy Administrator) at the organization level.
Trap: the Organization Administrator role does NOT include permission to edit org
policies. But if you are an Org Admin, you can grant roles/orgpolicy.policyAdmin to
yourself (because you hold setIamPolicy on the org).
Steps to turn the policy off in the Console:
- In the resource picker, select the Organization.
- Go to IAM -> Grant access -> add the
Organization Policy Administratorrole for your email. - Wait 1-2 minutes for the role to take effect.
- Go to IAM & Admin -> Organization Policies -> open
iam.disableServiceAccountKeyCreation. - Click Edit -> Override parent's policy -> rule Not enforced -> Set policy.
For least-privilege, override it at the project level rather than for the whole org.
This removes a security guardrail that Google turned on deliberately. Think it through, and consider re-enabling the policy once you have the key.
429 RESOURCE_EXHAUSTED and Dynamic Shared Quota (DSQ)
The newer Gemini models (2.x, 3.x) on Vertex run under Dynamic Shared Quota:
- There is no fixed RPM/TPM limit, and you cannot request a quota increase (there is no quota left to raise). Google's documentation says DSQ removes quotas entirely, so quota increase requests are no longer needed.
- A 429 does NOT mean you exceeded your own limit - it means the shared pool is congested at that moment. 429s are more likely with large multimodal inputs (audio/video/images).
How to reduce 429s:
- Exponential backoff + retry (Google's number one recommendation) - wait 1s, 2s, 4s... plus jitter.
- Lower the number of parallel requests (concurrent workers).
- Use the global endpoint (
location=global) - it routes to the region with the most spare capacity. - Batch Prediction API for offline bulk jobs (separate pool, no fixed limit) - a good fit for transcription/bulk processing.
- Provisioned Throughput (prepaid GSUs) if production needs guaranteed capacity - expensive, not worth it for a one-off job.
In practice free-trial/new accounts may get throttled earlier (new projects tend to have low priority in the pool), but Google publishes no official numbers - don't treat it as a guarantee.
404 "Publisher Model ... was not found" with Gemini 3.x
{ "error": { "code": 404,
"message": "Publisher Model `.../locations/us-central1/publishers/google/models/gemini-3.1-pro-preview` was not found",
"status": "NOT_FOUND" } }Cause: The Gemini 3.x preview models (gemini-3.1-pro-preview,
gemini-3-flash-preview...) only run on the global endpoint; they do not exist on
regional ones (us-central1...).
Fix: change location to global:
# WRONG - regional endpoints don't have the 3.x preview models
client = genai.Client(vertexai=True, project='<project-id>', location='us-central1')
# RIGHT
client = genai.Client(vertexai=True, project='<project-id>', location='global')REST: the host is
https://aiplatform.googleapis.com/v1/projects/<project-id>/locations/global/... (NOT
global-aiplatform... or <region>-aiplatform...).
SDK requirement: google-genai >= 1.51.0 for Gemini 3.x.
Regional vs Global endpoint
| Criteria | Regional (us-central1...) | Global |
|---|---|---|
| Pros | Data residency; supports tuning, batch prediction, context caching | More capacity (fewer 429s); has all the 3.x preview models |
| Cons | Some preview models are missing | No control over the processing region; does NOT support tuning / batch / context caching |
| 3.x preview models | No (returns 404) | Yes (works) |
Recommendation: use global unless you need data residency or
batch/tuning/context-caching.
Gemini models on Vertex AI (as of around July 2026)
| Model | Status | Notes |
|---|---|---|
gemini-2.5-flash | GA | Stable, recommended for most use cases |
gemini-2.5-pro | GA | Higher quality, pricier, hits 429 more easily |
gemini-2.5-flash-lite | GA | Cheapest, for simple tasks |
gemini-3-flash-preview | Preview | Global endpoint required |
gemini-3.1-pro-preview | Preview | Global endpoint required |
gemini-3.1-flash-lite | Preview | Global endpoint required |
gemini-3-pro-preview | Retired (around March 2026) | Replaced by gemini-3.1-pro-preview - don't use the old ID |
Models change very fast. Gemini 3 is not GA yet (3.1 Pro is still Preview as of July 2026), and the Gemini 3.5 family is starting to appear. Before hardcoding a model ID, check the live model list.
Script to test which models work
import os
from dotenv import load_dotenv
from google import genai
from google.genai import types
load_dotenv()
# Keyless: no credentials= passed. Use global so the 3.x preview models are tested too.
client = genai.Client(
vertexai=True,
project=os.getenv('GOOGLE_CLOUD_PROJECT'),
location='global',
)
models_to_test = [
'gemini-2.5-flash',
'gemini-2.5-pro',
'gemini-2.5-flash-lite',
'gemini-3-flash-preview',
'gemini-3.1-pro-preview',
]
for m in models_to_test:
try:
client.models.generate_content(
model=m, contents='Hello',
config=types.GenerateContentConfig(max_output_tokens=20),
)
print(f'[OK] {m}')
except Exception as e:
print(f'[FAIL] {m}: {str(e)[:80]}')AI Studio vs Vertex AI
| Criteria | AI Studio | Vertex AI |
|---|---|---|
| Auth | API Key (GOOGLE_API_KEY) | ADC / Service Account (keyless or key) |
| Endpoint | generativelanguage.googleapis.com | aiplatform.googleapis.com |
| Billing | Billed directly / free tier | Google Cloud billing (the $300 credit applies) |
| Management | Simple | Full IAM, logging, monitoring |
| Best for | Prototypes, personal use | Production, teams, VPS/CI |
Common errors
| Error | Cause | Fix |
|---|---|---|
Service account key creation is disabled | Org policy blocks key creation | Go Keyless, or turn off iam.disableServiceAccountKeyCreation (see Troubleshooting) |
403 PERMISSION_DENIED | The SA lacks a role, or the project differs | Grant Vertex AI User; check that GOOGLE_CLOUD_PROJECT matches the SA's project |
Vertex AI API has not been enabled | The API is not enabled | gcloud services enable aiplatform.googleapis.com |
Could not automatically determine credentials | No ADC and no key | Run gcloud auth application-default login, or set GOOGLE_APPLICATION_CREDENTIALS |
404 ... Publisher Model was not found | A 3.x model called on a regional endpoint | Change to location=global |
429 RESOURCE_EXHAUSTED | The DSQ pool is congested | Backoff + retry, fewer concurrent requests, global endpoint (see Troubleshooting) |
FAILED_PRECONDITION | First time using Vertex | Wait 2-3 minutes while Google sets up the service agents |