Vũ Văn HảiFull-stack · AI-native
GuidesBlog
Discuss a project

© 2026 Vu Van Hai · Written from real deployment experience.

HomeGuidesBlogRSS
  1. Guides
  2. /Integrations
  3. /Set up Vertex AI for the Gemini API (keyless and service account)

Set up Vertex AI for the Gemini API (keyless and service account)

Set up Google Cloud Vertex AI (now Gemini Enterprise Agent Platform) to call Gemini models: enable the API, create a service account, authenticate keyless with ADC or with a JSON key, install the google-genai SDK, and fix org policy, 429 and 404 errors.

Updated: Sep 21, 202614 min read
GeminiVertex AIGoogle CloudSecurity
On this page
  • When to use Vertex AI instead of AI Studio
  • What the $300 credit actually covers
  • Quick reference
  • Step 1: Enable the Vertex AI API
  • Step 2: Create a Service Account and grant roles
  • Step 3: Choose an authentication method
  • Option A - User ADC (fast for dev)
  • Option B - Attached Service Account (production)
  • Option C - Impersonation (VPS/CI, no key)
  • Option D - JSON key (when you have no choice)
  • Step 4: Configure environment variables
  • Step 5: Install the SDK
  • Step 6: Sample code
  • Python - Keyless (recommended, no key file)
  • Python - Using a JSON key (option D)
  • REST API (cURL) - Global endpoint
  • Troubleshooting
  • Org policy blocks key creation: "Service account key creation is disabled"
  • 429 RESOURCE_EXHAUSTED and Dynamic Shared Quota (DSQ)
  • 404 "Publisher Model ... was not found" with Gemini 3.x
  • Regional vs Global endpoint
  • Gemini models on Vertex AI (as of around July 2026)
  • Script to test which models work
  • AI Studio vs Vertex AI
  • Common errors
  • References

This guide sets up Google Cloud Vertex AI so you can call the Gemini models (2.5, 3.x...) over its API. Use it when your environment cannot reach the Google AI Studio API (VPS, CI/CD, production), or when you want to spend the $300 Google Cloud Free Trial Credit. It covers enabling the API, creating a service account, picking an auth method (keyless first), sample code, and fixes for the common errors.

Throughout, replace <project-id> with your Project ID, and <username> and <server-ip> with your server user and IP. my-ai-worker is an example service account name and service-account-key.json is an example key file name - rename them freely.

Rename (2026): Google renamed "Vertex AI" to "Gemini Enterprise Agent Platform" (announced at Google Cloud Next 2026, rolled out by around late May 2026). The API endpoint did NOT change - it is still aiplatform.googleapis.com, and existing code keeps working. Docs under cloud.google.com/vertex-ai/... now 301-redirect to .../gemini-enterprise-agent-platform/...; this guide keeps saying "Vertex AI" because that is the familiar name.

When to use Vertex AI instead of AI Studio

  • Your environment blocks the AI Studio API endpoint (generativelanguage.googleapis.com).
  • You want to use the $300 Free Trial Credit (see the limits right below).
  • You need proper IAM, logging, monitoring and billing.

What the $300 credit actually covers

This matters because many people get it wrong:

How you call the modelCovered by the $300 credit?Notes
Gemini through Vertex AIYesGoogle's own models, billed via aiplatform.googleapis.com
Gemini API through AI Studio (generativelanguage.googleapis.com)NoExcluded by Google since around March 2026
Partner models / model-as-a-service in Model Garden (Claude, Llama, Mistral...)NoThe credit does not apply to third-party models

Quick reference

There are two ways to authenticate. Prefer Keyless (ADC) - it is safer and sidesteps the org policy that blocks key creation (see Troubleshooting).

# === 0. Enable the API (the service name did NOT change with the rebrand) ===
gcloud services enable aiplatform.googleapis.com
gcloud services enable storage.googleapis.com   # if you need to upload large files via GCS

# === 1. Create a Service Account + grant roles ===
gcloud iam service-accounts create my-ai-worker --display-name="AI Worker"

gcloud projects add-iam-policy-binding <project-id> \
    --member="serviceAccount:my-ai-worker@<project-id>.iam.gserviceaccount.com" \
    --role="roles/aiplatform.user"
# If you use GCS: add the storage.objectAdmin role the same way

# === 2A. KEYLESS AUTH (recommended) - no key file needed ===
gcloud auth application-default login                       # login = your own identity (dev machine)
gcloud auth application-default set-quota-project <project-id>
# On GCE / Cloud Run: attach the SA to the machine/service -> ADC just works, no command needed

# === 2B. Or use a JSON key (when forced to, e.g. a VPS that cannot log in) ===
gcloud iam service-accounts keys create service-account-key.json \
    --iam-account=my-ai-worker@<project-id>.iam.gserviceaccount.com
# NOTE: may be blocked by the org policy iam.disableServiceAccountKeyCreation -> see Troubleshooting

# === 3. Environment variables ===
export GOOGLE_CLOUD_PROJECT=<project-id>
export GOOGLE_CLOUD_LOCATION=global    # "global" is required for Gemini 3.x preview models
export GOOGLE_APPLICATION_CREDENTIALS=/path/to/service-account-key.json   # ONLY needed for option 2B

# === 4. Install the SDK (use google-genai, NOT the old vertexai SDK - removed on June 24, 2026) ===
pip install "google-genai>=1.51.0" google-auth

# === 5. Quick test (keyless - no credentials= passed -> the SDK uses ADC) ===
python3 -c "
from google import genai
client = genai.Client(vertexai=True, project='<project-id>', location='global')
r = client.models.generate_content(model='gemini-2.5-flash', contents='Hello')
print(r.text)
"

Step 1: Enable the Vertex AI API

The most reliable way (no hunting for the name in the UI):

gcloud services enable aiplatform.googleapis.com

Or open the Library page of the exact service directly (replace <project-id>):

https://console.cloud.google.com/apis/library/aiplatform.googleapis.com?project=<project-id>

After the rebrand, if you go to APIs & Services -> Library, search for "Vertex AI API" and find nothing (only "Vertex AI Search for commerce API"...), don't panic. Usually the API is already enabled so it is hidden from the Browse area, or the UI changed with the rename. Search for aiplatform (the service ID always matches) or use the URL/gcloud command above.

If you need to upload large files through Cloud Storage:

gcloud services enable storage.googleapis.com

Step 2: Create a Service Account and grant roles

gcloud iam service-accounts create my-ai-worker --display-name="AI Worker"

Grant the minimum roles:

RolePurpose
roles/aiplatform.user (Vertex AI User)Call the Gemini API through Vertex AI
roles/storage.objectAdmin (Storage Object Admin)Upload/delete files on GCS (only if needed)
gcloud projects add-iam-policy-binding <project-id> \
    --member="serviceAccount:my-ai-worker@<project-id>.iam.gserviceaccount.com" \
    --role="roles/aiplatform.user"

For a quick test you can grant Owner, but production should always follow least-privilege (only the 2 roles above).

With Keyless using your personal identity (Step 3, option A), you don't have to create an SA at all - use your own account if you are already Owner/Editor of the project. An SA is only needed for attached-SA, impersonation, or a key file.

Step 3: Choose an authentication method

MethodWhen to useKey file needed?Notes
A. User ADCDeveloping on your own machineNogcloud auth application-default login
B. Attached SARunning on GCE / Cloud Run / GKENoAttach the SA to the machine -> ADC just works. Best for production
C. ImpersonationVPS/CI that needs the SA identity without a keyNoLog in, then impersonate the SA
D. JSON keyEnvironments where an ADC login is not possibleYesMay be blocked by org policy (see Troubleshooting)

Google recommends Keyless (A/B/C) over a key file (D) - long-lived JSON keys are the most common source of credential leaks.

Option A - User ADC (fast for dev)

gcloud auth application-default login
gcloud auth application-default set-quota-project <project-id>

ADC is stored at ~/.config/gcloud/application_default_credentials.json. The SDK picks it up on its own, no need to pass credentials=.

If your .env has GOOGLE_API_KEY / GEMINI_API_KEY, delete it or leave it empty - otherwise the SDK may prefer that key over ADC.

Option B - Attached Service Account (production)

Attach the SA my-ai-worker@... to the GCE VM / Cloud Run service when you create it (or edit it in the settings). No command, no file - ADC fetches credentials from the metadata server.

Option C - Impersonation (VPS/CI, no key)

gcloud auth application-default login \
    --impersonate-service-account=my-ai-worker@<project-id>.iam.gserviceaccount.com

Your account needs the roles/iam.serviceAccountTokenCreator role on that SA.

Option D - JSON key (when you have no choice)

gcloud iam service-accounts keys create service-account-key.json \
    --iam-account=my-ai-worker@<project-id>.iam.gserviceaccount.com

If you get "Service account key creation is disabled" (iam.disableServiceAccountKeyCreation), an org policy is blocking it. See the Troubleshooting section below, and prefer switching to Keyless (A/B/C).

Push the key to the server (if needed) and make it readable by its owner only:

scp service-account-key.json <username>@<server-ip>:/path/to/project/service-account-key.json
ssh <username>@<server-ip> "chmod 600 /path/to/project/service-account-key.json"

NEVER commit the key file to Git. Add it to .gitignore as soon as you create it.

echo "service-account-key.json" >> .gitignore
# Check that it is ignored:
git check-ignore service-account-key.json

Step 4: Configure environment variables

# === VERTEX AI CONFIG ===
GOOGLE_CLOUD_PROJECT=<project-id>
GOOGLE_CLOUD_LOCATION=global
# Only needed with a JSON key (option D); REMOVE this line when going Keyless:
GOOGLE_APPLICATION_CREDENTIALS=/path/to/service-account-key.json

# Model
GEMINI_MODEL=gemini-2.5-flash
VariableDescriptionExample
GOOGLE_CLOUD_PROJECTProject ID (must match the project of the SA/key)<project-id>
GOOGLE_CLOUD_LOCATIONEndpoint locationglobal or us-central1
GOOGLE_APPLICATION_CREDENTIALSPath to the key file - option D only/app/service-account-key.json
GEMINI_MODELModel namegemini-2.5-flash

Common trap: GOOGLE_CLOUD_PROJECT must be the project the SA/credential belongs to. If the projects differ, the SA has no permission -> a 403 error.

Step 5: Install the SDK

pip install "google-genai>=1.51.0" google-auth
# google-cloud-storage if you upload files to GCS
Package (Python)Purpose
google-genaiThe main SDK for calling Gemini through Vertex AI (the replacement)
google-authService Account / ADC authentication
google-cloud-storageUpload files to GCS (if needed)

The old SDK is gone: the generative-AI modules of vertexai / google-cloud-aiplatform (vertexai.generative_models, .language_models...) were deprecated on June 24, 2025 and removed on June 24, 2026. Use google-genai; Gemini 3.x needs google-genai >= 1.51.0.

Node.js: npm install @google/genai (the old @google-cloud/vertexai package is also being replaced by @google/genai).

Step 6: Sample code

Python - Keyless (recommended, no key file)

import os
from dotenv import load_dotenv
from google import genai
from google.genai import types

load_dotenv()

# No credentials= passed -> the SDK uses Application Default Credentials (ADC)
# (from 'gcloud auth application-default login', an attached SA, or impersonation)
client = genai.Client(
    vertexai=True,
    project=os.getenv('GOOGLE_CLOUD_PROJECT'),
    location=os.getenv('GOOGLE_CLOUD_LOCATION', 'global'),
)

response = client.models.generate_content(
    model=os.getenv('GEMINI_MODEL', 'gemini-2.5-flash'),
    contents='Hello, please introduce yourself.',
    config=types.GenerateContentConfig(max_output_tokens=1024, temperature=0.7),
)
print(response.text)

Python - Using a JSON key (option D)

import os
from dotenv import load_dotenv
from google import genai
from google.oauth2 import service_account

load_dotenv()

creds = service_account.Credentials.from_service_account_file(
    os.getenv('GOOGLE_APPLICATION_CREDENTIALS'),
    scopes=['https://www.googleapis.com/auth/cloud-platform'],
)

client = genai.Client(
    vertexai=True,
    project=os.getenv('GOOGLE_CLOUD_PROJECT'),
    location=os.getenv('GOOGLE_CLOUD_LOCATION', 'global'),
    credentials=creds,
)
# ... call generate_content as above

REST API (cURL) - Global endpoint

ACCESS_TOKEN=$(gcloud auth print-access-token)

curl -X POST \
  "https://aiplatform.googleapis.com/v1/projects/<project-id>/locations/global/publishers/google/models/gemini-2.5-flash:generateContent" \
  -H "Authorization: Bearer $ACCESS_TOKEN" \
  -H "Content-Type: application/json" \
  -d '{
    "contents": [{"role": "user", "parts": [{"text": "Hello"}]}]
  }'

Troubleshooting

Org policy blocks key creation: "Service account key creation is disabled"

An Organization Policy that blocks service accounts key creation has been enforced
on your organization.  (iam.disableServiceAccountKeyCreation)

Cause: Google enables the iam.disableServiceAccountKeyCreation constraint by default, under "Secure by Default", for every organization created on or after May 3, 2024 (including the org auto-created for a personal account). It blocks JSON key creation.

Fixes (in recommended order):

  1. Switch to Keyless (Step 3, option A/B/C) - avoids the problem entirely, no policy change needed, and it is safer.
  2. Turn the policy off (only if you truly need a key). You need the roles/orgpolicy.policyAdmin role (Organization Policy Administrator) at the organization level.

Trap: the Organization Administrator role does NOT include permission to edit org policies. But if you are an Org Admin, you can grant roles/orgpolicy.policyAdmin to yourself (because you hold setIamPolicy on the org).

Steps to turn the policy off in the Console:

  1. In the resource picker, select the Organization.
  2. Go to IAM -> Grant access -> add the Organization Policy Administrator role for your email.
  3. Wait 1-2 minutes for the role to take effect.
  4. Go to IAM & Admin -> Organization Policies -> open iam.disableServiceAccountKeyCreation.
  5. Click Edit -> Override parent's policy -> rule Not enforced -> Set policy.

For least-privilege, override it at the project level rather than for the whole org.

This removes a security guardrail that Google turned on deliberately. Think it through, and consider re-enabling the policy once you have the key.

429 RESOURCE_EXHAUSTED and Dynamic Shared Quota (DSQ)

The newer Gemini models (2.x, 3.x) on Vertex run under Dynamic Shared Quota:

  • There is no fixed RPM/TPM limit, and you cannot request a quota increase (there is no quota left to raise). Google's documentation says DSQ removes quotas entirely, so quota increase requests are no longer needed.
  • A 429 does NOT mean you exceeded your own limit - it means the shared pool is congested at that moment. 429s are more likely with large multimodal inputs (audio/video/images).

How to reduce 429s:

  • Exponential backoff + retry (Google's number one recommendation) - wait 1s, 2s, 4s... plus jitter.
  • Lower the number of parallel requests (concurrent workers).
  • Use the global endpoint (location=global) - it routes to the region with the most spare capacity.
  • Batch Prediction API for offline bulk jobs (separate pool, no fixed limit) - a good fit for transcription/bulk processing.
  • Provisioned Throughput (prepaid GSUs) if production needs guaranteed capacity - expensive, not worth it for a one-off job.

In practice free-trial/new accounts may get throttled earlier (new projects tend to have low priority in the pool), but Google publishes no official numbers - don't treat it as a guarantee.

404 "Publisher Model ... was not found" with Gemini 3.x

{ "error": { "code": 404,
  "message": "Publisher Model `.../locations/us-central1/publishers/google/models/gemini-3.1-pro-preview` was not found",
  "status": "NOT_FOUND" } }

Cause: The Gemini 3.x preview models (gemini-3.1-pro-preview, gemini-3-flash-preview...) only run on the global endpoint; they do not exist on regional ones (us-central1...).

Fix: change location to global:

# WRONG - regional endpoints don't have the 3.x preview models
client = genai.Client(vertexai=True, project='<project-id>', location='us-central1')
# RIGHT
client = genai.Client(vertexai=True, project='<project-id>', location='global')

REST: the host is https://aiplatform.googleapis.com/v1/projects/<project-id>/locations/global/... (NOT global-aiplatform... or <region>-aiplatform...).

SDK requirement: google-genai >= 1.51.0 for Gemini 3.x.

Regional vs Global endpoint

CriteriaRegional (us-central1...)Global
ProsData residency; supports tuning, batch prediction, context cachingMore capacity (fewer 429s); has all the 3.x preview models
ConsSome preview models are missingNo control over the processing region; does NOT support tuning / batch / context caching
3.x preview modelsNo (returns 404)Yes (works)

Recommendation: use global unless you need data residency or batch/tuning/context-caching.

Gemini models on Vertex AI (as of around July 2026)

ModelStatusNotes
gemini-2.5-flashGAStable, recommended for most use cases
gemini-2.5-proGAHigher quality, pricier, hits 429 more easily
gemini-2.5-flash-liteGACheapest, for simple tasks
gemini-3-flash-previewPreviewGlobal endpoint required
gemini-3.1-pro-previewPreviewGlobal endpoint required
gemini-3.1-flash-litePreviewGlobal endpoint required
gemini-3-pro-previewRetired (around March 2026)Replaced by gemini-3.1-pro-preview - don't use the old ID

Models change very fast. Gemini 3 is not GA yet (3.1 Pro is still Preview as of July 2026), and the Gemini 3.5 family is starting to appear. Before hardcoding a model ID, check the live model list.

Script to test which models work

import os
from dotenv import load_dotenv
from google import genai
from google.genai import types

load_dotenv()

# Keyless: no credentials= passed. Use global so the 3.x preview models are tested too.
client = genai.Client(
    vertexai=True,
    project=os.getenv('GOOGLE_CLOUD_PROJECT'),
    location='global',
)

models_to_test = [
    'gemini-2.5-flash',
    'gemini-2.5-pro',
    'gemini-2.5-flash-lite',
    'gemini-3-flash-preview',
    'gemini-3.1-pro-preview',
]

for m in models_to_test:
    try:
        client.models.generate_content(
            model=m, contents='Hello',
            config=types.GenerateContentConfig(max_output_tokens=20),
        )
        print(f'[OK]   {m}')
    except Exception as e:
        print(f'[FAIL] {m}: {str(e)[:80]}')

AI Studio vs Vertex AI

CriteriaAI StudioVertex AI
AuthAPI Key (GOOGLE_API_KEY)ADC / Service Account (keyless or key)
Endpointgenerativelanguage.googleapis.comaiplatform.googleapis.com
BillingBilled directly / free tierGoogle Cloud billing (the $300 credit applies)
ManagementSimpleFull IAM, logging, monitoring
Best forPrototypes, personal useProduction, teams, VPS/CI

Common errors

ErrorCauseFix
Service account key creation is disabledOrg policy blocks key creationGo Keyless, or turn off iam.disableServiceAccountKeyCreation (see Troubleshooting)
403 PERMISSION_DENIEDThe SA lacks a role, or the project differsGrant Vertex AI User; check that GOOGLE_CLOUD_PROJECT matches the SA's project
Vertex AI API has not been enabledThe API is not enabledgcloud services enable aiplatform.googleapis.com
Could not automatically determine credentialsNo ADC and no keyRun gcloud auth application-default login, or set GOOGLE_APPLICATION_CREDENTIALS
404 ... Publisher Model was not foundA 3.x model called on a regional endpointChange to location=global
429 RESOURCE_EXHAUSTEDThe DSQ pool is congestedBackoff + retry, fewer concurrent requests, global endpoint (see Troubleshooting)
FAILED_PRECONDITIONFirst time using VertexWait 2-3 minutes while Google sets up the service agents

References

  • Gemini Enterprise Agent Platform (formerly Vertex AI)
  • Model list + lifecycle
  • Locations / Global endpoint
  • Dynamic Shared Quota
  • Handling 429 errors
  • Keyless auth / disable SA keys
  • Google Gen AI SDK
PreviousCheck and refresh the OG image cache on social platforms

Related articles

  • Configure Google OAuth for a new domain (Next.js and .NET API)

    Create a new OAuth Client ID in Google Cloud Console for a new domain, declare the right JavaScript origins and redirect URIs, then update the environment variables of the Next.js frontend and the .NET API backend.

    Integrations

    Integrations
  • Debug and bypass SSL pinning on an Android emulator

    Inspect the API traffic of an Android app that uses SSL pinning with LDPlayer 9 and HTTP Toolkit: enable root, connect over ADB, auto-bypass SSL, and handle advanced protections.

    Dev tools

    Dev tools
  • Set up PostgreSQL on a VPS securely

    Install and configure PostgreSQL on an Ubuntu/Debian VPS for production: create a database and user, open remote access safely (SSL + pg_hba + UFW), and understand transaction isolation.

    Database

    Database

Written by Vu Van Hai

I'm Hai, a full-stack developer based in Ho Chi Minh City. These guides come from systems I built and run myself. Need to build or untangle something similar? Get in touch.

Discuss a projectMore guides

Spot a mistake or a command that no longer works? Let me know

On this page

  • When to use Vertex AI instead of AI Studio
  • What the $300 credit actually covers
  • Quick reference
  • Step 1: Enable the Vertex AI API
  • Step 2: Create a Service Account and grant roles
  • Step 3: Choose an authentication method
  • Option A - User ADC (fast for dev)
  • Option B - Attached Service Account (production)
  • Option C - Impersonation (VPS/CI, no key)
  • Option D - JSON key (when you have no choice)
  • Step 4: Configure environment variables
  • Step 5: Install the SDK
  • Step 6: Sample code
  • Python - Keyless (recommended, no key file)
  • Python - Using a JSON key (option D)
  • REST API (cURL) - Global endpoint
  • Troubleshooting
  • Org policy blocks key creation: "Service account key creation is disabled"
  • 429 RESOURCE_EXHAUSTED and Dynamic Shared Quota (DSQ)
  • 404 "Publisher Model ... was not found" with Gemini 3.x
  • Regional vs Global endpoint
  • Gemini models on Vertex AI (as of around July 2026)
  • Script to test which models work
  • AI Studio vs Vertex AI
  • Common errors
  • References