Tamarind Bio
Protein and molecular design
- Category
- AI
- Primary Subcategory
- Biomedical & Pharma Research Intelligence
Integration details
Description
Connect ChatGPT to your Tamarind Bio workspace for protein and molecular science. Validate, estimate, submit, and monitor jobs across 300+ tools — structure prediction, docking, protein and antibody design, molecular dynamics, and property prediction. Chain tools into pipelines, batch many inputs at once, and pull back structures, results, and logs — all from chat.
- Integration type
- Plugin
- Verification status
- Not applicable
- Platform
- ChatGPT
- Primary Subcategory
- Biomedical & Pharma Research Intelligence
- Secondary Subcategories
- None listed
- Brand
- Tamarind Bio
- Access
- Account required
- First tracked
- 2026-07-29
- Tool count
- 36
- Geography
- US
The Primary Subcategory used for this profile’s headline score.
Other Subcategories where the Integration is listed.
Get alerts for Tamarind Bio
Get updates when Tamarind Bio’s Discoverability Score or category rank changes.
ChatGPT Plugin Discovery Score
ChatGPT Plugin discovery is coming soon
ChatGPT can surface a Plugin when it matches a user's request.Your Plugin Discovery Score measures how often yours appears.
No spam. Unsubscribe any time.
What discovery looks like

Competing in ChatGPT Biomedical & Pharma Research Intelligence
View Category36 tools agents can invoke
Soft-stop every member of a batch or pipeline in one call. Flips JobStatus to "Stopped" on all batch children, all pipeline stages, and the pipeline parent row (if one exists). Doesn't delete rows or outputs. Use this when the user wants to stop a batch or pipeline as a unit. Don't call cancelJob N times — this fans out server-side via the Batch-index GSI in one round-trip. Pipeline support is for LEGACY pipelines only: their stages carry `Batch=<pipeline_name>`, so a pipeline name resolves through the same index a batch name does, and the parent row (Type='pipelines') is looked up separately via JobName-index and cancelled too. This does NOT stop a run created by submitPipeline. Those are Type='pipeline-run', a deliberately different job type that neither index reaches — use stopPipelineRun(runId) for them. Only cancels jobs in "In Queue", "Running", or "Pending" status — already-terminal jobs are skipped. **Responses:** - **200** (Success): Batch cancellation summary - Content-Type: `application/json` - **Response Properties:** - **cancelledCount**: Number of rows successfully flipped to Stopped. - **totalChildren**: Total batch/pipeline children found (any status). - **cancellableChildren**: Children in In Queue / Running / Pending — the subset eligible for cancel. - **parentCancelled**: True if a pipeline parent row was found and cancelled too. - **Example:** ```json { "success": true, "message": "Cancelled 12 jobs in batch \"my-batch\"", "batchName": "string" } ``` - **400**: batchName missing or invalid - **401**: Unauthorized - **404**: No batch or pipeline with that name found for this user
cancelBatch
Soft-stop a job — flips JobStatus to "Stopped" (same value the website cancel button writes) without deleting the row or outputs. Only works on jobs in "In Queue", "Running", or "Pending" status; jobs already in a terminal state are rejected with 400. Accepts either `jobName` (preferred — looked up via JobName-index GSI server-side, no need for callers to know the Id) or `jobId` (the DDB Id, if already known). At least one is required. Different from deleteJob, which permanently removes the row + outputs. cancelJob preserves both. **Responses:** - **200** (Success): Job cancelled - Content-Type: `application/json` - **Response Properties:** - **Example:** ```json { "success": true, "message": "Job cancelled successfully", "jobId": "string" } ``` - **400**: Job not in a cancellable state (already Completed / Failed / Cancelled) - **401**: Unauthorized - **403**: Access denied (job belongs to another user) - **404**: Job not found
cancelJob
Create a project, so a job can be tagged with it. CALL listProjects FIRST. An organization that requires project tags is usually curating that list deliberately, and a near-duplicate project is worse than reusing one — it fragments the job history the requirement exists to keep together. If nothing on the list fits, ask the user before creating rather than inventing a name for them. Args: name: Project name, unique within the scope. Required. description: Optional free text. scope: "org" (default) creates it for your whole organization, which is the ONLY scope that satisfies a required-project policy. "personal" creates a private one, which does NOT satisfy it and is only meaningful if you belong to an organization. Returns: JSON { created: {id, name, scope}, hint }. Pass `id` (or `name`) as `projectTag` on your next submitJob / submitBatch.
createProject
Delete a workspace file (or a whole folder) from your Tamarind account. Pass `filePath` to delete one file, or `folder` to delete every file under a folder. After a successful delete the cached file listing is invalidated, so the file stops appearing in getFiles/getFileStats immediately. Args: filePath: Name/path of the file to delete (e.g. "path/to/myFile.txt"). folder: Folder path — deletes all files under it. Returns: JSON confirmation, e.g. {"message": "File deleted successfully"}.
deleteFile
Permanently delete a job by name (if it's a batch, all of its subjobs are deleted too). Unlike cancelJob (which only stops a running job without removing it), this permanently removes the job row and its outputs. Returns a plain-text confirmation; raises a clean error if the job doesn't exist or isn't yours. Args: jobName: Name of the job (or batch) to delete. Returns: Plain-text confirmation, e.g. "myJob deleted".
deleteJob
Deploy your organization's own code as a Tamarind tool: create it if needed, upload the source, build an image, run a build as a test, and publish it. Calling this again with new source IS the update path — there is no separate update call. Identical source returns `unchanged` without rebuilding, and a code-only change returns `reuse_image` and skips the image build. Exactly one mode per call: - BUILD — pass `files` (and `binaryFiles`). Creates the tool if it does not exist, uploads, and starts a build. - TEST — pass `testVersion` (and `testSettings`) to run a Complete build as a job in the tool's Test history on the website, without publishing it. It runs the tool, so it uses compute. An identical test runs once; calling again reports that run. - PUBLISH — pass `publishVersion` alone to promote an already-built version. This is also how you roll back: publish an older completed version. - CANCEL — pass `cancelVersion` alone to stop an in-progress build. SOURCE FORMAT. `files` maps archive-relative path to text content; `Dockerfile`, `run.sh`, and `config.json` belong at its ROOT. Binary members go in `binaryFiles` as base64. At runtime the working directory is `/app`, scalar inputs arrive as environment variables named after `config.json`'s `inputs[]`, file inputs arrive as absolute paths under `/app/inputs/` (which is writable, so a tool whose code wants its input at a fixed path can stage it there), results must be written to `/app/out/`, and THE CONTAINER HAS NO NETWORK — bake packages and model weights into the image in the Dockerfile, which does have network. Bake them OUTSIDE `/app` (use `/opt/<tool>/`): the uploaded source replaces `/app` at runtime, so anything the Dockerfile fetches there is missing when the job runs, even though it passed its build-time checksum. PUBLISHING IS ORGANIZATION-WIDE. It changes what every member gets when they submit this tool, so only pass `publish=True` (or `publishVersion`) once the user has authorized that exact promotion. Args: name: Organization-unique tool name. files: {"Dockerfile": "...", "run.sh": "...", "config.json": "..."} — text content. binaryFiles: Same shape for binary members, base64-encoded. displayName: Human-readable name. Used only when the tool is created. description: One-line description. Used only when the tool is created. publish: Publish the new version once it builds. Requires waitSeconds > 0, so it only lands when no image build was needed. Otherwise publish separately with publishVersion once the build is Complete. waitSeconds: Poll the build for up to this long (max 45 — MCP clients abort longer calls). 0 returns immediately. An image build takes minutes, so expect to poll with getCustomTool rather than wait here. validateOnly: Check and hash the archive without creating or uploading anything. publishVersion: PUBLISH mode — the version.id returned by build or getCustomTool to promote. cancelVersion: CANCEL mode — the version.id returned by build or getCustomTool to cancel. testVersion: TEST mode — the version.id returned by build or getCustomTool to run. testSettings: TEST mode — the run's inputs, named as in config.json `inputs[]`. projectTag: TEST mode — the organization project id to file the run under (listProjects). Orgs that require a project tag refuse a test without one. idempotencyKey: Optional stable key for retrying the same build after an ambiguous response. Reuse only for the same source and runtime request; a different request conflicts. In TEST mode, a new key runs an identical test again. expectNew: Refuse to update an existing tool. Pass True whenever you believe the name is free — a listing is a snapshot, and another member can claim the name between your check and this call. Returns: JSON. Build: {"action": "build"|"reuse_image"|"unchanged", "version", "sourceDigest", "published"}. `action: "unchanged"` means identical source already has a version and nothing was rebuilt. Poll with getCustomTool until the version is terminal. Test: {"action": "test", "version", "job": {"jobName", "id", "status"}}; follow the job with getJobs(jobName=...). `alreadySubmitted` means this exact test already ran and nothing new started.
deployCustomTool
Estimate how LONG a job will run and how many WEIGHTED (compute) HOURS it will cost — BEFORE submitting anything. Use it to plan a run, compare tools, warn the user about a slow or expensive job, or budget a design loop. The number returned is the SAME estimate the submit path's budget gate uses, so it matches what submitting this job would quote. No job is created. Returns JSON: { jobType, numJobs, perJobRuntimeSeconds, perJobRuntimeHuman, tier, weightedHours?, note? } - perJobRuntimeSeconds / perJobRuntimeHuman: wall-clock estimate for ONE job. - numJobs: how many jobs these settings expand into (e.g. multiple sequences). - weightedHours: total compute-hour cost across the whole submission (the billing unit). Returned for PREMIUM users only; free users get `note` instead, mirroring the website (compute cost is shown to Premium users only). Args: type: Tool name — see getAvailableTools(). settings: Job settings object — see getJobSchema(jobType=type). Estimate one job at a time (a single settings object, not a list). jobName: Optional; does not create a job and does not affect the numbers. Leave as the default. atomCount: Optional atom count for MD tools (gromacs/openmm) whose runtime scales with system size; a typical default is used when omitted.
estimateTime
Get summary statistics about available files (counts by type). Use this BEFORE getFiles() to understand what file types are available and decide what to filter for. Returns: JSON with file counts by type and total count
getFileStats
Inspect your organization's custom tools, or poll one tool's build. Call with no `name` to list the custom tools your organization owns, including Drafts that are not yet published. Published custom tools are also visible as ordinary tools through getAvailableTools(custom=True), with their submit schema through getJobSchema. Call with a `name` to get that tool plus one version's build status and logs — this is the polling call after deployCustomTool. Add `listVersions=True` to get the tool's version history instead, newest first; that is where a rollback target comes from, because publishing an older Complete version IS the rollback. Stop when the version's `terminal` is true; statuses are Queued, Running, Complete, and Stopped, and a terminal build can carry a structured `error`. Build-log cursors stay live while a build runs: carry `logs.nextCursor` into the next poll and sleep between polls. A REPEATED non-null cursor means "no new logs yet", not "drain another page immediately". It goes null once the terminal log stream is exhausted. Args: name: The tool's name. Omit to list every custom tool in the organization. version: The version.id returned by build, or "latest" (default) for the newest. listVersions: With `name`, list that tool's versions newest-first instead of reading one. This is how you find a rollback target — an older version whose status is Complete. Paginate with `cursor`. logCursor: `logs.nextCursor` from the previous poll, to resume the log stream. cursor: Listing mode only — `nextCursor` from the previous page. Follow it until null. Returns: JSON. Without `name`: {"items": [...], "nextCursor": ...}. With `name`: {"tool", "version", "logs"}. With `listVersions`: {"tool", "versions": [...], "nextCursor"}.
getCustomTool
Download job output files — ONE call, ONE link, however many files. Give it the "jobName/fileName" paths you want. Those are exactly the `s3Path` values listJobFiles() returns, so pass them straight through: getJobFile(files=["my-job/results.csv"]) # one file getJobFile(files=["job-a/x.csv", "job-b/y.csv"]) # any number Use listJobFiles() first to see what exists — this tool does not search, it fetches exactly what you name. A whole batch is therefore two calls: found = listJobFiles("my-batch-*", "*-scores.csv")["files"] getJobFile(files=[f["s3Path"] for f in found]) Do NOT call this once per file. One call takes the whole set; a loop is what makes a batch download slow and floods the context. You get back a short-lived `downloadUrl` plus a `files` manifest of what it covers. Fetch everything with the single command in `instructions`: curl -fsS -o results.zip "<downloadUrl>" && unzip -o results.zip -d <DEST> One file streams raw; many files stream as a zip. As a convenience, a SINGLE text file under 256 KB comes back as `content` inline instead, so reading one log or CSV stays a single call. Pass maxSizeMB=0 to decline that and always get a `downloadUrl` — do that when you want the file ON DISK, since a 256 KB log inlined is ~65K tokens of context spent on bytes you meant to save. When you want only PART of what a glob matches — "batch 28 and onwards", "the ten largest" — filter the listJobFiles() result in code and pass the paths you kept. IF curl FAILS with "could not resolve host" / "network unreachable" / HTTP 000, your MCP client's sandbox is blocking outbound network. The MCP tools still work (they go through the client, not your shell) — only curl needs egress to the downloadUrl's host, which is named in `instructions`. Tell the USER to allow that host (Claude Desktop/Claude.ai: Settings -> Capabilities -> "Allow network egress"; other clients: their network allowlist), then re-run the curl. Args: files: List of "jobName/fileName" paths — the `s3Path` values from listJobFiles(), e.g. ["my-job/results.csv"] maxSizeMB: How many megabytes a SINGLE file may ride the tool channel as inline `content`. Default 256 KB. RAISE it (max 50) when your shell has no network access and curl is unavailable; set it to 0 to switch inlining OFF and always receive a `downloadUrl`, which is what you want when the file is headed for disk. jobName, fileName: DEPRECATED single-file form, equivalent to files=["<jobName>/<fileName>"]. Still accepted so clients holding a cached tool list keep working. Prefer `files`. Returns: JSON: {"downloadUrl", "downloadType", "count", "totalBytes", "expiresInSeconds", "instructions", "files": [{"s3Path", "size"}]} — or, for one small text file, {"content", "jobName", "fileName", "s3Path", "size", ...} CHECK "error" FIRST here too: a refusal (unknown or unauthorised job, a path that does not exist, a selection over the limits) returns {"error", "hint", ...} with NO "downloadUrl", "count" or "files". For a zip, "zipLayout" says how the members are named: "flat" when every file's path WITHIN its job is unique — any subdirectories are kept, only the "<jobName>/" prefix is dropped — otherwise every member is nested under "<jobName>/". Unzip into a directory (-d) rather than the cwd either way. Security: - Every entry is checked for path traversal, and matching is confined to your own files - Every job named is checked against the backend before anything is served — owning the prefix is not by itself authorization
getJobFile
Fetch the output logs for a specific job from S3. This is useful for debugging job failures, checking progress, or understanding what happened during job execution. Args: jobName: The name of the job (e.g., "03-11_10-proteinmpnn-9h0O-canary") maxLines: Maximum number of lines to return from the end of the log (default 500) Returns: The output log contents (last N lines) or error message Examples: - getJobLogs("03-11_10-proteinmpnn-9h0O-canary") - getJobLogs("my-alphafold-job", maxLines=100)
getJobLogs
Get the detailed parameter schema for a specific job type. Checks both standard tools (DynamoDB) and custom tools (API) to find the schema. Returns the complete schema including all parameter definitions with their types, descriptions, validation rules, and metadata. Args: jobType: Job type name (e.g., "proteinmpnn", "alphafold", "rfdiffusion", or custom tool name) Returns: Complete JSON schema for the job type including: - parameters: full parameter definitions (name, type, description, required, etc.) - displayName: human-readable tool name - description: tool description - categories: tool categories - tags: tool tags - outputTypes: what this tool declares it produces — molecular (pdb, sequence, sdf, smiles) and otherwise (csv, score, cif, ...); a few are compound, like "pdb-list", so match by CONTAINMENT. Chain on the molecular ones: they must cover what the next tool takes as input. Absent means the tool declares nothing (UNKNOWN), not that it produces nothing. - outputs: what the results look like — `taskType` (`generate` vs `score` is what tells you whether the tool creates a molecule or just annotates one), `produces` (molecular representations the output carries — as a column of the result table OR as a file written alongside it — INCLUDING any echoed from the input, so a scorer can list one), `mainCSV`, and `columns` (name, type, description, units). The column names are the metrics to filter or rank results on. Absent when the tool declares no contract. For a tool whose output depends on the task it runs, `outputs.byTask` gives the per-task contract and is AUTHORITATIVE for the task you set — read it instead of the scalar fields, which summarize across tasks and are NOT guaranteed to describe the task selector's default. Each entry is complete (a task inherits the top-level `mainCSV`/`taskType` when it doesn't redeclare them) and carries `generates`: what that task creates fresh, which is empty for a scoring task whose table still echoes its input. A column carrying a `tasks` list appears only in those tasks; an UNTAGGED column is simply not task-gated in the declaration, which does not promise every task fills it — read `generates` for what a task actually produces. `outputs.byTaskNote` restates this on the response. - custom: true if this is a custom tool (only present for custom tools) Examples: - getJobSchema("alphafold") - getJobSchema("proteinmpnn") - getJobSchema("protein_properties_wmo6g") # custom tool
getJobSchema
Get one pipeline run's status and per-step progress. Poll this to follow a run submitted with submitPipeline. Each step reports its own status, how many of its jobs are complete, and `outputGroup` — the molecules group that step produced. Use getPipelineRunResults(runId) to read those outputs. Args: runId: The run id (the `id` returned by submitPipeline). Returns: JSON PublicRun: { id, templateId, templateVersion, name, source, status, startedAt, completedAt, inputs, nodeRuns: [{id, nodeId, label, nodeType, status, startedAt, completedAt, outputCount, jobsTotal, jobsComplete, outputGroup}] }. This tool returns the API payload verbatim, so `nodeRuns` is the key you index — a step's own id is `id`, and the stable pipeline node id you pass to getPipelineRunResults(node=...) is `nodeId`. A backend that predates the node-run rename answers with `steps: [{..., node}]` instead; read `nodeRuns or steps` and `nodeId or node` if you must work against both. getPipelineRunResults normalizes this for you — only this passthrough varies.
getPipelineRun
Read a pipeline run's OUTPUTS — the molecules each step produced, with their scores — rather than just its status. Returns every step's molecules, with their scores — which is what you rank and compare on. A step's molecules come from one of two API resources, and WHERE the scores sit differs between them, so read whichever shape you get: * group molecules — id in `id`, scores under `metadata`; * the step's own outputs — id in `complexId`, scores at top level in `scores`. Passed through as they come rather than reshaped here, so the fields match what the docs for each resource say. A step with an `outputGroup` is normally read from the group, but falls back to its own outputs when the group returns nothing — so do not assume a step's shape from whether `outputGroup` is set. `outputGroup` is the molecules group a step MINTED, and a step that produced plenty may still have none: one that enriches its inputs in place (scoring, structure prediction) leaves its molecules in the group they came from, and a filter step's survivors exist only as step outputs. Those steps report `outputGroup: null` and their molecules are read from the step itself, so a null group no longer means "nothing here". Results appear per step as it finishes, so this is useful before the whole run is complete. A step reports an empty `molecules` list when it genuinely has not produced anything yet. To page through a step that produced more than `limit` molecules: call again with `node` set to that step and `cursor` set to its `nextCursor`. Args: runId: The run id. node: Only this step's outputs — the `nodeId` value from getPipelineRun's `nodeRuns`. Omit for every step that has produced output. limit: Molecules per step, 1-100 (default 25). Keep it small — molecules carry their sequences, files and full score history. cursor: A `nextCursor` from a previous call, to fetch the next page. Requires `node`, since a cursor belongs to one step's group. Returns: JSON: { runId, status, steps: [{node, label, status, outputGroup, molecules[], nextCursor, error?}] }.
getPipelineRunResults
Get the pipeline IR schema and the full authoring guide. Call this BEFORE writing a `pipeline` graph for submitPipeline/validatePipeline — the IR has rules that are not guessable, and authoring blind reliably produces graphs the API rejects. Returns the complete, uncompressed JSON Schema plus the prose guide. Both are fetched live from the app host, so they always match the deployed validator. The guide covers, among other things: the four node kinds; that there is NO edge list (topology lives inside each node's `inputs`); that every molecule `user_input` needs a reference group; that a `filter` node must either set `intersect: true` or carry at least one rule; and that residue selections are supplied per-run in the BINDING (`residuesByChain`), never as a node setting. Returns: JSON: { schemaUrl, guideUrl, schema, guide }. On failure, `error` plus the two URLs so you can read them another way.
getPipelineSchema
Get one pipeline template: its graph, its bindable inputs, and its versions. The `inputs` array is the important part — it is the manifest of what you must pass as `bindings` to submitPipeline/validatePipeline. Each entry gives the input node id to key the binding by, whether it takes molecules or a file, the molecule type it expects, its reference chain labels, and any `residueFields`. A non-empty `residueFields` means the run MAY take a `residuesByChain` selection in this input's binding — it does NOT mean one is required. This said REQUIRES, and it is not true: production runs of the rfdiffusion, bindcraft and antibody-boltzgen templates whose manifests declare residueFields finished with no `residuesByChain` at all, and the entries carry no per-field `required` flag to tell the two apart. Believing it forces a choice between refusing a runnable template and inventing a selection — and an invented one is not a no-op: it reaches the tool as designed residues / binder hotspots / a binding site, constraining the design to the wrong residues and spending GPU on scientifically wrong output that still looks successful. Do not guess. Call validatePipeline with the bindings you have: if a selection is genuinely required, it comes back as `required-field-unset` naming the field. Note also that `targetsChains` is empty on some entries and names a chain outside the input's own `chains` on others (a de-novo chain the graph creates), so it is not always a usable key for `residuesByChain`. Args: templateId: The template id (from listPipelineTemplates). version: A version handle like "v2". Omit for the version a default run would use (published, else latest). Returns: JSON PublicTemplate: { id, name, description, isPublished, publishedVersion, version, pipeline (the IR), inputs[], versions[], createdAt, updatedAt }.
getPipelineTemplate
Retrieve results for a completed job **Responses:** - **200** (Success): Job results - Content-Type: `application/json` - **Example:** ```json "https://s3.amazonaws.com/bucket/job-results.zip" ``` - **400**: Bad request - **401**: Unauthorized - **404**: Job not found
getResult
Get list of all available tools/job types in the Tamarind platform. Use this tool BEFORE submitting jobs to discover what tools are available. The catalog is large, so FILTER rather than listing everything: narrow by `modality` (the molecule type) and/or `function` (what the tool does), or by a free-text `search`. Call listModalities() and listTags() first to see the valid values (with descriptions) for those two filters. Args: modality: Filter by molecule type (e.g. "antibody", "protein", "peptide", "small-molecule", "nucleic-acid"). See listModalities(). function: Filter by what the tool does (e.g. "structure-prediction", "binder-design", "protein-ligand-docking"). See listTags(). category: Deprecated alias for `modality` (still honored). tag: Deprecated alias for `function` (still honored). search: Search in tool name or description custom: If true, return only custom tools; if false/None, return standard tools from database The catalog drifts as tools are added/removed, so don't guess values — every response includes the exact live `availableCategories` and `availableTags` (a.k.a. modalities and functions) computed from the current catalog, or call listModalities()/listTags() for the same vocabulary with descriptions. Returns: JSON with `tools` (name, displayName, categories, outputTypes[, tags, description]), plus `availableCategories` and `availableTags` — the live facet lists to filter against. `outputTypes` is what a tool declares it produces — molecular (pdb, sequence, sdf, smiles) and otherwise (csv, score, cif, ...). A few tools use compound values such as "pdb-list" or "sequence-list", so match by CONTAINMENT, not equality. Use it to plan a multi-step run: a molecular type here must cover what the next tool takes as input. It is OMITTED when the tool declares nothing, which means UNKNOWN — not "produces nothing", so confirm with getJobSchema rather than ruling the tool out. A type here is also not proof the tool GENERATED it: a scorer can echo its input sequence alongside the scores. getJobSchema's `taskType` is what distinguishes generating from scoring, along with the full result contract — the results CSV and its columns, i.e. the metric names to filter or rank on. Examples: - Find antibody tools: modality="antibody" - Find structure prediction tools: function="structure-prediction" - Narrow to antibody structure prediction: modality="antibody", function="structure-prediction" - Search for AlphaFold variants: search="alphafold" - Find custom tools: custom=true
getAvailableTools
List files available in the Tamarind workspace with filtering. IMPORTANT: This returns a LIMITED subset of files to avoid context overflow. Use 'types' and 'search' parameters to filter before listing. Use 'offset' to paginate through results. Args: types: Comma-separated file types to filter (e.g., "pdb,cif"). HIGHLY RECOMMENDED to avoid getting too many results. search: Search term to filter filenames (e.g., "output", "job123") limit: Maximum number of files to return (default 100, max 500) offset: Number of files to skip for pagination (default 0) includeMetadata: Include file size and modification time Returns: JSON with files array, count, hasMore flag, and total count Examples: - List PDB files: types="pdb" - Search for specific job: search="job_12345" - Get next page: offset=100, limit=100
getFiles
Retrieve a list of folders in the user's storage **Query Parameters:** - **limit**: Maximum number of folders to return - **offset**: Number of folders to skip - **loadAll**: Load all folders without pagination **Responses:** - **200** (Success): List of folders - Content-Type: `application/json` - **Response Properties:** - **hasMore**: Whether there are more folders available - **total**: Total number of folders - **offset**: Current offset in the results - **limit**: Maximum number of folders returned - **Example:** ```json { "folders": [ "string" ], "hasMore": true, "total": 1 } ``` - **401**: Unauthorized - **500**: Internal server error - Content-Type: `application/json` - **Response Properties:** - **Example:** ```json { "error": "Failed to fetch folders" } ```
getFolders
List the FUNCTIONS (a.k.a. tags) you can filter tools by. A function is what a tool DOES — structure-prediction, binder-design, protein-ligand-docking, binding-affinity, etc. This is one of the two axes for narrowing the tool catalog (the other is modality — see listModalities()). Call this BEFORE getAvailableTools to pick a function, then filter with getAvailableTools(function="<value>"). Combine with modality= to narrow more. Returns: JSON: { functions: [{ value, label, description?, toolCount }], ... }. Pass `value` (e.g. "structure-prediction") as the function filter. `toolCount` is how many tools you can currently see with that function.
listTags
List job output files — across ONE job or MANY, without downloading anything. Answers "what exists?" cheaply: names, sizes and `s3Path` only — no file contents and no download URLs. This is always the first step, because getJobFile does not search; it fetches the paths you hand it. listJobFiles("my-job") # every file in one job listJobFiles("my-job", "*.csv") # just the CSVs in one job listJobFiles("my-batch-*", "*-scores.csv") # one file across many jobs BOTH arguments accept globs (`*`, `?`, `[...]`) and match the WHOLE name. `fileName` defaults to "*", so a bare job name lists everything in it. Pass the `s3Path` values you want to getJobFile(files=[...]) — one call for the whole set. When your criterion cannot be written as a glob ("the newest ten", "everything after run 28"), filter this result in code and pass the paths you kept. IMPORTANT: the returned s3Path also works directly in submitJob to chain jobs — no need to download and re-upload. Args: jobName: A job name, or a glob over job names (e.g. "my-batch-*"). Put as much LITERAL text as you can before the first wildcard — that prefix is what makes the search cheap. A pattern starting with a wildcard ("*-scores") has to walk every job you own and will usually time out. fileName: Glob over files within each job (default "*" = everything; e.g. "*-scores.csv"). Matches the whole name, and `*` spans "/" — "*.csv" also matches "seqs/out.csv". Returns: JSON {"count", "bytesListed", "files": [{"jobName", "path", "s3Path", "size", "sizeHuman", "lastModified"}], "hints"}. One job additionally returns "jobName" and "s3Prefix". No contents, no download URLs. CHECK "error" FIRST. A job name with no wildcard is an exact lookup, so an unknown or unauthorised one comes back as {"error", "hint"} with NO "count" and NO "files" — code that reaches straight for "count" dies on a mistyped name. A job name WITH a wildcard is a search and never errors that way; it answers "count": 0. When the request succeeded but matched nothing, "count" is 0 and a singular "hint" replaces "hints". A cross-job search may stop early on a cap; then "truncated" says why, and "count"/"bytesListed" describe only what is listed — narrow jobName and search again. A single job is always listed in full. Examples: - One job: listJobFiles("03-11_10-proteinmpnn-9h0O-canary") - A sweep: listJobFiles("my-batch-*", "*-scores.csv") - Then fetch: getJobFile(files=["job-a/x.csv", "job-b/y.csv"])
listJobFiles
List your jobs (status, timing, score), or fetch one job by name. By default each job's bulky input JSON is OMITTED: for a big batch it dominates the payload and overflows the context. The input blob is named `Sequence` on the production API and `Settings` on the legacy API, so both are dropped (whichever a job carries). Set includeSequences=true to keep them, or use getResult()/getJobFile() for a single job's full inputs/outputs. The `Score` field is retained as a ranking signal. Args: jobName: Return just this one job. batch: Return all jobs in this batch. startKey: Pagination cursor from a previous call (to page past `limit`). limit: Max jobs to return (default 1000); a startKey is returned if more remain. organization: If true, return all jobs across your organization. includeSubjobs: If true, include batch subjobs (only top-level jobs by default). jobEmail: Return jobs for another member of your organization. includeSequences: If true, keep each job's bulky input JSON (the `Sequence`/`Settings` fields, omitted by default). Returns: JSON: {jobs: [...], startKey?, statuses?} for listings, or a single job object when jobName is given.
getJobs
List the biological MODALITIES (molecule types) you can filter tools by. A modality is what KIND of molecule a tool works on — protein, antibody, peptide, small-molecule, nucleic-acid, etc. This is one of the two axes for narrowing the tool catalog (the other is function — see listTags()). Call this BEFORE getAvailableTools to pick a modality, then filter with getAvailableTools(modality="<value>"). Combine with function= to narrow more. Returns: JSON: { modalities: [{ value, label, description?, toolCount }], ... }. Pass `value` (e.g. "antibody") as the modality filter. `toolCount` is how many tools you can currently see in that modality.
listModalities
List molecule groups — the way to FIND an existing group to feed a pipeline. A group id from here is the same id a binding takes ({"group": "<id>"}) and the same id a run reports as a step's `outputGroup`, so a group produced by one run can be fed straight into the next. Rows are summaries: id, name, type, moleculeCount, source, tags, timestamps — no molecules. Use searchMolecules(group=<id>) to see what is inside one. Args: search: Case-insensitive LITERAL substring of the group name — `_` and `%` match themselves, not as wildcards. scope: "mine" (default — groups you created) or "org" (your whole organization's). sort: "recent" (default), "name", or "size". dir: "asc" or "desc". limit: Page size, 1-200 (default 50). Keep it small and follow `nextCursor` — a full 200-row page costs roughly 19k tokens of context. cursor: `nextCursor` from a previous page. Returns: JSON: { items: [{id, name, type, moleculeCount, source, tags, status, createdAt, updatedAt}], nextCursor }.
listMoleculeGroups
List pipeline runs. Note that pipeline runs do NOT appear in getJobs() — that endpoint covers single jobs and batches only, so this is the way to see them. Args: status: Filter by run status. templateId: Only runs of this template. owner: "mine" (default) or "org". limit: Page size, 1-200 (default 50). cursor: `nextCursor` from a previous page. Returns: JSON: { items: [{id, templateId, templateVersion, name, source, status, startedAt, completedAt}], nextCursor }.
listPipelineRuns
List saved pipeline templates you can run. Use this to find an existing pipeline instead of authoring a graph from scratch — then read it with getPipelineTemplate() and run it with submitPipeline(templateId=...). Rows are summaries only (no graph). Call getPipelineTemplate for the IR and the list of inputs you must bind. Args: search: Match templates by name. owner: "org" (the DEFAULT — your whole organization's templates, so results include ones you did not create) or "mine" to narrow to your own. isPublished: True for published templates only, False for unpublished. limit: Page size, 1-200 (default 50). cursor: `nextCursor` from a previous page. Returns: JSON: { items: [{id, name, description, isPublished, versionCount, runCount, createdBy, createdAt, updatedAt}], nextCursor }.
listPipelineTemplates
List the projects you can tag jobs with, to resolve a `projectTag`. Call this when a submission is refused for a missing or invalid project tag and you need more than the capped `availableProjects` the refusal carries, or when the user names a project and you want its id before submitting. ONLY a project with scope "org" satisfies an organization's "require a project tag" policy. A "personal" project is your own and a "shared" one is someone else's private project you were invited to; submitting with either is refused exactly as if you had sent no tag at all. Returns: JSON { projects: [{ id, name, scope }], isInOrg }. Pass a project's `id` as `projectTag` on submitJob / submitBatch, or as `project` on submitPipeline. `id` ALWAYS works; a `name` works only for "org" and "personal" scopes, because a "shared" project lives in another user's partition where names are not unique and are therefore not resolved — sending one comes back "not found". Prefer the id and the distinction stops mattering. When `isInOrg` is false you have no organization, so every project is reported as "org" and no required-project policy applies.
listProjects
Query molecules across every group and run — filter and rank by TOOL SCORE, not just by name. This is the cross-run view. getPipelineRunResults(runId) answers "what did THIS run produce, step by step"; this answers "across everything I have ever run, which molecules score best" — and takes a run's `outputGroup` as `group` when you want to narrow back down to one step. Filtering is the point. `filter` takes repeatable `field:operator:value` predicates, AND-ed, over metadata fields AND tool scores written as `tool.metric`: filter=["alphafold.plddt:gt:0.8", "round:eq:2"] Operators: gt, gte, lt, lte, eq, neq, and `field:exists` (metadata only). `sortBy` takes the same `tool.metric` form, so you can rank by a score directly. Molecules are heavy — they carry chains, files and score history — so keep `limit` small and follow `nextCursor`. Args: group: Only this group's molecules (a group id, or a step's `outputGroup`). jobName: Only molecules a job or batch produced or consumed, by name. Names are not unique; every visible match is included. jobId: The same, by id. Pass a batch id, not a child job id. type: Restrict to molecule kinds, e.g. ["protein"], ["small_molecule"]. search: A term. With mode="default" (the default) it matches molecule NAMES. `matchedOn` says which field hit, and in practice it is always ["name"]: the documented metadata-key and score-metric matching does not fire, so searching for an annotation key returns an empty page with a NULL `nextCursor` — which per this tool's own contract means "finished, no matches", not "not implemented". To find molecules by an annotation, use `filter` instead (`filter=["my_key:exists"]`), which does match metadata. With mode="sequence" it is an amino-acid subsequence. mode: "default" or "sequence". filter: Repeatable `field:operator:value` predicates (max 32), AND-ed. `neq` also excludes molecules where the field is absent, so it is not the complement of `eq`. scope: "mine" (only groups YOU created) or "org" (your whole organization). Defaults to "mine" — EXCEPT when you pass `group`, where it defaults to "org", because a group id you were given is often a coworker's and "mine" would answer with an empty page. Pass "mine" explicitly to override that. sortBy: "id", "added", "type", "groupName", or a metadata field / tool score (`tool.metric`, on that tool's best run). WARNING: an unknown field here is IGNORED rather than rejected — the page comes back 200 in the default order with nothing saying the ranking was dropped. Spell it exactly as `filter` would take it, and sanity-check that the first item really does lead on that metric before acting on the order. dir: "asc" or "desc". limit: Molecules per page, 1-100 (default 25). cursor: `nextCursor` from the previous page. Returns: JSON: { items[], nextCursor, mode, scanProgress }. EITHER mode may return an EMPTY page while `nextCursor` is set — that means "keep paging", not "no results"; only a null `nextCursor` ends the search. In sequence mode `scanProgress` (0.0-1.0) tracks the scan.
searchMolecules
Stop a running pipeline. Steps that already finished keep their outputs; work still in flight is cancelled. Use this rather than cancelBatch — cancelBatch only understands the older pipeline job layout and will not stop one of these runs. Args: runId: The run id to stop. Returns: JSON PublicRun, reflecting the run's state after the stop.
stopPipelineRun
Submit MANY jobs as a SINGLE batch (one tool call, not one per job). Submits the settings you construct to the Tamarind API (POST https://app.tamarind.bio/api/submit-batch). Requests only ever go to the Tamarind platform; the endpoint is fixed and not caller-controlled. ALWAYS prefer this over calling submitJob in a loop. If you are about to submit more than one job of the same `type`, use submitBatch instead — it is one request, groups the jobs under a shared batch name, and is far cheaper than N separate submissions. But when `settings` holds exactly ONE entry, call submitJob instead. A one-job batch is pure overhead: the platform writes a batch parent row beside the single child and runs an aggregation pass to wrap it, and nothing is being grouped. Stay here even for a single job when the request needs something submitJob cannot express — it takes only jobName, type and settings: - weightedHoursBudget, a runaway-cost cap. It is enforced off the batch parent row, so a single job without that row has no cap at all. - maxRuntimeSeconds, a per-job timeout. - fromJob / fromFile (+ topN, sequenceField, sharedSettings). Reading the source, parsing it and ranking it all happen HERE, so a chained request cannot be restated as a submitJob call — including when the source turns out to hold a single sequence. Three ways to provide the per-job inputs: 1. fromJob (chaining — RECOMMENDED for ProteinMPNN -> AlphaFold/Chai): Pass a COMPLETED sequence-design job name. This reads that job's generated sequences from its output (filtered_sequences.fa / seqs/*.fa / metrics.csv), and folds EACH sequence as its own job in one batch. Example: submitBatch(batchName="verify-mpnn", type="alphafold", fromJob="proteinmpnn-example-7v3te") 2. fromFile (a fasta or csv of sequences): Pass a workspace filename or a job s3Path (from listJobFiles) to a .fasta/.fa or .csv containing one sequence per record. Each sequence becomes one job. Example: submitBatch(batchName="fold-designs", type="alphafold", fromFile="designs.fasta") 3. settings + jobNames (explicit, full control): Provide an array of per-job settings objects and a matching array of job names. Use this when each job needs different parameters. NAMING RULE. Never reuse `batchName` for one of the `jobNames`. The batch is itself a job under that name, so a job sharing it collides and the batch never finishes; the server rejects the submission with 400. The comparison is against the CLEANED name — characters outside [A-Za-z0-9._-] are stripped and whitespace becomes "_", so a batch sent as "my batch" is stored as "my_batch", and that is the value no job may equal. Args: batchName: Unique name for the batch (also the prefix for auto job names). type: Tool to run for EVERY job in the batch (e.g. "alphafold", "chai"). All jobs in a batch use the same tool. Use getJobSchema(type) for params. settings: (mode 3) Array of per-job settings objects. jobNames: (mode 3) Array of job names, same length as settings, and none equal to batchName (see NAMING RULE). For fromJob/fromFile you can omit this — names are auto-generated as "{batchName}-1", "{batchName}-2", ... fromJob: (mode 1) Completed job whose generated sequences to fold. fromFile: (mode 2) Workspace path / job s3Path to a fasta or csv of sequences. sequenceField: Settings field each sequence is placed into (default "sequence" — correct for alphafold/chai/boltz/esmfold). topN: Keep only the top-N sequences (ranked by MPNN/overall_confidence score when available). Omit, or pass <=0, to fold ALL (default). sharedSettings: Settings applied to every generated job (mode 1 & 2), e.g. {"useMSA": true, "numModels": "5"}. maxRuntimeSeconds: Optional per-job timeout (integer seconds, >=1). weightedHoursBudget: Optional budget for the whole batch (number, >=0). projectTag: Organization project to file EVERY job in this batch under — its project id or its name. REQUIRED for orgs that have turned on "require a project tag"; those reject an untagged batch with 400 and a PERSONAL project does not count. The 400 carries `availableProjects` (id + name), so retry with one of those rather than looking them up first. Ask the user which project when more than one could fit. Note: sequences are NOT de-duplicated — each record in the source becomes its own job (folding an identical sequence twice is left to the caller). Returns: JSON summary: batchName, type, count, jobNames, source, and the backend response. On problems, an "error" field with a hint.
submitBatch
Submit ONE job for protein analysis using one of the available tools. Submits the settings you construct to the Tamarind API (POST https://app.tamarind.bio/api/submit-job). Requests only ever go to the Tamarind platform; the endpoint is fixed and not caller-controlled. IMPORTANT: Use this only for a single job. If you are about to submit more than one job (e.g. folding several sequences, scanning variants, or chaining a design job's output into structure prediction), call submitBatch ONCE instead of calling submitJob in a loop. Looping submitJob is slower, ungroupable, and rate-limited. For ProteinMPNN -> AlphaFold/Chai, use submitBatch(fromJob="<mpnn job>", type="alphafold"). **Responses:** - **200** (Success): Job submitted successfully - **400**: Bad request - invalid parameters - **401**: Unauthorized - invalid API key - **429**: Too many requests - rate limit exceeded
submitJob
Run a multi-tool pipeline. This SPENDS COMPUTE — call validatePipeline on the same graph and bindings first, and only submit once it reports valid:true. Two modes. Pass `pipeline` (a full IR graph) to author and run in one call: the server saves a template you own, unpublished, and runs it. Or pass `templateId` to run a template you already have, which is how you re-run without re-authoring the graph. Poll the returned run with getPipelineRun(runId), and read its outputs with getPipelineRunResults(runId). Args: name: Name for the pipeline. In inline mode it names the created template too. Required. bindings: One entry per input node, keyed by the input node's id. Give molecules exactly one way: {"group": "<groupId>"}, or raw values — {"sequences": [...]}, {"smiles": [...]}, {"pdbs": [...]}, {"sdfs": [...]} — and the server creates the group for you (and rolls it back if the submit fails). A file input takes {"file": "<path>"}. Add "residuesByChain": {"A": "1-76"} where a tool needs a residue selection, and "chainMapping" only when your chain IDs differ from the template's reference chains. pipeline: INLINE mode — a full pipeline IR. Call getPipelineSchema() first. Mutually exclusive with templateId. templateId: REFERENCE mode — an existing template to run. version: REFERENCE mode only — a version handle like "v2". Omit for published-then-latest. settings: REFERENCE mode only — per-node overrides {nodeId: {settingKey: value}}, limited to the template's editable settings. runName: An explicit name for THIS run. Omit to derive a unique one from `name`. project: Organization project id to stamp on the jobs this run creates. idempotencyKey: A client key that makes a retried submit safe. Returns: JSON PublicRun: { id, templateId, templateVersion, name, status, startedAt, completedAt, inputs, steps[] }. `id` is the runId.
submitPipeline
Upload a local file to your Tamarind file store (up to 200 MB). RECOMMENDED (up to 200 MB) — call with just `filename` to get a short-lived `uploadUrl` on mcp.tamarind.bio, then stream your file to it with ONE shell command: curl -T /path/to/<filename> "<uploadUrl>" The bytes go straight from your machine to your store. Afterwards the file appears in getFiles(), the web Database tab, and can be referenced by its BARE `filename` in submitJob/submitBatch file fields. IF curl FAILS with "could not resolve host" / "network unreachable" / HTTP 000, your MCP client's sandbox is blocking outbound network. The MCP tools keep working (they go through the client, not your shell) — only the upload curl needs egress to the uploadUrl's host (mcp.tamarind.bio in production). The exact host to allow is named in the uploadUrl/instructions you get back; tell the USER to allow that host, then re-run the curl: - Claude Desktop / Claude.ai: Settings -> Capabilities -> "Allow network egress" -> under "Additional allowed domains" add that host - Other MCP clients (e.g. Codex): add that host to the client's network / egress allowlist. FALLBACK (no shell available, or a small file you can't enable egress for): pass the file's bytes inline via `content` (text, or base64 for binary), up to the 4 MB inline cap. The server writes them directly — no curl, no egress: uploadFile(filename="target.pdb", content=<file text>) uploadFile(filename="model.bcif", content=<base64>, encoding="base64") FLAT FILES ONLY: pass a bare name like "target.pdb" (no "/"). Args: filename: Bare destination name, e.g. "target.pdb". content: FALLBACK only — the file's bytes inline (no curl, up to 4 MB). Leave EMPTY to get an uploadUrl to stream to (recommended). encoding: "text" (default; for text like PDB/CIF/FASTA) or "base64" (for binary or non-UTF8 inline content). contentType: MIME type to store (optional). Returns: JSON. Default: {"uploadUrl", "filename", "instructions", ...}. Inline fallback: {"success": true, "filename", "bytes", "hint"}.
uploadFile
Validate a job's settings BEFORE submitting, so you can catch missing or invalid fields and fix them (or ask the user) instead of submitting and surfacing a 400. Returns JSON: { valid, normalized?, error?, ... }. HTTP 200 is used for both outcomes — read the `valid` field. When valid is false, `error` explains the first problem. IMPORTANT: only submit settings that came back valid:true WITHOUT a `mutatedFields` warning. A mutated input means the validator silently dropped disallowed characters (e.g. non-standard residues), so submitting as-is would run the job on ALTERED input — fix the input and re-validate first. Args: jobName: Unique name for the job (letters, digits, '_' and '-'). type: Tool name — see getAvailableTools(). settings: Job settings object — see getJobSchema(jobType=type).
validateJob
Validate a pipeline run WITHOUT executing or saving anything. ALWAYS call this before submitPipeline — a pipeline run costs real compute, and most authoring mistakes are caught here. Takes the arguments submitPipeline validates: the graph (`pipeline` or `templateId`+`version`), `bindings`, `settings` and `name`. submitPipeline's remaining arguments — `runName`, `project`, `idempotencyKey` — name and tag the run rather than describing it, and the validator does not read them, so passing them here would change nothing. If `validationUnavailable` comes back, the validator could not be reached and your pipeline was NOT judged — retry a couple of times rather than editing the graph, and if it fails identically every time, report it instead of looping. Returns { valid, errors[] } at HTTP 200 for both outcomes; read `valid`. Each error carries a stable `code`, the `node` it belongs to, the `field` at fault where there is one, and a `message` telling you how to fix it. Codes you will actually hit include `required-field-unset` (often a residue selection that belongs in the binding, not in settings), `input-unbound`, `tool-unknown`, `setting-invalid`, `param-out-of-range`, `chain-incompatible` and `molecule-class-incompatible` (the upstream tool does not produce what the downstream tool consumes). `valid: true` is NOT a guarantee that submitPipeline will accept the run. Known gaps, measured: it does not confirm the tools exist, and it does not check that a bound molecule group exists or has anything in it. So if submitPipeline rejects a graph this approved, believe submitPipeline — re-validating will not reproduce the error, and the fix is in what it names (usually a tool name or a group id), not in the graph's shape. Fix the reported errors and re-validate until `valid` is true. Args: name: A name for the pipeline/run. Required. bindings: One entry per input node, keyed by the input node's id. Give molecules exactly one way: {"group": "<groupId>"} for an existing molecule group, or raw values — {"sequences": [...]}, {"smiles": [...]}, {"pdbs": ["<uploadedPath>"]}, {"sdfs": [...]} — and the server creates the group for you. A file input takes {"file": "<path>"}. Add "residuesByChain": {"A": "1-76"} where the template's residueFields ask for one, and "chainMapping" only when your chain IDs differ from the template's reference chains. pipeline: INLINE mode — a full pipeline IR. Call getPipelineSchema() first. Mutually exclusive with templateId. templateId: REFERENCE mode — an existing template to run. version: REFERENCE mode only — a version handle like "v2". settings: REFERENCE mode only — per-node overrides {nodeId: {settingKey: value}}, limited to the settings the template's author marked editable. Returns: JSON: { valid: bool, errors: [{code, severity, node, field, message}] }.
validatePipeline
How do I improve a ChatGPT Plugin's discoverability?
The levers are the listing surface agents actually read: names, descriptions, keywords, tool metadata, and registry health. Which lever matters depends on where discovery breaks, which is what continuous measurement shows.
What are Tamarind Bio alternatives on ChatGPT?
As of 2026-09-29, Tamarind Bio competes with Amass, Anida Clinical Trials, CAIRO Medicines, Convoke, DailyMed, DrugBank, Glow Research Docs, Korea Drug Info, Life Analytics IAS, NyquistAI, openFDA, Rhizome AI, RxNorm in ChatGPT Biomedical & Pharma Research Intelligence, ranked by public Discoverability Score.
Where is this profile measured?
This profile uses the geography attached to the latest public registry snapshot: US. Locale tags are intentionally omitted.