Build. Create. Grow.

Preparing your workspace

Skip to content
SONARA IndustriesSONARA One
Menu
Research Lab

Governed Hugging Face model and dataset catalog

Curated models, datasets, specifications, licensing posture, runtime placement, and staged SONARA adoption decisions.

Curated scope

33 Hugging Face models, datasets, and platform specifications are classified for SONARA.

Adoption state

4 resources are pilot-ready; 2 are blocked from commercial use; all execution remains disabled by default.

Supply-chain boundary

Model files, dataset files, custom code, and remote installers are never executed by this public catalog. Production use requires an isolated worker, pinned revision, license review, and human approval.

BGE Small English v1.5: pilot ready

Type: model. Task: text-embeddings. License: MIT. Commercial use: permissive. Runtime: embedding worker. Product fit: All studios, Files & Records, Search. Capabilities: semantic search, similarity, RAG indexing. Next: Pilot on a CPU-isolated Text Embeddings Inference worker against SONARA search evaluations.

Qwen3 Embedding 0.6B: evaluation

Type: model. Task: text-embeddings. License: Apache-2.0. Commercial use: permissive. Runtime: embedding worker. Product fit: All studios, Multilingual search, Research Lab. Capabilities: multilingual embeddings, long-text retrieval, semantic search. Next: Evaluate only if multilingual retrieval quality materially exceeds the smaller BGE pilot.

BGE Reranker v2 M3: pilot ready

Type: model. Task: text-ranking. License: Apache-2.0. Commercial use: permissive. Runtime: reranker worker. Product fit: Search, Research Lab, Files & Records. Capabilities: query-document reranking, multilingual relevance scoring. Next: Run offline retrieval benchmarks after the embedding pilot; do not place it in the synchronous web process.

Qwen3 4B: evaluation

Type: model. Task: text-generation. License: Apache-2.0. Commercial use: permissive. Runtime: text generation worker. Product fit: Business Builder, Creator Studio, Growth Studio, Private model mode. Capabilities: draft generation, structured extraction, reasoning assistance. Next: Evaluate in an isolated model gateway with deterministic SONARA validation and no production write tools.

Phi-4 Mini Instruct: review required

Type: model. Task: text-generation. License: MIT. Commercial use: permissive. Runtime: text generation worker. Product fit: Private model mode, Internal assistance, Multilingual drafting. Capabilities: instruction following, long-context drafting, multilingual text. Next: Security-review custom code before any isolated evaluation; otherwise prefer a standard-architecture model.

Granite Docling 258M: pilot ready

Type: model. Task: document-parsing. License: Apache-2.0. Commercial use: permissive. Runtime: document worker. Product fit: Business Builder, Files & Records, Invoices, Reports. Capabilities: OCR, layout parsing, tables, formulas, document structure. Next: Pilot receipt, invoice, menu, and contract extraction in a queue-backed document worker.

Whisper Large v3 Turbo: pilot ready

Type: model. Task: automatic-speech-recognition. License: MIT. Commercial use: permissive. Runtime: speech worker. Product fit: Creator Studio, Meeting imports, Voice interface, Accessibility. Capabilities: speech transcription, multilingual ASR, timestamps. Next: Pilot asynchronous transcription on user-uploaded audio with consent records and deletion controls.

Kokoro 82M: review required

Type: model. Task: text-to-speech. License: Apache-2.0. Commercial use: permissive. Runtime: tts worker. Product fit: Voice interface, Accessibility, Creator previews. Capabilities: text-to-speech, multilingual voices, low-cost local synthesis. Next: Prefer an ONNX or reviewed safe-weight build and add voice-consent and provenance controls before pilot use.

GLiNER Multilingual PII v1: review required

Type: model. Task: pii-detection. License: Apache-2.0. Commercial use: permissive. Runtime: pii worker. Product fit: Files & Records, Support, Meeting imports, Privacy. Capabilities: PII entity detection, multilingual redaction assistance. Next: Evaluate as a redaction assistant with mandatory rule-based checks and human review.

SigLIP 2 Base 224: evaluation

Type: model. Task: image-text-retrieval. License: Apache-2.0. Commercial use: permissive. Runtime: vision worker. Product fit: Creator Studio, Asset search, Catalog tagging. Capabilities: image-text embeddings, zero-shot image classification, asset retrieval. Next: Evaluate for creator-asset search and duplicate discovery, not identity or eligibility decisions.

CLAP HTSAT Unfused: review required

Type: model. Task: audio-text-embeddings. License: Apache-2.0. Commercial use: permissive. Runtime: audio embedding worker. Product fit: Creator Studio, Audio search, Sound tagging. Capabilities: audio-text similarity, audio embeddings, semantic sound search. Next: Evaluate on SONARA-owned audio only after safe serialization or isolated loading review.

FLUX.1 Schnell: review required

Type: model. Task: text-to-image. License: Apache-2.0. Commercial use: permissive. Runtime: image generation worker. Product fit: Creator Studio, Marketing concepts, Draft artwork. Capabilities: text-to-image generation, rapid concept rendering. Next: Use a managed or isolated GPU endpoint only after gated-access, cost, rights, and moderation review.

Phi-4 Multimodal Instruct: review required

Type: model. Task: multimodal-assistance. License: MIT. Commercial use: permissive. Runtime: multimodal worker. Product fit: Research Lab, Document and audio experiments, Accessibility. Capabilities: speech recognition, speech summarization, visual question answering, multimodal text generation. Next: Keep research-only until custom-code review shows a material advantage over specialized workers.

Stable Audio Open 1.0: license review

Type: model. Task: text-to-audio. License: Stability AI Community License. Commercial use: conditional. Runtime: audio generation worker. Product fit: Creator Studio, Sound-effect concepts. Capabilities: text-to-audio, short music and sound generation. Next: Do not enable until counsel approves the current license and required product attribution.

MusicGen Small: blocked commercial

Type: model. Task: text-to-music. License: CC-BY-NC-4.0 model weights. Commercial use: noncommercial. Runtime: audio generation worker. Product fit: Research Lab only. Capabilities: text-to-music, audio-conditioned music generation. Next: Retain only as a research comparison; select a commercially compatible alternative for production.

MERT v1 95M: blocked commercial

Type: model. Task: music-understanding. License: CC-BY-NC-4.0. Commercial use: noncommercial. Runtime: audio embedding worker. Product fit: Research Lab only. Capabilities: music feature extraction, music representation learning. Next: Use only for offline comparison on authorized audio, or replace with a permissively licensed model.

BANKING77: evaluation

Type: dataset. Task: intent-classification. License: CC-BY-4.0. Commercial use: permissive with attribution. Runtime: dataset evaluation. Product fit: Business Builder, Support, Voice interface. Capabilities: fine-grained customer-service intent evaluation. Next: Map a subset of intents to SONARA support taxonomy and use it as a benchmark, not production customer data.

MInDS-14: evaluation

Type: dataset. Task: spoken-intent-detection. License: CC-BY-4.0. Commercial use: permissive with attribution. Runtime: dataset evaluation. Product fit: Voice interface, Business Builder, Accessibility. Capabilities: multilingual spoken-intent evaluation, ASR and intent testing. Next: Use selected language subsets to benchmark voice intent routing and error rates.

FLEURS: evaluation

Type: dataset. Task: multilingual-speech-evaluation. License: CC-BY-4.0. Commercial use: permissive with attribution. Runtime: dataset evaluation. Product fit: Voice interface, Accessibility, Research Lab. Capabilities: multilingual ASR evaluation, language coverage testing. Next: Use small streamed validation subsets for supported-language ASR benchmarking.

MMLU: evaluation

Type: dataset. Task: general-knowledge-evaluation. License: MIT. Commercial use: permissive. Runtime: dataset evaluation. Product fit: Research Lab, Model gateway evaluation. Capabilities: multi-domain multiple-choice evaluation. Next: Add a small reproducible evaluation slice for candidate text models.

GSM8K: evaluation

Type: dataset. Task: reasoning-evaluation. License: MIT. Commercial use: permissive. Runtime: dataset evaluation. Product fit: Research Lab, Formula system evaluation. Capabilities: multi-step arithmetic reasoning evaluation. Next: Use test-only evaluation to compare model reasoning with SONARA deterministic formulas.

FineWeb-Edu: research only

Type: dataset. Task: language-model-training-corpus. License: ODC-By-1.0 plus Common Crawl terms. Commercial use: conditional. Runtime: dataset evaluation. Product fit: Research Lab only. Capabilities: large-scale educational web corpus. Next: Catalog only; do not ingest until a separate data-governance and training program exists.

FUNSD LayoutLMv2 Variant: blocked license review

Type: dataset. Task: form-understanding-evaluation. License: Unverified in retrieved dataset card. Commercial use: unknown. Runtime: dataset evaluation. Product fit: Document worker evaluation. Capabilities: annotated form layout evaluation. Next: Verify the authoritative source and license before any use.

Model Cards: spec adopted

Type: spec. Task: governance. License: Documentation/reference specification. Commercial use: metadata only. Runtime: platform spec. Product fit: Platform governance. Capabilities: intended-use documentation, limitations, training data, evaluation metadata. Next: Require a SONARA model card for every approved model revision.

Dataset Cards: spec adopted

Type: spec. Task: data-governance. License: Documentation/reference specification. Commercial use: metadata only. Runtime: platform spec. Product fit: Platform governance. Capabilities: license metadata, bias documentation, dataset structure, responsible-use context. Next: Require a dataset card and license record before ingestion.

Safetensors: spec adopted

Type: spec. Task: model-serialization. License: Documentation/reference specification. Commercial use: metadata only. Runtime: platform spec. Product fit: Platform governance. Capabilities: non-pickle tensor loading, zero-copy reads, partial tensor access. Next: Prefer Safetensors and reject unreviewed pickle weights in production workers.

Hub Security Scanning: spec adopted

Type: spec. Task: supply-chain-security. License: Documentation/reference specification. Commercial use: metadata only. Runtime: platform spec. Product fit: Platform governance. Capabilities: malware scanning, pickle import scanning, secret scanning, signed commits. Next: Record scan state, source owner, revision, file hashes, and signature status before deployment.

Hub OpenAPI and REST API: spec adopted

Type: spec. Task: metadata-integration. License: Documentation/reference specification. Commercial use: metadata only. Runtime: platform spec. Product fit: Platform governance. Capabilities: model metadata, dataset metadata, repository revisions, webhooks. Next: Use read-only metadata calls first; keep write scopes disabled.

Transformers.js: spec adopted

Type: spec. Task: browser-inference. License: Documentation/reference specification. Commercial use: metadata only. Runtime: platform spec. Product fit: Platform governance. Capabilities: ONNX browser inference, NLP, vision, audio. Next: Allow only small reviewed ONNX models; never ship customer secrets or unrestricted model downloads to the browser.

Text Embeddings Inference: spec adopted

Type: spec. Task: embedding-serving. License: Documentation/reference specification. Commercial use: metadata only. Runtime: platform spec. Product fit: Platform governance. Capabilities: dynamic batching, Safetensors loading, OpenTelemetry, Prometheus metrics. Next: Use TEI as the preferred isolated embedding service for the pilot.

Inference Endpoints: spec adopted

Type: spec. Task: managed-inference. License: Documentation/reference specification. Commercial use: metadata only. Runtime: platform spec. Product fit: Platform governance. Capabilities: managed engines, autoscaling, scale-to-zero, private endpoints. Next: Compare managed endpoint cost and security against self-hosted workers before production.

Datasets Streaming: spec adopted

Type: spec. Task: dataset-access. License: Documentation/reference specification. Commercial use: metadata only. Runtime: platform spec. Product fit: Platform governance. Capabilities: iterable streaming, partial exploration, large-dataset access. Next: Stream bounded evaluation subsets and enforce sample-count and byte limits.

Croissant Dataset Metadata: spec adopted

Type: spec. Task: dataset-interoperability. License: Documentation/reference specification. Commercial use: metadata only. Runtime: platform spec. Product fit: Platform governance. Capabilities: JSON-LD metadata, schema.org dataset description, Parquet distributions. Next: Store Croissant metadata snapshots for approved dataset revisions.

Command
Experience settings