# SEA-LION

*Built for Southeast Asia, by Southeast Asia*

SEA-LION (Southeast Asian Languages in One Network) is a family of open-source Large Language Models (LLMs) that better understands Southeast Asia’s (SEA) diverse contexts, languages, and cultures.

It is an open-source project anchored by the Products Pillar of AI Singapore. Our work in SEA-LION aims to create LLMs that cater to under-represented population groups and low resource languages in the SEA region. You can [read more about our motivations for SEA-LION here](/overview/readme/why_sea-lion).

This site provides information and resources on SEA-LION, including how to access the models, hosting, and how-to guides.

## Key Features of SEA-LION

| Model Collection                                | Size       | Context Length | Training Strategy                                      | Available in                              |
| ----------------------------------------------- | ---------- | -------------- | ------------------------------------------------------ | ----------------------------------------- |
| [**SEA-LION v4.5**](/models/sea-lion-v4.5)      | E2B        | 128K           | SFT¹ of gemma-4-E2B-it (Latest)                        | Instruct, GGUF                            |
|                                                 | 27B        | 262K           | SFT¹ of Qwen3.6-27B (Latest)                           | Instruct, SpecDecoder, GGUF               |
| [**SEA-LION v4**](/models/sea-lion-v4)          | 4B         | 128K           | SFT¹ of Gemma 3 4B IT                                  | VLM                                       |
|                                                 | 8B         | 65K            | SFT¹ of Apertus-8B-Instruct-2509                       | Instruct, Reasoning                       |
|                                                 | 4B, 8B     | 256K           | SFT¹ of Qwen3-VL-4B-Instruct, Qwen3-VL-8B-Instruct     | VLM                                       |
|                                                 | 27B        | 128K           | CPT² of Gemma 3 27B IT                                 | Base, Instruct, GGUF, NVFP4, FP8\_Dynamic |
|                                                 | 27B        | 128K           | SFT¹ of Gemma-SEA-LION-v4-27B-IT                       | VLM                                       |
|                                                 | 32B        | 128K           | SFT¹ of Qwen3-32B                                      | Instruct, GPTQ-4BIT, GPTQ-8BIT            |
| [**SEA-LION v3.5**](/models/sea-lion-v3.5)      | 8B         | 128K           | SFT¹ of Llama-SEA-LION-v3-8B-IT                        | Reasoning, GGUF                           |
|                                                 | 70B        | 128K           | SFT¹ of Llama-SEA-LION-v3-70B-IT                       | Reasoning, GGUF                           |
| [**SEA-LION v3**](/models/sea-lion-v3)          | 9B         | 8192           | CPT² of Gemma2                                         | Base, Instruct, GGUF                      |
|                                                 | 8B         | 128K           | CPT² of Llama 3.1 8B                                   | Base, Instruct, GGUF                      |
|                                                 | 70B        | 128K           | CPT² of Llama 3.1 70B                                  | Base, Instruct, GGUF                      |
| [**SEA-LION v2**](/models/sea-lion-v2)          | 8B         | 8192           | CPT² of Llama3                                         | Base, Instruct, GGUF                      |
| [**SEA-LION v1**](/models/sea-lion-v1)          | 3B         | 2048           | Pre-training from scratch                              | Base                                      |
|                                                 | 7B         | 2048           | Pre-training from scratch                              | Instruct                                  |
| [**SEA-LION Embedding**](/models/sea-embedding) | 300M, 600M | 8K             | ModernBERT from scratch                                | Base-Embedding, Checkpoints               |
|                                                 | 300M, 600M | 8K             | SFT¹ of SEA-LION-ModernBERT-Embedding                  | Tuned-Embedding                           |
|                                                 | 600M       | 512            | SFT¹ of E5-Large                                       | Tuned-Embedding                           |
| [**SEA-GUARD**](/models/sea-guard)              | 4B, 8B     | 128K           | SFT¹ of Qwen-SEA-LION-v4-4B-VL, Qwen-SEA-LION-v4-8B-VL | VLM                                       |
|                                                 | 8B         | 128K           | SFT¹ of Llama-SEA-LION-v3-8B-IT                        | Instruct                                  |
|                                                 | 12B        | 128K           | SFT¹ of Gemma 3 12B IT                                 | VLM                                       |

¹ Supervised Fine-Tuning

² Continued Pre-Training

## Performance and Benchmarks

SEA-LION has seen:

* In v1, ability to outperform most models based on SEA-HELM (Southeast Asian Holistic Evaluation of Language Models) when it was released
* In v2, outperformance for SEA tasks, while retaining credible performance on standard (English) benchmarks
* In v2.1, key improvements in conversational abilities across SEA languages, while providing more helpful and contextually appropriate responses to user prompts
* In v3, outperforms similar sized open source models, and even some larger models in both general and SEA capabilities
* In v3.5, ability to handle reasoning tasks, with the versatility of handling general tasks as well while maintaining similar performance with state-of-the-art models.
* In v4, our first multimodal SEA-LION models, extending capabilities beyond text to handle image + text inputs with massive 256K native context windows and specialized regional OCR, while continuing our focus on Southeast Asian languages, culture, and use cases.
* In v4.5, rapid specialization of state-of-the-art open foundation models via knowledge distillation and model merging, delivering high-capacity reasoning and agentic tool-use capabilities alongside speed-optimized booster configurations for low-latency production deployment.

SEA-LION-Embedding: The Vector Foundation

The SEA-LION-Embedding suite provides the semantic foundation for the ecosystem. Released in March 2026, these models represent a significant leap in regional retrieval performance. Tested on the SEA-BED (Southeast Asia Embedding Benchmark), which utilizes human-curated native data rather than machine translations, our embeddings consistently set new state-of-the-art records for 10 regional languages across retrieval, reranking, and semantic textual similarity tasks.

SEA-Guard: The Protector

Building on the sophisticated reasoning and multimodal foundation laid by SEA-LION v4, we are proud to introduce the first generation of SEA-Guard. Released on 4 Feb 2026, SEA-Guard is the dedicated safety counterpart to the SEA-LION family. As our foundational models gained the power to "see" (Multimodality in v4) and "think" (Reasoning in v3.5), the need for a culturally attuned safety layer became paramount.

We use a holistic approach to evaluation, including not just traditional Natural Language Processing (NLP) benchmarking tasks (such as sentiment analysis and question answering) but also [meticulously handcrafted linguistic and cultural diagnostic tests tailored to Southeast Asia](https://arxiv.org/abs/2309.06085v2).

Visit our [Leaderboard](https://leaderboard.sea-lion.ai/) for more detailed breakdown on:

1. How SEA-LION compares to other available models along different metrics
2. What SEA-HELM is and the four key capabilities it is evaluated on: English performance, Proficiency in SEA chat, Instruction-following and Linguistic tasks
3. What each of these globally recognized metrics mean under SEA-HELM

## Licensing

**Transparent and Open Source**

We have benefited greatly from the open-source community and believe that efforts to better represent our region will similarly be well served by open-source efforts.

All SEA-LION releases will therefore embrace an open-source ethos under the MIT license as much as possible; however, the exact licensing terms may vary depending on the underlying base model’s restrictions or requirements. For instance, if the model leverages Meta’s Llama3 codebase, it may be bound by the [Llama3 License](https://huggingface.co/meta-llama/Meta-Llama-3-8B/blob/main/LICENSE), which places certain restrictions on commercial use. Similarly, the Gemma-based variants may carry different terms. Users should always refer to the Hugging Face model card of each specific SEA-LION model for the most accurate, up-to-date license information.

SEA-LION will also be open and transparent in the following areas throughout this guide:

1. Pre-Training data
2. Model training code
3. Fine-Tuning data
4. Evaluation benchmarks

## Community

We welcome contributions to SEA-LION! Check out the [contributing guide](/overview/readme/contributions) to get started.

Some ways to contribute:

* Report bugs and issues
* Enhance the documentation
* Add more model evaluation tasks and metrics
* Train versions of the model in more SEA languages

Check out our [collaborations guide](/overview/readme/contributions) also, for possible ways to further enhance and expand the capabilities of SEA-LION together.

## To Cite SEA-LION

If you use SEA-LION in your work, please cite it as:

```bibtex
@misc{sea_lion_2024,
  title={SEA-LION (Southeast Asian Languages In One Network): A Family of Large Language Models for Southeast Asia},
  author={AI Singapore},
  year={2024},
  howpublished={\url{https://github.com/aisingapore/sealion}}
}
```

If you are using SEA-LION v3 for your work, please cite it as:

```bibtex
@misc{2504.05747,
      title={SEA-LION: Southeast Asian Languages in One Network},
      author={Raymond Ng and Thanh Ngan Nguyen and Yuli Huang and Ngee Chia Tai and Wai Yi Leong and Wei Qi Leong and Xianbin Yong and Jian Gang Ngui and Yosephine Susanto and Nicholas Cheng and Hamsawardhini Rengarajan and Peerat Limkonchotiwat and Adithya Venkatadri Hulagadri and Kok Wai Teng and Yeo Yeow Tong and Bryan Siow and Wei Yi Teo and Wayne Lau and Choon Meng Tan and Brandon Ong and Zhi Hao Ong and Jann Railey Montalan and Adwin Chan and Sajeban Antonyrex and Ren Lee and Esther Choa and David Ong Tat-Wee and Bing Jie Darius Liu and William Chandra Tjhi and Erik Cambria and Leslie Teo},
      year={2025},
      eprint={2504.05747},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2504.05747},
}
```

## Acknowledgements

AI Singapore is a national programme supported by the National Research Foundation, Singapore and hosted by the National University of Singapore. Any opinion, finding, conclusion or recommendation expressed in this material are those of the author(s) and do not reflect the views of National Research Foundation, Singapore, or the National University of Singapore.

We also grateful for the support of the Infocomm Media Development Authority (IMDA) of Singapore.

SEA-LION would not be possible without a growing list of Singapore, regional, and international collaborators. Please see our website for more details.

## Contact

If you have questions, comments, or issues, please open a GitHub issue or contact us via <sealion@aisingapore.org>.


# Motivations

Large Language Models (LLMs) are a type of artificial intelligence model designed to understand and generate human language. Recent developments in LLMs have showcased remarkable capabilities in understanding and generating human language, with applications spanning translation, summarization, coding assistance, question answering, and more.

Many existing LLMs, however, are trained upon massive internet-based datasets, which often has disproportionately large influences from western, industrialized, rich, educated, and democratic (WIRED) societies, as people from non-WIRED societies are less likely to be literate, to use the Internet, and to have their output easily accessed. Such an imbalance in the training data can lead to model outputs that\
display strong bias in terms of cultural values, political beliefs and social attitudes.

LLMs trained on predominantly WIRED-centric content risk neglecting the linguistic and cultural diversity inherent in non-WIRED populations. These biases become evident not only in examples like cultural references or local idioms, but also in more critical domains such as decision-making, social attitudes and moral reasoning, and which can vary significantly across global communities. By overlooking these variations, mainstream LLMs may inadvertently perpetuate inaccurate assumptions or exclude large segments of the global population.

Our work in SEA-LION, now part of Singapore’s National Multi-Modal Large Language Model project, aims to address these disparities by creating LLMs that cater to under-represented population groups and low resource languages in the SEA region.

SEA-LION is trained on more content produced in Southeast Asian languages like Thai, Vietnamese and Bahasa Indonesia to ensure better representation in data and alignment compared to Western or Chinese models. SEA-LION models understand nuances in SEA languages and demonstrate greater awareness of cultural context specific to the region.

This lowers the bar for governments, industries, and academia that seek LLM solutions that fit local languages and reflect local cultural norms, since WIRED-centric models can pose langauge barriers and misalign with local sensibilities in the SEA region.


# Contributing

Launched in May 2017, AI Singapore is a national R\&D programme in AI – the very first of its kind on a national level. Through SEA-LION, we aim to support **under-represented population groups and low-resource languages** in Southeast Asia. We welcome collaborations that further enhance our work and help our shared mission of democratizing AI.

We invite researchers, developers, language enthusiasts, and organizations to collaborate in **enhancing and expanding** the capabilities of our models. This page outlines the various ways you can work with us—whether it’s **joint model building, research partnerships, code contributions, documentation improvements**, or **data sharing**.

We appreciate your interest in contributing to SEA-LION, and we look forward to collaborating with you!

## 1. Collaboration Opportunities

We offer various partnership modalities to suit different needs and interests, with a focus on open-sourcing results to benefit the wider community. For collaborations other than a self-serve model, an NDA/MOU/CA may have to be signed by both parties.

### Self-Service Model Usage

*Using our model (self-service)*

SEA-LION models are open-source so do feel free to download them for your own needs. If SEA-LION manages to be helpful to you/your team, we would love to hear about it and share your successes with our wider commnunity. Send us a note at <sealion@aisingapore.org>.

### Joint Model Building

*Building models together, Data contributions*

Have a unique dataset or domain expertise? Contribute data that we can incorporate into future SEA-LION releases. Collaborate with us to train specialized or improved versions of SEA-LION that will be open-sourced, benefiting both you and the broader community.

### Research Partnerships

If there are areas of research topics where you think that AISG and yourself can work together to work on a common research agenda that can further the work of SEA-LION, please feel free to reach out to us as well. We will be open to joint publications arising from these collaborations too.

If you are interested in any form of collaboration, please feel free to reach out to <sealion@aisingapore.org>

## 2. Community Contributions

In addition to formal partnerships, there are many individual and community-driven ways to help SEA-LION grow. Whether it’s fixing bugs, improving documentation, or proposing new features, as a developer, researcher, or just an enthusiast, your input matters.

Before you begin, please take a moment to review this guide, which outlines the contribution process, code of conduct, and how to get help if needed.

### Getting Started

Before you start contributing, please ensure you have the following:

* A GitHub account: If you don't have one, you can [create an account here](https://github.com/join).
* Familiarity with Git: You'll need to know the basics of Git for version control.

### Reporting Bugs

If you encounter any issues, bugs, or unexpected behavior while using SEA-LION, please help us by [reporting them](https://github.com/aisingapore/sealion/issues).

To report a bug:

1. Check if the issue has already been reported by searching the [GitHub Issues](https://github.com/aisingapore/sealion/issues) page.
2. If not, create a new issue with a descriptive title and detailed description of the problem you encountered.
3. Include relevant information such as your operating system, Python version, and any error messages.

### Suggesting Enhancements

We appreciate your suggestions for improving SEA-LION. If you have an idea for an enhancement or new feature, follow these steps:

1. Check if your suggestion has already been proposed in the [GitHub Issues](https://github.com/aisingapore/sealion/issues) section.
2. If not, create a new issue with a clear and concise title and a detailed description of your suggestion.
3. Include any relevant context or examples to illustrate the enhancement's value.

### Code of Conduct

Please review and adhere to our [Code of Conduct](/overview/readme/code_of_conduct). We expect all contributors and community members to treat each other with respect and kindness.

### Get Help

If you have questions, need assistance, or want to discuss contributions further, please feel free to contact us or open an issue for discussion.


# Code of Conduct

## Our Pledge

We, the community of contributors and users of SEA-LION, pledge to create a welcoming and inclusive environment for everyone. We are committed to fostering a respectful and harassment-free space where diverse ideas and perspectives can thrive.

## Expected Behavior

To contribute to a positive and inclusive atmosphere, we expect all participants, including contributors, users, and maintainers, to:

1. Be respectful and considerate: Treat others with kindness, respect, and empathy. Recognize and embrace diversity in backgrounds, experiences, and opinions.
2. Be inclusive: Welcome and support people of all backgrounds, identities, and abilities. Avoid any form of discrimination or exclusionary behavior.
3. Listen actively: Pay attention to others' ideas, experiences, and feedback. Be open to constructive criticism and different points of view.
4. Show empathy: Understand that people may have different cultural norms, communication styles, and perspectives. Be patient and considerate when engaging with others.
5. Resolve conflicts constructively: Disagreements and conflicts are natural, but we encourage participants to address them in a respectful and solution-oriented manner. Avoid personal attacks and name-calling.
6. Use clear and inclusive language: Use language that is respectful, inclusive, and considerate of all participants. Avoid offensive, derogatory, or discriminatory language.

## Unacceptable Behavior

The following behaviors are considered unacceptable and will not be tolerated within the SEA-LION community:

1. Harassment: Any form of harassment, including but not limited to offensive comments, slurs, intimidation, or unwelcome advances, is strictly prohibited.
2. Discrimination: Discriminatory actions or comments based on race, ethnicity, nationality, gender, gender identity, sexual orientation, disability, religion, age, or any other characteristic will not be tolerated.
3. Hate speech: Hate speech, promoting violence, or advocating harm towards individuals or groups based on their identity is not allowed.
4. Personal attacks: Engaging in personal attacks, insults, or trolling of others within the community is unacceptable.
5. Disruptive behavior: Deliberate disruption of discussions, events, or community activities is discouraged.

## Reporting Violations

If you witness or experience behavior that violates this code of conduct, please report it promptly to the project maintainers by contacting [sealion@aisingapore.org](https://github.com/aisingapore/sealion/blob/main/overview/sealion@aisingapore.org)

All reports will be treated with confidentiality, and the project maintainers will take appropriate action as necessary to address violations. We are committed to providing a safe and welcoming environment for all participants.

## Enforcement

Enforcement of this code of conduct will be carried out in a fair and just manner. Depending on the severity and frequency of violations, consequences may include warnings, temporary or permanent bans from the community, or other appropriate actions.

## Attribution

This code of conduct is adapted from the [Contributor Covenant](https://www.contributor-covenant.org), version 2.0, available at <https://www.contributor-covenant.org/version/2/0/code\\_of\\_conduct.html>.


# FAQ

Common questions we receive at the SEA-LION inbox, grouped by topic. If your question is not covered here, write to us at <sealion@aisingapore.org> or use the Collaborate form on [sea-lion.ai](https://sea-lion.ai/).

## Getting started

### What is SEA-LION?

SEA-LION (Southeast Asian Languages In One Network) is a family of open-source large language models built for Southeast Asia's languages and cultural contexts. The models are developed by AI Singapore.

### How do I try SEA-LION?

You have a few options:

* Test the models directly in our [Playground](https://sea-lion.ai/).
* Chat with SEA-LION on our [Telegram bot](https://t.me/sealion_ai_bot).
* Download the models from [Hugging Face](https://huggingface.co/aisingapore).
* Inference the models through major cloud platforms (see "Running SEA-LION in production" below).

### Is SEA-LION free to use? Can I use it commercially?

Yes. SEA-LION is open-source and free to use, including for commercial purposes. You can download and self-host the models from Hugging Face, or access them through cloud platform partners.

## API access and rate limits

### How do I get an API key?

You can generate an API key through the SEA-LION Playground. Documentation is available at [docs.sea-lion.ai](https://docs.sea-lion.ai/).

### What are the API rate limits?

Our free APIs are intended for proof-of-concept (POC) development and are rate-limited (10 requests per minute). They are not designed for production workloads.

### Can I get a rate limit increase?

We are not able to raise the limits on the free POC tier. If you need higher throughput or production-grade reliability, inference SEA-LION through one of our cloud platform partners, where you can scale usage according to your needs.

### My API key expired, or I am hitting timeouts. What should I do?

The free API is rate-limited and meant for POC testing, so heavy or sustained use may hit limits. For continued or production use, move to a cloud platform partner. If you believe you are seeing an actual outage rather than a rate limit, write to us with the details.

### How do I run SEA-LION in production?

For production, inference SEA-LION through major cloud platforms rather than the free POC API. The models are available on hyperscalers including AWS (Bedrock) and Google Cloud (Vertex AI), among others. See our [inferencing guide](https://docs.sea-lion.ai/guides/inferencing) for the current list and setup steps.

### Which model and hardware should I use?

If you are running on limited GPU resources, use the quantised versions of the models. The model card on Hugging Face lists the available variants and their requirements.

## Languages and capabilities

### Which languages does SEA-LION support?

SEA-LION is trained on and supports the national languages of Southeast Asia. This includes (among others) Indonesian, Thai, Vietnamese, Malay, Tamil, Khmer, Burmese, Lao and Filipino. For the authoritative and most current list, see the model documentation.

### Does SEA-LION support \[Khmer / Tamil / Lao / Burmese]?

Yes. These are among the Southeast Asian languages SEA-LION is trained on. We are continually working with local and international partners to strengthen performance across all supported languages.

### Does SEA-LION support Tetum?

We are actively improving Tetum capability, working together with local and international partners. It is an area of ongoing development.

### Does SEA-LION support dialects such as Cantonese or Hokkien?

No. We evaluate and support the national languages of Southeast Asia. This does not include dialects such as Cantonese, Hokkien, Teochew, Hakka or other Singaporean dialects.

### What about Singlish, or speech-to-text and text-to-speech?

SEA-LION currently supports text and image (vision) inputs, and does not yet have a voice or audio model, though we do have plans to release text-to-speech (TTS).

### Does SEA-LION have a voice or audio model?

Not yet. SEA-LION today supports text and vision, and does not currently have a voice or audio model. We do have plans to release text-to-speech (TTS). Follow our website and AI Singapore's social channels to hear when it is released.

### What is the difference between SEA-LION and MERaLiON?

SEA-LION and MERaLiON are complementary models that address different needs and work in tandem to strengthen Singapore's national AI foundation. SEA-LION, built by AI Singapore, is an open-source, SEA-relevant large language model. It is multilingual, multicultural and multimodal (handling text and vision), and is built for general tasks such as instruction following, tool use and reasoning. MERaLiON is a speech-to-text model that prioritises empathetic learning and understanding, starting with Singlish and Singapore's multicultural context and extending to regional languages. It also supports speech-related tasks such as emotion recognition and spoken dialogue summarisation. MERaLiON is developed by A\*STAR's Institute for Infocomm Research (I²R).

The two leverage shared data, compute and research to build complementary text and speech models.

## Benchmarks and evaluation

### Where can I see how SEA-LION performs?

Our benchmark evaluations are published on the [SEA-LION Leaderboard](https://leaderboard.sea-lion.ai/).

### I am running SEA-HELM and seeing inconsistencies. What should I check?

Make sure you are using the latest version of the SEA-HELM evaluation code base, which is updated as we add languages and metrics. We recommend running with the default SEA-HELM configuration.

## Partnerships and collaboration

### How can I partner or collaborate with SEA-LION?

We welcome collaboration across the public sector, private sector, academia and the non-profit space. The best first step is to submit the Collaborate form on [sea-lion.ai](https://sea-lion.ai/) or email us. To help us route your request quickly, please include a short overview of your organisation, your use case or proposed collaboration, and the languages or capabilities you are interested in.

### I would like to contribute language data or annotation. Is that useful?

We are always glad to hear from people who can help strengthen SEA-LION's coverage and quality. Share details of the data or capability you can offer, and we will let you know if it aligns with our current priorities. We do review fit case by case, so not every offer will match an active need at a given time.

### Can AI Singapore/SEA-LION fund my project or research?

AI Singapore/SEA-LION is a non-profit initiative funded by Singapore's National Research Foundation (NRF). We are not able to provide grants or funding for external projects. You are very welcome to use the open-source models freely for your work.

### We are a cloud or hosting provider interested in hosting SEA-LION. Who do we talk to?

Reach out via the Collaborate form or email. Note that SEA-LION is already available on several major platforms, so do let us know what you have in mind and we can explore fit.

## Programmes

### What is the Pinnacle AI Industry Programme (PAIP) and how do I apply?

PAIP is a company-sponsored programme that helps organisations build applied AI capability within their teams. Because it is tailored to each organisation, we usually start with a short call to understand your team, your goals and the number of participants before sharing detailed materials. If you are interested, email us with your organisation and a brief description of your needs.

## Media and speaking

### I am a journalist with a media or interview request.

Please email us with details of your outlet, the focus of the story and your timeline. Media requests are handled together with AI Singapore's Marketing Communications team, and we will follow up to confirm.

### I would like to invite SEA-LION to speak at an event.

We are glad to receive speaking and panel invitations. Email us with the event details, format, date and audience, and we will assess fit and timing.

## Security

### How do I report a security vulnerability?

If you have identified a potential security issue, please email us at <sealion@aisingapore.org> with the details so we can route it to the right team.

## Staying in touch

### How do I keep up with new releases and events?

Follow AI Singapore on LinkedIn and our other social channels, and keep an eye on the SEA-LION website. That is where we announce new model releases, events and opportunities.


# SEA-LION v4.5 (Latest)

SEA-LION version 4.5, released in May 2026, is our latest collection of foundational, agentic, and multimodal models optimized for Southeast Asia. Utilizing advanced post-training methodologies—such as knowledge distillation and model merging—this suite delivers state-of-the-art regional fluency, precise tool use, and high computational efficiency.

## Gemma-SEA-LION-v4.5 (E2B Series)

Our highly efficient 4-billion parameter series built on Gemma 4. It is optimized for precise function-calling, structured JSON outputs, and autonomous agentic tool-use with minimal memory overhead.

* [Gemma-SEA-LION-v4.5-E2B-IT](/models/sea-lion-v4.5/gemma-sea-lion-v4.5)

  Refer to the detailed page to access model cards, usage scripts, and deployment configurations for the entire Gemma family:

  * Core Agentic Foundations

## Qwen-SEA-LION-v4.5 (27B Series)

Our high-capacity flagship causal and vision-language series built on Qwen3.6. It features a native 262K context window, robust multi-turn reasoning, and repository-level coding adapted for regional linguistic and cultural contexts.

* [Qwen-SEA-LION-v4.5-27B-IT](/models/sea-lion-v4.5/qwen-sea-lion-v4.5)

  Refer to the detailed page to explore full benchmarks, technical specifications, and configuration details for the entire Qwen family:

  * Base & Multimodal Foundations
  * High-Throughput Booster (SpecDecoder version)


# Gemma-SEA-LION-v4.5

*Last update: 2026-05-19*

SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

**Gemma-SEA-LION-v4.5-E2B-IT** built upon the gemma-4-E2B-it architecture with 2.3B effective (5.1B with embeddings). To ensure deep domain adaptation, the model underwent distillation from google/gemma-4-31B-it on an updated aisingapore/SEA-Instruct-2602, instilling multilingual and multicultural fluency across English and key SEA languages: Burmese, Indonesian, Filipino (Tagalog), Malay, Tamil, Thai, and Vietnamese.

Gemma-SEA-LION-v4.5-E2B-IT inherits the following features from Gemma 4:

* **Native System Prompt Support** – Gemma 4 introduces native support for the `system` role, enabling more structured and controllable conversations.
* **Reasoning** – Highly capable reasoning model, with configurable thinking modes.
* **Extended Multimodalities** – Processes Text, Image with variable aspect ratio and resolution support (all models), Video, and Audio (featured natively).
* **Optimized for On-Device** – designed for efficient local execution on laptops and mobile devices.
* **Enhanced Coding & Agentic Capabilities** – Achieves notable improvements in coding benchmarks alongside native function-calling support, powering highly capable autonomous agents.

## Model Details

### Model Description

SEA-LION stands for Southeast Asian Languages In One Network.

We performed post-training in English and SEA languages on `gemma-4-E2B-it`, a multimodal learning model using the Gemma 4 architecture, to create Gemma-SEA-LION-v4.5-E2B-IT.

For tokenization, the model employs the default tokenizer used in `gemma-4-E2B-it`.

* **Developed by:** AI Products Pillar, AI Singapore
* **Funded by:** Singapore NRF
* **Shared by:** AI Products Pillar, AI Singapore
* **Model type:** Transformer Decoder with Vision and Audio Encoder
* **Training Stage:** Post-Training (Logit Distillation & Model Merging)
* **Context length:** 128k
* **Language(s):** fine-tuned on Burmese, Indonesian, Filipino, Malay, Tamil, Thai, and Vietnamese
* **License:** [Apache-2.0](https://ai.google.dev/gemma/apache_2)
* **Finetuned from model:** <https://huggingface.co/google/gemma-4-E4B-it>

### Model Sources

* **Repository:** [SEA-LION v4.5 - an aisingapore Collection](https://huggingface.co/collections/aisingapore/sea-lion-v45)

## How to Get Started with the Model

### Download the Models

Gemma-SEA-LION-v4.5-E2B-IT models are available for download via the following channels: 🤗[HuggingFace SEA-LION v4.5 Collection](https://github.com/aisingapore/sealion/blob/main/models/sea-lion-v4.5/\(https:/huggingface.co/collections/aisingapore/sea-lion-v45\)/README.md)

| Model                           | Download                                                                                                                                               |
| ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Gemma-SEA-LION-v4.5-E2B-IT      | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-LION-v4.5-E2B-IT), [Ollama](https://ollama.com/aisingapore/Gemma-SEA-LION-v4.5-E2B-IT)      |
| Gemma-SEA-LION-v4.5-E2B-IT-GGUF | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-LION-v4.5-E2B-IT-GGUF), [Ollama](https://ollama.com/aisingapore/Gemma-SEA-LION-v4.5-E2B-IT) |

Use the code below to get started with the model with 🤗 Transformers libraries.

```
pip install -U transformers torch accelerate
```

```
from transformers import AutoProcessor, AutoModelForCausalLM
MODEL_ID = "aisingapore/Gemma-SEA-LION-v4.5-E2B-IT"
# Load model
processor = AutoProcessor.from_pretrained(MODEL_ID)
model = AutoModelForCausalLM.from_pretrained(
    MODEL_ID,
    dtype="auto",
    device_map="auto"
)
# Prompt
messages = [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Write a short joke about HDB flat in Singapore."},
]
# Process input
text = processor.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=False
)
inputs = processor(text=text, return_tensors="pt").to(model.device)
input_len = inputs["input_ids"].shape[-1]
# Generate output
outputs = model.generate(**inputs, max_new_tokens=1024)
response = processor.decode(outputs[0][input_len:], skip_special_tokens=False)
# Parse output
processor.parse_response(response)
```

## Training Details

The instruction fine-tuning text dataset comprises a collection of OSS & synthetic data, including approximately 8.54 million instruction-text pairs and advanced tool-calling instruction pairs specifically curated for the region.

### Training Data

🤗[aisingapore/SEA-Instruct-2602](https://huggingface.co/datasets/aisingapore/SEA-Instruct-2602)

### Training Regime

To maintain high efficiency, our post-training pipeline is centred entirely on knowledge distillation.

## Evaluation

### Testing Data, Factors & Metrics

We evaluated Gemma-SEA-LION-v4.5-E2B-IT on general language, multi-turn chat, instruction-following capabilities, and vision-language benchmarks.

#### Testing Data

**General language capabilities**

For the evaluation of general language capabilities, we employed the [SEA-HELM evaluation benchmark](https://arxiv.org/abs/2502.14301) across a variety of tasks. These tasks include Question Answering (QA), Sentiment Analysis (Sentiment), Toxicity Detection (Toxicity), Translation in both directions (Eng>Lang & Lang>Eng), Abstractive Summarisation (Abssum), Causal Reasoning (Causal), Natural Language Inference (NLI), Linguistic Diagnostics (LINDSEA), Cultural Knowledge (Kalahi) and Global MMLU Lite/Thai Exam.

**Instruction-following and Multi-turn Chat**

We evaluated the models on Instruction-following and Multi-turn Chat capabilities with SEA-IFEval (based on [IFEval](https://arxiv.org/abs/2311.07911)) and SEA-MTBench (based on [MT-Bench](https://arxiv.org/abs/2306.05685)) respectively. The two datasets were originally in English, the linguists and native speakers in the team worked together to filter, localise and translate the datasets into the respective target languages to ensure that the examples remained reasonable, meaningful and natural.

#### Factors

All evaluations were run with the model specific generation parameters defined in the model config. Each evaluation comprised of 8 runs with different seeds and the final results were averaged across these runs.

For all tasks, the model was expected to provide an answer tag from which the answer was automatically extracted. For tasks where options were provided, the answer should comprise one of the pre-defined options.

The evaluation was done **zero-shot** with native prompts on a sample of 100-1000 instances for each dataset.

* **SEA-IFEval:** Evaluates a model's ability to adhere to constraints provided in the prompt, for example beginning a response with a specific word/phrase or answering with a certain number of sections. Additionally, accuracy is normalised by the proportion of responses in the correct language (if the model performs the task correctly but responds in the wrong language, it is judged to have failed the task).
* **SEA-MTBench:** Evaluates a model's ability to engage in Multi-turn (2 turns) conversations and respond in ways that align with human needs. We use `gpt-oss-120b` as the judge model and compare against `gpt-oss-120b` as the baseline model. The metric used is the weighted win rate against the baseline model (i.e. average win rate across each category: Math, Reasoning, STEM, Humanities, Roleplay, Writing, Extraction).

#### Metrics

The following metrics were used for text capabilities:

| **Task**                       | **Metric**                        |
| ------------------------------ | --------------------------------- |
| Sentiment Analysis             | Accuracy                          |
| Extractive QA (ID, VI, TH, TA) | ChrF++                            |
| MCQ-QA (TL, MY, MS)            | Accuracy                          |
| Metaphor                       | Accuracy                          |
| Abstractive Summarisation      | Rouge-L                           |
| Translations                   | MetricX-24 score (with reference) |
| Causal Reasoning               | Accuracy                          |
| Natural Language Inference     | Accuracy                          |
| LINDSEA                        | Accuracy                          |
| Global MMLU Lite               | Accuracy                          |
| ThaiExam                       | Accuracy                          |
| Kalahi                         | Accuracy                          |
| SEA-IFEval                     | Accuracy                          |
| SEA-MTBench                    | Win rate against a reference      |

#### Results

For details on Gemma-SEA-LION-v4.5-E2B-IT performance, please refer to the <https://leaderboard.sea-lion.ai/>.

## Technical Specifications

### Model Architecture

The architecture is based on the dense model architecture of Gemma4 E2B, can be found in the [gemma-4-E2B-it Model Card](https://huggingface.co/google/gemma-4-E2B-it#dense-models).

## Uses

### Out-of-Scope Use

The model has not been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

### Bias, Risks, and Limitations

The model was not tested for robustness against adversarial prompting. It is important for users to be aware that our model exhibits certain limitations that warrant consideration. Like many LLMs, the model can hallucinate and occasionally generates irrelevant content, introducing fictional elements that are not grounded in the provided context. Users should also exercise caution in interpreting and validating the model's responses due to the potential inconsistencies.


# Qwen-SEA-LION-v4.5

Last update: 2026-05-19

**SEA-LION** is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

The **Qwen-SEA-LION-v4.5-27B-IT** sub-collection — comprising the standard high-fidelity model and its speed-optimized companion, the **27B-IT-SpecDecoder** — is built upon the Qwen3.6-27B dense architecture, a 27-billion parameter model featuring a hybrid Linear and Full Attention design. To ensure deep domain adaptation, both models underwent extensive distillation from Qwen/Qwen3.5-397B-A17B on the updated aisingapore/SEA-Instruct-2602 dataset. This instills native multilingual and multicultural fluency across English and key SEA languages (Burmese, Indonesian, Filipino, Malay, Tamil, Thai, and Vietnamese), with the SpecDecoder variant specifically engineered to maximize throughput and minimize inference latency in production environments.

**Qwen-SEA-LION-v4.5-27B-IT-SpecDecoder** is a draft model using speculative decoding method to employ a lightweight **block diffusion** model to draft multiple tokens in parallel trained from **Qwen-SEA-LION-v4.5-27B-IT**. This is the drafter model, which must be paired with aisingapore/Qwen-SEA-LION-v4.5-27B-IT.

Qwen-SEA-LION-v4.5-27B-IT inherits the following features from Qwen3.6:

* **Context Window (262K):** Reduce if facing OOM errors but keep ≥128K to preserve full reasoning capabilities.
* **Unified Vision-Language:** Early fusion training delivers good performance across multimodal reasoning, coding, and visual tasks.
* **Scalable RL:** Trained in million-agent environments for robust, real-world SEA adaptability.
* **Broad Linguistic Coverage:** Deeply specialized in SEA cultural nuances while supporting 201 languages globally.
* **Advanced Infrastructure:** Utilizes highly efficient multimodal training and asynchronous RL frameworks.
* **Agentic Coding:** High-precision handling of repository-level reasoning and frontend workflows.
* **Thinking Preservation:** Retains historical reasoning context to streamline iterative development and reduce compute overhead.

## Model Details

### Model Description

SEA-LION stands for Southeast Asian Languages In One Network.

We performed post-training in English and SEA languages on Qwen3.6-27B, a multimodal learning model using the Qwen3.6 architecture, to create Qwen-SEA-LION-v4.5.

For tokenization, the model employs the default tokenizer used in Qwen3.6.

* **Developed by:** AI Products Pillar, AI Singapore
* **Funded by:** Singapore NRF
* **Shared by:** AI Products Pillar, AI Singapore
* **Model type:** Causal Language Model with Vision Encoder
* **Training Stage:** Post-Training (Logit Distillation & Model Merging))
* **Context length:** 262k
* **Language(s):** fine-tuned on Burmese, Indonesian, Filipino, Malay, Tamil, Thai, and Vietnamese
* **License:** [Apache-2.0](https://github.com/QwenLM/Qwen3.6/blob/main/LICENSE)
* **Finetuned from model:** <https://huggingface.co/Qwen/Qwen3.6-27B>

SpecDecoder was Finedtuned from [z-lab/Qwen3.6-27B-DFlash](https://huggingface.co/z-lab/Qwen3.6-27B-DFlash) targeted to [Qwen-SEA-LION-v4.5-27B-IT](https://github.com/aisingapore/sealion/blob/main/models/sea-lion-v4.5/\(https:/huggingface.co/aisingapore/Qwen-SEA-LION-v4.5-27B-IT\))

### Model Sources

**Qwen-SEA-LION-v4.5-27B-IT** models are available for download via the following channels:

[HuggingFace SEA-LION v4.5 Collection](https://huggingface.co/collections/aisingapore/sea-lion-v45)

| Model                                 | Download                                                                                                                                             |
| ------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- |
| Qwen-SEA-LION-v4.5-27B-IT             | [HuggingFace](https://huggingface.co/aisingapore/Qwen-SEA-LION-v4.5-27B-IT), [Ollama](https://ollama.com/aisingapore/Qwen-SEA-LION-v4.5-27B-IT)      |
| Qwen-SEA-LION-v4.5-27B-IT-SpecDecoder | [HuggingFace](https://huggingface.co/aisingapore/Qwen-SEA-LION-v4.5-27B-IT-SpecDecoder)                                                              |
| Qwen-SEA-LION-v4.5-27B-IT-GGUF        | [HuggingFace](https://huggingface.co/aisingapore/Qwen-SEA-LION-v4.5-27B-IT-GGUF), [Ollama](https://ollama.com/aisingapore/Qwen-SEA-LION-v4.5-27B-IT) |

## How to Get Started with the Model

Use the code below to get started with the model with 🤗 Transformers libraries.

```
pip install "transformers>=4.57.0" accelerate vllm
```

```python
# ============================================================
# TEXT-ONLY INFERENCE example
# ============================================================

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

MODEL_ID = "aisingapore/Qwen-SEA-LION-v4.5-27B-IT"

# ── Load tokenizer ──
tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)

# ── Load model in bfloat16 across all available GPUs ──
# attn_implementation="sdpa" is safer for the hybrid DeltaNet arch;
# flash_attention_2 compatibility depends on your transformers version
model = AutoModelForCausalLM.from_pretrained(
    MODEL_ID,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    attn_implementation="sdpa",  # use sdpa for DeltaNet hybrid layers
)

# ── Message: text-only, same Malay query from original snippet ──
messages = [
    {
        "role": "user",
        "content": "Tolong carikan flat 4-bilik dekat Tampines, bajet bawah $500,000. "
                   "Nak tahu juga berapa anggaran pinjaman bulanan."
    }
]

# ── Apply chat template — text-only, thinking disabled ──
# enable_thinking=False → instruct/non-thinking mode
# Qwen3.6 does NOT support /no_think soft switch unlike Qwen3
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=False,      # hard-disable CoT thinking blocks
)

# ── Tokenize ──
inputs = tokenizer(text, return_tensors="pt").to(model.device)

# ── Generate — non-thinking mode params ──
# presence_penalty=1.5 is important for Qwen3.6 non-thinking mode
# to suppress repetition; not available in model.generate() directly,
# so use do_sample=True with the temperature/top_p/top_k trio
generated_ids = model.generate(
    **inputs,
    max_new_tokens=512,
    do_sample=True,
    temperature=0.7,    # non-thinking instruct mode
    top_p=0.80,
    top_k=20,
    # Note: presence_penalty requires vLLM/SGLang for full effect;
    # in transformers use repetition_penalty as a proxy
    repetition_penalty=1.1,
)

# ── Decode only newly generated tokens ──
output_ids = generated_ids[0][inputs["input_ids"].shape[1]:]
response = tokenizer.decode(output_ids, skip_special_tokens=True).strip()
print(response)
```

#### Tool Calling example

```python
# ============================================================
# TOOL CALLING (Transformers, local)
# ============================================================

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

MODEL_ID = "aisingapore/Qwen-SEA-LION-v4.5-27B-IT"

tokenizer = AutoTokenizer.from_pretrained(MODEL_ID)

model = AutoModelForCausalLM.from_pretrained(
    MODEL_ID,
    torch_dtype=torch.bfloat16,
    device_map="auto",
    attn_implementation="sdpa",
)

messages = [
    {
        "role": "user",
        "content": "Tolong carikan flat 4-bilik dekat Tampines, bajet bawah $500,000. "
                   "Nak tahu juga berapa anggaran pinjaman bulanan."
    }
]

tools = [
    {
        "type": "function",
        "function": {
            "name": "search_hdb_listings",
            "description": "Search for HDB flats available for sale",
            "parameters": {
                "type": "object",
                "properties": {
                    "location": {
                        "type": "string",
                        "description": "Town or area name"
                    },
                    "flat_type": {
                        "type": "string",
                        "description": "Flat type e.g. 3-room, 4-room, 5-room"
                    },
                    "max_price": {
                        "type": "number",
                        "description": "Maximum price in SGD"
                    }
                },
                "required": ["location", "flat_type"]
            }
        }
    },
    {
        "type": "function",
        "function": {
            "name": "calculate_mortgage",
            "description": "Calculate estimated monthly mortgage payment",
            "parameters": {
                "type": "object",
                "properties": {
                    "loan_amount": {
                        "type": "number",
                        "description": "Loan amount in SGD"
                    },
                    "interest_rate": {
                        "type": "number",
                        "description": "Annual interest rate as percentage"
                    },
                    "loan_tenure_years": {
                        "type": "integer",
                        "description": "Loan period in years"
                    }
                },
                "required": ["loan_amount"]
            }
        }
    }
]

# ============================================================
# apply_chat_template returns BatchEncoding with keys:
#   input_ids, attention_mask (and sometimes token_type_ids)
# ============================================================

inputs = tokenizer.apply_chat_template(
    messages,
    tools=tools,
    return_tensors="pt",
    return_dict=True,          # ← returns BatchEncoding dict with attention_mask
    add_generation_prompt=True,
    enable_thinking=False,     # disable CoT for structured tool call output
).to(model.device)

# ── Unpack BatchEncoding dict with ** — fixes the AttributeError ──
generated_ids = model.generate(
    **inputs,                  # ← unpack: passes input_ids + attention_mask
    max_new_tokens=512,
    do_sample=False,
)

# ── Decode only new tokens — slice off the prompt portion ──
output_ids = generated_ids[0][inputs["input_ids"].shape[1]:]
response = tokenizer.decode(output_ids, skip_special_tokens=True).strip()
print(response)

```

**Agentic Example:**

```python
# ============================================================
# NO-VLLM AGENTIC LOOP
# ============================================================

import os
import json
import re
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
from dotenv import load_dotenv

load_dotenv()

MODEL_ID = "aisingapore/Qwen-SEA-LION-v4.5-27B-IT"

print("Loading tokenizer...")
tokenizer = AutoTokenizer.from_pretrained(
    MODEL_ID,
    token=os.getenv("HF_TOKEN"),
)

print("Loading model across GPUs...")
model = AutoModelForCausalLM.from_pretrained(
    MODEL_ID,
    token=os.getenv("HF_TOKEN"),
    torch_dtype=torch.bfloat16,
    device_map="auto",
    attn_implementation="sdpa",
)

device_info = getattr(model, "hf_device_map", None) or str(model.device)
print(f"Model loaded. Device: {device_info}")

TOOLS = [
    {
        "type": "function",
        "function": {
            "name": "search_hdb_listings",
            "description": "Search for HDB flats available for sale",
            "parameters": {
                "type": "object",
                "properties": {
                    "location": {"type": "string", "description": "Town or area name"},
                    "flat_type": {"type": "string", "description": "e.g. 4-room"},
                    "max_price": {"type": "number", "description": "Max price in SGD"},
                },
                "required": ["location", "flat_type"],
            },
        },
    },
    {
        "type": "function",
        "function": {
            "name": "calculate_mortgage",
            "description": "Calculate estimated monthly mortgage payment",
            "parameters": {
                "type": "object",
                "properties": {
                    "loan_amount": {"type": "number", "description": "Loan amount SGD"},
                    "interest_rate": {"type": "number", "description": "Annual rate %"},
                    "loan_tenure_years": {"type": "integer", "description": "Loan years"},
                },
                "required": ["loan_amount"],
            },
        },
    },
]

def execute_tool(name: str, arguments: dict) -> str:
    """Mock tool executor — replace with real API calls."""
    if name == "search_hdb_listings":
        return json.dumps({
            "listings": [
                {
                    "address": "Blk 472 Tampines St 43",
                    "flat_type": arguments.get("flat_type"),
                    "resale_price": 488000,
                    "floor_area_sqm": 93,
                    "remaining_lease": "67 years",
                },
                {
                    "address": "Blk 512 Tampines Ave 4",
                    "flat_type": arguments.get("flat_type"),
                    "resale_price": 475000,
                    "floor_area_sqm": 89,
                    "remaining_lease": "62 years",
                },
            ]
        })
    elif name == "calculate_mortgage":
        principal = arguments["loan_amount"]
        r = (arguments.get("interest_rate", 2.6) / 100) / 12
        n = arguments.get("loan_tenure_years", 25) * 12
        monthly = principal * (r * (1 + r) ** n) / ((1 + r) ** n - 1)
        return json.dumps({
            "loan_amount": principal,
            "monthly_repayment_sgd": round(monthly, 2),
        })
    return json.dumps({"error": f"Unknown tool: {name}"})

def generate_response(messages: list) -> str:
    """
    Single model.generate() call.
    Returns the raw decoded string (may contain tool call JSON).
    """
    # ── Render chat template to string first ──
    text = tokenizer.apply_chat_template(
        messages,
        tools=TOOLS,
        tokenize=False,
        add_generation_prompt=True,
        enable_thinking=False,      # no  blocks for tool calling
 )

 # ── Tokenize separately ──
 inputs = tokenizer(text, return_tensors="pt").to(model.device)

 # ── Generate ──
 with torch.no_grad(): # saves memory during inference
 generated_ids = model.generate(
 **inputs,
 max_new_tokens=512,
 do_sample=False, # greedy for deterministic tool JSON
 )

 # ── Decode new tokens only ──
 output_ids = generated_ids[0][inputs["input_ids"].shape[1]:]
 return tokenizer.decode(output_ids, skip_special_tokens=True).strip()

def parse_tool_calls(response_text: str) -> list:
 """
 Parse Hermes-style tool call JSON from model output.
 Qwen3.6 emits tool calls wrapped in ... tags.
 Returns list of {"name": ..., "arguments": {...}} dicts.
 Falls back to empty list if no tool calls found.
 """
 import re
 tool_calls = []

 # ── Match {...} blocks ──
 pattern = r"(.*?)"
 matches = re.findall(pattern, response_text, re.DOTALL)

 for match in matches:
 try:
 call = json.loads(match.strip())
 tool_calls.append(call)
 except json.JSONDecodeError:
 print(f" [WARN] Could not parse tool call JSON: {match[:100]}")

 return tool_calls

def run_agent(user_query: str, max_steps: int = 10) -> str:
 """
 Transformers-native agentic loop — no vLLM or API server needed.

 Loop:
 1. Generate response
 2. Parse tool calls from output
 3. Execute tools, append results
 4. Repeat until no tool calls in response
 """
 messages = [
 {
 "role": "system",
 "content": (
 "You are a helpful Singapore housing assistant. "
 "Always call the relevant tools to get accurate data before answering. "
 "Give a clear, concise summary after gathering all information."
 ),
 },
 {"role": "user", "content": user_query},
 ]

 print(f"\n{'='*60}")
 print(f"USER: {user_query}")
 print(f"{'='*60}")

 for step in range(max_steps):
 print(f"\n[Step {step + 1}] Generating...")

 response_text = generate_response(messages)
 print(f" Raw output: {response_text[:200]}...")

 # ── Try to parse tool calls from the response ──
 tool_calls = parse_tool_calls(response_text)

 if tool_calls:
 print(f" → Found {len(tool_calls)} tool call(s)")

 # ── Append assistant turn with raw response ──
 messages.append({
 "role": "assistant",
 "content": response_text,
 })

 # ── Execute each tool and append results ──
 for call in tool_calls:
 fn_name = call.get("name", "")
 fn_args = call.get("arguments", {})

 # ── arguments may be a string or dict depending on model output ──
 if isinstance(fn_args, str):
 fn_args = json.loads(fn_args)

 print(f" • {fn_name}({json.dumps(fn_args, ensure_ascii=False)})")
 result = execute_tool(fn_name, fn_args)
 print(f" ↳ {result[:150]}")

 # ── Append tool result as tool role message ──
 messages.append({
 "role": "tool",
 "name": fn_name,
 "content": result,
 })

 continue # loop back for next generation

 # ── No tool calls — this is the final answer ──
 print(f"\n{'='*60}")
 print(f"AGENT FINAL ANSWER:\n{response_text}")
 print(f"{'='*60}\n")
 return response_text

 return "[Agent stopped: exceeded maximum steps]"

# ── Run examples ──
if __name__ == "__main__":

 run_agent(
 "Tolong carikan flat 4-bilik dekat Tampines, bajet bawah $500,000. "
 "Nak tahu juga berapa anggaran pinjaman bulanan."
 )
```

Output

```
============================================================
AGENT FINAL ANSWER:

Tampines

4-room

500000

============================================================
```

Use the code below to get aisingapore/Qwen-SEA-LION-v4.527B-IT-SpecDecoder booster with vLLM.

```
CUDA_VISIBLE_DEVICES=0 vllm serve aisingapore/Qwen-SEA-LION-v4.5-27B-IT \
--speculative-config '{"method": "dflash", "model": "aisingapore/Qwen-SEA-LION-v4.5-27B-IT-SpecDecoder", "num_speculative_tokens": 16}' \
--attention-backend flash_attn \
--max-num-batched-tokens 32768 \
--gdn-prefill-backend triton
```

## Training Details

### Training Data

🤗[aisingapore/SEA-Instruct-2602](https://huggingface.co/datasets/aisingapore/SEA-Instruct-2602)

### Training Regime

Our post-training workflow consists solely of distillation and model merging.

## Evaluation

### Testing Data, Factors & Metrics

We evaluated Qwen-SEA-LION-v4.5 on general language, multi-turn chat and instruction-following capabilities.

#### Results

For details on Qwen-SEA-LION-v4.5-27B-IT performance, please refer to the [SEA-LION Leaderboard](https://leaderboard.sea-lion.ai/).

## Technical Specifications

### Model Architecture

The architecture is based on the highly efficient Qwen3.6 foundation. The detailed architecture can be found at <https://huggingface.co/Qwen/Qwen3.6-27B#model-overview>.

## Uses

### Out-of-Scope Use

The model has not been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

### Bias, Risks, and Limitations

The model was not tested for robustness against adversarial prompting. It is important for users to be aware that our model exhibits certain limitations that warrant consideration. Like many LLMs, the model can hallucinate and occasionally generates irrelevant content, introducing fictional elements that are not grounded in the provided context. Users should also exercise caution in interpreting and validating the model's responses due to the potential inconsistencies.


# SEA-LION v4

SEA-LION version 4, released in August 2025, represents our first collection of multimodal models trained on Southeast Asian text. Each model offers unique strengths tailored to specific needs:

## Vision Language Models (VLMs)

* [**Gemma-SEA-LION-v4-4B-VL**](/models/sea-lion-v4/gemma-sea-lion-v4-4b-vl) A lightweight model optimized for mobile and edge devices, bringing robust SEA language understanding to environments with strict memory and latency constraints.
* [**Gemma-SEA-LION-v4-27B-VL**](/models/sea-lion-v4/gemma-sea-lion-v4-27b-vl) Our powerful vision-language model, expertly trained to interpret both images and text with a deep, nuanced understanding of Southeast Asian cultural contexts.
* [**Qwen-SEA-LION-v4-4B-VL and 8B-VL**](/models/sea-lion-v4/qwen-sea-lion-v4-vl) Our latest specialized vision-language models, featuring a native 256K context window and superior OCR capabilities for Indonesian, Thai, and Vietnamese. The **4B model** is optimized for efficient edge deployment, while the **8B model** is engineered for complex multi-modal reasoning.

## Text Generation LLMs

* [**Apertus-SEA-LION-v4-8B**](/models/sea-lion-v4/apertus-sea-lion-v4-8b) A versatile, compact model designed for high throughput and general-purpose SEA language tasks, offering an optimal trade-off between size and capability for developers.
* [**Gemma-SEA-LION-v4-27B**](/models/sea-lion-v4/gemma-sea-lion-v4-27b) Suited for translation, abstractive summarisation, natural language inference, causal reasoning, metaphor understanding, question answering, paraphrasing, and sentiment analysis—applications where regional language support and advanced reasoning are critical.
* [**Gemma-SEA-LION-v4-27B-IT**](/models/sea-lion-v4/gemma-sea-lion-v4-27b#training-procedure) Suited for knowledge-intensive tasks and high-demand contexts where comprehensive language comprehension is essential.
  * *Available Variants:* [**Gemma-SEA-LION-v4-27B-IT-GGUF, NVFP4, FP8\_Dynamic**](/models/sea-lion-v4/gemma-sea-lion-v4-27b#gemma-sea-lion-v4-27b-it-quantized-version) support inference on a range of consumer-grade GPUs and are compatible with various inference engines.
* [**Qwen-SEA-LION-v4-32B-IT**](/models/sea-lion-v4/qwen-sea-lion-v4-32b) Our flagship instruction-tuned model, designed for maximum performance in general SEA context tasks.
  * *Available Variants:* [**Qwen-SEA-LION-v4-32B-IT-4BIT / 8BIT**](/models/sea-lion-v4/qwen-sea-lion-v4-32b#available-versions) offer a near-perfect balance of performance and efficiency, making the power of SEA-LION accessible on resource-constrained hardware like consumer-grade GPUs.

SEA-LION v4 continues our mission to create language models that understand and respond with greater cultural awareness and depth across Southeast Asia. In addition, SEA-LION v4 has the ability to handle both image and text input as a multimodal model.

For detailed information of each of the SEA-LION v4 models, please refer to their individual documentation pages via the links above.


# Apertus-SEA-LION-v4-8B

Last updated: 2026-02-05

**SEA-LION** is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

## Introduction

SEA-LION stands for *Southeast Asian Languages In One Network*.

**Apertus-SEA-LION-v4-8B-IT** is a 8-billion parameter model built upon the Apertus-8B-Instruct architecture. To ensure **domain adaptation** for the region, the model underwent rigorous post-training on a curated dataset of approximately **8.54 million** instruction-text pairs.

This extensive post-training instills **multilingual** and **multicultural** fluency, covering key SEA languages such as Burmese, Malay, Tagalog and Tamil. The dataset also includes 347,000 tool-calling instruction-text pairs to impart these capabilities, in addition to linguistic fluency.

Apertus-SEA-LION-v4-8B-IT is designed as a fully open model; to align with this core philosophy, we have released the datasets used for post-training, as well as the evaluation codes and datasets used to evaluate the model.

These resources can be accessed via the link below.

* [Open post-training datasets](#Training-Data) we used.
* [SEA-HELM Evaluation codes and datasets](https://github.com/aisingapore/SEA-HELM)

## Model Details

### Model Description

SEA-LION stands for *Southeast Asian Languages In One Network*.

We performed post-training in English and SEA languages on Apertus-8B-Instruct-2509, a decoder model using the Apertus architecture, and post-training to create Apertus-SEA-LION-v4-8B-IT.

For tokenization, the model employs the default tokenizer used in Apertus-8B-Instruct-2509.

* **Developed by:** AI Products Pillar, AI Singapore
* **Funded by:** Singapore NRF
* **Shared by:** AI Products Pillar, AI Singapore
* **Model type:** Decoder
* **Context length:** 65k
* **Language(s):** fine-tuned on English, Burmese, Tagalog, Malay and Tamil
* **License:** [Apache-2.0](https://choosealicense.com/licenses/apache-2.0/)
* **Finetuned from model:** [Apertus-8B-Instruct](https://huggingface.co/swiss-ai/Apertus-8B-Instruct-2509)

### Model Sources

* **Repository:** <https://huggingface.co/collections/aisingapore/sea-lion-v4>

## Uses

### Out-of-Scope Use

The model has not been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

### Bias, Risks, and Limitations

*The models were not tested for robustness against adversarial prompting.* It is important for users to be aware that our models exhibit certain limitations that warrant consideration. Like many LLMs, the models can hallucinate and occasionally generates irrelevant content, introducing fictional elements that are not grounded in the provided context. Users should also exercise caution in interpreting and validating the model's responses due to the potential inconsistencies.

## How to Get Started with the Model

Use the code below to get started with the model with 🤗 Transformers libraries.

```python
pip install transformers>=4.56.0

```

```python
# The code is adopted from Apertus example
from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "aisingapore/Apertus-SEA-LION-v4-8B-IT"
device = "cuda"  # for GPU usage or "cpu" for CPU usage

# load the tokenizer and the model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
).to(device)

# prepare the model input
prompt = "Explain the concept of 'Hari Raya Puasa' in simple terms."
messages_think = [
    {"role": "user", "content": prompt}
]

text = tokenizer.apply_chat_template(
    messages_think,
    tokenize=False,
    add_generation_prompt=True,
)
model_inputs = tokenizer([text], return_tensors="pt", add_special_tokens=False).to(model.device)

# Generate the output
generated_ids = model.generate(**model_inputs, max_new_tokens=32768)

# Get and decode the output
output_ids = generated_ids[0][len(model_inputs.input_ids[0]) :]
print(tokenizer.decode(output_ids, skip_special_tokens=True))
```

### Tool Calling

The prompt in the example is in Malay and translates to “Please help me find a 4-room flat near Tampines, budget under $500,000. I also want to know the estimated monthly loan payment.”

```python
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained('aisingapore/gemma-sealion-4b-v4-cand5', trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    "aisingapore/gemma-sealion-4b-v4-cand5",
    torch_dtype=torch.bfloat16,
    device_map="auto"
)

messages = [
    {"role": "user", "content": "Tolong carikan flat 4-bilik dekat Tampines, bajet bawah $500,000. Nak tahu juga berapa anggaran pinjaman bulanan."}
]

tools = [
    {
        "type": "function",
        "function": {
            "name": "search_hdb_listings",
            "description": "Search for HDB flats available for sale",
            "parameters": {
                "type": "object",
                "properties": {
                    "location": {"type": "string", "description": "Town or area name"},
                    "flat_type": {"type": "string", "description": "Flat type e.g. 3-room, 4-room, 5-room"},
                    "max_price": {"type": "number", "description": "Maximum price in SGD"}
                },
                "required": ["location", "flat_type"]
            }
        }
    },
    {
        "type": "function",
        "function": {
            "name": "calculate_mortgage",
            "description": "Calculate estimated monthly mortgage payment",
            "parameters": {
                "type": "object",
                "properties": {
                    "loan_amount": {"type": "number", "description": "Loan amount in SGD"},
                    "interest_rate": {"type": "number", "description": "Annual interest rate as percentage"},
                    "loan_tenure_years": {"type": "integer", "description": "Loan period in years"}
                },
                "required": ["loan_amount"]
            }
        }
    }
]

input_ids = tokenizer.apply_chat_template(
    messages,
    tools=tools,
    return_tensors="pt",
    add_generation_prompt=True
).to(model.device)

generated_ids = model.generate(
    input_ids,
    max_new_tokens=512,
    do_sample=False,
)

response = tokenizer.decode(
    generated_ids[0][input_ids.shape[1]:],
    skip_special_tokens=False,
).replace("", "").strip()

print(response)
```

## Training Details

### Training Data

The instruction fine-tuning text dataset comprises of a collection of OSS & synthetic data. The datasets used for post-training can be accessed via the link below.

**Datasets for Instruction Fine Tuning**:

* 🤗[aisingapore/SEA-Instruct-2602](https://huggingface.co/datasets/aisingapore/SEA-Instruct-2602)

**Datasets for Tool-calling**:

* 🤗[allenai/Dolci-Instruct-SFT-Tool-Use](https://huggingface.co/datasets/allenai/Dolci-Instruct-SFT-Tool-Use)
* 🤗[Agent-Ark/Toucan-1.5M](https://huggingface.co/datasets/Agent-Ark/Toucan-1.5M)

**Datasets for Reinforcement Learning:**

* 🤗[nvidia/Nemotron-Cascade-RL-Instruction-Following](https://huggingface.co/datasets/nvidia/Nemotron-Cascade-RL-Instruction-Following)
* 🤗[openai/gsm8k](https://huggingface.co/datasets/openai/gsm8k)

### Training Procedure

#### Training Hyperparameters

* **Training regime:** Our post-training workflow consists of instruction fine-tuning and model merging.
* **Training hyperparameters:** The following hyperparameters were used during training:

| Category         | Hyperparameter                | Value                                           |
| ---------------- | ----------------------------- | ----------------------------------------------- |
| **Optimization** | Optimizer                     | `ADAMW_TORCH_FUSED` (β1=0.9, β2=0.999, ε=1e-08) |
| **Batch Size**   | Train Batch Size (per device) | `1`                                             |
|                  | Eval Batch Size (per device)  | `1`                                             |
| **Hardware**     | Distributed Type              | `multi-GPU`                                     |
|                  | Number of Devices             | `64`                                            |
| **Schedule**     | LR Scheduler Type             | `constant_with_warmup`                          |
|                  | LR Scheduler Warmup Steps     | `269`                                           |
| **Other**        | Training Steps                | `5397`                                          |
|                  | Seed                          | `42`                                            |

## Evaluation

### Testing Data, Factors & Metrics

We evaluated Apertus-SEA-LION-v4-8B-IT on general language capabilities and LLM-specific capabilities using SEA-HELM.

### Results

For details on Apertus-SEA-LION-v4-8B-IT performance, please refer to the [Leaderboard results on SEA-HELM](https://leaderboard.sea-lion.ai/).

## Download the Models

The Apertus-SEA-LION-v4-8B models are available for download via the 🤗 [HuggingFace Apertus-SEA-LION-v4-8B-IT](https://huggingface.co/aisingapore/Apertus-SEA-LION-v4-8B-IT) repository. You can also explore more models in the same collection at 🤗 [HuggingFace SEA-LION v4 Collection](https://huggingface.co/collections/aisingapore/sea-lion-v4).

## More Information

This is the repository for the commercial instruction-tuned model. The models have *not* been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

For more info, please contact us at [SEA-LION Inquiry Form](https://forms.gle/sLCUVb95wmGf43hi6) or <sealion@aisingapore.org>


# Gemma-SEA-LION-v4-4B-VL

Last updated: 2026-02-05

SEA-LION is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

**Gemma-SEA-LION-v4-4B-VL** is a 4-billion parameter Vision-Language Model (VLM) built upon the gemma-3-4b-it architecture. To ensure **domain adaptation** for the region, the model underwent rigorous post-training on a curated dataset of approximately **8.54 million** instruction-text pairs. This curated dataset included various SEA Languages such as Burmese, Malay, Tagalog and Tamil. This extensive post-training instills **multilingual** and **multicultural** fluency for both SEA languages and English.

Gemma-SEA-LION-v4-4B-VL inherits the image and text capabilities from gemma-3-4b-it alongside its large context length of 128K tokens. Additionally, beyond extending the multilingual capabilities of the original gemma model for SEA languages, we experimented with:

1. Adding function calling to the model to allow for this model to be used in tool calling applications
2. The visual parsing capabilities in Thai, Chinese and English

## Model Details

### Model Description

We performed post-training in English and SEA languages on Gemma-SEA-LION-v4-27B-IT, a decoder model using the Gemma 3 architecture, to create Gemma-SEA-LION-v4-4B-VL.

For tokenization, the model employs the default tokenizer used in gemma-3-4b-it.

* **Developed by:** AI Products Pillar, AI Singapore
* **Funded by:** Singapore NRF
* **Shared by:** AI Products Pillar, AI Singapore
* **Model type:** Decoder
* **Context length:** 128k tokens
* **Language(s) (NLP):** fine-tuned on Burmese, English, Khmer, Malay, Tagalog and Tamil
* **License:** [Gemma Terms of Use](https://ai.google.dev/gemma/terms)
* **Finetuned from model:** [gemma-3-4b-it](https://huggingface.co/google/gemma-3-4b-it)

As of Feb 2026, Gemma-SEA-LION-v4-4B-VL outperforms other open models in the small parameter class on SEA tasks, achieving performance comparable to larger, proprietary models. For detailed rankings, please refer to the [leaderboard](https://leaderboard.sea-lion.ai/).

## Uses

### Out-of-Scope Use

The model has not been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

## Bias, Risks, and Limitations

*The model was not tested for robustness against adversarial prompting.* It is important for users to be aware that our model exhibits certain limitations that warrant consideration. Like many LLMs, the model can hallucinate and occasionally generates irrelevant content, introducing fictional elements that are not grounded in the provided context. Users should also exercise caution in interpreting and validating the model's responses due to the potential inconsistencies.

**Limitations**

In terms of text capability, Gemma-SEA-LION-v4-4B-VL has been trained and fine-tuned exclusively on the vision-text backend. As a result, its text capabilities are expected to be comparable to those of Gemma-SEA-LION-v4-27B-IT, and may not exhibit significant improvements or differences in this area.

## How to Get Started with the Model

Use the code below to get started with the model using the 🤗 Transformers library.

```python
pip install transformers>=4.50.0

```

```python
from transformers import pipeline
import torch

pipe = pipeline(
    "image-text-to-text",
    model="aisingapore/Gemma-SEA-LION-v4-4B-VL",
    device="cuda",
    torch_dtype=torch.bfloat16
)
messages = [
  {
      "role": "system",
      "content": [{"type": "text", "text": "You are a helpful assistant."}]
  },
  {
      "role": "user",
      "content": [
          {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
          {"type": "text", "text": "Response in Indonesian. What animal is on the candy? Describe in Indonesian and English."}
      ]
  }
]

output = pipe(text=messages, max_new_tokens=200)
print(output[0]["generated_text"][-1]["content"])
# Okay, let's take a look!
# Based on the image, the animal on the candy is a **turtle**.
# You can see the shell shape and the head and legs.

```

## Training Details

### Training Data

The instruction fine-tuning text dataset comprises of a collection of OSS & synthetic data. The datasets used for post-training can be accessed via the link below.

**Datasets for Instruction Fine Tuning**:

* 🤗[aisingapore/SEA-Instruct-2602](https://huggingface.co/datasets/aisingapore/SEA-Instruct-2602)

**Datasets for Tool-calling**:

* 🤗[allenai/Dolci-Instruct-SFT-Tool-Use](https://huggingface.co/datasets/allenai/Dolci-Instruct-SFT-Tool-Use)
* 🤗[Agent-Ark/Toucan-1.5M](https://huggingface.co/datasets/Agent-Ark/Toucan-1.5M)

**Datasets for Reinforcement Learning:**

* 🤗[nvidia/Nemotron-Cascade-RL-Instruction-Following](https://huggingface.co/datasets/nvidia/Nemotron-Cascade-RL-Instruction-Following)
* 🤗[openai/gsm8k](https://huggingface.co/datasets/openai/gsm8k)

### Training Procedure

#### Training Hyperparameters

* **Training regime:** Our post-training workflow consists of Instruction Fine-tuning, Model Merging and Reinforcement Learning.

For details on Gemma-SEA-LION-v4-4B-VL performance, please refer to the SEA-HELM leaderboard, <https://leaderboard.sea-lion.ai/>.

## Hardware

* **Hardware Type:** Nvidia H200 140GB GPUs
* **Cloud Provider:** SMC H200
* **Compute Region:** Singapore

## More Information

This is the repository for the commercial instruction-tuned model. The model has not been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

AI Singapore is a national programme supported by the National Research Foundation, Singapore and hosted by the National University of Singapore. Any opinions, findings and conclusions or recommendations expressed in this material are those of the author(s) and do not reflect the views of the National Research Foundation or the National University of Singapore.


# Gemma-SEA-LION-v4-27B

Last updated: 2025-08-25

**SEA-LION** is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

## Introduction

SEA-LION stands for *Southeast Asian Languages In One Network*.

Gemma-SEA-LION-v4-27B is based on Gemma 3 (which supports over 100 languages) and is a multilingual model that has undergone continued pre-training on approximately **500B tokens**, sampled from a pool of 1 trillion tokens across 11 SEA languages: Burmese, English, Indonesia, Khmer, Lao, Malay, Mandarin, Tagalog, Tamil, Thai and Vietnamese, to create *Gemma-SEA-LION-v4-27B*. Subsequently, we performed post-training on Gemma-SEA-LION-v4-27B using multiple RL algorithms to produce *Gemma-SEA-LION-v4-27B-IT*.

Gemma-SEA-LION-v4-27B inherits Gemma 3’s:

* Large 128K context length
* Image and text understanding capabilities, including document comprehension, visual Q\&A, and image-grounded reasoning
* Advanced function calling and structured outputs to allow for seamless integration into larger systems

For tokenization, the model employs the default tokenizer used in Gemma 3 27B IT.

At a glance:

* **Developed by:** Products Pillar, AI Singapore
* **Funded by:** Singapore NRF
* **Shared by:** Products Pillar, AI Singapore
* **Model type:** Decoder
* **Context length:** 128k
* **Language(s):** Burmese, English, Indonesia, Khmer, Lao, Malay, Mandarin, Tagalog, Tamil, Thai and Vietnamese
* **License:** [Gemma Terms of Use](https://ai.google.dev/gemma/terms)
* **Continued pretrained and finetuned from model:** [Gemma-3-27B-IT](https://huggingface.co/google/gemma-3-27b-it)

As of 25 Aug 2025, Gemma-SEA-LION-v4-27B-IT excels at Southeast Asian (SEA) tasks when compared to other open models with fewer than 200 billion parameters and demonstrates performance comparable to that of larger and top closed models. For detailed rankings, please refer to the [leaderboard](https://leaderboard.sea-lion.ai/).

## Training Details

### Training Data

The dataset comprises Burmese, English, Indonesia, Khmer, Lao, Malay, Mandarin, Tagalog, Tamil, Thai and Vietnamese languages, collected from a mixture of sources including web data, code, open-source datasets, and synthetically generated datasets, amounting to a total of 500 billion tokens.

The 500 billion tokens are sampled from a much larger pool of 1 trillion tokens from open-sourced datasets with the optimal datamix shown below determined by our experiments.

| Language                  | Dataset Name               | Total Tokens (B) | Percentage (%) | Total percentage (%) |
| ------------------------- | -------------------------- | ---------------- | -------------- | -------------------- |
| Code                      | StarCoder (OLMo 2 Version) | 50B              | 10             | 10                   |
| EN                        | Fineweb-Edu                | 80B              | 16             | 40                   |
|                           | DCLM-OLMo2-HQ              | 80B              | 16             |                      |
|                           | Non-CC-EN                  | 40B              | 8              |                      |
| ZH                        | SEA-LION Pile v1           | 13.5B            | 2.7            | 9                    |
|                           | Fineweb2                   | 13.5B            | 2.7            |                      |
|                           | Fineweb2-HQ                | 4.5B             | 0.9            |                      |
| VI                        | SEA-LION Pile v1           | 4.25B            | 0.85           | 8.5                  |
|                           | SEA-LION Pile v2           | 12.75B           | 2.55           |                      |
|                           | Fineweb2                   | 8.5B             | 1.7            |                      |
|                           | Non-CC-VI                  | 17B              | 3.4            |                      |
| ID                        | SEA-LION Pile v1           | 5.66B            | 1.13           | 8.5                  |
|                           | SEA-LION Pile v2           | 17B              | 3.4            |                      |
|                           | Fineweb2                   | 11.33B           | 2.27           |                      |
|                           | Non-CC-ID                  | 8.5B             | 1.7            |                      |
| TH                        | SEA-LION Pile v1           | 3.035B           | 0.61           | 8.5                  |
|                           | SEA-LION Pile v2           | 9.107B           | 1.82           |                      |
|                           | Fineweb2                   | 3.035B           | 0.61           |                      |
|                           | WangChanBERTa              | 3.035B           | 0.61           |                      |
|                           | Dolmav1                    | 3.035B           | 0.61           |                      |
|                           | Non-CC-TH                  | 21.25B           | 4.25           |                      |
| TL, TA, MS, KM, LO and MY | ALL\_LANG                  | 77.5B            | 15.5           | 15.5                 |

Note:

* All token counts are counted using Gemma 3 tokenizer.
* Pre-training was conducted with batches of 8k token lengths.
* SEA-Pile v1 is processed from Common Crawl WET, which is published [here](https://huggingface.co/datasets/aisingapore/sea-lion-pile). The main proportion is from mC4 dataset (corpus [link](https://huggingface.co/datasets/bertin-project/mc4-sampling)). The cutoff date of this version is September 2020.
* SEA-Pile v2 is processed from Common Crawl WARC from October 2020 to April 2024.
* Tamil news is sourced with permission from [Seithi](https://seithi.mediacorp.sg/)
* We utilized 0.5% of synthetically generated datasets for the low-resource language, Khmer.

### Training Procedure

**Training regime:**

| Hyperparameter    | Gemma-SEA-LION-v4-27B |
| ----------------- | --------------------- |
| Precision         | bfloat16              |
| Optimizer         | decoupled\_adamw      |
| Scheduler         | CosineAnnealing       |
| Learning Rate     | 4.00E-08              |
| Global Batch Size | 1024                  |

Gemma-SEA-LION-v4-27B has undergone post-training using a QA pairs dataset in Burmese, English, Indonesia, Khmer, Lao, Malay, Tagalog, Tamil, Thai and Vietnamese, comprising approximately 10M samples in total, to create *Gemma-SEA-LION-v4-27B-IT*.

The instruction fine-tuning dataset combines our SEA-Instruct, Infinity-Instruct, and OpenMath-Instruct 2 with open-source datasets. For the Online RL datasets, open sourced datasets such as nvidia/Llama-Nemotron-Post-Training-Dataset (RL set) and zwhe99/DeepMath-103K were used. For alignment, rejected-chosen pairs are generated from the target model, with the chosen responses obtained by rewriting and improving upon the rejected outputs. Prompt sampling is guided by a gradient-based analysis process.

Our post-training workflow consists of multiple stages: instruction fine-tuning, model merging, online RL for both instruction following and math using DRGPPO, and then followed by on-policy alignment via APO.

## Evaluation

We evaluated Gemma-SEA-LION-v4-27B models on both general language capabilities and instruction-following capabilites.

### Testing Data, Factors & Metrics

#### Testing Data

We evaluated Gemma-SEA-LION-v4-27B models on both general language capabilities and instruction-following capabilites.

General language capabilities

For the evaluation of general language capabilities, we employed the [SEA-HELM evaluation benchmark](https://arxiv.org/abs/2502.14301) across a variety of tasks. These tasks include Question Answering (QA), Sentiment Analysis (Sentiment), Toxicity Detection (Toxicity), Translation in both directions (Eng>Lang & Lang>Eng), Abstractive Summarisation (Abssum), Causal Reasoning (Causal), Natural Language Inference (NLI), Linguistic Diagnostics (LINDSEA), Cultural Knowledge (Kalahi) and Global MMLU Lite.

Instruction-following and Multi-turn Chat

We evaluated the models on instruction-following and multi-turn chat capabilities with SEA-IFEval (based on [IFEval](https://arxiv.org/abs/2311.07911)) and SEA-MTBench (based on [MT-Bench](https://arxiv.org/abs/2306.05685)) respectively. The two datasets were originally in English, the linguists and native speakers in the team worked together to filter, localise and translate the datasets into the respective target languages to ensure that the examples remained reasonable, meaningful and natural.

#### Factors

All evaluations were run with the model specific generation parameters defined in the model config. Each evaluation comprised of 8 runs with different seeds and the final results were averaged across these runs.

For all tasks, the model was expected to provide an answer tag from which the answer was automatically extracted. For tasks where options were provided, the answer should comprise one of the pre-defined options.

The evaluation was done zero-shot with native prompts on a sample of 100-1000 instances for each dataset.

SEA-IFEval

SEA-IFEval evaluates a model's ability to adhere to constraints provided in the prompt, for example beginning a response with a specific word/phrase or answering with a certain number of sections. Additionally, accuracy is normalised by the proportion of responses in the correct language (if the model performs the task correctly but responds in the wrong language, it is judged to have failed the task).

SEA-MTBench

SEA-MTBench evaluates a model's ability to engage in multi-turn (2 turns) conversations and respond in ways that align with human needs. We use `gpt-4.1-2025-04-14` as the judge model and compare against `gpt-4.1-2025-04-14` as the baseline model. The metric used is the weighted win rate against the baseline model (i.e. average win rate across each category: Math, Reasoning, STEM, Humanities, Roleplay, Writing, Extraction).

#### Metrics

The following metrics were used:

| Task                           | Metric                            |
| ------------------------------ | --------------------------------- |
| Sentiment Analysis             | Accuracy                          |
| Extractive QA (ID, VI, TH, TA) | ChrF++                            |
| MCQ-QA (TL, MY, MS)            | Accuracy                          |
| Metaphor                       | Accuracy                          |
| Abstractive Summarisation      | Rouge-L                           |
| Translations                   | MetricX-24 score (with reference) |
| Causal Reasoning               | Accuracy                          |
| Natural Language Inference     | Accuracy                          |
| LINDSEA                        | Accuracy                          |
| Global MMLU Lite               | Accuracy                          |
| Kalahi                         | Accuracy                          |
| SEA-IFEval                     | Accuracy                          |
| SEA-MTBench                    | Win rate against a reference      |
| Toxicity Detection             | Accuracy                          |

### Results

For details on Gemma-SEA-LION-v4-27B model performances, please refer to the SEA-HELM leaderboard, [Leaderboard results on SEA-HELM](https://leaderboard.sea-lion.ai/).

## Gemma-SEA-LION-v4-27B-IT Quantized Version

The following quantized versions of our Gemma-SEA-LION-v4-27B-IT model are available:

* Gemma-SEA-LION-v4-27B-IT-Q4\_K\_M
* Gemma-SEA-LION-v4-27B-IT-Q8\_0
* Gemma-SEA-LION-v4-27B-IT-BF16
* Gemma-SEA-LION-v4-27B-IT-NVFP4
* Gemma-SEA-LION-v4-27B-IT-FP8-Dynamic

Please refer to our [Download the Models](#download-the-models) section for more details on how to access them.

## How to Get Started with the Model

## Download the Models

Gemma-SEA-LION-v4-27B models are available for download via the following channels: 🤗[HuggingFace SEA-LION v4 Collection](https://github.com/aisingapore/sealion/blob/main/models/sea-lion-v4/\(https:/huggingface.co/collections/aisingapore/sea-lion-v4-68aa7bb8061d497a4f9f2fec\)/README.md)

| Model                                | Download                                                                                                                                           |
| ------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------- |
| Gemma-SEA-LION-v4-27B                | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-LION-v4-27B)                                                                            |
| Gemma-SEA-LION-v4-27B-IT             | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-LION-v4-27B-IT)                                                                         |
| Gemma-SEA-LION-v4-27B-IT-GGUF        | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-LION-v4-27B-IT-GGUF), [Ollama](https://ollama.com/aisingapore/Gemma-SEA-LION-v4-27B-IT) |
| Gemma-SEA-LION-v4-27B-IT-NVFP4       | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-LION-v4-27B-IT-NVFP4)                                                                   |
| Gemma-SEA-LION-v4-27B-IT-FP8-Dynamic | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-LION-v4-IT-27B-FP8-Dynamic)                                                             |

## Usage

Use the code below to get started with the model using the 🤗 Transformers library.

```python
from transformers import pipeline
import torch

pipe = pipeline(
    "text-generation",
    model="aisingapore/Gemma-SEA-LION-v4-27B-IT",
    device="cuda",
    torch_dtype=torch.bfloat16
)

messages = [
    {
        "role": "system",
        "content": [{"type": "text", "text": "You are a helpful assistant."}]
    },
    {
        "role": "user",
        "content": [
            {"type": "text", "text": "Write a poem on southeast asian countries in Indonesian."}
        ]
    }
]

output = pipe(text=messages, max_new_tokens=200)
print(output[0]["generated_text"][-1]["content"])
```

## Disclaimer

The models have not been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

### Bias, Risks, and Limitations

*The models were not tested for robustness against adversarial prompting.* It is important for users to be aware that our models exhibit certain limitations that warrant consideration. Like many LLMs, the models can hallucinate and occasionally generates irrelevant content, introducing fictional elements that are not grounded in the provided context. Users should also exercise caution in interpreting and validating the model's responses due to the potential inconsistencies.

**Limitations**

In terms of vision capability, Gemma-SEA-LION-v4-27B has been trained and fine-tuned exclusively on the text back-end. As a result, its vision capabilities are expected to be comparable to those of Gemma 3 IT 27B, and may not exhibit significant improvements or differences in this area. [🤗 google/gemma-3-27b-it](https://huggingface.co/google/gemma-3-27b-it)

## More Information

This is the repository for the commercial instruction-tuned model. The models have *not* been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

For more info, please contact us at [SEA-LION Inquiry Form](https://forms.gle/sLCUVb95wmGf43hi6) or <sealion@aisingapore.org>


# Gemma-SEA-LION-v4-27B-VL

Last updated: 2025-10-17

**SEA-VLM** is an instruct-tuned vision-text model for the Southeast Asia (SEA) region.

Gemma-SEA-LION-v4-27B-VL has undergone post-training using instruction-image pairs datasets in Burmese, English, Indonesian, Khmer, Lao, Malay, Mandarin, Tagalog, Tamil, Thai and Vietnamese, comprising approximately 540k samples in total, to create *Gemma-SEA-LION-v4-27B-VL*.

Gemma-SEA-LION-v4-27B-VL inherits Gemma 3's:

* Large 128K context length
* Image and text understanding capabilities, including document comprehension, visual Q\&A, and image-grounded reasoning
* Advanced function calling and structured outputs to allow for seamless integration into larger systems

## Model Details

### Model Description

We performed post-training in English and SEA languages on Gemma-SEA-LION-v4-27B-IT, a decoder model using the Gemma 3 architecture, to create Gemma-SEA-LION-v4-27B-VL.

For tokenization, the model employs the default tokenizer used in Gemma 3 27B IT.

* **Developed by:** SEACrowd and Products Pillar, AI Singapore
* **Funded by:** Singapore NRF
* **Shared by:** SEACrowd and Products Pillar, AI Singapore
* **Model type:** Decoder
* **Context length:** 128k tokens
* **Language(s) (NLP):** Burmese, English, Indonesian, Khmer, Lao, Malay, Mandarin, Tagalog, Tamil, Thai and Vietnamese
* **License:** [Gemma Terms of Use](https://ai.google.dev/gemma/terms)
* **Finetuned from model:** [Gemma-SEA-LION-v4-27B-IT](https://huggingface.co/aisingapore/Gemma-SEA-LION-v4-27B-IT)

As of 15 October 2025, Gemma-SEA-LION-v4-27B-VL excels at Southeast Asian (SEA) tasks when compared to other open models with fewer than 200 billion parameters and demonstrates performance comparable to that of larger and top closed models. For detailed rankings, please refer to the [leaderboard](https://leaderboard.sea-lion.ai/).

## Uses

### Out-of-Scope Use

The model has not been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

## Bias, Risks, and Limitations

*The model was not tested for robustness against adversarial prompting.* It is important for users to be aware that our model exhibits certain limitations that warrant consideration. Like many LLMs, the model can hallucinate and occasionally generates irrelevant content, introducing fictional elements that are not grounded in the provided context. Users should also exercise caution in interpreting and validating the model's responses due to the potential inconsistencies.

**Limitations**

In terms of text capability, Gemma-SEA-LION-v4-27B-VL has been trained and fine-tuned exclusively on the vision-text backend. As a result, its text capabilities are expected to be comparable to those of Gemma-SEA-LION-v4-27B-IT, and may not exhibit significant improvements or differences in this area.

## How to Get Started with the Model

Use the code below to get started with the model using the 🤗 Transformers library.

```python
from transformers import pipeline
import torch

pipe = pipeline(
    "image-text-to-text",
    model="aisingapore/Gemma-SEA-LION-v4-27B-VL",
    device="cuda",
    torch_dtype=torch.bfloat16
)

messages = [
    {
        "role": "system",
        "content": [{"type": "text", "text": "You are a helpful assistant."}]
    },
    {
        "role": "user",
        "content": [
            {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"},
            {"type": "text", "text": "What animal is on the candy?"}
        ]
    }
]

output = pipe(text=messages, max_new_tokens=200)
print(output[0]["generated_text"][-1]["content"])

```

## Training Details

### Training Data

The dataset comprises vision-text paired in Burmese, English, Indonesian, Khmer, Lao, Malay, Mandarin, Tagalog, Tamil, Thai and Vietnamese languages, collected from a mixture of sources including web data, code, open-source datasets.

### Training Procedure

#### Training Hyperparameters

* **Training regime:**

We perform SFT using 540k of vision-text samples written in 10 languages. Then, we perform model merging with Gemma3-27B-IT to preserve general vision-text knowledge.

For details on Gemma-SEA-LION-v4-27B-VL performance, please refer to the SEA-LION.ai blogpost, [SEA-LION v4 VL new members](https://sea-lion.ai/sea-lion-v4-VL-new-members/).

## Environmental Impact

Carbon emissions can be estimated using the [Machine Learning Impact calculator](https://mlco2.github.io/impact#compute) presented in [Lacoste et al. (2019)](https://arxiv.org/abs/1910.09700).

* **Hardware Type:** Nvidia H200 140GB GPUs
* **Hours used:** 13 hrs
* **Cloud Provider:** SMC H200
* **Compute Region:** Singapore
* **Carbon Emitted:** appx. 27 kg CO2 e

## More Information

This is the repository for the commercial instruction-tuned model. The model has not been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

AI Singapore is a national programme supported by the National Research Foundation, Singapore and hosted by the National University of Singapore. Any opinions, findings and conclusions or recommendations expressed in this material are those of the author(s) and do not reflect the views of the National Research Foundation or the National University of Singapore.


# Qwen-SEA-LION-v4-32B-IT

Last updated: 2025-10-17

*SEA-LION* is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

**Qwen-SEA-LION-v4-32B-IT** is based on Qwen3, which provides a strong foundation with support for over 100 languages and advanced reasoning capabilities. The model underwent continued pre-training on approximately **100B tokens** sampled from the SEA-Pile v2 pretraining corpus of over one trillion tokens across 7 SEA languages: Burmese, Indonesian, Malay, Filipino, Tamil, Thai, and Vietnamese. Finally, it was post-trained on a high-quality dataset of approximately **8 million question-and-answer pairs** to create the final instruction-tuned model.

Qwen-SEA-LION-v4-32B-IT inherits the following features from Qwen3-32B:

* 32,768 of context length natively

## Model Details

### Model Description

SEA-LION stands for *Southeast Asian Languages In One Network*.

We performed continued pre-training in English and SEA languages on Qwen3-32B, a decoder model using the Gemma 3 architecture, and post-training to create Qwen-SEA-LION-v4-32B-IT.

For tokenization, the model employs the default tokenizer used in Qwen3-32B.

* **Developed by:** Products Pillar, AI Singapore
* **Funded by:** Singapore NRF
* **Shared by:** Products Pillar, AI Singapore
* **Model type:** Decoder
* **Context Length:** 128k tokens
* **Language(s) (NLP):** Burmese, English, Indonesian, Khmer, Lao, Malay, Mandarin, Tagalog, Tamil, Thai, and Vietnamese
* **License:** [Qwen Terms of Service](https://qwen.ai/termsservice) / [Qwen Usage Policy](https://qwen.ai/usagepolicy)
* **Continue pretrained from model:** [Qwen-3-32B](https://huggingface.co/Qwen/Qwen3-32B)

## Uses

### Out-of-Scope Use

The model has *not* been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

## Bias, Risks, and Limitations

### Caveats || Risks

*The model was not tested for robustness against adversarial prompting.* It is important for users to be aware that our model exhibits certain limitations that warrant consideration. Like many LLMs, the model can hallucinate and occasionally generates irrelevant content, introducing fictional elements that are not grounded in the provided context. Users should also exercise caution in interpreting and validating the model's responses due to the potential inconsistencies.

## Limitations

In terms of vision capability, Qwen-SEA-LION-v4-32B-IT has been trained and fine-tuned exclusively on the text back-end. As a result, its vision capabilities are expected to be comparable to those of Qwen3-32B, and may not exhibit significant improvements or differences in this area. (<https://huggingface.co/Qwen/Qwen3-32B> )

## How to Get Started with the Model

Use the code below to get started with the model using the 🤗 Transformers library.

> The model defaults to non-thinking mode. To enable thinking mode, please use `enable_thinking=True`.

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "aisingapore/Qwen-SEA-LION-v4-32B"

# load the tokenizer and the model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype="auto",
    device_map="auto"
)

# prepare the model input
prompt = "Create a poem that captures the seaside scenery across Southeast Asian countries, including transcriptions in their respective languages."
messages = [
    {"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=True # Switches between thinking and non-thinking modes. Default is True.
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)

# conduct text completion
generated_ids = model.generate(
    **model_inputs,
    max_new_tokens=32768
)
output_ids = generated_ids[0][len(model_inputs.input_ids[0]):].tolist()

# parsing thinking content
try:
    # rindex finding 151668 ()
    index = len(output_ids) - output_ids[::-1].index(151668)
except ValueError:
    index = 0

thinking_content = tokenizer.decode(output_ids[:index], skip_special_tokens=True).strip("\n")
content = tokenizer.decode(output_ids[index:], skip_special_tokens=True).strip("\n")

print("thinking content:", thinking_content)
print("content:", content)

```

## Training Details

**Training Datasets**

The instruction fine-tuning dataset combines our SEA-Instruct, Infinity-Instruct, and OpenMath-Instruct 2 with open-source datasets. For the Online RL datasets, open sourced datasets such as nvidia/Llama-Nemotron-Post-Training-Dataset (RL set) and zwhe99/DeepMath-103K were used.

#### Training Procedure

**Training regime**

Our post-training workflow consists of multiple stages: instruction fine-tuning, model merging, online RL for both instruction following and math using DRGPPO, and then followed by on-policy alignment via APO.

## Uses

### Available Versions

* [Qwen-SEA-LION-v4-32B-IT](https://huggingface.co/aisingapore/Qwen-SEA-LION-v4-32B-IT)
* [Qwen-SEA-LION-v4-32B-IT-4BIT](https://huggingface.co/aisingapore/Qwen-SEA-LION-v4-32B-IT-4BIT)
* [Qwen-SEA-LION-v4-32B-IT-8BIT](https://huggingface.co/aisingapore/Qwen-SEA-LION-v4-32B-IT-8BIT)

### Resource Metrics

| Quantized Variant | Model Size (GB) | VRAM Required (GB) | Time to First Token (s) | Tokens per Second |
| ----------------- | --------------- | ------------------ | ----------------------- | ----------------- |
| BF16              | 65.57           | 61.03              | 0.34                    | 58.84             |
| 8-bit (GPTQ)      | 34.34           | 32.04              | 0.35                    | 68.85             |
| 4-bit (GPTQ)      | 19.93           | 19.43              | 0.34                    | 78.20             |

*Additional Remarks:*

* TTFT and Toks per Sec: measured with vLLM on localhost and concurrency = 1.
* Reported results are the median (p50) values, calculated across 10 requests. (11 requests were run and the first result was dropped, to eliminate cold-start delays)
* Model size taken from vLLM upon loading
* Input size 4K, output 1K
* Tests conducted on a system with an NVIDIA H200 GPU

## Evaluation

### Testing Data, Factors & Metrics

We evaluated Qwen-SEA-LION-v4-32B-IT on general language, multi-turn chat and instruction-following capabilities.

#### Testing Data

General language capabilities

For the evaluation of general language capabilities, we employed the [SEA-HELM evaluation benchmark](https://arxiv.org/abs/2502.14301) across a variety of tasks. These tasks include Question Answering (QA), Sentiment Analysis (Sentiment), Toxicity Detection (Toxicity), Translation in both directions (Eng>Lang & Lang>Eng), Abstractive Summarisation (Abssum), Causal Reasoning (Causal), Natural Language Inference (NLI), Linguistic Diagnostics (LINDSEA), Cultural Knowledge (Kalahi) and Global MMLU Lite.

Instruction-following and Multi-turn Chat

We evaluated the models on instruction-following and multi-turn chat capabilities with SEA-IFEval (based on [IFEval](https://arxiv.org/abs/2311.07911)) and SEA-MTBench (based on [MT-Bench](https://arxiv.org/abs/2306.05685)) respectively. The two datasets were originally in English, the linguists and native speakers in the team worked together to filter, localise and translate the datasets into the respective target languages to ensure that the examples remained reasonable, meaningful and natural.

#### Factors

All evaluations were run with the model specific generation parameters defined in the model config. Each evaluation comprised of 8 runs with different seeds and the final results were averaged across these runs.

For all tasks, the model was expected to provide an answer tag from which the answer was automatically extracted. For tasks where options were provided, the answer should comprise one of the pre-defined options.

The evaluation was done **zero-shot** with native prompts on a sample of 100-1000 instances for each dataset.

SEA-IFEval

SEA-IFEval evaluates a model's ability to adhere to constraints provided in the prompt, for example beginning a response with a specific word/phrase or answering with a certain number of sections. Additionally, accuracy is normalised by the proportion of responses in the correct language (if the model performs the task correctly but responds in the wrong language, it is judged to have failed the task).

SEA-MTBench

SEA-MTBench evaluates a model's ability to engage in multi-turn (2 turns) conversations and respond in ways that align with human needs. We use `gpt-4.1-2025-04-14` as the judge model and compare against `gpt-4.1-2025-04-14` as the baseline model. The metric used is the weighted win rate against the baseline model (i.e. average win rate across each category: Math, Reasoning, STEM, Humanities, Roleplay, Writing, Extraction).

#### Metrics

The following metrics were used:

| **Task**                       | **Metric**                        |
| ------------------------------ | --------------------------------- |
| Sentiment Analysis             | Accuracy                          |
| Extractive QA (ID, VI, TH, TA) | ChrF++                            |
| MCQ-QA (TL, MY, MS)            | Accuracy                          |
| Metaphor                       | Accuracy                          |
| Abstractive Summarisation      | Rouge-L                           |
| Translations                   | MetricX-24 score (with reference) |
| Causal Reasoning               | Accuracy                          |
| Natural Language Inference     | Accuracy                          |
| LINDSEA                        | Accuracy                          |
| Global MMLU Lite               | Accuracy                          |
| Kalahi                         | Accuracy                          |
| SEA-IFEval                     | Accuracy                          |
| SEA-MTBench                    | Win rate against a reference      |
| Toxicity Detection             | Accuracy                          |

For details on Qwen-SEA-LION-v4-27B-IT performance, please refer to the SEA-HELM leaderboard, <https://leaderboard.sea-lion.ai/> .

## More Information

This is the repository for the commercial instruction-tuned model. The model has *not* been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

For more info, please contact us using this <sealion@aisingapore.org>


# Qwen-SEA-LION-v4-VL

Last update: 2025-12-1

**SEA-LION** is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

**Qwen-SEA-LION-v4-VL** contains 4/8-billion parameter Vision-Language Models (VLM) built upon the Qwen3-VL-4B/8B-Instruct architecture. To ensure **domain adaptation** for the region, the model underwent rigorous supervised fine-tuning (SFT) on a curated dataset of approximately **9 million** instruction-text pairs. This extensive post-training instills **multilingual** and **multicultural** fluency, covering English and 7 key SEA languages: Burmese, Indonesian, Filipino, Malay, Tamil, Thai, and Vietnamese.

Qwen-SEA-LION-v4-4B/8B-VL inherits the following features from Qwen3-VL:

* Long-Context Multimodal Architecture (Native 256K context window)
* Edge-Optimized Inference (Resource Efficient)
* Enhanced Vision-Language Capabilities
* Tool Use

## Introduction

SEA-LION stands for *Southeast Asian Languages In One Network*.

Qwen-SEA-LION-v4-4B/8B-VL contains 4/8-billion parameter Vision-Language Models (VLM) built upon the Qwen3-VL-4B/8B-Instruct architecture. To ensure **domain adaptation** for the region, the model underwent rigorous supervised fine-tuning (SFT) on a curated dataset of approximately **9 million** instruction-text pairs. This extensive post-training instills **multilingual** and **multicultural** fluency, covering English and 7 key SEA languages: Burmese, Indonesian, Filipino, Malay, Tamil, Thai, and Vietnamese.

Qwen-SEA-LION-v4-4B/8B-VL inherits the following features from Qwen3-VL:

* Long-Context Multimodal Architecture (Native 256K context window)
* Edge-Optimized Inference (Resource Efficient)
* Enhanced Vision-Language Capabilities
* Tool Use

For tokenization, the model employs the default tokenizer used in Qwen3-VL.

* **Developed by:** AI Products Pillar, AI Singapore
* **Funded by:** Singapore NRF
* **Shared by:** AI Products Pillar, AI Singapore
* **Model type:** Decoder
* **Context length:** 256k
* **Language(s):** fine-tuned on Burmese, Indonesian, Filipino, Malay, Tamil, Thai, and Vietnamese
* **License:** [Apache-2.0](https://choosealicense.com/licenses/apache-2.0/)
* **Finetuned from model:** [Qwen3-VL-4B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct), [Qwen3-VL-8B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-8B-Instruct)

## Training Details

### Training Data

The instruction fine-tuning text dataset comprises of a collection of OSS & synthetic data.

### Training Procedure

#### Training Hyperparameters

* **Training regime:** Our workflow consists of instruction fine-tuning and model merging.

## Evaluation

### Testing Data, Factors & Metrics

#### Testing Data

We evaluated Qwen-SEA-LION-v4-4B/8B-VL on general language capabilities.

*General language capabilities*

For the evaluation of general language capabilities, we employed the [SEA-HELM evaluation benchmark](https://arxiv.org/abs/2502.14301) across a variety of tasks. These tasks include Question Answering (QA), Sentiment Analysis (Sentiment), Toxicity Detection (Toxicity), Translation in both directions (Eng>Lang & Lang>Eng), Abstractive Summarisation (Abssum), Causal Reasoning (Causal), Natural Language Inference (NLI), Linguistic Diagnostics (LINDSEA), Cultural Knowledge (Kalahi) and Global MMLU Lite.

*Instruction-following and Multi-turn Chat*

We evaluated the models on instruction-following and multi-turn chat capabilities with SEA-IFEval (based on [IFEval](https://arxiv.org/abs/2311.07911)) and SEA-MTBench (based on [MT-Bench](https://arxiv.org/abs/2306.05685)) respectively. The two datasets were originally in English, the linguists and native speakers in the team worked together to filter, localise and translate the datasets into the respective target languages to ensure that the examples remained reasonable, meaningful and natural.

#### Factors

All evaluations were run with the model specific generation parameters defined in the model config. Each evaluation comprised of 8 runs with different seeds and the final results were averaged across these runs.

For all tasks, the model was expected to provide an answer tag from which the answer was automatically extracted. For tasks where options were provided, the answer should comprise one of the pre-defined options.

The evaluation was done **zero-shot** with native prompts on a sample of 100-1000 instances for each dataset.

*SEA-IFEval*

SEA-IFEval evaluates a model's ability to adhere to constraints provided in the prompt, for example beginning a response with a specific word/phrase or answering with a certain number of sections. Additionally, accuracy is normalised by the proportion of responses in the correct language (if the model performs the task correctly but responds in the wrong language, it is judged to have failed the task).

*SEA-MTBench*

SEA-MTBench evaluates a model's ability to engage in multi-turn (2 turns) conversations and respond in ways that align with human needs. We use `gpt-4.1-2025-04-14` as the judge model and compare against `gpt-4.1-2025-04-14` as the baseline model. The metric used is the weighted win rate against the baseline model (i.e. average win rate across each category: Math, Reasoning, STEM, Humanities, Roleplay, Writing, Extraction).

#### Metrics

The following metrics were used for text capabilities:

| **Task**                       | **Metric**                        |
| ------------------------------ | --------------------------------- |
| Sentiment Analysis             | Accuracy                          |
| Extractive QA (ID, VI, TH, TA) | ChrF++                            |
| MCQ-QA (TL, MY, MS)            | Accuracy                          |
| Metaphor                       | Accuracy                          |
| Abstractive Summarisation      | Rouge-L                           |
| Translations                   | MetricX-24 score (with reference) |
| Causal Reasoning               | Accuracy                          |
| Natural Language Inference     | Accuracy                          |
| LINDSEA                        | Accuracy                          |
| Global MMLU Lite               | Accuracy                          |
| ThaiExam                       | Accuracy                          |
| Kalahi                         | Accuracy                          |
| SEA-IFEval                     | Accuracy                          |
| SEA-MTBench                    | Win rate against a reference      |

### Retaining VL Capabilities

We also evaluated our models on two types of tasks using datasets specifically focused on Southeast Asian examples to benchmark and compared our models' performances against the original base models (Qwen3-VL-4B/8B).

* Visual Question Answering (VQA): We utilised Multiple Choice Question (MCQ) style tasks, including MARVL, CVQA, and WorldCuisines.
* Image Captioning: We employed the XM3600 dataset, evaluating strictly on examples relevant to the SEA region.

Key Insight: Despite our fine-tuning process focusing primarily on text data (approximately 8 million regional Q\&A and instruction pairs), our evaluations confirm that Qwen-SEA-LION-v4 (4B/8B) successfully retains the high-performance vision-language capabilities of the original base models.

#### Factors

The evaluation was done **zero-shot** with native prompts.

#### Metrics

The following metrics were used to measure performance:

* **Normalised accuracy** was the primary metric for the VQA tasks (CVQA, MARVL, and WorldCuisines).
* **RefCLIP Score** was used for the XM3600 image captioning task.

### Results

For details on Qwen-SEA-LION-v4-VL performances, please refer to the SEA-HELM leaderboard, <https://leaderboard.sea-lion.ai/>.

## Download the Models

Qwen-SEA-LION-v4-VL models are available for download via the following channels: 🤗[HuggingFace SEA-LION v4 Collection](https://github.com/aisingapore/sealion/blob/main/models/sea-lion-v4/\(https:/huggingface.co/collections/aisingapore/sea-lion-v4\)/README.md)

| Model                  | Download                                                                 |
| ---------------------- | ------------------------------------------------------------------------ |
| Qwen-SEA-LION-v4-4B-VL | [HuggingFace](https://huggingface.co/aisingapore/Qwen-SEA-LION-v4-4B-VL) |
| Qwen-SEA-LION-v4-8B-VL | [HuggingFace](https://huggingface.co/aisingapore/Qwen-SEA-LION-v4-8B-VL) |

## Uses

## How to Get Started with the Model

Use the code below to get started with the model with 🤗 Transformers libraries.

```bash
  pip install transformers>=4.57.0 
```

```python
from transformers import Qwen3VLForConditionalGeneration, AutoProcessor

# default: Load the model on the available device(s)
model = Qwen3VLForConditionalGeneration.from_pretrained(
    "aisingapore/Qwen-SEA-LION-v4-8B-VL", dtype="auto", device_map="auto"
)

# We recommend enabling flash_attention_2 for better acceleration and memory saving, especially in multi-image and video scenarios.
# model = Qwen3VLForConditionalGeneration.from_pretrained(
#     "aisingapore/Qwen-SEA-LION-v4-8B-VL",
#     dtype=torch.bfloat16,
#     attn_implementation="flash_attention_2",
#     device_map="auto",
# )

processor = AutoProcessor.from_pretrained("Qwen/Qwen3-VL-4B-Instruct")

messages = [
    {
        "role": "system",
        "content": [{"type": "text", "text": "You are a helpful assistant."}]
    },
    {
        "role": "user",
        "content": [
            {"type": "text", "text": "Write a poem on southeast asian countries in Indonesian."}
        ],
    }
]

# Preparation for inference
inputs = processor.apply_chat_template(
    messages,
    tokenize=True,
    add_generation_prompt=True,
    return_dict=True,
    return_tensors="pt"
)
inputs = inputs.to(model.device)

# Inference: Generation of the output
generated_ids = model.generate(**inputs, max_new_tokens=128)
generated_ids_trimmed = [
    out_ids[len(in_ids) :] for in_ids, out_ids in zip(inputs.input_ids, generated_ids)
]
output_text = processor.batch_decode(
    generated_ids_trimmed, skip_special_tokens=True, clean_up_tokenization_spaces=False
)
print(output_text)  
```

### Disclaimer

The model has not been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

## Bias, Risks, and Limitations

*The model was not tested for robustness against adversarial prompting.* It is important for users to be aware that our model exhibits certain limitations that warrant consideration. Like many LLMs, the model can hallucinate and occasionally generates irrelevant content, introducing fictional elements that are not grounded in the provided context. Users should also exercise caution in interpreting and validating the model's responses due to the potential inconsistencies.

## More Information

This is the repository for the commercial instruction-tuned model. The model has *not* been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

For more info, please contact us at [SEA-LION Inquiry Form](https://forms.gle/sLCUVb95wmGf43hi6) or <sealion@aisingapore.org>


# SEA-LION v3.5

SEA-LION version 3.5, released in Apr 2025, is our first set of hybrid reasoning models trained on Southeast Asian data, each with their unique strengths:

* [Llama-SEA-LION-v3.5-8B-R](/models/sea-lion-v3.5/llama-sea-lion-v3.5-8b) is suited for cost-efficient, low-latency applications where regional language support and advanced reasoning is critical.
* [Llama-SEA-LION-v3.5-70B-R](/models/sea-lion-v3.5/llama-sea-lion-v3.5-70b) is suited for knowledge-intensive tasks, and for high-demand environments where advanced reasoning and comprehensive language comprehension are essential.

Trained off our SEA-LION v3 models, SEA-LION v3.5 is explicitly enhanced for reasoning tasks with the inclusion of thinking blocks, continuing our mission to create language models that understand and respond with greater cultural awareness and depth across Southeast Asia.

A distinctive feature of SEA-LION v3.5 is its dynamic reasoning toggle - Mode selection is managed through customizable chat template configurations, offering versatile functionality in handling both complex reasoning tasks and general text generation.

For detailed information of each of the SEA-LION v3.5 models, please refer to their individual documentation pages via the links above.


# Llama-SEA-LION-v3.5-8B

## Introduction

Llama-SEA-LION-v3.5-8B-R is a hybrid model offering versatile functionality, handling both complex reasoning tasks and general text generation, with mode selection managed through the tokenizer's chat template.

We performed instruction tuning in English and also in SEA languages such as Filipino, Indonesian, Tamil, Thai and Vietnamese on our [continued pre-trained Llama-SEA-LION-v3-8B-IT](/models/sea-lion-v3/llama-sea-lion-v3-8b), a decoder model using the Llama 3.1 architecture, to create Llama-SEA-LION-v3.5-8B-R.

By leveraging on SEA-LION v3’s strong foundation, Llama-SEA-LION-v3.5-8B-R ensures broader accessibility and usability, empowering diverse communities and use cases throughout the region. It is particularly suited for cost-efficient, low-latency applications where regional language support and advanced reasoning is critical.

For tokenisation, the model employs the default tokenizer used in Llama 3.1 70B Instruct. The model has a context length of 128k.

At a glance:

* **Model type:** Decoder
* **Tokenizer**: Default tokenizer used in Llama 3.1 8B Instruct
* **Context Length**: 128K
* **Available Formats**:
  * Reasoning (Llama-SEA-LION-v3.5-8B-R)
  * GGUF (Llama-SEA-LION-v3.5-8B-R-GGUF)
* **Supported Languages:**
  1. Burmese
  2. Chinese
  3. English
  4. Filipino
  5. Indonesia
  6. Javanese
  7. Khmer
  8. Lao
  9. Malay
  10. Sundanese
  11. Tamil
  12. Thai
  13. Vietnamese
* **License:** [Llama3.1 Community License](https://huggingface.co/meta-llama/Llama-3.1-70B-Instruct/blob/main/LICENSE)

## Llama-SEA-LION-v3.5-8B-R

### Post-Training

Llama-SEA-LION-v3.5-8B-R was trained with an additional series of supervised fine-tuning atop our existing Llama-SEA-LION-v3-8B-IT models across multiple stages, culminating in a final tune with distilled reasoning data of **1.5M** traces from Deepseek-R1 across multiple SEA languages such as Indonesian Tamil, Thai, Tagalog and Vietnamese.

A distinctive feature of Llama-SEA-LION-v3.5-8B-R is its dynamic reasoning toggle. By default, the model operates in a detailed reasoning mode, thoughtfully guiding users through step-by-step solutions. Users retain full control, easily switching reasoning mode off using customizable chat template configurations, allowing concise interactions suitable for straightforward queries. During the tuning process, reasoning and non-reasoning data were simultaneously incorporated, resulting in a versatile model adaptable to varied user needs.

We also scaled up our instruction set to **30M** instructions (across a training time of a month on a single node for the 70B), incorporating the latest in open-source alongside multiple rounds of synthetic aggregation and rewrite, improving the quality of its responses and leaning the model towards accounting for our region's unique cultural diversity and history. It comprises a mix of curated publicly available open source data, synthetic generations from stronger models and handwritten instructions centered around Southeast Asian culture (particularly from Project SEALD), general multilingual instruction-following and chat prompt-response pairs.

Llama-SEA-LION-v3.5-8B-R training uniquely emphasizes region-specific data aggregation and synthetic instruction generation, undergoing multiple refinement cycles and model merging to enhance multilingual proficiency and reasoning capabilities, ensuring exceptional performance across both complex and general-purpose tasks. This ensures that Llama-SEA-LION-v3.5-8B-R maintains its superior performance while mitigating issues like catastrophic forgetting.

### Benchmark Performance

We evaluated Llama-SEA-LION-v3.5-8B-R on both general language capabilities and instruction-following capabilities.

#### General Language Capabilities

For the evaluation of general language capabilities, we employed the [SEA-HELM (also known as BHASA) evaluation benchmark](https://arxiv.org/abs/2309.06085v2) across a variety of tasks.\
These tasks include Question Answering (QA), Sentiment Analysis (Sentiment), Toxicity Detection (Toxicity), Translation in both directions (Eng>Lang & Lang>Eng), Abstractive Summarisation (Abssum), Causal Reasoning (Causal), Natural Language Inference (NLI), and linguistic diagnostics (LINDSEA).

Note: SEA-HELM is implemented using prompts to elicit answers in a strict format. For all tasks, the model is expected to provide an answer tag from which the answer is automatically extracted. For tasks where options are provided, the answer should comprise one of the pre-defined options. The scores for each task is normalised to account for baseline performance due to random chance.

The evaluation was done **zero-shot** with native prompts on a sample of 100-1000 instances for each dataset.

#### Instruction-following Capabilities

Since Llama-SEA-LION-v3.5-8B-R is an instruction-following model, we also evaluated it on instruction-following capabilities with two datasets, SEA-IFEval (based on [IFEval](https://arxiv.org/abs/2311.07911)) and SEA-MTBench (based on [MT-Bench](https://arxiv.org/abs/2306.05685)).

As these two datasets were originally in English, the linguists and native speakers in the team worked together to filter, localise and translate the datasets into the respective target languages to ensure that the examples remained reasonable, meaningful and natural.

**SEA-IFEval**

SEA-IFEval evaluates a model's ability to adhere to constraints provided in the prompt, for example beginning a response with a specific word/phrase or answering with a certain number of sections. Additionally, accuracy is normalised by the proportion of responses in the correct language (if the model performs the task correctly but responds in the wrong language, it is judged to have failed the task).

**SEA-MTBench**

SEA-MTBench evaluates a model's ability to engage in multi-turn (2 turns) conversations and respond in ways that align with human needs. We use `gpt-4-1106-preview` as the judge model and compare against `gpt-3.5-turbo-0125` as the baseline model. The metric used is the weighted win rate against the baseline model (i.e. average win rate across each category: Math, Reasoning, STEM, Humanities, Roleplay, Writing, Extraction). A tie is given a score of 0.5.

For more details on Llama-SEA-LION-v3.5-8B-R benchmark performance, please refer to the [SEA-HELM leaderboard](https://leaderboard.sea-lion.ai/).

<br>

## Llama-SEA-LION-v3.5-8B-R-GGUF

The following quantized GGUF formats of our Llama-SEA-LION-v3.5-8B-R model are available:

* Llama-SEA-LION-v3.5-8B-R-F16
* Llama-SEA-LION-v3.5-8B-R-Q2\_K
* Llama-SEA-LION-v3.5-8B-R-Q3\_K\_M
* Llama-SEA-LION-v3.5-8B-R-Q4\_0
* Llama-SEA-LION-v3.5-8B-R-Q4\_K\_M
* Llama-SEA-LION-v3.5-8B-R-Q5\_0
* Llama-SEA-LION-v3.5-8B-R-Q5\_K\_M
* Llama-SEA-LION-v3.5-8B-R-Q6\_K
* Llama-SEA-LION-v3.5-8B-R-Q8\_0

Please refer to our [Download the Model(s)](#download-the-model-s) section for more details on how to access them.

<br>

## Download the Model(s)

Llama-SEA-LION-v3.5-8B-R models are available for download via the following channels:

[HuggingFace SEA-LION v3.5 Collection](https://huggingface.co/collections/aisingapore/sea-lion-v35-67fc3ab84300d7e6088fa32c)

| Model                         | Download                                                                                                                                           |
| ----------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
| Llama-SEA-LION-v3.5-8B-R      | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v3.5-8B-R)                                                                         |
| Llama-SEA-LION-v3.5-8B-R-GGUF | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v3.5-8B-R-GGUF), [Ollama](https://ollama.com/aisingapore/Llama-SEA-LION-v3.5-8B-R) |

<br>

## Usage

Llama-SEA-LION-v3.5-8B-R can be run using the 🤗 Transformers library

```python
import transformers
import torch

model_id = "aisingapore/Llama-SEA-LION-v3.5-8B-R"

pipeline = transformers.pipeline(
    "text-generation",
    model=model_id,
    model_kwargs={"torch_dtype": torch.bfloat16},
    device_map="auto",
)
messages = [
    {"role": "user", "content": "Apa sentimen dari kalimat berikut ini?\nKalimat: Buku ini sangat membosankan.\nJawaban: "},
]

outputs = pipeline(
    messages,
    max_new_tokens=256,
)
print(outputs[0]["generated_text"][-1])
```

### Thinking Mode Toggle

Llama-SEA-LION-v3.5-8B-R defaults to reasoning with `thinking_mode="on"` passed to the chat template. To use non-thinking mode ie. standard generations, pass `thinking_mode="off"` to the chat template instead.

```python
import transformers
import torch

model_id = "aisingapore/Llama-SEA-LION-v3.5-70B-R"

pipeline = transformers.pipeline(
    "text-generation",
    model=model_id,
    model_kwargs={"torch_dtype": torch.bfloat16},
    device_map="auto",
)

tokenizer = pipeline.tokenizer

messages = [
    {"role": "user", "content": "Apa sentimen dari kalimat berikut ini?\nKalimat: Buku ini sangat membosankan.\nJawaban: "},
]

prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False, thinking_mode="off")

outputs = pipeline(
    prompt,
    max_new_tokens=256,
)

print(outputs[0]["generated_text"])
```

## Disclaimer

It is important for users to be aware that our models exhibits certain limitations that warrant consideration:

1. The model can hallucinate and occasionally generates irrelevant content, introducing fictional elements that are not grounded in the provided context. Users should also exercise caution in interpreting and validating the model's responses due to the potential inconsistencies in its reasoning.
2. The model has not been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.


# Llama-SEA-LION-v3.5-70B

## Introduction

Llama-SEA-LION-v3.5-70B-R is a hybrid model offering versatile functionality, handling both complex reasoning tasks and general text generation, with mode selection managed through the tokenizer's chat template.

We performed instruction tuning in English and also in SEA languages such as Filipino, Indonesian, Tamil, Thai and Vietnamese on our [continued pre-trained Llama-SEA-LION-v3-70B-IT](/models/sea-lion-v3/llama-sea-lion-v3-70b), a decoder model using the Llama 3.1 architecture, to create Llama-SEA-LION-v3.5-70B-R.

By leveraging on SEA-LION v3’s strong foundation, Llama-SEA-LION-v3.5-70B-R ensures broader accessibility and usability, empowering diverse communities and use cases throughout the region. It is particularly suited for knowledge-intensive tasks, and for high-demand environments where advanced reasoning and comprehensive language comprehension are essential.

For tokenisation, the model employs the default tokenizer used in Llama 3.1 70B Instruct. The model has a context length of 128k.

At a glance:

* **Model type:** Decoder
* **Tokenizer**: Default tokenizer used in Llama 3.1 70B Instruct
* **Context Length**: 128K
* **Available Formats**:
  * Reasoning (Llama-SEA-LION-v3.5-70B-R)
  * GGUF (Llama-SEA-LION-v3.5-70B-R-GGUF)
* **Supported Languages:**
  1. Burmese
  2. Chinese
  3. English
  4. Filipino
  5. Indonesia
  6. Javanese
  7. Khmer
  8. Lao
  9. Malay
  10. Sundanese
  11. Tamil
  12. Thai
  13. Vietnamese
* **License:** [Llama3.1 Community License](https://huggingface.co/meta-llama/Llama-3.1-70B-Instruct/blob/main/LICENSE)

## Llama-SEA-LION-v3.5-70B-R

### Post-Training

Llama-SEA-LION-v3.5-70B-R was trained with an additional series of supervised fine-tuning atop our existing Llama-SEA-LION-v3-70B-IT models across multiple stages, culminating in a final tune with distilled reasoning data of **1.5M** traces from Deepseek-R1 across multiple SEA languages such as Indonesian Tamil, Thai, Tagalog and Vietnamese.

A distinctive feature of Llama-SEA-LION-v3.5-70B-R is its dynamic reasoning toggle. By default, the model operates in a detailed reasoning mode, thoughtfully guiding users through step-by-step solutions. Users retain full control, easily switching reasoning mode off using customizable chat template configurations, allowing concise interactions suitable for straightforward queries. During the tuning process, reasoning and non-reasoning data were simultaneously incorporated, resulting in a versatile model adaptable to varied user needs.

We also scaled up our instruction set to **30M** instructions (across a training time of a month on a single node for the 70B), incorporating the latest in open-source alongside multiple rounds of synthetic aggregation and rewrite, improving the quality of its responses and leaning the model towards accounting for our region's unique cultural diversity and history. It comprises a mix of curated publicly available open source data, synthetic generations from stronger models and handwritten instructions centered around Southeast Asian culture (particularly from Project SEALD), general multilingual instruction-following and chat prompt-response pairs.

Llama-SEA-LION-v3.5-70B-R training uniquely emphasizes region-specific data aggregation and synthetic instruction generation, undergoing multiple refinement cycles and model merging to enhance multilingual proficiency and reasoning capabilities, ensuring exceptional performance across both complex and general-purpose tasks. This ensures that Llama-SEA-LION-v3.5-70B-R maintains its superior performance while mitigating issues like catastrophic forgetting.

### Benchmark Performance

We evaluated Llama-SEA-LION-v3.5-70B-R on both general language capabilities and instruction-following capabilities.

#### General Language Capabilities

For the evaluation of general language capabilities, we employed the [SEA-HELM (previously known as BHASA) evaluation benchmark](https://arxiv.org/abs/2309.06085v2) across a variety of tasks.\
These tasks include Question Answering (QA), Sentiment Analysis (Sentiment), Toxicity Detection (Toxicity), Translation in both directions (Eng>Lang & Lang>Eng), Abstractive Summarisation (Abssum), Causal Reasoning (Causal), Natural Language Inference (NLI), and linguistic diagnostics (LINDSEA).

Note: SEA-HELM is implemented using prompts to elicit answers in a strict format. For all tasks, the model is expected to provide an answer tag from which the answer is automatically extracted. For tasks where options are provided, the answer should comprise one of the pre-defined options. The scores for each task is normalised to account for baseline performance due to random chance.

The evaluation was done **zero-shot** with native prompts on a sample of 100-1000 instances for each dataset.

#### Instruction-following Capabilities

Since Llama-SEA-LION-v3.5-70B-R is an instruction-following model, we also evaluated it on instruction-following capabilities with two datasets, SEA-IFEval (based on [IFEval](https://arxiv.org/abs/2311.07911)) and SEA-MTBench (based on [MT-Bench](https://arxiv.org/abs/2306.05685)).

As these two datasets were originally in English, the linguists and native speakers in the team worked together to filter, localise and translate the datasets into the respective target languages to ensure that the examples remained reasonable, meaningful and natural.

**SEA-IFEval**

SEA-IFEval evaluates a model's ability to adhere to constraints provided in the prompt, for example beginning a response with a specific word/phrase or answering with a certain number of sections. Additionally, accuracy is normalised by the proportion of responses in the correct language (if the model performs the task correctly but responds in the wrong language, it is judged to have failed the task).

**SEA-MTBench**

SEA-MTBench evaluates a model's ability to engage in multi-turn (2 turns) conversations and respond in ways that align with human needs. We use `gpt-4-1106-preview` as the judge model and compare against `gpt-3.5-turbo-0125` as the baseline model. The metric used is the weighted win rate against the baseline model (i.e. average win rate across each category: Math, Reasoning, STEM, Humanities, Roleplay, Writing, Extraction). A tie is given a score of 0.5.

For more details on Llama-SEA-LION-v3.5-70B-R benchmark performance, please refer to the [SEA-HELM leaderboard](https://leaderboard.sea-lion.ai/).

<br>

## Llama-SEA-LION-v3.5-70B-R-GGUF

The following quantized GGUF formats of our Llama-SEA-LION-v3.5-70B-R model are available:

* Llama-SEA-LION-v3.5-70B-R-F16
* Llama-SEA-LION-v3.5-70B-R-Q2\_K
* Llama-SEA-LION-v3.5-70B-R-Q3\_K\_M
* Llama-SEA-LION-v3.5-70B-R-Q4\_0
* Llama-SEA-LION-v3.5-70B-R-Q4\_K\_M
* Llama-SEA-LION-v3.5-70B-R-Q5\_0
* Llama-SEA-LION-v3.5-70B-R-Q5\_K\_M
* Llama-SEA-LION-v3.5-70B-R-Q6\_K
* Llama-SEA-LION-v3.5-70B-R-Q8\_0

Please refer to our [Download the Model(s)](#download-the-model-s) section for more details on how to access them.

<br>

## Download the Model(s)

Llama-SEA-LION-v3.5-70B-R models are available for download via the following channels:

[HuggingFace SEA-LION v3.5 Collection](https://huggingface.co/collections/aisingapore/sea-lion-v35-67fc3ab84300d7e6088fa32c)

| Model                          | Download                                                                                                                                             |
| ------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------- |
| Llama-SEA-LION-v3.5-70B-R      | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v3.5-70B-R)                                                                          |
| Llama-SEA-LION-v3.5-70B-R-GGUF | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v3.5-70B-R-GGUF), [Ollama](https://ollama.com/aisingapore/Llama-SEA-LION-v3.5-70B-R) |

<br>

## Usage

Llama-SEA-LION-v3.5-70B-R can be run using the 🤗 Transformers library

```python
import transformers
import torch

model_id = "aisingapore/Llama-SEA-LION-v3.5-70B-R"

pipeline = transformers.pipeline(
    "text-generation",
    model=model_id,
    model_kwargs={"torch_dtype": torch.bfloat16},
    device_map="auto",
)
messages = [
    {"role": "user", "content": "Apa sentimen dari kalimat berikut ini?\nKalimat: Buku ini sangat membosankan.\nJawaban: "},
]

outputs = pipeline(
    messages,
    max_new_tokens=256,
)
print(outputs[0]["generated_text"][-1])
```

### Thinking Mode Toggle

Llama-SEA-LION-v3.5-70B-R defaults to reasoning with `thinking_mode="on"` passed to the chat template. To use non-thinking mode ie. standard generations, pass `thinking_mode="off"` to the chat template instead.

```python
import transformers
import torch

model_id = "aisingapore/Llama-SEA-LION-v3.5-70B-R"

pipeline = transformers.pipeline(
    "text-generation",
    model=model_id,
    model_kwargs={"torch_dtype": torch.bfloat16},
    device_map="auto",
)

tokenizer = pipeline.tokenizer

messages = [
    {"role": "user", "content": "Apa sentimen dari kalimat berikut ini?\nKalimat: Buku ini sangat membosankan.\nJawaban: "},
]

prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True, tokenize=False, thinking_mode="off")

outputs = pipeline(
    prompt,
    max_new_tokens=256,
)

print(outputs[0]["generated_text"])
```

## Disclaimer

It is important for users to be aware that our models exhibits certain limitations that warrant consideration:

1. The model can hallucinate and occasionally generates irrelevant content, introducing fictional elements that are not grounded in the provided context. Users should also exercise caution in interpreting and validating the model's responses due to the potential inconsistencies in its reasoning.
2. The model has not been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

<br>


# SEA-LION v3

SEA-LION version 3, released in Dec 2024, is a collection of 3 models (and their variants), each with their unique strengths:

* [Gemma-SEA-LION-v3-9B](/models/sea-lion-v3/gemma-sea-lion-v3-9b), based on Gemma2, is our best performing SEA-LION v3 model on SEA-HELM benchmarks for similar sized models
* [Llama-SEA-LION-v3-8B](/models/sea-lion-v3/llama-sea-lion-v3-8b), based on Llama 3.1 8B, now provides a larger context length of 128K
* [Llama-SEA-LION-v3-70B](/models/sea-lion-v3/llama-sea-lion-v3-70b), based on Llama 3.1 70B, is our largest model to date, also providing a 128K context length

With enhanced natural language reasoning (NLR) abilities and superior instruction-following in SEA languages, SEA-LION v3 sets new standards in multilingual AI for the region.

For detailed information of each of the SEA-LION v3 models, please refer to their individual documentation pages via the links above.


# Gemma-SEA-LION-v3-9B

## Introduction

Our Gemma-SEA-LION-v3-9B models have been continued pre-trained on top of the Gemma2 base model that is 9 billion parameters in size, and has a **context length of 8192**.

The training data for Gemma-SEA-LION-v3-9B comprises approximately **200B tokens** of Burmese, Chinese, English, Filipino, Indonesia, Khmer, Lao, Malay, Tamil, Thai and Vietnamese. The training data is sampled from the Dolma dataset, the original SEA-LION pretraining data, and Wiki sources for Burmese, Chinese, English, Filipino, Khmer, Lao, Malay, Indonesian, Tamil, Thai and Vietnamese. In addition, the training data for SFT for the Instruct model includes Javanese and Sudanese.

Gemma-SEA-LION-v3-9B benefits from the strong Gemma2 performance in Southeast Asian (SEA) languages, allowing it to significantly outperform its predecessors across SEA evaluation metrics, indicating improved language capabilities for the SEA region.

Due to its pre-training and fine-tuning data mix, Gemma-SEA-LION-v3-9B exhibits a boost in natural language reasoning (NLR) abilities in Indonesian, Tamil, Thai and Vietnamese, and achieves state-of-the-art performances in SEA instruction-following and multi-turn chat, while retaining Gemma2’s general abilities.

At a glance:

* **Model type:** Decoder
* **Tokenizer**: Default tokenizer used in Gemma2 9B
* **Training Data Size**: 200B tokens of SEA data
* **Context Length**: 8192
* **Available Formats**:
  * Base (Gemma-SEA-LION-v3-9B)
  * Instruct (Gemma-SEA-LION-v3-9B-IT)
  * GGUF (Gemma-SEA-LION-v3-9B-IT-GGUF)
* **Supported Languages:**
  1. Burmese
  2. Chinese
  3. English
  4. Filipino
  5. Indonesia
  6. Khmer
  7. Lao
  8. Malay
  9. Tamil
  10. Thai
  11. Vietnamese
* **License:** [Gemma Community License](https://ai.google.dev/gemma/terms)

## Gemma-SEA-LION-v3-9B

### Training Infrastructure

Gemma-SEA-LION-v3-9B was trained using [MosaicML Composer](https://github.com/mosaicml/composer) on the following hardware:

| Training Details     | Gemma-SEA-LION-v3-9B |
| -------------------- | :------------------: |
| SingTel HGX-100      |      8 instances     |
| Nvidia H100 80GB GPU |          64          |
| Training Duration    |        10 days       |

**Configuration**

| HyperParameter    |  Gemma-SEA-LION-v3-9B |
| ----------------- | :-------------------: |
| Precision         |        bfloat16       |
| Optimizer         |    decoupled\_adamw   |
| Scheduler         | weight\_stable\_decay |
| Learning Rate     |         1.0e-5        |
| Global Batch Size |          512          |
| Micro Batch Size  |           1           |

### Tokenizer

For tokenisation, the model employs the default tokenizer used in Gemma2 9B.

### Training Data

Gemma-SEA-LION-v3-9B base model was continued pre-trained on 200B tokens of the following data:

| Language                 | Source           | Total Tokens (B) | Percentage (%) | Total percentage (%) |
| ------------------------ | ---------------- | ---------------- | -------------- | -------------------- |
| Code                     | Stackv2          | 40               | 20             | 20                   |
| English                  | Dolma            | 37.5             | 18.75          | 25                   |
|                          | Fineweb-Edu      | 7.5              | 3.75           |                      |
|                          | Others           | 5                | 2.5            |                      |
| Chinese                  | SEA-LION Pile v1 | 12               | 6              | 13                   |
|                          | Others           | 14               | 7              |                      |
| Vietnamese               | SEA-LION Pile v1 | 8.4              | 4.2            | 13                   |
|                          | VinBigData       | 16               | 8              |                      |
|                          | Others           | 1.6              | 0.8            |                      |
| Indonesian               | SEA-LION Pile v1 | 7                | 3.5            | 13                   |
|                          | SEA-LION Pile v2 | 7                | 3.5            |                      |
|                          | Others           | 12               | 6              |                      |
| Thai                     | SEA-LION Pile v1 | 10.7             | 5.35           | 10                   |
|                          | WangChanBERTa    | 8.5              | 4.25           |                      |
|                          | Others           | 0.8              | 0.4            |                      |
| Filipino - Malay - Tamil | SEA-LION Pile v1 | 4.28             | 2.14           | 3                    |
|                          | Others           | 1.72             | 0.86           |                      |
| Khmer - Lao - Burmese    | SEA-LION Pile v1 | 5.2              | 2.6            | 3                    |
|                          | Others           | 0.8              | 0.4            |                      |

Note:

* All token counts are counted using Gemma2 9B tokenizer
* SEA-LION Pile v1 is processed from Common Crawl WET, which is published [here](https://huggingface.co/datasets/aisingapore/sea-lion-pile). The cutoff date of this version is September 2020.
* SEA-LION Pile v2 is processed from Common Crawl WARC from October 2020 to April 2024.
* Tamil news is sourced with permission from [Seithi](https://seithi.mediacorp.sg/)

### Benchmark Performance

We evaluated Gemma-SEA-LION-v3-9B base model on general language capabilities.

#### General Language Capabilities

For the evaluation of general language capabilities, we employed the [SEA-HELM](/benchmarks/sea-helm) (also known as BHASA) evaluation benchmark across a variety of tasks.\
These tasks include Question Answering (QA), Sentiment Analysis (Sentiment), Toxicity Detection (Toxicity), Translation in both directions (Eng>Lang & Lang>Eng), Abstractive Summarization (Summ), Causal Reasoning (Causal) and Natural Language Inference (NLI).

Note: SEA-HELM is implemented using prompts to elicit answers in a strict format. For all tasks, the model is expected to provide an answer tag from which the answer is automatically extracted. For tasks where options are provided, the answer should comprise one of the pre-defined options. The scores for each task is normalised to account for baseline performance due to random chance.

The evaluation was done **five-shot** with native prompts on a sample of 100-1000 instances for each dataset.

For more details on Gemma-SEA-LION-v3-9B base benchmark performance, please refer to the SEA-HELM leaderboard, <https://leaderboard.sea-lion.ai/>

<br>

## Gemma-SEA-LION-v3-9B-IT

Gemma-SEA-LION-v3-9B-IT is a multilingual instruction-following model which has been fine-tuned with around **500,000 English instruction-completion pairs** alongside a larger pool of around **1,000,000 instruction-completion pairs** from other ASEAN languages, such as Indonesian, Thai and Vietnamese.

### Fine-Tuning Methodology

Gemma-SEA-LION-v3-9B-IT was built using a combination of a full parameter fine-tune, on-policy alignment, and model merges of the best performing checkpoints. The training process for fine-tuning was approximately 15 hours, with alignment taking 2 hours, both on 8x H100-80GB GPUs.

### Fine-Tuning Data

Gemma-SEA-LION-v3-9B-IT was trained on a wide range of synthetic instructions, alongside publicly available instructions hand-curated by the team with the assistance of native speakers. In addition, special care was taken to ensure that the datasets used had commercially permissive licenses through verification with the original data source.

#### Indonesian, Javanese & Sudanese Specific SEA-LION

Our partners at GoTo have continued pretrained and instruction tuned a variant of Gemma-SEA-LION-v3-9B, specifically enhancing its capabilities for Indonesian, Javanese, and Sundanese languages. Find the continued pretrained model at [Gemma2 9B CPT SahabatAIv1 Base](https://huggingface.co/GoToCompany/gemma2-9b-cpt-sahabatai-v1-base), and its corresponding instructioned tuned version at [Gemma2 9B CPT SahabatAIv1 Instruct](https://huggingface.co/GoToCompany/gemma2-9b-cpt-sahabatai-v1-instruct).

### Benchmark Performance

We evaluated Gemma-SEA-LION-v3-9B-IT on both general language capabilities and instruction-following capabilities.

#### General Language Capabilities

For the evaluation of general language capabilities, we employed the [SEA-HELM](/benchmarks/sea-helm) (also known as BHASA) evaluation benchmark across a variety of tasks.\
These tasks include Question Answering (QA), Sentiment Analysis (Sentiment), Toxicity Detection (Toxicity), Translation in both directions (Eng>Lang & Lang>Eng), Abstractive Summarization (Summ), Causal Reasoning (Causal) and Natural Language Inference (NLI).

Note: SEA-HELM is implemented using prompts to elicit answers in a strict format. For all tasks, the model is expected to provide an answer tag from which the answer is automatically extracted. For tasks where options are provided, the answer should comprise one of the pre-defined options. The scores for each task is normalised to account for baseline performance due to random chance.

The evaluation was done **zero-shot** with native prompts on a sample of 100-1000 instances for each dataset.

#### Instruction-following Capabilities

Since Gemma-SEA-LION-v3-9B-IT is an instruction-following model, we also evaluated it on instruction-following capabilities with two datasets, [IFEval](https://arxiv.org/abs/2311.07911) and [MT-Bench](https://arxiv.org/abs/2306.05685).

As these two datasets were originally in English, the linguists and native speakers in the team worked together to filter, localize and translate the datasets into the respective target languages to ensure that the examples remained reasonable, meaningful and natural.

**IFEval**

IFEval evaluates a model's ability to adhere to constraints provided in the prompt, for example beginning a response with a specific word/phrase or answering with a certain number of sections. Additionally, accuracy is normalized by the proportion of responses in the correct language (if the model performs the task correctly but responds in the wrong language, it is judged to have failed the task).

**MT-Bench**

MT-Bench evaluates a model's ability to engage in multi-turn (2 turns) conversations and respond in ways that align with human needs. We use `gpt-4-1106-preview` as the judge model and compare against `gpt-3.5-turbo-0125` as the baseline model. The metric used is the weighted win rate against the baseline model (i.e. average win rate across each category: Math, Reasoning, STEM, Humanities, Roleplay, Writing, Extraction). A tie is given a score of 0.5.

For more details on Gemma-SEA-LION-v3-9B-IT benchmark performance, please refer to the SEA-HELM leaderboard, <https://leaderboard.sea-lion.ai/>

<br>

## Gemma-SEA-LION-v3-9B-IT-GGUF

The following quantized GGUF formats of our Gemma-SEA-LION-v3-9B-IT model are available:

* Gemma-SEA-LION-v3-9B-IT-F16
* Gemma-SEA-LION-v3-9B-IT-Q2\_K
* Gemma-SEA-LION-v3-9B-IT-Q3\_K\_M
* Gemma-SEA-LION-v3-9B-IT-Q4\_0
* Gemma-SEA-LION-v3-9B-IT-Q4\_K\_M
* Gemma-SEA-LION-v3-9B-IT-Q5\_0
* Gemma-SEA-LION-v3-9B-IT-Q5\_K\_M
* Gemma-SEA-LION-v3-9B-IT-Q6\_K
* Gemma-SEA-LION-v3-9B-IT-Q8\_0

Please refer to our [Download the Model(s)](#download-the-model-s) section for more details on how to access them.

## Download the Model(s)

Gemma-SEA-LION-v3-9B models are available for download via the following channels:

[HuggingFace SEA-LION v3 Collection](https://huggingface.co/collections/aisingapore/sea-lionv3-672589a39cdadd6a5b199581)

| Model                        | Download                                                                                                                                                          |
| ---------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Gemma-SEA-LION-v3-9B         | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-LION-v3-9B), [Kaggle](https://www.kaggle.com/models/ai-singapore/gemma2-9b-cpt-sea-lionv3-base)        |
| Gemma-SEA-LION-v3-9B-IT      | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-LION-v3-9B-IT), [Kaggle](https://www.kaggle.com/models/ai-singapore/gemma2-9b-cpt-sea-lionv3-instruct) |
| Gemma-SEA-LION-v3-9B-IT-GGUF | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-LION-v3-9B-IT-GGUF), [Ollama](https://ollama.com/aisingapore/Gemma-SEA-LION-v3-9B-IT)                  |

<br>

## Usage

**NOTE** This model has not been trained to use a system prompt or to use tool calling.

Gemma-SEA-LION-v3-9B-IT can be run using the 🤗 Transformers library

```python
# Please use transformers==4.45.2

import transformers
import torch

model_id = "aisingapore/Gemma-SEA-LION-v3-9B-IT"

pipeline = transformers.pipeline(
    "text-generation",
    model=model_id,
    model_kwargs={"torch_dtype": torch.bfloat16},
    device_map="auto",
)
messages = [
    {"role": "user", "content": "Apa sentimen dari kalimat berikut ini?\nKalimat: Buku ini sangat membosankan.\nJawaban: "},
]

outputs = pipeline(
    messages,
    max_new_tokens=256,
)
print(outputs[0]["generated_text"][-1])
```

## Disclaimer

It is important for users to be aware that our models exhibits certain limitations that warrant consideration:

1. The model can hallucinate and occasionally generates irrelevant content, introducing fictional elements that are not grounded in the provided context. Users should also exercise caution in interpreting and validating the model's responses due to the potential inconsistencies in its reasoning.
2. Current SEA-LION models, including this commercially permissive release, have not been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

<br>

## References

### Thai Pre-Training Data Reference

```bibtex
@misc{lowphansirikul2021wangchanberta,
    title={WangchanBERTa: Pretraining transformer-based Thai Language Models},
    author={Lalita Lowphansirikul and Charin Polpanumas and Nawat Jantrakulchai and Sarana Nutanong},
    year={2021},
    eprint={2101.09635},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}
```


# Llama-SEA-LION-v3-8B

## Introduction

Our Llama-SEA-LION-v3-8B models have been continued pre-trained on top of [Llama 3.1 8B Instruct](https://huggingface.co/meta-llama/Meta-Llama-3.1-8B-Instruct) that is 8 billion parameters in size, with **context length of 128K tokens**, making it one of the SEA-LION models with the longest context length to date. It achieves state-of-the-art performance on regional benchmarks like SEA-HELM and outperforms models such as Llama 3.3 70B Instruct on key metrics.

Llama-SEA-LION-v3-8B was trained on data comprised of approximately **200B tokens** across 11 SEA languages: Burmese, Chinese, English, Filipino, Indonesia, Khmer, Lao, Malay, Tamil, Thai and Vietnamese.

Llama-SEA-LION-v3-8B-IT was fine-tuned in two stages on approximately **12.3M English instruction-completion pairs** alongside a pool of **4.5M Southeast Asian instruction-completion pairs** from SEA languages such as Indonesian, Javanese, Sundanese, Tamil, Thai and Vietnamese.

At a glance:

* **Model type:** Decoder
* **Tokenizer**: Default tokenizer used in Llama 3.1 8B Instruct
* **Context Length**: 128K
* **Available Formats**:
  * Base (Llama-SEA-LION-v3-8B)
  * Instruct (Llama-SEA-LION-v3-8B-IT)
  * GGUF (Llama-SEA-LION-v3-8B-IT-GGUF)
* **Supported Languages:**
  1. Burmese
  2. Chinese
  3. English
  4. Filipino
  5. Indonesia
  6. Javanese (Instruct/GGUF only)
  7. Khmer
  8. Lao
  9. Malay
  10. Sundanese (Instruct/GGUF only)
  11. Tamil
  12. Thai
  13. Vietnamese
* **License:** [Llama3.1 Community License](https://github.com/meta-llama/llama-models/blob/main/models/llama3_1/LICENSE)

## Llama-SEA-LION-v3-8B

### Training Infrastructure

Llama-SEA-LION-v3-8B was trained using [MosaicML Composer](https://github.com/mosaicml/composer) on the following hardware:

| Training Details      | Llama-SEA-LION-v3-8B |
| --------------------- | :------------------: |
| AWS p5e.48xlarge      |      8 instances     |
| Nvidia H200 140GB GPU |          64          |
| Training Duration     |       136 Hours      |

**Configuration**

| HyperParameter    |  Llama-SEA-LION-v3-8B |
| ----------------- | :-------------------: |
| Precision         |        bfloat16       |
| Optimizer         |    decoupled\_adamw   |
| Scheduler         | weight\_stable\_decay |
| Learning Rate     |         1.0e-5        |
| Global Batch Size |          512          |

### Tokenizer

For tokenisation, the model employs the default tokenizer used in Llama 3.1 8B Instruct.

### Training Data

Llama-SEA-LION-v3-8B base model was continued pre-trained on 200B tokens of the following data:

| Language                 | Source                               | Total Tokens (B) | Percentage (%) | Total percentage (%) |
| ------------------------ | ------------------------------------ | ---------------- | -------------- | -------------------- |
| Code                     | Stackv2                              | 40               | 20             | 20                   |
| English                  | Dolma                                | 37.5             | 18.75          | 25                   |
|                          | Fineweb-Edu                          | 7.5              | 3.75           |                      |
|                          | Others                               | 5                | 2.5            |                      |
| Chinese                  | SEA-LION Pile v1                     | 12               | 6              | 13                   |
|                          | Others                               | 14               | 7              |                      |
| Vietnamese               | SEA-LION Pile v1                     | 8.4              | 4.2            | 13                   |
|                          | VinBigData                           | 16               | 8              |                      |
|                          | Others                               | 1.6              | 0.8            |                      |
| Indonesian               | SEA-LION Pile v1                     | 7                | 3.5            | 13                   |
|                          | SEA-LION Pile v2                     | 7                | 3.5            |                      |
|                          | Others                               | 12               | 6              |                      |
| Thai                     | SEA-LION Pile v1                     | 10.7             | 5.35           | 10                   |
|                          | WangChanBERTa                        | 8.5              | 4.25           |                      |
|                          | Others                               | 0.8              | 0.4            |                      |
| Filipino - Malay - Tamil | SEA-LION Pile v1, AI4Bharat Sangraha | 4.28             | 2.14           | 3                    |
|                          | Others                               | 1.72             | 0.86           |                      |
| Khmer - Lao - Burmese    | SEA-LION Pile v1                     | 5.2              | 2.6            | 3                    |
|                          | Others                               | 0.8              | 0.4            |                      |

Note:

* All token counts are counted using Llama 3.1 8B Instruct tokenizer
* SEA-LION Pile v1 is processed from Common Crawl WET, which is published [here](https://huggingface.co/datasets/aisingapore/sea-lion-pile). The cutoff date of this version is September 2020.
* SEA-LION Pile v2 is processed from Common Crawl WARC from October 2020 to April 2024.
* Tamil data from Sangraha is published [here](https://huggingface.co/datasets/ai4bharat/sangraha). The paper can be found [here](https://arxiv.org/abs/2403.06350).
* Tamil news is sourced with permission from [Seithi](https://seithi.mediacorp.sg/)

### Benchmark Performance

We evaluated Llama-SEA-LION-v3-8B base model on general language capabilities and constraint-following behaviour.

#### General Language Capabilities and Constraint-following Behaviour

For the evaluation of general language capabilities, we employed the [SEA-HELM](/benchmarks/sea-helm) (also known as BHASA) evaluation benchmark across a variety of tasks.\
These tasks include Question Answering (QA), Sentiment Analysis (Sentiment), Toxicity Detection (Toxicity), Translation in both directions (Eng>Lang & Lang>Eng), Abstractive Summarisation (Abssum), Causal Reasoning (Causal) and Natural Language Inference (NLI).

Note: SEA-HELM is implemented using prompts to elicit answers in a strict format. For all tasks, the model is expected to provide an answer tag from which the answer is automatically extracted. For tasks where options are provided, the answer should comprise one of the pre-defined options. The scores for each task is normalised to account for baseline performance due to random chance.

The evaluation was done **five-shot** with native prompts on a sample of 100-1000 instances for each dataset.

Following the implementation of IFEval in OpenLLM leaderboard, we also implement SEA-IFEval to provide a comparison of the ability of the model to follow specific constraints in English and in SEA languages.

**SEA-IFEval**

Based on [IFEval](https://arxiv.org/abs/2311.07911), the linguists and native speakers in the team worked together to filter, localise and translate the datasets into the respective target languages to ensure that the examples remained reasonable, meaningful and natural.

SEA-IFEval evaluates a model's ability to adhere to constraints provided in the prompt, for example beginning a response with a specific word/phrase or answering with a certain number of sections. Additionally, accuracy is normalised by the proportion of responses in the correct language (if the model performs the task correctly but responds in the wrong language, it is judged to have failed the task).

For more details on Llama-SEA-LION-v3-8B base benchmark performance, please refer to the SEA-HELM leaderboard, <https://leaderboard.sea-lion.ai/>.

## Llama-SEA-LION-v3-8B-IT

### Fine-Tuning Methodology

Llama-SEA-LION-v3-8B-IT is a multilingual instruction-following model that has been fine-tuned in two stages on approximately 12.3M English instruction-completion pairs alongside a pool of 4.5M Southeast Asian instruction-completion pairs from SEA languages such as Indonesian, Javanese, Sundanese, Tamil, Thai and Vietnamese.

We performed instruction tuning in English and also in SEA languages such as Indonesian, Javanese, Sundanese, Tamil, Thai and Vietnamese on our continued pre-trained Llama-SEA-LION-v3-8B, to create Llama-SEA-LION-v3-8B-IT.

### Fine-Tuning Data

Llama-SEA-LION-v3-8B-IT was trained on a wide range of synthetic instructions, alongside publicly available instructions hand-curated by the team with the assistance of native speakers. In addition, special care was taken to ensure that the datasets used had commercially permissive licenses through verification with the original data source.

### Benchmark Performance

We evaluated Llama-SEA-LION-v3-8B Instruct on both general language capabilities and instruction-following capabilities.

#### General Language Capabilities

For the evaluation of general language capabilities, we employed the [SEA-HELM](/benchmarks/sea-helm) (also known as BHASA) evaluation benchmark across a variety of tasks.\
These tasks include Question Answering (QA), Sentiment Analysis (Sentiment), Toxicity Detection (Toxicity), Translation in both directions (Eng>Lang & Lang>Eng), Abstractive Summarisation (Abssum), Causal Reasoning (Causal) and Natural Language Inference (NLI).

Note: SEA-HELM is implemented using prompts to elicit answers in a strict format. For all tasks, the model is expected to provide an answer tag from which the answer is automatically extracted. For tasks where options are provided, the answer should comprise one of the pre-defined options. The scores for each task is normalised to account for baseline performance due to random chance.

The evaluation was done **zero-shot** with native prompts on a sample of 100-1000 instances for each dataset.

#### Instruction-following Capabilities

Since Llama-SEA-LION-v3-8B-IT is an instruction-following model, we also evaluated it on instruction-following capabilities with two datasets, SEA-IFEval (based on [IFEval](https://arxiv.org/abs/2311.07911)) and SEA-MTBench (based on [MT-Bench](https://arxiv.org/abs/2306.05685)).

As these two datasets were originally in English, the linguists and native speakers in the team worked together to filter, localise and translate the datasets into the respective target languages to ensure that the examples remained reasonable, meaningful and natural.

**SEA-IFEval**

SEA-IFEval evaluates a model's ability to adhere to constraints provided in the prompt, for example beginning a response with a specific word/phrase or answering with a certain number of sections. Additionally, accuracy is normalised by the proportion of responses in the correct language (if the model performs the task correctly but responds in the wrong language, it is judged to have failed the task).

**SEA-MTBench**

SEA-MTBench evaluates a model's ability to engage in multi-turn (2 turns) conversations and respond in ways that align with human needs. We use `gpt-4-1106-preview` as the judge model and compare against `gpt-3.5-turbo-0125` as the baseline model. The metric used is the weighted win rate against the baseline model (i.e. average win rate across each category: Math, Reasoning, STEM, Humanities, Roleplay, Writing, Extraction). A tie is given a score of 0.5.

For more details on Llama-SEA-LION-v3-8B-IT benchmark performance, please refer to the SEA-HELM leaderboard, <https://leaderboard.sea-lion.ai/>.

## Llama-SEA-LION-v3-8B-IT-GGUF

The following quantized GGUF formats of our Llama-SEA-LION-v3-8B-IT model are available:

* Llama-SEA-LION-v3-8B-IT-F16
* Llama-SEA-LION-v3-8B-IT-Q2\_K
* Llama-SEA-LION-v3-8B-IT-Q3\_K\_M
* Llama-SEA-LION-v3-8B-IT-Q4\_0
* Llama-SEA-LION-v3-8B-IT-Q4\_K\_M
* Llama-SEA-LION-v3-8B-IT-Q5\_0
* Llama-SEA-LION-v3-8B-IT-Q5\_K\_M
* Llama-SEA-LION-v3-8B-IT-Q6\_K
* Llama-SEA-LION-v3-8B-IT-Q8\_0

Please refer to our [Download the Model(s)](#download-the-model-s) section for more details on how to access them.

<br>

## Download the Model(s)

Llama-SEA-LION-v3-8B models are available for download via the following channels:

[HuggingFace SEA-LION v3 Collection](https://huggingface.co/collections/aisingapore/sea-lionv3-672589a39cdadd6a5b199581)

| Model                        | Download                                                                                                                                                            |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Llama-SEA-LION-v3-8B         | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v3-8B), [Kaggle](https://www.kaggle.com/models/ai-singapore/llama3.1-8b-cpt-sea-lionv3-base)        |
| Llama-SEA-LION-v3-8B-IT      | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v3-8B-IT), [Kaggle](https://www.kaggle.com/models/ai-singapore/llama3.1-8b-cpt-sea-lionv3-instruct) |
| Llama-SEA-LION-v3-8B-IT-GGUF | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v3-8B-IT-GGUF), [Ollama](https://ollama.com/aisingapore/Llama-SEA-LION-v3-8B-IT)                    |

<br>

## Usage

Llama-SEA-LION-v3-8B-IT can be run using the 🤗 Transformers library

```python
import transformers
import torch

model_id = "aisingapore/Llama-SEA-LION-v3-8B-IT"

pipeline = transformers.pipeline(
    "text-generation",
    model=model_id,
    model_kwargs={"torch_dtype": torch.bfloat16},
    device_map="auto",
)
messages = [
    {"role": "user", "content": "Apa sentimen dari kalimat berikut ini?\nKalimat: Buku ini sangat membosankan.\nJawaban: "},
]

outputs = pipeline(
    messages,
    max_new_tokens=256,
)
print(outputs[0]["generated_text"][-1])
```

## Disclaimer

It is important for users to be aware that our models exhibits certain limitations that warrant consideration:

1. The model can hallucinate and occasionally generates irrelevant content, introducing fictional elements that are not grounded in the provided context. Users should also exercise caution in interpreting and validating the model's responses due to the potential inconsistencies in its reasoning.
2. The model has not been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

<br>

## References

### Thai Pre-Training Data Reference

```bibtex
@misc{lowphansirikul2021wangchanberta,
    title={WangchanBERTa: Pretraining transformer-based Thai Language Models},
    author={Lalita Lowphansirikul and Charin Polpanumas and Nawat Jantrakulchai and Sarana Nutanong},
    year={2021},
    eprint={2101.09635},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}
```


# Llama-SEA-LION-v3-70B

## Introduction

Our Llama-SEA-LION-v3-70B models have been continued pre-trained on top of [Llama 3.1 70B Instruct](https://huggingface.co/meta-llama/Llama-3.1-70B-Instruct) that is 70 billion parameters in size. Similar to our Llama-SEA-LION-v3-8B model, our Llama-SEA-LION-v3-70B also has a **context length of 128K tokens**, making them our SEA-LION models with the longest context length to date.

Llama-SEA-LION-v3-70B was trained on data comprised of approximately **200B tokens** across 11 SEA languages: Burmese, Chinese, English, Filipino, Indonesia, Khmer, Lao, Malay, Tamil, Thai and Vietnamese.

Llama-SEA-LION-v3-70B-IT was fine-tuned in two stages on approximately **12.3M English instruction-completion pairs** alongside a pool of **4.5M Southeast Asian instruction-completion pairs** from SEA languages such as Indonesian, Javanese, Sundanese, Tamil, Thai and Vietnamese.

At a glance:

* **Model type:** Decoder
* **Tokenizer**: Default tokenizer used in Llama 3.1 70B Instruct
* **Available Formats**:
  * Base (Llama-SEA-LION-v3-70B)
  * Instruct (Llama-SEA-LION-v3-70B-IT)
  * GGUF (Llama-SEA-LION-v3-70B-IT-GGUF)
* **Languages supported:**
  1. Burmese
  2. Chinese
  3. English
  4. Filipino
  5. Indonesia
  6. Javanese (Instruct/GGUF only)
  7. Khmer
  8. Lao
  9. Malay
  10. Sundanese (Instruct/GGUF only)
  11. Tamil
  12. Thai
  13. Vietnamese
* **License:** [Llama3.1 Community License](https://github.com/meta-llama/llama-models/blob/main/models/llama3_1/LICENSE)

## Llama-SEA-LION-v3-70B

### Training Infrastructure

Llama-SEA-LION-v3-70B was trained in two stages using [MosaicML Composer](https://github.com/mosaicml/composer) on the following hardware:

| Stage        | Training Details      |    Llama-SEA-LION-v3-70B    |
| ------------ | --------------------- | :-------------------------: |
| First Stage  | AWS p5e.48xlarge      |         8 instances         |
|              | Nvidia H200 140GB GPU |              64             |
|              | Training Duration     |   200 hrs (step 0 - 9000)   |
| Second Stage | SingTel HGX-100       |         16 instances        |
|              | Nvidia H100 80GB GPU  |             128             |
|              | Training Duration     | 495 hrs (step 9000 - 47684) |

### Configuration

| HyperParameter    | Llama-SEA-LION-v3-70B |
| ----------------- | :-------------------: |
| Precision         |        bfloat16       |
| Optimizer         |    decoupled\_adamw   |
| Scheduler         | weight\_stable\_decay |
| Learning Rate     |         1.0e-5        |
| Global Batch Size |          512          |

### Tokenizer

For tokenisation, the model employs the default tokenizer used in Llama 3.1 70B Instruct.

### Training Data

Llama-SEA-LION-v3-70B base model was continued pre-trained on 200B tokens of the following data:

| Language                 | Source                               | Total Tokens (B) | Percentage (%) | Total percentage (%) |
| ------------------------ | ------------------------------------ | ---------------- | -------------- | -------------------- |
| Code                     | Stackv2                              | 40               | 20             | 20                   |
| English                  | Dolma                                | 37.5             | 18.75          | 25                   |
|                          | Fineweb-Edu                          | 7.5              | 3.75           |                      |
|                          | Others                               | 5                | 2.5            |                      |
| Chinese                  | SEA-LION Pile v1                     | 12               | 6              | 13                   |
|                          | Others                               | 14               | 7              |                      |
| Vietnamese               | SEA-LION Pile v1                     | 8.4              | 4.2            | 13                   |
|                          | VinBigData                           | 16               | 8              |                      |
|                          | Others                               | 1.6              | 0.8            |                      |
| Indonesian               | SEA-LION Pile v1                     | 7                | 3.5            | 13                   |
|                          | SEA-LION Pile v2                     | 7                | 3.5            |                      |
|                          | Others                               | 12               | 6              |                      |
| Thai                     | SEA-LION Pile v1                     | 10.7             | 5.35           | 10                   |
|                          | WangChanBERTa                        | 8.5              | 4.25           |                      |
|                          | Others                               | 0.8              | 0.4            |                      |
| Filipino - Malay - Tamil | SEA-LION Pile v1, AI4Bharat Sangraha | 4.28             | 2.14           | 3                    |
|                          | Others                               | 1.72             | 0.86           |                      |
| Khmer - Lao - Burmese    | SEA-LION Pile v1                     | 5.2              | 2.6            | 3                    |
|                          | Others                               | 0.8              | 0.4            |                      |

Note:

* All token counts are counted using Llama 3.1 70B Instruct tokenizer
* SEA-LION Pile v1 is processed from Common Crawl WET, which is published [here](https://huggingface.co/datasets/aisingapore/sea-lion-pile). The cutoff date of this version is September 2020.
* SEA-LION Pile v2 is processed from Common Crawl WARC from October 2020 to April 2024.
* Tamil data from Sangraha is published [here](https://huggingface.co/datasets/ai4bharat/sangraha). The paper can be found [here](https://arxiv.org/abs/2403.06350).
* Tamil news is sourced with permission from [Seithi](https://seithi.mediacorp.sg/)

### Benchmark Performance

We evaluated Llama-SEA-LION-v3-70B base model on general language capabilities and constraint-following behaviour.

#### General Language Capabilities and Constraint-following Behaviour

For the evaluation of general language capabilities, we employed the [SEA-HELM](/benchmarks/sea-helm) (also known as BHASA) evaluation benchmark across a variety of tasks.\
These tasks include Question Answering (QA), Sentiment Analysis (Sentiment), Toxicity Detection (Toxicity), Translation in both directions (Eng>Lang & Lang>Eng), Abstractive Summarisation (Abssum), Causal Reasoning (Causal) and Natural Language Inference (NLI).

Note: SEA-HELM is implemented using prompts to elicit answers in a strict format. For all tasks, the model is expected to provide an answer tag from which the answer is automatically extracted. For tasks where options are provided, the answer should comprise one of the pre-defined options. The scores for each task is normalised to account for baseline performance due to random chance.

The evaluation was done **five-shot** with native prompts on a sample of 100-1000 instances for each dataset.

Following the implementation of IFEval in OpenLLM leaderboard, we also implement SEA-IFEval to provide a comparison of the ability of the model to follow specific constraints in English and in SEA languages.

**SEA-IFEval**

Based on [IFEval](https://arxiv.org/abs/2311.07911), the linguists and native speakers in the team worked together to filter, localise and translate the datasets into the respective target languages to ensure that the examples remained reasonable, meaningful and natural.

SEA-IFEval evaluates a model's ability to adhere to constraints provided in the prompt, for example beginning a response with a specific word/phrase or answering with a certain number of sections. Additionally, accuracy is normalised by the proportion of responses in the correct language (if the model performs the task correctly but responds in the wrong language, it is judged to have failed the task).

For more details on Llama-SEA-LION-v3-70B base benchmark performance, please refer to the SEA-HELM leaderboard, <https://leaderboard.sea-lion.ai/>.

## Llama-SEA-LION-v3-70B-IT

### Fine-Tuning Methodology

Llama-SEA-LION-v3-70B-IT is a multilingual instruction-following model that has been tuned using a combination of a full parameter fine-tune, on-policy alignment, and model merges of the best performing checkpoints. The training process for fine-tuning was approximately 3200 GPU hours, on a single node of 8x H100-80GB GPUs.

### Fine-Tuning Data

Llama-SEA-LION-v3-70B-IT was trained on a wide range of synthetic instructions, alongside publicly available instructions hand-curated by the team with the assistance of native speakers. In addition, special care was taken to ensure that the datasets used had commercially permissive licenses through verification with the original data sources.

### Benchmark Performance

We evaluated Llama-SEA-LION-v3-70B-IT on both general language capabilities and instruction-following capabilities.

#### General Language Capabilities

For the evaluation of general language capabilities, we employed the [SEA-HELM](/benchmarks/sea-helm) (also known as BHASA) evaluation benchmark across a variety of tasks.\
These tasks include Question Answering (QA), Sentiment Analysis (Sentiment), Toxicity Detection (Toxicity), Translation in both directions (Eng>Lang & Lang>Eng), Abstractive Summarisation (Abssum), Causal Reasoning (Causal) and Natural Language Inference (NLI).

Note: SEA-HELM is implemented using prompts to elicit answers in a strict format. For all tasks, the model is expected to provide an answer tag from which the answer is automatically extracted. For tasks where options are provided, the answer should comprise one of the pre-defined options. The scores for each task is normalised to account for baseline performance due to random chance.

The evaluation was done **zero-shot** with native prompts on a sample of 100-1000 instances for each dataset.

#### Instruction-following Capabilities

Since Llama-SEA-LION-v3-70B-IT is an instruction-following model, we also evaluated it on instruction-following capabilities with two datasets, SEA-IFEval (based on [IFEval](https://arxiv.org/abs/2311.07911)) and SEA-MTBench (based on [MT-Bench](https://arxiv.org/abs/2306.05685)).

As these two datasets were originally in English, the linguists and native speakers in the team worked together to filter, localise and translate the datasets into the respective target languages to ensure that the examples remained reasonable, meaningful and natural.

**SEA-IFEval**

SEA-IFEval evaluates a model's ability to adhere to constraints provided in the prompt, for example beginning a response with a specific word/phrase or answering with a certain number of sections. Additionally, accuracy is normalised by the proportion of responses in the correct language (if the model performs the task correctly but responds in the wrong language, it is judged to have failed the task).

**SEA-MTBench**

SEA-MTBench evaluates a model's ability to engage in multi-turn (2 turns) conversations and respond in ways that align with human needs. We use `gpt-4-1106-preview` as the judge model and compare against `gpt-3.5-turbo-0125` as the baseline model. The metric used is the weighted win rate against the baseline model (i.e. average win rate across each category: Math, Reasoning, STEM, Humanities, Roleplay, Writing, Extraction). A tie is given a score of 0.5.

For more details on Llama-SEA-LION-v3-70B-IT benchmark performance, please refer to the SEA-HELM leaderboard, <https://leaderboard.sea-lion.ai/>.

## Llama-SEA-LION-v3-70B-IT-GGUF

The following quantized GGUF formats of our Llama-SEA-LION-v3-70B-IT model are available:

* Llama-SEA-LION-v3-70B-IT-F16
* Llama-SEA-LION-v3-70B-IT-Q2\_K
* Llama-SEA-LION-v3-70B-IT-Q3\_K\_M
* Llama-SEA-LION-v3-70B-IT-Q4\_0
* Llama-SEA-LION-v3-70B-IT-Q4\_K\_M
* Llama-SEA-LION-v3-70B-IT-Q5\_0
* Llama-SEA-LION-v3-70B-IT-Q5\_K\_M
* Llama-SEA-LION-v3-70B-IT-Q6\_K
* Llama-SEA-LION-v3-70B-IT-Q8\_0

Please refer to our [Download the Model(s)](#download-the-model-s) section for more details on how to access them.

<br>

## Download the Model(s)

Llama-SEA-LION-v3-70B models are available for download via the following channels:

[HuggingFace SEA-LION v3 Collection](https://huggingface.co/collections/aisingapore/sea-lionv3-672589a39cdadd6a5b199581)

| Model                         | Download                                                                                                                                           |
| ----------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
| Llama-SEA-LION-v3-70B         | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v3-70B)                                                                            |
| Llama-SEA-LION-v3-70B-IT      | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v3-70B-IT)                                                                         |
| Llama-SEA-LION-v3-70B-IT-GGUF | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v3-70B-IT-GGUF), [Ollama](https://ollama.com/aisingapore/Llama-SEA-LION-v3-70B-IT) |

<br>

## Usage

Llama-SEA-LION-v3-70B-IT can be run using the 🤗 Transformers library

```python
import transformers
import torch

model_id = "aisingapore/Llama-SEA-LION-v3-70B-IT"

pipeline = transformers.pipeline(
    "text-generation",
    model=model_id,
    model_kwargs={"torch_dtype": torch.bfloat16},
    device_map="auto",
)
messages = [
    {"role": "user", "content": "Apa sentimen dari kalimat berikut ini?\nKalimat: Buku ini sangat membosankan.\nJawaban: "},
]

outputs = pipeline(
    messages,
    max_new_tokens=256,
)
print(outputs[0]["generated_text"][-1])
```

## Disclaimer

It is important for users to be aware that our models exhibits certain limitations that warrant consideration:

1. The model can hallucinate and occasionally generates irrelevant content, introducing fictional elements that are not grounded in the provided context. Users should also exercise caution in interpreting and validating the model's responses due to the potential inconsistencies in its reasoning.
2. The model has not been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

<br>

## References

### Thai Pre-Training Data Reference

```bibtex
@misc{lowphansirikul2021wangchanberta,
    title={WangchanBERTa: Pretraining transformer-based Thai Language Models},
    author={Lalita Lowphansirikul and Charin Polpanumas and Nawat Jantrakulchai and Sarana Nutanong},
    year={2021},
    eprint={2101.09635},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}
```


# SEA-LION v2

## Introduction

SEA-LION version 2, released in July 2024, has been continued-pretrained on top of the [Llama 3 8B Instruct model](https://huggingface.co/meta-llama/Meta-Llama-3-8B-Instruct) that is 8 billion parameters in size, with **context length of 8192 tokens**.

Using continued-pretraining let us leverage the powerful capabilities of the Llama3 base model and build a stronger model with far fewer resources than pre-training from scratch. Compared to the 980B tokens used in for SEA-LION v1, approximately **48B** tokens across 5 SEA languages (English, Indonesia, Tamil, Thai and Vietnamese) was used for the continued pre-training of SEA-LION v2.

At a glance:

* **Model type:** Decoder
* **Tokenizer**: Default tokenizer used in Llama 3 8B Instruct
* **Training Data Size**: 48B tokens of SEA data
* **Context Length**: 8192
* **Available Formats**:
  * Base (Llama-SEA-LION-v2-8B)
  * Instruct (Llama-SEA-LION-v2-8B-IT)
  * GGUF (Llama-SEA-LION-v2-8B-IT-GGUF)
* **Supported Languages:**
  1. English
  2. Indonesian
  3. Thai
  4. Vietnamese
  5. Tamil
* **License:** [Llama3 Community License](https://huggingface.co/meta-llama/Meta-Llama-3-8B/blob/main/LICENSE)

## Llama-SEA-LION-v2-8B

### Training Infrastructure

Llama-SEA-LION-v2-8B was trained using [MosaicML Composer](https://github.com/mosaicml/composer) on the following hardware:

| Training Details     | Llama-SEA-LION-v2-8B |
| -------------------- | :------------------: |
| AWS EC2 p5d.24xlarge |      8 instances     |
| Nvidia H100 80GB GPU |          64          |
| Training Duration    |        2 days        |

**Configuration**

| HyperParameter    |  Llama-SEA-LION-v2-8B |
| ----------------- | :-------------------: |
| Precision         |        bfloat16       |
| Optimizer         |    decoupled\_adamw   |
| Scheduler         | weight\_stable\_decay |
| Learning Rate     |         1.0e-5        |
| Global Batch Size |          512          |
| Micro Batch Size  |           2           |

### Tokenizer

For tokenisation, the model employs the default tokenizer used in Llama 3 8B Instruct.

### Training Data

The Llama-SEA-LION-v2-8B base model was continued pre-trained on 48B tokens of the following data:

| Data Source                | Unique Tokens (B) | Multiplier | Total Tokens (B) | Percentage (%) |
| -------------------------- | :---------------: | :--------: | :--------------: | :------------: |
| Dolma RefinedWeb - English |       7.650       |      1     |       7.650      |      15.90     |
| Dolma C4 - English         |       1.160       |      1     |       1.16       |      9.21      |
| Dolma Reddit - English     |       1.339       |      1     |       1.339      |      2.42      |
| Dolma Semantic Scholar     |       0.959       |      1     |       0.959      |      2.79      |
| Dolma arXiv                |       0.469       |      1     |       0.469      |      1.99      |
| Dolma StarCoder            |       4.422       |      1     |       4.422      |      0.98      |
| SEA-LION Pile - Indonesian |        3.4        |      2     |        6.8       |      14.17     |
| Wiki\* - Indonesian        |        0.3        |      4     |        1.2       |      2.50      |
| SEA-LION Pile - Tamil      |        5.6        |      1     |        5.6       |      11.67     |
| Wiki\* + News - Tamil      |        0.6        |      4     |        2.4       |      5.00      |
| SEA-LION Pile - Thai       |        2.28       |      1     |       2.28       |      4.75      |
| WangChanBERTa - Thai       |         5         |      1     |         5        |      10.42     |
| Wiki\* - Thai              |        0.18       |      4     |       0.72       |      1.50      |
| SEA-LION Pile - Vietnamese |        6.76       |      1     |       6.76       |      14.08     |
| Wiki\* - Vietnamese        |        0.31       |      4     |       1.24       |      2.58      |

Note:

* All token counts are counted using Llama3 tokenizer
* wiki\* sources includes Wikipedia, Wiki Books, Wiki Source and Wiki Voyage
* Tamil news is sourced with permission from [Seithi](https://seithi.mediacorp.sg/)

### Benchmark Performance

We evaluated Llama-SEA-LION-v2-8B base model on general language capabilities.

#### General Language Capabilities

For the evaluation of general language capabilities in SEA languages, we employed the [BHASA evaluation benchmark](https://arxiv.org/abs/2309.06085v2) across a variety of tasks.\
These tasks include Question Answering (QA), Sentiment Analysis (Sentiment), Toxicity Detection (Toxicity), Translation in both directions (Eng>Lang & Lang>Eng), Abstractive Summarization (Summ), Causal Reasoning (Causal) and Natural Language Inference (NLI).

The evaluation was done **five-shot** with native prompts and only a sample of 100-1000 instances for each dataset was used as per the setting described in the paper.

For more details on Llama-SEA-LION-v2-8B benchmark performance, please refer to the SEA HELM leaderboard, <https://leaderboard.sea-lion.ai/>

<br>

## Llama-SEA-LION-v2-8B-IT

Llama-SEA-LION-v2-8B-IT is a multilingual instruction-following model which has been fine-tuned with around 100,000 English instruction-completion pairs alongside a smaller pool of around 50,000 instruction-completion pairs from other ASEAN languages, such as Indonesian, Thai and Vietnamese.

These instructions have been carefully curated and rewritten to ensure the model was trained on truly open, commercially permissive and high quality datasets.

### Fine-Tuning Methodology

The Llama-SEA-LION-v2-8B-IT model was fine-tuned using 8x A100-40GB using parameter efficient fine tuning in the form of LoRA.

### Fine-Tuning Data

Llama-SEA-LION-v2-8B-IT was trained on a wide range of instructions that were manually and stringently verified by our team. A large portion of the effort was dedicated to ensuring that each instruction-completion pair that the model sees is of high quality and any errors were corrected and rewritten by native speakers or else dropped from our mix.

In addition, special care was taken to ensure that the datasets used had commercially permissive licenses through verification with the original data source.

### Benchmark Performance

We evaluated Llama-SEA-LION-v2-8B-IT on both general language capabilities and instruction-following capabilities.

#### General Language Capabilities

For the evaluation of general language capabilities, we employed the [BHASA evaluation benchmark](https://arxiv.org/abs/2309.06085v2) across a variety of tasks. These tasks include Question Answering (QA), Sentiment Analysis (Sentiment), Toxicity Detection (Toxicity), Translation in both directions (Eng>Lang & Lang>Eng), Abstractive Summarization (Summ), Causal Reasoning (Causal) and Natural Language Inference (NLI).

Note: BHASA is implemented following a strict answer format, and only spaces and punctuations are cleaned. For tasks where options are provided, the answer should only include one of the pre-defined options, nothing else. If the model continues to generate more tokens (e.g. to explain its answer), it will be considered to be a wrong response. For the F1 score metric (as used in Sentiment Analysis and Toxicity Detection), all answers that do not fall under the pre-defined labels will be treated as a separate label (to mark it as a wrong answer) and included in the calculations so that the model is penalized for not generating one of the pre-defined labels.

The evaluation was done **zero-shot** with native prompts and only a sample of 100-1000 instances for each dataset was used as per the setting described in the paper.

#### Instruction-following Capabilities

Since Llama-SEA-LION-v2-8B-IT is an instruction-following model, we also evaluated it on instruction-following capabilities with two datasets, [IFEval](https://arxiv.org/abs/2311.07911) and [MT-Bench](https://arxiv.org/abs/2306.05685).

As these two datasets were originally in English, the linguists and native speakers in the team worked together to filter, localize and translate the datasets into the respective target languages to ensure that the examples remained reasonable, meaningful and natural.

**IFEval**

IFEval evaluates a model's ability to adhere to constraints provided in the prompt, for example beginning a response with a specific word/phrase or answering with a certain number of sections. The metric used is accuracy normalized by language (if the model performs the task correctly but responds in the wrong language, it is judged to have failed the task).

**MT-Bench**

MT-Bench evaluates a model's ability to engage in multi-turn (2 turns) conversations and respond in ways that align with human needs. We use `gpt-4-1106-preview` as the judge model and compare against `gpt-3.5-turbo-0125` as the baseline model. The metric used is the weighted win rate against the baseline model (i.e. average win rate across each category (Math, Reasoning, STEM, Humanities, Roleplay, Writing, Extraction)). A tie is given a score of 0.5.

<br>

## Llama-SEA-LION-v2-8B-IT-GGUF

The following quantized GGUF formats of our Llama-SEA-LION-v2-8B-IT model are available:

* Llama-SEA-LION-v2-8B-IT-Q2\_K
* Llama-SEA-LION-v2-8B-IT-Q3\_K\_M
* Llama-SEA-LION-v2-8B-IT-Q4\_0
* Llama-SEA-LION-v2-8B-IT-Q4\_K\_M
* Llama-SEA-LION-v2-8B-IT-Q5\_0
* Llama-SEA-LION-v2-8B-IT-Q5\_K\_M
* Llama-SEA-LION-v2-8B-IT-Q6\_K
* Llama-SEA-LION-v2-8B-IT-Q8\_0

Please refer to our [Download the Model(s)](#download-the-model-s) section for more details on how to access them.

<br>

## Download the Model(s)

SEA-LION v2 models are available for download via the following channels:

[HuggingFace SEA-LION v2 Collection](https://huggingface.co/collections/aisingapore/sea-lionv2-672589c4c7ea47e4174d3e7f)

| Model                        | Download                                                                                                                                         |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ |
| Llama-SEA-LION-v2-8B         | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v2-8B)                                                                           |
| Llama-SEA-LION-v2-8B-IT      | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v2-8B-IT)                                                                        |
| Llama-SEA-LION-v2-8B-IT-GGUF | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v2-8B-IT-GGUF), [Ollama](https://ollama.com/aisingapore/Llama-SEA-LION-v2-8B-IT) |

<br>

## Usage

Llama-SEA-LION-v2-8B-IT can be run using the 🤗 Transformers library

```python
# Please use transformers==4.43.2

import transformers
import torch

model_id = "aisingapore/Llama-SEA-LION-v2-8B-IT"

pipeline = transformers.pipeline(
    "text-generation",
    model=model_id,
    model_kwargs={"torch_dtype": torch.bfloat16},
    device_map="auto",
)
messages = [
    {"role": "user", "content": "Apa sentimen dari kalimat berikut ini?\nKalimat: Buku ini sangat membosankan.\nJawaban: "},
]

outputs = pipeline(
    messages,
    max_new_tokens=256,
)
print(outputs[0]["generated_text"][-1])
```

<br>

## Disclaimer

It is important for users to be aware that our models exhibits certain limitations that warrant consideration:

1. The model can hallucinate and occasionally generates irrelevant content, introducing fictional elements that are not grounded in the provided context. Users should also exercise caution in interpreting and validating the model's responses due to the potential inconsistencies in its reasoning.
2. The model has not been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

<br>

## References

### Thai Pre-Training Data Reference

```bibtex
@misc{lowphansirikul2021wangchanberta,
    title={WangchanBERTa: Pretraining transformer-based Thai Language Models},
    author={Lalita Lowphansirikul and Charin Polpanumas and Nawat Jantrakulchai and Sarana Nutanong},
    year={2021},
    eprint={2101.09635},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}
```


# SEA-LION v1

## Introduction

SEA-LION version 1, released in December 2023, was our first collection of Large Language Models (LLMs) that were specifically **pretrained** and **instruct-tuned** for the Southeast Asia (SEA) region, making a significant leap forward in the field of Natural Language Processing in understanding the SEA regional context.

SEA-LION v1 comes in two model sizes – one with 3 billion parameters (SEA-LION-v1-3B) and another with 7 billion parameters (SEA-LION-v1-7B). Both variants are built on the robust MPT architecture and utilise a vocabulary size of 256K, with **context length of 2048 tokens**.

Our SEA-LION-v1-7B model was then further instruct-tuned to produce SEA-LION-v1-7B-IT.

At a glance:

* **Model type:** Decoder
* **Tokenizer**: Custom SEABPETokenizer
* **Available Formats**:
  * 3B Base (SEA-LION-v1-3B)
  * 7B Base (SEA-LION-v1-7B)
  * 7B Instruct (SEA-LION-v1-7B-IT)
  * 7B GGUF (SEA-LION-v1-7B-IT-GGUF)
* **Languages:**
  1. English
  2. Chinese
  3. Indonesian
  4. Malay
  5. Thai
  6. Vietnamese
  7. Filipino
  8. Tamil
  9. Burmese
  10. Khmer
  11. Lao
* **License:** MIT

## SEA-LION-v1-3B / SEA-LION-v1-7B

### Model Architecture

SEA-LION-v1-3B and SEA-LION-v1-7B are both decoder models built on the robust MPT architecture:

| Parameter       | SEA-LION-v1-3B | SEA-LION-v1-7B |
| --------------- | :------------: | :------------: |
| Layers          |       32       |       32       |
| d\_model        |      2560      |      4096      |
| head\_dim       |       20       |       32       |
| Vocabulary      |     256000     |     256000     |
| Sequence Length |      2048      |      2048      |

### Training Infrastructure

SEA-LION v1 was trained using [MosaicML Composer](https://github.com/mosaicml/composer) on the following hardware:

| Training Details     | SEA-LION-v1-3B | SEA-LION-v1-7B |
| -------------------- | :------------: | :------------: |
| AWS EC2 p4d.24xlarge |  30 instances  |  32 instances  |
| Nvidia A100 40GB GPU |       240      |       256      |
| Training Duration    |     14 days    |     22 days    |

**Configuration:**

| HyperParameter    |    SEA-LION-v1-3B    |    SEA-LION-v1-7B    |
| ----------------- | :------------------: | :------------------: |
| Precision         |       bfloat16       |       bfloat16       |
| Optimizer         |   decoupled\_adamw   |   decoupled\_adamw   |
| Scheduler         | cosine\_with\_warmup | cosine\_with\_warmup |
| Learning Rate     |        1.6e-4        |        6.0e-5        |
| Global Batch Size |         1200         |         2048         |
| Micro Batch Size  |           5          |           4          |

For full details on our SEA-LION v1 training infrastructure and configuration, please refer to our [SEA-LION Pre-Training Setup Guide](https://github.com/aisingapore/sealion/blob/main/models/sea-lion-v1/pre-training/README-PRE-TRAINING.md)

### Tokenizer

For tokenization, both SEA-LION-v1-3B and SEA-LION-v1-7B employed our custom SEABPETokenizer, which is specially tailored for SEA languages, ensuring optimal model performance.

SEABPETokenizer was trained by sampling 20M lines from the model training data, using the SentencePiece framework. The tokenizer type is Byte-Pair Encoding (BPE).

### Training Data

SEA-LION-v1-3B and SEA-LION-v1-7B were trained on 980B tokens of text data from 11 languages spoken across SEA:

* English
* Chinese
* Indonesian
* Malay
* Thai
* Vietnamese
* Filipino
* Tamil
* Burmese
* Khmer
* Lao

These 980B tokens comprised of the following data mix:

| Data Source               | Unique Tokens | Multiplier | Total Tokens | Percentage |
| ------------------------- | ------------: | ---------: | -----------: | ---------: |
| RefinedWeb - English      |        571.3B |          1 |       571.3B |     58.20% |
| mC4 - Chinese             |         91.2B |          1 |        91.2B |      9.29% |
| mC4 - Indonesian          |         3.68B |          4 |        14.7B |      1.50% |
| mC4 - Malay               |         0.72B |          4 |         2.9B |      0.29% |
| mC4 - Filipino            |         1.32B |          4 |         5.3B |      0.54% |
| mC4 - Burmese             |          1.2B |          4 |         4.9B |      0.49% |
| mC4 - Vietnamese          |         63.4B |          1 |        63.4B |      6.46% |
| mC4 - Thai                |          5.8B |          2 |        11.6B |      1.18% |
| WangChanBERTa - Thai      |            5B |          2 |          10B |      1.02% |
| mC4 - Lao                 |         0.27B |          4 |         1.1B |      0.12% |
| mC4 - Khmer               |         0.97B |          4 |         3.9B |      0.40% |
| mC4 - Tamil               |         2.55B |          4 |        10.2B |      1.04% |
| the Stack - Python        |         20.9B |          2 |        41.8B |      4.26% |
| the Stack - Javascript    |         55.6B |          1 |        55.6B |      5.66% |
| the Stack - Shell         |         1.2B5 |          2 |         2.5B |      0.26% |
| the Stack - SQL           |          6.4B |          2 |        12.8B |      1.31% |
| the Stack - Markdown      |         26.6B |          1 |        26.6B |      2.71% |
| RedPajama - StackExchange |         21.2B |          1 |        21.2B |      2.16% |
| RedPajama - ArXiv         |         30.6B |          1 |        30.6B |      3.12% |

The dataset is available here: [SEA-LION-PILE](https://huggingface.co/datasets/aisingapore/sea-lion-pile).

### Performance

SEA-LION-v1-3B and SEA-LION-v1-7B base models had an average performance on general tasks in English (as measured by Hugging Face's LLM Leaderboard):

| Model          |  ARC  | HellaSwag |  MMLU | TruthfulQA | Average |
| -------------- | :---: | :-------: | :---: | :--------: | :-----: |
| SEA-LION-v1-3B | 36.26 |   64.59   | 24.07 |    36.46   |  40.35  |
| SEA-LION-v1-7B | 39.93 |   68.51   | 26.87 |    35.09   |  42.60  |

For up-to-date comparison of SEA-LION performance against other latest models, please refer to our [SEA-LION Leaderboard](https://leaderboard.sea-lion.ai)

## SEA-LION-v1-7B-IT

SEA-LION-v1-7B-IT is a multilingual model which has been fine-tuned with thousands of English and Indonesian instruction-completion pairs alongside a smaller pool of instruction-completion pairs from other ASEAN languages, using our pre-trained SEA-LION-v1-7B as base.

These instructions have been carefully curated and rewritten to ensure the model was trained on truly open, commercially permissive and high quality datasets.

### Fine-Tuning Methodology

The SEA-LION-v1-7B-IT was fine-tuned using 8x A100-40GB using parameter efficient fine tuning in the form of LoRA.

To perform similar fine-tuning on our SEA-LION-v1-7B base model using the HuggingFace TRL library, you can refer to sample configurations provided in our [SEA-LION QLoRA Fine-Tuning Guide.](https://github.com/aisingapore/sealion/blob/main/models/sea-lion-v1/fine-tuning/README.md)

### Fine-Tuning Data

SEA-LION-v1-7B-IT was trained on a wide range of instructions that were manually and stringently verified by our team. A large portion of the effort was dedicated to ensuring that each instruction-completion pair that the model sees is of a high quality and any errors were corrected and rewritten by native speakers or else dropped from our mix.

In addition, special care was taken to ensure that the datasets used had commercially permissive licenses through verification with the original data source.

### Benchmarks

We evaluated SEA-LION-v1-7B-IT on the BHASA benchmark ([arXiv](https://arxiv.org/abs/2309.06085v2) and [GitHub](https://github.com/aisingapore/bhasa)) across a variety of tasks.

BHASA stands out amongst other evaluations for SEA languages for its holistic approach to evaluation, including not just traditional Natural Language Processing (NLP) benchmarking tasks (such as sentiment analysis and question answering), but also linguistic and cultural diagnostic tests which are meticulously handcrafted.

The evaluation was done zero-shot with Indonesian prompts and only a sample of 100-1000 instances for each dataset was used as per the setting described in the BHASA paper.

* For Natural Language Understanding (NLU) tasks, we tested the model on Sentiment Analysis (Sentiment) using the NusaX dataset, Question Answering (QA) using the TyDiQA dataset, and Toxicity Detection (Toxicity) using the Indonesian Multi-Label Hate Speech Detection dataset. The metrics used are F1 scores for all three tasks.
* For Natural Language Generation (NLG) tasks, we tested the model on Machine Translation from English to Indonesian (Eng>Indo) and from Indonesian to English (Indo>Eng) using the FLORES-200 dataset, and Abstractive Summarization (Summary) using the XLSum dataset. The metrics used for Machine Translation and Abstractive Summarization are ChrF++ and ROUGE-L respectively.
* For Natural Language Reasoning (NLR) tasks, we tested the model on Natural Language Inference (NLI) using the IndoNLI lay dataset and on Causal Reasoning (Causal) using the XCOPA dataset. The metrics are based on accuracy for both tasks.

### Performance

SEA-LION v1 models achieved better or competitive performances on tasks in regional languages at the time of release:

| Model                         | QA (F1) | Sentiment (F1) | Toxicity (F1) | Eng>Indo (ChrF++) | Indo>Eng (ChrF++) | Summary (ROUGE-L) | NLI (Acc) | Causal (Acc) |
| ----------------------------- | ------- | -------------- | ------------- | ----------------- | ----------------- | ----------------- | --------- | ------------ |
| SEA-LION-7B-Instruct-Research | 24.86   | 76.13          | 24.45         | 52.50             | 46.82             | 15.44             | 33.20     | 23.80        |
| SEA-LION-7B-Instruct          | 68.41   | 91.45          | 17.98         | 57.48             | 58.04             | 17.54             | 53.10     | 60.80        |
| SeaLLM 7B v1                  | 30.96   | 56.29          | 22.60         | 62.23             | 41.55             | 14.03             | 26.50     | 56.60        |
| SeaLLM 7B v2                  | 44.40   | 80.13          | 55.24         | 64.01             | 63.28             | 17.31             | 43.60     | 82.00        |
| Sailor-7B                     | 65.43   | 59.48          | 20.48         | 64.27             | 60.68             | 8.69              | 15.10     | 38.40        |
| Llama 2 7B Chat               | 11.12   | 52.32          | 0.00          | 44.09             | 57.58             | 9.24              | 0.00      | 0.00         |
| Mistral 7B Instruct v0.1      | 38.85   | 74.38          | 20.83         | 30.60             | 51.43             | 15.63             | 28.60     | 50.80        |
| GPT-4                         | 73.60   | 74.14          | 63.96         | 69.38             | 67.53             | 18.71             | 83.20     | 96.00        |

For up-to-date comparison of SEA-LION performance against other latest models, please refer to our [SEA-LION Leaderboard](https://leaderboard.sea-lion.ai)

## SEA-LION-7B-Instruct-GGUF

Support for SEA-LION-v1-IT-GGUF was merged into `llama.cpp` as of 4th Apr 2024.

SEA-LION can be run using the `llama.cpp` library from commit id [bb43cf7](https://github.com/ggerganov/llama.cpp/commit/bb43cf7e9d86d69ffd9c7f008f75db890a35b45a) or later.

**Prompt Template:**

```
### USER:
{{prompt}}

### RESPONSE:

```

**Recommended `llama.cpp` command:**

```
./main -m sea-lion-7b-instruct-Q4_0.gguf --temp 0 --repeat-penalty 1.2 -e -ngl 32 -p "### USER:\nwhat is a sea lion?\n\n### RESPONSE:\n"
```

**To convert & quantize your own SEA-LION model:**

```
python convert-hf-to-gguf.py {{model path}}

./quantize ggml-model-f16.gguf {{Quant Type}}
```

For other parameters and how to use them, please refer to [llama.cpp documentation.](https://github.com/ggerganov/llama.cpp/blob/master/examples/main/README.md)

The following quantized GGUF formats of our SEA-LION-v1-7B-IT model are available:

* sea-lion-7b-instruct-Q2\_K
* sea-lion-7b-instruct-Q3\_K\_M
* sea-lion-7b-instruct-Q4\_0
* sea-lion-7b-instruct-Q4\_K\_M
* sea-lion-7b-instruct-Q5\_0
* sea-lion-7b-instruct-Q5\_K\_M
* sea-lion-7b-instruct-Q6\_K
* sea-lion-7b-instruct-Q8\_0

Please refer to our [Download the Model(s)](#download-the-model-s) section for more details on how to access them.

## Download the Model(s)

SEA-LION v1 models are available for download via the following channels:

[HuggingFace SEA-LION v1 Collection](https://huggingface.co/collections/aisingapore/sea-lionv1-672589cd29a1781afa6be35e)

| Model                  | Download                                                                 |
| ---------------------- | ------------------------------------------------------------------------ |
| SEA-LION-v1-3B         | [HuggingFace](https://huggingface.co/aisingapore/SEA-LION-v1-3B)         |
| SEA-LION-v1-7B         | [HuggingFace](https://huggingface.co/aisingapore/SEA-LION-v1-7B)         |
| SEA-LION-v1-7B-IT      | [HuggingFace](https://huggingface.co/aisingapore/SEA-LION-v1-7B-IT)      |
| SEA-LION-v1-7B-IT-GGUF | [HuggingFace](https://huggingface.co/aisingapore/SEA-LION-v1-7B-IT-GGUF) |

## Usage

SEA-LION-v1-7B-IT can be run using the 🤗 Transformers library

```python
# Please use transformers==4.37.2
from transformers import AutoModelForCausalLM, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("aisingapore/SEA-LION-v1-7B-IT", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("aisingapore/SEA-LION-v1-7B-IT", trust_remote_code=True)

prompt_template = "### USER:\n{human_prompt}\n\n### RESPONSE:\n"
prompt = """Apa sentimen dari kalimat berikut ini?
Kalimat: Buku ini sangat membosankan.
Jawaban: """
full_prompt = prompt_template.format(human_prompt=prompt)

tokens = tokenizer(full_prompt, return_tensors="pt")
output = model.generate(tokens["input_ids"], max_new_tokens=20, eos_token_id=tokenizer.eos_token_id)
print(tokenizer.decode(output[0], skip_special_tokens=True))
```

## Prompting Guide

A basic prompting guide for the SEALION v1 models is provided [here](https://github.com/aisingapore/sealion/blob/main/models/sea-lion-v1/sea-lion-v1_promptguide.md)

## Disclaimer

It is important for users to be aware that our models exhibits certain limitations that warrant consideration:

1. The model can hallucinate and occasionally generates irrelevant content, introducing fictional elements that are not grounded in the provided context. Users should also exercise caution in interpreting and validating the model's responses due to the potential inconsistencies in its reasoning.
2. The model has not been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.
3. It should be noted that the model has not been optimized for multi-turn dialogue interactions, which may result in reduced effectiveness in extended conversations.

## References

Thai Pre-Training Data Reference

```
@misc{lowphansirikul2021wangchanberta,
    title={WangchanBERTa: Pretraining transformer-based Thai Language Models},
    author={Lalita Lowphansirikul and Charin Polpanumas and Nawat Jantrakulchai and Sarana Nutanong},
    year={2021},
    eprint={2101.09635},
    archivePrefix={arXiv},
    primaryClass={cs.CL}
}
```


# SEA-LION Foundation Family

The SEA-LION models have served as a foundation for developing localized AI solutions tailored to specific linguistic and cultural needs. Over multiple iterations, other regions have built upon SEA-LION’s architecture to create specialized versions that enhance language understanding for their respective regions. These models leverage SEA-LION’s robust multilingual capabilities while being fine-tuned with localized datasets, ensuring better performance in regional contexts.

## SEA-LION Model Tree

<table><thead><tr><th width="144" valign="top">Base</th><th>Localized Model</th></tr></thead><tbody><tr><td valign="top">SEA-LION v1</td><td>→ <a href="https://huggingface.co/airesearch/WangchanLion7B">WangchanLion 7B (Thai)</a>: WangchanLion 7B is a multilingual instruction-following model developed by PyThaiNLP and the VISTEC-depa AI Research Institute of Thailand. Fine-tuned on SEA-LION-v1-7B, it incorporates approximately 500,000 samples from open-source, commercially permissible datasets, with a focus on Thai and English languages.</td></tr><tr><td valign="top">SEA-LION v2</td><td>→ <a href="https://huggingface.co/GoToCompany/llama3-8b-cpt-sahabatai-v1-instruct">Llama3 8B CPT Sahabat-AI v1 Instruct (Indonesian)</a>: Llama3 8B CPT Sahabat-AI v1 Instruct model, co-developed by GoTo Group and AI Singapore, is an Indonesian-focused adaptation fine-tuned with 448,000 Indonesian instruction-completion pairs, along with 96,000 Javanese and 98,000 Sundanese pairs. It supports Indonesian, Javanese, Sundanese, and English, making it a significant advancement in AI for the Indonesian linguistic landscape.</td></tr><tr><td valign="top">SEA-LION v3</td><td><p>→ <a href="https://huggingface.co/aisingapore/Gemma2-9b-WangchanLIONv2-instruct">Gemma2 9B WangchanLIONv2 (Thai)</a>: The Gemma2 9B WangchanLIONv2 Instruct model is a collaborative effort between VISTEC and AI Singapore. It has been fine-tuned with approximately 3,760,000 Thai instruction-completion pairs derived from human-annotated instructions, FLAN-style automatic data construction, and synthetic samples. This multilingual model supports both Thai and English languages.</p><p>→ <a href="https://huggingface.co/GoToCompany/gemma2-9b-cpt-sahabatai-v1-instruct">Gemma2 9B CPT Sahabat-AI (Indonesian)</a>: The Gemma2 9B CPT Sahabat-AI v1 Instruct model, co-developed by GoTo Group and AI Singapore, has been fine-tuned with approximately 448,000 Indonesian instruction-completion pairs, along with 96,000 in Javanese, 98,000 in Sundanese, and an additional 129,000 in English. This multilingual model supports Indonesian, Javanese, Sundanese, and English.</p></td></tr><tr><td valign="top">SEA-LION v4</td><td><p>→ <a href="https://huggingface.co/aisingapore/Gemma-SEA-LION-v4-27B-VL">Gemma 3 27B/4B SEA-LION v4 (Multimodal)</a>: The first multimodal release in the SEA-LION family, based on the Gemma 3 architecture. Co-developed with Google, these models support a 128K context window. It underwent post-training on ~10M samples across 11 SEA languages and features vision-text capabilities for document understanding and visual Q&#x26;A. This model supports text and image understanding with a commercially permissive license. It is designed to handle Southeast Asian cultural nuances and visual contexts.</p><p>→ <a href="https://huggingface.co/aisingapore/Qwen-SEA-LION-v4-4B-VL">Qwen-SEA-LION-v4-4B/8B-IT</a>: The lightweight flagship multimodal model built on the Qwen3-VL framework, optimized for regional linguistic and cultural nuances.</p><p>→ <a href="/pages/AOV90iS5faoHqt3NN1S3">Apertus 8B SEA-LION v4</a>: A fully open regional adaptation based on the Swiss AI "Apertus" architecture, focusing on highly efficient, transparent, and community-driven multilingual support. Features 15T token pre-training across 1,000+ languages with a focus on deep multilingual depth and transparency.</p></td></tr><tr><td valign="top">SEA-Guard</td><td><p>A specialized collection of safety-focused LLMs built on the SEA-LION family released on 4 Feb 2026.</p><p>→ <a href="https://huggingface.co/aisingapore/Qwen-SEA-Guard-4B-040226">Qwen-SEA-Guard-4B/8B (Image-to-Text)</a>: The 4B variant is a lightweight visual guardrail optimized for edge applications, while the 8B offers balanced visual moderation with stronger reasoning capabilities.</p><p>→ <a href="https://huggingface.co/aisingapore/Llama-SEA-Guard-8B-040226">Llama-SEA-Guard-8B (Text Generation)</a>: A text-only safety specialist optimized for chat moderation and policy enforcement, ensuring safe interactions in dialogue systems.</p><p>→ <a href="https://huggingface.co/aisingapore/Gemma-SEA-Guard-12B-040226">Gemma-SEA-Guard-12B (Multimodal)</a>: The high-capacity safety flagship capable of interpreting complex relationships between visual and textual data for deep content analysis.</p></td></tr><tr><td valign="top">SEA-LION Embedding</td><td><p>A suite of high-performance encoder-only models specifically architected for Southeast Asian languages, providing state-of-the-art vector representations for RAG and semantic search.</p><p>→ <a href="https://huggingface.co/collections/aisingapore/sea-lion-modernbert-and-embedding">SEA-LION-ModernBERT-Embedding (Encoder)</a>: Our flagship encoder line utilizing the ModernBERT architecture. It offers superior efficiency and long-context handling, specifically tuned for the unique scripts and linguistic structures of the SEA region.</p><p>→ <a href="hhttps://huggingface.co/aisingapore/SEA-LION-E5-Embedding-600M">SEA-LION-E5-Embedding (Encoder)</a>: A high-precision semantic encoder line fine-tuned from the E5-large foundation, optimized for maximum accuracy in cross-lingual retrieval and semantic matching across regional and global languages.</p></td></tr></tbody></table>

## Impact and Future Directions

By leveraging SEA-LION’s architecture, these localized models provide AI solutions that align more closely with native language requirements. As SEA-LION continues to evolve, more localized versions are expected to emerge, further expanding the reach and effectiveness of AI in Southeast Asian languages.


# SEA-LION Embedding

[**SEA-LION Embedding**](https://huggingface.co/collections/aisingapore/sea-lion-modernbert-and-embedding), released in March 2026, is a suite of high-performance encoder models specifically architected for Southeast Asian languages. The collection features two primary product lines:

* [SEA-LION-ModernBERT (Efficiency & Context)](/models/sea-embedding/sea-modernbert): Built on the ModernBERT architecture, these models feature a native 8,192 token context window and are optimized for high-throughput production environments and long-document RAG.
* [SEA-LION-Embedding-E5 (Semantic Precision))](/models/sea-embedding/sea-e5): Fine-tuned from the E5-large foundation, these models are designed for maximum semantic accuracy in retrieval and similarity tasks across the region's diverse linguistic landscape.

For detailed information on specific models and checkpoints, please refer to their individual documentation pages.


# SEA-LION-ModernBERT and Embedding

Last update: 2026-03-16

**SEA-LION** is a collection of Large Language Models (LLMs) and encoders which have been pretrained and fine-tuned for the Southeast Asia (SEA) region.

## Introduction

SEA-LION stands for *Southeast Asian Languages In One Network*.

This encoder-only model leverages the advanced **ModernBERT** architecture combined with the Gemma 3 SentencePiece tokenizer. The adoption of the **Gemma 3 tokenizer** with ModernBERT allows the model to achieve highly efficient and culturally nuanced text processing. This combination significantly improves the tokenization fertility and compression rates for complex regional scripts and diverse Southeast Asian languages, enabling the model to handle longer context windows and cross-lingual tasks with greater computational efficiency.

To achieve this level of performance, the model was developed through a rigorous, multi-stage training pipeline. The foundation was established through extensive **pre-training on 2 Trillion (2T) tokens**, followed by a mid-training phase on an additional 1 Trillion (1T) tokens. Both of these massive training phases comprehensively covered code alongside 13 specific languages: Burmese, Chinese, English, Filipino, Indonesian, Javanese, Khmer, Lao, Malay, Sundanese, Tamil, Thai, and Vietnamese.

## Model Details

### Model Description

The SEA-LION-ModernBERT-based models are built on the ModernBERT architecture and has a vocabulary size of 262K.

For tokenization, the model employs our custom [Gemma3](https://storage.googleapis.com/deepmind-media/gemma/Gemma3Report.pdf) tokenizer, which has excellent performance for SEA languages, ensuring optimal model performance.

* **Developed by:** AI Products Pillar, AI Singapore
* **Funded by:** Singapore NRF
* **Shared by:** AI Products Pillar, AI Singapore
* **Model type:** Encoder
* **Context length:** 8k
* **Languages:** Burmese, Chinese, English, Filipino, Indonesian, Javanese, Khmer, Lao, Malay, Sundanese, Tamil, Thai, and Vietnamese
* **License:** [MIT](https://tlo.mit.edu/understand-ip/exploring-mit-open-source-license-comprehensive-guide)

### Model Sources

* **Repository:** The weights for this model and its various training stages are being released to support transparency, research, and diverse downstream applications. [**Link to HF Repo**](https://huggingface.co/collections/aisingapore/sea-lion-modernbert-and-embedding)

## Uses

This model card details one of the variants available within this ModernBERT-based collection.

| Model Variant                     | Model Repository                                                                                                                                                                                                                                                                                                                                                                                            | Suggesting Applications & Use Cases                                                                                           |
| --------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------- |
| **Fine-tuned Embedding Models**   | <p>- <a href="https://huggingface.co/aisingapore/SEA-LION-E5-Embedding-600M">aisingapore/SEA-LION-E5-Embedding-600M</a><br>- <a href="https://huggingface.co/aisingapore/SEA-LION-ModernBERT-Embedding-300M">aisingapore/SEA-LION-ModernBERT-Embedding-300M</a><br>- <a href="https://huggingface.co/aisingapore/SEA-LION-ModernBERT-Embedding-600M">aisingapore/SEA-LION-ModernBERT-Embedding-600M</a></p> | <p>- Retrieval-Augmented Generation (RAG)<br>- Information retrieval, and search<br>- Similarity comparisons</p>              |
| **Pre-trained Encoder Models**    | <p>- <a href="https://huggingface.co/aisingapore/SEA-LION-ModernBERT-300M">aisingapore/SEA-LION-ModernBERT-300M</a><br>- <a href="https://huggingface.co/aisingapore/SEA-LION-ModernBERT-600M">aisingapore/SEA-LION-ModernBERT-600M</a><br></p>                                                                                                                                                             | <p>- Fill mask<br>- Text classification<br>- Fine-tuning for downstream tasks (e.g., sentiment analysis, classification).</p> |
| **Pre-trained Model Checkpoints** | <p>- <a href="https://huggingface.co/aisingapore/SEA-LION-ModernBERT-300M-checkpoints">aisingapore/SEA-LION-ModernBERT-300M-checkpoints</a><br>- <a href="https://huggingface.co/aisingapore/SEA-LION-ModernBERT-600M-checkpoints">aisingapore/SEA-LION-ModernBERT-600M-checkpoints</a><br></p>                                                                                                             | <p>- Continued Pre-Training (CPT)<br>- Fine-tuning for downstream tasks (e.g., sentiment analysis, classification).</p>       |

The checkpoints repository contains available of model variants.

| Model Variant                                    | Suggesting Applications & Use Cases                                                                                                                                                             |
| ------------------------------------------------ | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| stage1-pre-training/SEA-LION-PT-300M.pt          | Composer checkpoint from the **Pre-Training Stage** suitable for continued pre-training (CPT).                                                                                                  |
| stage1-pre-training/SEA-LION-PT-300M             | Folder for the HuggingFace checkpoints from the **Pre-Training Stage**, suitable for continued pre-training or fine tuning.                                                                     |
| stage2-mid-training/SEA-LION-MT-300M-w-decay.pt  | Composer checkpoint from the **Mid-Training stage** with learning rate annealing suitable for fine tuning with learning rate warmup.                                                            |
| stage2-mid-training/SEA-LION-MT-300M-wo-decay.pt | Composer checkpoint from the **Mid-Training stage** without learning rate annealing suitable for continued pre-training (CPT) and fine tuning without learning rate warmup.                     |
| stage2-mid-training/SEA-LION-MT-300M-wo-decay    | Folder for the HuggingFace checkpoints from the **Mid-Training stage** without learning rate annealing. suitable for continued pre-training (CPT) and fine tuning without learning rate warmup. |

Note: For stage2-mid-train*ing checkpoints with learning rate annealing, please refer to* [*aisingapore/SEA-LION-ModernBERT-300M*](https://huggingface.co/aisingapore/SEA-LION-ModernBERT-300M) and [*aisingapore/SEA-LION-ModernBERT-600M*](https://huggingface.co/aisingapore/SEA-LION-ModernBERT-600M)

*Note: If you are deploying our models for your specific use case, we would love to hear from you! Please feel free to* [*contact us*](mailto:sealion@aisingapore.org) *to share your experience or explore potential collaborations.*

### Bias, Risks, and Limitations

The model was not tested for robustness against adversarial usage. It is important for users to be aware that our model exhibits certain limitations that warrant consideration. Users should also exercise caution in continue-implementing and validating the model's responses due to the potential inconsistencies.

### Recommendations

Users (both direct and downstream) should be made aware of the risks, biases and limitations of the model.

## How to Get Started with the Model

Use this code snippet below to run inference with the model via the API.

```
from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.sea-lion.ai/v1/embeddings")

result = client.embeddings.create(
    model="aisingapore/SEA-LION-ModernBERT-Embedding-600M",
    input=[
        "Singapore is a tropical island city-state.",
        "The Lion City sits at the tip of the Malay Peninsula.",
    ],
)

for item in result.data:
    print(f"index {item.index}: dim={len(item.embedding)}")
```

Use the code below to download the model locally.

```
pip install -U transformers>=4.48.0
```

```
#########################
## Download checkpoints locally for continued pre-training or fine tuning
#########################
from huggingface_hub import snapshot_download

# Download the stage-1-pre-training Huggingface checkpoint
snapshot_download(
  "aisingapore/SEA-LION-ModernBERT-300M-checkpoints",
  repo_type="model",
  allow_patterns=["stage1-pre-training/SEA-LION-PT-300M/*"],
  local_dir="checkpoints"
)

# Download the stage-1-pre-training Composer checkpoint
snapshot_download(
  "aisingapore/SEA-LION-ModernBERT-300M-checkpoints",
  repo_type="model",
  allow_patterns=["stage1-pre-training/SEA-LION-PT-300M.pt"],
  local_dir="checkpoints"
)
```

```
import torch
from transformers import pipeline

pipeline = pipeline(
    task="fill-mask",
    model="checkpoints/stage1-pre-training/SEA-LION-PT-300M", # loading from local folder
    dtype=torch.float16,
    device=0
)
pipeline("Plants create  through a process known as photosynthesis.")
```

*Note: To get started with Continued Pre-Training of the Composer checkpoints, we recommend refering to this* [*guide*](https://huggingface.co/blog/thomas-sounack/bioclinical-modernbert-tutorial)*.*

***

## Training Details

The models are pre-trained from scratch through a two-phase pipeline, beginning with an extensive initial stage on 2 trillion tokens, followed by a mid-training phase on an additional 1 trillion tokens. Both phases incorporated a diverse dataset covering programming code and 13 languages: Burmese, Chinese, English, Filipino, Indonesian, Javanese, Khmer, Lao, Malay, Sundanese, Tamil, Thai, and Vietnamese.

### Training Data

The pre-trained checkpoints were pre-trained from scratch on a number of trillion tokens corpus with the following linguistic and thematic distribution:

| Data Source     | Percentage |
| --------------- | ---------- |
| code            | 10%        |
| EN - English    | 35%        |
| ID - Indonesian | 8%         |
| JV - Javanese   | 0.5%       |
| KM - Khmer      | 1.5%       |
| LO - Lao        | 0.5%       |
| MS - Malay      | 4.75%      |
| MY - Burmese    | 1.75%      |
| SU - Sundanese  | 0.5%       |
| TA - Tamil      | 4.5%       |
| TH - Thai       | 8%         |
| TL - Filipino   | 2.5%       |
| VI - Vietnamese | 8.5%       |
| ZH - Chinese    | 14%        |

## Evaluation

### Testing Data, Factors & Metrics

#### Testing Data

The model is evaluated across three primary benchmark suites to provide a comprehensive assessment of embedding quality across Southeast Asian, Chinese, and English contexts:

* [**SEA-BED (Southeast Asia Embedding Benchmark)**](https://arxiv.org/pdf/2508.12243): The primary testing suite, consisting of 169 datasets across 10 Southeast Asian languages (Burmese, Filipino, Indonesian, Khmer, Malay, Lao, Tamil, Tetum, Thai, and Vietnamese). Notably, 71% of these datasets are native-authored or human-curated to preserve regional linguistic properties.
* **CMTEB (Chinese Massive Text Embedding Benchmark)**: A specialised subset of MTEB focused on Chinese language tasks, used to evaluate performance in one of the region's most prominent scripts.
* **MTEB (Massive Text Embedding Benchmark)**: The industry-standard global benchmark used to gauge general-purpose English embedding performance across a wide array of tasks.

## Results

For details on Performance comparison of embedding models on SEA-BED, please refer to the [SEA-HELM](https://leaderboard.sea-lion.ai/embedding/SEA).

## Environmental Impact

Carbon emission was estimated using the fact sheet from TRG [Datacenters](https://www.trgdatacenters.com/resource/h200-power-consumption/).

* **Hardware Type:** Nvidia H200 140GB GPUs
* **Hours used:** 1,825 GPU hours
* **Cloud Provider:** SMC H200
* **Compute Region:** Singapore
* **Carbon Emitted:** appx. 513.27 kg CO2 e

## Technical Specifications

### Model Architecture and Objective

SEA-LION-ModernBERT-300M is an encoder model using the ModernBERT architecture.

| Parameter       | SEA-LION-ModernBERT |
| --------------- | ------------------- |
| Layers          | 22                  |
| d\_model        | 768                 |
| head\_dim       | 12                  |
| Vocabulary      | 262144              |
| Sequence Length | 8k                  |

## More Information

This is the repository for the commercial fine-tuned model. The model has *not* been aligned for safety. Developers and users should perform their own safety fine-tuning and related security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

For more info, please contact us at <sealion@aisingapore.org>


# SEA-LION-Embedding-E5

Last update: 2026-03-16

**SEA-LION** is a collection of Large Language Models (LLMs) and encoders which have been pretrained and fine-tuned for the Southeast Asia (SEA) region.

## Introduction

SEA-LION stands for *Southeast Asian Languages In One Network*.

The **SEA-LION-Embedding-E5-600M** model is a Sentence Transformer optimised for 11 Southeast Asian languages. It has been fine-tuned from the [multilingual-e5-large](https://huggingface.co/intfloat/multilingual-e5-large) base, mapping sentences and paragraphs to a 1024-dimensional dense vector space. This model is designed for high-accuracy semantic textual similarity, semantic search, paraphrase mining, text classification, clustering, and RAG (Retrieval-Augmented Generation) workflows. It leverages the robust XLM-RoBERTa architecture pretrained on 100 languages, optimised here for 11 Southeast Asian languages: Burmese, Chinese, English, Filipino, Indonesian, Khmer, Lao, Malay, Tamil, Thai, and Vietnamese.

## Model Details

### Model Description

The SEA-LION-E5-Embedding-600M model is a Sentence Transformer built on the [Multilingual E5 Text Embeddings](https://arxiv.org/pdf/2402.05672) which was initialised from [xlm-roberta-large](https://huggingface.co/xlm-roberta-large) architecture.

* **Model Type:** Sentence Transformer
* **Base Architecture:** E5 (Transformer Encoder)
* **Developed by:** AI Products Pillar, AI Singapore
* **Funded by:** Singapore NRF
* **Shared by:** AI Products Pillar, AI Singapore
* **Context length:** 512
* **Languages:** Burmese, Chinese, English, Filipino, Indonesian, Khmer, Lao, Malay, Tamil, Thai, and Vietnamese
* **License:** [MIT](https://tlo.mit.edu/understand-ip/exploring-mit-open-source-license-comprehensive-guide)
* **Finetuned from model:** [multilingual-e5-large](https://huggingface.co/intfloat/multilingual-e5-large)

### Model Sources

* **Documentation:** [Sentence Transformers Documentation](https://www.sbert.net/)
* **Repository:** [aisingapore/SEA-LION-E5-Embedding-600M](https://huggingface.co/aisingapore/SEA-LION-E5-Embedding-600M)

## Uses

SEA-LION-E5-Embedding-600M details one of the variants available within this collection. If you are deploying our models for your specific use case, we would love to hear from you! Please feel free to [contact us](mailto:sealion@aisingapore.org) to share your experience or explore potential collaborations.

| Model Variant                     | Model Repository                                                                                                                                                                                                                                                                                                                                                                                            | Suggesting Applications & Use Cases                                                                                           |
| --------------------------------- | ----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | ----------------------------------------------------------------------------------------------------------------------------- |
| **Fine-tuned Embedding Models**   | <p>- <a href="https://huggingface.co/aisingapore/SEA-LION-E5-Embedding-600M">aisingapore/SEA-LION-E5-Embedding-600M</a><br>- <a href="https://huggingface.co/aisingapore/SEA-LION-ModernBERT-Embedding-300M">aisingapore/SEA-LION-ModernBERT-Embedding-300M</a><br>- <a href="https://huggingface.co/aisingapore/SEA-LION-ModernBERT-Embedding-600M">aisingapore/SEA-LION-ModernBERT-Embedding-600M</a></p> | <p>- Retrieval-Augmented Generation (RAG)<br>- Information retrieval, and search<br>- Similarity comparisons</p>              |
| **Pre-trained Encoder Models**    | <p>- <a href="https://huggingface.co/aisingapore/SEA-LION-ModernBERT-300M">aisingapore/SEA-LION-ModernBERT-300M</a><br>- <a href="https://huggingface.co/aisingapore/SEA-LION-ModernBERT-600M">aisingapore/SEA-LION-ModernBERT-600M</a><br></p>                                                                                                                                                             | <p>- Fill mask<br>- Text classification<br>- Fine-tuning for downstream tasks (e.g., sentiment analysis, classification).</p> |
| **Pre-trained Model Checkpoints** | <p>- <a href="https://huggingface.co/aisingapore/SEA-LION-ModernBERT-300M-checkpoints">aisingapore/SEA-LION-ModernBERT-300M-checkpoints</a><br>- <a href="https://huggingface.co/aisingapore/SEA-LION-ModernBERT-600M-checkpoints">aisingapore/SEA-LION-ModernBERT-600M-checkpoints</a><br></p>                                                                                                             | <p>- Continued Pre-Training (CPT)<br>- Fine-tuning for downstream tasks (e.g., sentiment analysis, classification).</p>       |

## Direct Usage (Sentence Transformers)

First install the Sentence Transformers library:

```
pip install -U sentence_transformers>=2.2.2
```

Then you can load this model and run inference.

```
from sentence_transformers import SentenceTransformer

model = SentenceTransformer(
    "aisingapore/SEA-LION-E5-Embedding-600M",
    prompts={
        "STS": "Instruct: Retrieve semantically similar text.\nQuery: ",
        "Clustering": "Instruct: Classify text into its appropriate category\nQuery: ",
        "Classification": "Instruct: Classify text into its appropriate category\nQuery: ",
        "Retrieval": "Instruct: Given a passage that is guaranteed to contain the answer, retrieve relevant passages that answer the query.\nQuery: ",
        "BitextMining": "Instruct: Retrieve parallel sentences.\nQuery: ",
        "PairClassification": "Instruct: Retrieve semantically similar text.\nQuery: ",
        "Reranking": "Instruct: Retrieve semantically similar text.\nQuery: ",
        "InstructionRetrieval": "Instruct: Given a instruction and a output, retrieve the most relevant output that answer the instruction.\nQuery: ",
        "MultiLabelTextClassification": "Instruct: Classify the given text into its appropriate classes\nQuery: ",
        "QARetrieval": "Instruct: Given a passage that is guaranteed to contain the answer, retrieve relevant passages that answer the query.\nQuery: ",
        "Summarization": "Instruct: Summarize the given text.\nQuery: "
    },
)

sentences = [
    "The weather is lovely today.",
    "อากาศวันนี้ดีมาก",
    "Dia berkendara ke stadion.",
]
embeddings = model.encode(sentences, prompt_name="STS")
print(embeddings.shape)
# [3, 1024]

similarities = model.similarity(embeddings, embeddings)
print(similarities)
#tensor([[1.0000, 0.8594, 0.5534],
#        [0.8594, 1.0000, 0.6343],
#        [0.5534, 0.6343, 1.0000]])
```

## Training Details

### Training Data

This model was tuned using a multi-stage training pipeline with the following datasets:

* **Contrastive Pre-training:** 245 million text pairs (EN-EN and EN-SEA) to enhance cross-lingual alignment.
* **Fine-tuning:** 13 million diverse text pairs (spanning EN-EN, CN-CN, EN-SEA, and SEA-SEA) to create the final fine-tuned model.

| **Language** | Percentage |
| ------------ | ---------- |
| EN-EN        | 20%        |
| CN-CN        | 20%        |
| EN-SEA       | 10%        |
| SEA-SEA      | 50%        |

### Training Procedure

#### Preprocessing

Following the foundational training, the model's cross-lingual alignment was substantially enhanced by undergoing contrastive pre-training utilising 245 million text pairs, specifically focusing on English-to-English and English-to-Southeast Asian language mappings (EN-EN and EN-SEA). Finally, to ensure the model could effectively follow user instructions and handle complex interactions, it was instruction-tuned using a diverse dataset of 13 million text pairs spanning EN-EN, CN-CN, EN-SEA, and SEA-SEA, culminating in the final highly capable instruction-tuned model.

## Evaluation

### Testing Data, Factors & Metrics

#### Testing Data

The model is evaluated across three primary benchmark suites to provide a comprehensive assessment of embedding quality across Southeast Asian, Chinese, and English contexts:

* **SEA-BED (Southeast Asia Embedding Benchmark)** (<https://arxiv.org/pdf/2508.12243>): The primary testing suite, consisting of 169 datasets across 10 Southeast Asian languages (Burmese, Filipino, Indonesian, Khmer, Malay, Lao, Tamil, Tetum, Thai, and Vietnamese). Notably, 71% of these datasets are native-authored or human-curated to preserve regional linguistic properties.
* **CMTEB (Chinese Massive Text Embedding Benchmark)**: A specialised subset of MTEB focused on Chinese language tasks, used to evaluate performance in one of the region's most prominent scripts.
* **MTEB (Massive Text Embedding Benchmark)**: The industry-standard global benchmark used to gauge general-purpose English embedding performance across a wide array of tasks.

### Results

For details on Performance comparison of embedding models on SEA-BED, please refer to the [SEA-HELM](https://leaderboard.sea-lion.ai/embedding/SEA).

## Environmental Impact

Carbon emission was estimated using the fact sheet from TRG [Datacenters](https://www.trgdatacenters.com/resource/h200-power-consumption/).

* **Hardware Type:** Nvidia H200 140GB GPUs
* **Hours used:** 896 hrs
* **Cloud Provider:** SMC H200
* **Compute Region:** Singapore
* **Carbon Emitted:** appx. 252.13 kg CO2 e

## Technical Specifications

### Model Architecture and Objective

SEA-LION-E5-Embedding-600M is an encoder-only model based on XLM-R Large with E5-style contrastive pre-training and mean pooling.

| Parameter           | SEA-LION-E5-Embedding-600M        |
| ------------------- | --------------------------------- |
| **d\_model**        | 1024                              |
| **head\_dim**       | 16                                |
| **Vocabulary**      | 250,000 (SentencePiece)           |
| **Sequence Length** | 512                               |
| **Pooling Mode**    | Mean tokens (with attention mask) |

## Glossary

* **E5:** "EmbEddings from bidirEctional Encoder rEpresentations" – a weakly-supervised contrastive pre-training method for text embeddings.
* **SEA-BED:** Southeast Asia Embedding Benchmark – a comprehensive evaluation suite for embedding models on SEA languages.
* **Asymmetric Retrieval:** Retrieval tasks where query and document formulations differ; E5 uses prefixes to handle this.
* **Mean Pooling:** Aggregating token embeddings by averaging (weighted by attention mask) to produce a fixed-size sentence representation.

## More Information

While this model supports masked language modeling, it is primarily optimised via contrastive fine-tuning for downstream tasks such as sequence classification, token classification, or question answering. Please note that these weights have not been specifically aligned for safety; therefore, developers should implement their own safety evaluations and security measures. The authors disclaim all liability for any claims, damages, or other liabilities arising from the use of the released code or weights.

For more info, please contact us at <sealion@aisingapore.org>


# SEA-GUARD

**SEA-Guard**, released on **4 Feb 2026**, is our premier suite of safety and moderation models tailored for the Southeast Asian landscape. This collection focuses on robust visual and textual guardrails, ensuring AI deployments remain secure and culturally compliant across the region.

## The SEA-Guard Collection

The SEA-Guard suite currently consists of 4 specialized models:

* [**Qwen-SEA-Guard-4B**](/models/sea-guard/qwennllama-sea-guard) (`https://huggingface.co/aisingapore/Qwen-SEA-Guard-4B-040226`) is a **4B Image-to-Text** model serving as a lightweight visual safety guardrail, specifically optimized for efficient edge applications where low latency is critical.
* [**Qwen-SEA-Guard-8B**](/models/sea-guard/qwennllama-sea-guard) (`https://huggingface.co/aisingapore/Qwen-SEA-Guard-8B-040226`) is an **8B Image-to-Text** model that offers a balanced visual moderation solution with stronger reasoning capabilities for more nuanced content detection.
* [**Llama-SEA-Guard-8B**](/models/sea-guard/qwennllama-sea-guard) (`https://huggingface.co/aisingapore/Llama-SEA-Guard-8B-040226`) is an **8B Text Generation** model. This text-only safety specialist is optimized for chat moderation and policy enforcement, ensuring safe interactions in dialogue systems.
* [**Gemma-SEA-Guard-12B**](/models/sea-guard/gemma-sea-guard) (`https://huggingface.co/aisingapore/Gemma-SEA-Guard-12B-040226`) is a high-capacity **12B Image-Text-to-Text** multimodal safety model designed for complex content analysis, capable of interpreting intricate relationships between visual and textual data.

## Our Commitment to Safe Regional AI

The SEA-Guard collection advances our mission to build AI that is not only culturally intelligent but also inherently safe. By providing specialized guardrails for both text and vision, these models enable developers to deploy AI solutions across Southeast Asia with confidence in their safety and compliance standards.

> **Note:** For detailed integration guides and benchmark data for each SEA-Guard model, please refer to their individual documentation pages via the links above.


# Gemma-SEA-GUARD-12B

**SEA-Guard** is a collection of safety-focused Large Language Models (LLMs) designed specifically for the Southeast Asia (SEA) region. While the collection comprises four distinct models, we currently offer a single API endpoint that exclusively serves the [Gemma-based model](https://playground.sea-lion.ai/sea-guard). You can generate an API key to access this model at [sea-lion api key manager](https://playground.sea-lion.ai/key-manager).

## Model Details

### Model Description

SEA-LION stands for *Southeast Asian Languages In One Network* and is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

This model is a fine-tuned version of [Gemma 3 12B IT](https://huggingface.co/google/gemma-3-12b-it) on 1M instruction-following pairs. For more details on training data, please refer to the paper [SEA-Guard](https://arxiv.org/abs/2602.01618).

For tokenization, the model employs the default tokenizer used in Gemma 3.

* **Developed by:** AI Products Pillar, AI Singapore
* **Funded by:** Singapore NRF
* **Shared by:** AI Products Pillar, AI Singapore
* **Model type:** Decoder
* **Context length:** 128k tokens
* **Language(s) (text):** Burmese, English, Indonesian, Malay, Tagalog, Tamil, Thai, and Vietnamese
* **License:** [Gemma](https://ai.google.dev/gemma/terms)
* **Finetuned from model:** [Gemma 3 12B IT](https://huggingface.co/google/gemma-3-12b-it)

### Model Sources

* **Repository:** 🤗[HuggingFace SEA-Guard Collection](https://huggingface.co/collections/aisingapore/sea-guard)

## Intended Uses and Limitations

This model is optimized to return a binary classification in text form: \["safe", "unsafe"]. However, users must be aware that the model is subject to the limitations common to generative AI, including the potential to hallucinate or generate ungrounded, irrelevant text. Due to these inherent risks, human oversight is advised, and the model’s outputs should not be treated as absolute determinations without secondary verification.

## Uses

### Direct Use

The output of the model is only "safe" or "unsafe". Users can directly use it without any finetune or in-context learning since it is already trained with cultural safety for SEA contexts. We also release the API of this model at [sea-lion.ai](https://playground.sea-lion.ai/key-manager).

### Downstream Use

Users can also continue training this model further on the target tasks, e.g., vision-text safety datasets. Also, this model is supported by vLLM for fast inference.

## Training and Evaluation Data

For more details on training data, please refer to the paper [SEA-Guard](https://arxiv.org/abs/2602.01618).

## Training Procedure

We employ a supervised-finetuning technique (SFT) on Llama-factory with the following hyperparameters.

### Testing Data, Factors & Metrics

We use [SEA-SafeguardBench](https://github.com/aisingapore/sealion/blob/main/models/sea-guard/arxiv.org/abs/2512.05501) to evaluate our SEA-Guard. Note that we also evaluated the vision-text safety classification in [our research paper](https://arxiv.org/abs/2602.01618)

#### Metrics

AUPRC is the primary metric to evaluate the safety classification of our models.

## Citation

**BibTeX:**

```
@misc{tasawong2026seaguardculturallygroundedmultilingual,
      title={SEA-Guard: Culturally Grounded Multilingual Safeguard for Southeast Asia}, 
      author={Panuthep Tasawong and Jian Gang Ngui and Alham Fikri Aji and Trevor Cohn and Peerat Limkonchotiwat},
      year={2026},
      eprint={2602.01618},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2602.01618}, 
}
```

## More Information

This is the repository for the commercial instruction-tuned model. Notwithstanding the model's safety-aligned training, developers and users are advised to conduct their own safety fine-tuning and implement appropriate security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

AI Singapore is a national programme supported by the National Research Foundation, Singapore and hosted by the National University of Singapore. Any opinions, findings and conclusions or recommendations expressed in this material are those of the author(s) and do not reflect the views of the National Research Foundation or the National University of Singapore.

## Contact

<sealion@aisingapore.org>


# Qwen and Llama SEA-GUARD

**SEA-Guard** is a collection of safety-focused Large Language Models (LLMs) built upon the SEA-LION family, designed specifically for the Southeast Asia (SEA) region.

## Model Details

### Model Description

SEA-LION stands for *Southeast Asian Languages In One Network* and is a collection of Large Language Models (LLMs) which have been pretrained and instruct-tuned for the Southeast Asia (SEA) region.

This model is a fine-tuned version of [aisingapore/Qwen-SEA-LION-v4-4B-VL](https://huggingface.co/aisingapore/Qwen-SEA-LION-v4-4B-VL), [aisingapore/Qwen-SEA-LION-v4-8B-VL](https://huggingface.co/aisingapore/Qwen-SEA-LION-v4-8B-VL) and [aisingapore/Llama-SEA-LION-v3-8B-IT](https://huggingface.co/aisingapore/Llama-SEA-LION-v3-8B-IT) on 1M instruction-following pairs. For more details on training data, please refer to the paper [SEA-Guard](https://arxiv.org/abs/2602.01618).

For tokenization, the model employs the default tokenizer used in Qwen3-VL.

* **Developed by:** AI Products Pillar, AI Singapore
* **Funded by:** Singapore NRF
* **Shared by:** AI Products Pillar, AI Singapore
* **Model type:** Decoder
* **Context length:** 128k tokens
* **Language(s) (text):** Burmese, English, Indonesian, Malay, Tagalog, Tamil, Thai, and Vietnamese
* **License:**
  * Qwen : [Apache-2.0](https://choosealicense.com/licenses/apache-2.0/)
  * Llama : [Llama 3.1 Community License](https://huggingface.co/meta-llama/Llama-3.1-70B-Instruct/blob/main/LICENSE)
* **Finetuned from model:** [aisingapore/Qwen-SEA-LION-v4-4B-VL](https://huggingface.co/aisingapore/Qwen-SEA-LION-v4-4B-VL), [aisingapore/Qwen-SEA-LION-v4-8B-VL](https://huggingface.co/aisingapore/Qwen-SEA-LION-v4-8B-VL) and [aisingapore/Llama-SEA-LION-v3-8B-IT](https://huggingface.co/aisingapore/Llama-SEA-LION-v3-8B-IT)

## Download the Models

Qwen and Llama SEA-Guard models are available for download via the following channels:

* [Qwen-SEA-Guard-4B-040226](https://huggingface.co/aisingapore/Qwen-SEA-Guard-4B-040226)
* [Qwen-SEA-Guard-8B-040226](https://huggingface.co/aisingapore/Qwen-SEA-Guard-8B-040226)
* [Llama-SEA-Guard-8B-040226](https://huggingface.co/aisingapore/aisingapore/Llama-SEA-Guard-8B-040226)

**Repository:** 🤗[HuggingFace SEA-Guard Collection](https://huggingface.co/collections/aisingapore/sea-guard)

## Intended Uses and Limitations

This model is optimized to return a binary classification in text form: \["safe", "unsafe"]. However, users must be aware that the model is subject to the limitations common to generative AI, including the potential to hallucinate or generate ungrounded, irrelevant text. Due to these inherent risks, human oversight is advised, and the model’s outputs should not be treated as absolute determinations without secondary verification.

## Uses

### Direct Use

The output of the model is only "safe" or "unsafe". Users can directly use it without any finetune or in-context learning since it is already trained with cultural safety for SEA contexts.

### Downstream Use

Users can also continue training this model further on the target tasks, e.g., vision-text safety datasets. Also, this model is supported by vLLM for fast inference.

## Training and evaluation data

For more details on training data, please refer to the paper [SEA-Guard](https://arxiv.org/abs/2602.01618).

## Training procedure

We employ a supervised-finetuning technique (SFT) on Llama-factory with the following hyperparameters.

### Testing Data, Factors & Metrics

We use [SEA-SafeguardBench](https://arxiv.org/pdf/2512.05501) to evaluate our SEA-Guard. Note that we also evaluated the vision-text safety classification in [our research paper](https://arxiv.org/abs/2602.01618)

## Citation

**BibTeX:**

```
@misc{tasawong2026seaguardculturallygroundedmultilingual,
      title={SEA-Guard: Culturally Grounded Multilingual Safeguard for Southeast Asia}, 
      author={Panuthep Tasawong and Jian Gang Ngui and Alham Fikri Aji and Trevor Cohn and Peerat Limkonchotiwat},
      year={2026},
      eprint={2602.01618},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2602.01618}, 
}
```

## More Information

This is the repository for the commercial instruction-tuned model. Notwithstanding the model's safety-aligned training, developers and users are advised to conduct their own safety fine-tuning and implement appropriate security measures. In no event shall the authors be held liable for any claims, damages, or other liabilities arising from the use of the released weights and codes.

AI Singapore is a national programme supported by the National Research Foundation, Singapore and hosted by the National University of Singapore. Any opinions, findings and conclusions or recommendations expressed in this material are those of the author(s) and do not reflect the views of the National Research Foundation or the National University of Singapore.

## Contact

<sealion@aisingapore.org>


# Getting the models

SEA-LION models are available for download via the following channels:

## SEA-Lion v4.5 (Latest)

[Full HuggingFace SEA-LION v4.5 Collection](https://huggingface.co/collections/aisingapore/sea-lion-v45)

**Gemma-SEA-LION-v4.5-E2B-IT**

| Model                           | Download                                                                                                                                               |
| ------------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Gemma-SEA-LION-v4.5-E2B-IT      | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-LION-v4.5-E2B-IT), [Ollama](https://ollama.com/aisingapore/Gemma-SEA-LION-v4.5-E2B-IT)      |
| Gemma-SEA-LION-v4.5-E2B-IT-GGUF | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-LION-v4.5-E2B-IT-GGUF), [Ollama](https://ollama.com/aisingapore/Gemma-SEA-LION-v4.5-E2B-IT) |

<br>

**Qwen-SEA-LION-v4.5-27B-IT**

| Model                                 | Download                                                                                                                                             |
| ------------------------------------- | ---------------------------------------------------------------------------------------------------------------------------------------------------- |
| Qwen-SEA-LION-v4.5-27B-IT             | [HuggingFace](https://huggingface.co/aisingapore/Qwen-SEA-LION-v4.5-27B-IT), [Ollama](https://ollama.com/aisingapore/Qwen-SEA-LION-v4.5-27B-IT)      |
| Qwen-SEA-LION-v4.5-27B-IT-SpecDecoder | [HuggingFace](https://huggingface.co/aisingapore/Qwen-SEA-LION-v4.5-27B-IT-SpecDecoder)                                                              |
| Qwen-SEA-LION-v4.5-27B-IT-GGUF        | [HuggingFace](https://huggingface.co/aisingapore/Qwen-SEA-LION-v4.5-27B-IT-GGUF), [Ollama](https://ollama.com/aisingapore/Qwen-SEA-LION-v4.5-27B-IT) |

<br>

## SEA-Lion Embedding

[Full HuggingFace SEA-LION-Embedding Collection](https://huggingface.co/collections/aisingapore/sea-lion-modernbert-and-embedding)

| Model                                | Download                                                                               |
| ------------------------------------ | -------------------------------------------------------------------------------------- |
| SEA-LION-ModernBERT-300M-checkpoints | [HuggingFace](https://huggingface.co/aisingapore/SEA-LION-ModernBERT-300M-checkpoints) |
| SEA-LION-ModernBERT-600M-checkpoints | [HuggingFace](https://huggingface.co/aisingapore/SEA-LION-ModernBERT-600M-checkpoints) |
| SEA-LION-ModernBERT-300M             | [HuggingFace](https://huggingface.co/aisingapore/SEA-LION-ModernBERT-300M)             |
| SEA-LION-ModernBERT-600M             | [HuggingFace](https://huggingface.co/aisingapore/SEA-LION-ModernBERT-600M)             |
| SEA-LION-ModernBERT-Embedding-300M   | [HuggingFace](https://huggingface.co/aisingapore/ModernBERT-SEA-Embedding-300M)        |
| SEA-LION-ModernBERT-Embedding-600M   | [HuggingFace](https://huggingface.co/aisingapore/ModernBERT-SEA-Embedding-600M)        |
| SEA-LION-E5-Embedding-600M           | [HuggingFace](https://huggingface.co/aisingapore/SEA-LION-E5-Embedding-600M)           |

<br>

## SEA-LION v4

[Full HuggingFace SEA-LION v4 Collection](https://huggingface.co/collections/aisingapore/sea-lion-v4)

<br>

**Apertus-SEA-LION-v4-8B**

| Model                         | Download                                                                    |
| ----------------------------- | --------------------------------------------------------------------------- |
| **Apertus-SEA-LION-v4-8B-IT** | [HuggingFace](https://huggingface.co/aisingapore/Apertus-SEA-LION-v4-8B-IT) |

<br>

**Gemma-SEA-LION-v4-VL**

| Model                       | Download                                                                   |
| --------------------------- | -------------------------------------------------------------------------- |
| **Gemma-SEA-LION-v4-4B-VL** | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-LION-v4-4B-VL)  |
| Gemma-SEA-LION-v4-27B-VL    | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-LION-v4-27B-VL) |

<br>

**Gemma-SEA-LION-v4-27B**

| Model                                | Download                                                                                                                                           |
| ------------------------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------- |
| Gemma-SEA-LION-v4-27B                | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-LION-v4-27B)                                                                            |
| Gemma-SEA-LION-v4-27B-IT             | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-LION-v4-27B-IT)                                                                         |
| Gemma-SEA-LION-v4-27B-IT-GGUF        | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-LION-v4-27B-IT-GGUF), [Ollama](https://ollama.com/aisingapore/Gemma-SEA-LION-v4-27B-IT) |
| Gemma-SEA-LION-v4-27B-IT-NVFP4       | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-LION-v4-27B-IT-NVFP4)                                                                   |
| Gemma-SEA-LION-v4-27B-IT-FP8-Dynamic | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-LION-v4-27B-IT-FP8-Dynamic)                                                             |

<br>

**Qwen-SEA-LION-v4-VL**

| Model                  | Download                                                                 |
| ---------------------- | ------------------------------------------------------------------------ |
| Qwen-SEA-LION-v4-4B-VL | [HuggingFace](https://huggingface.co/aisingapore/Qwen-SEA-LION-v4-4B-VL) |
| Qwen-SEA-LION-v4-8B-VL | [HuggingFace](https://huggingface.co/aisingapore/Qwen-SEA-LION-v4-8B-VL) |

<br>

**Qwen-SEA-LION-v4-32B-IT**

| Model                        | Download                                                                       |
| ---------------------------- | ------------------------------------------------------------------------------ |
| Qwen-SEA-LION-v4-32B-IT      | [HuggingFace](https://huggingface.co/aisingapore/Qwen-SEA-LION-v4-32B-IT)      |
| Qwen-SEA-LION-v4-32B-IT-4BIT | [HuggingFace](https://huggingface.co/aisingapore/Qwen-SEA-LION-v4-32B-IT-4BIT) |
| Qwen-SEA-LION-v4-32B-IT-8BIT | [HuggingFace](https://huggingface.co/aisingapore/Qwen-SEA-LION-v4-32B-IT-8BIT) |

<br>

## SEA-LION v3.5

[Full HuggingFace SEA-LION v3.5 Collection](https://huggingface.co/collections/aisingapore/sea-lion-v35-67fc3ab84300d7e6088fa32c)

**Llama-SEA-LION-v3.5-8B**

| Model                         | Download                                                                                                                                           |
| ----------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
| Llama-SEA-LION-v3.5-8B-R      | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v3.5-8B-R)                                                                         |
| Llama-SEA-LION-v3.5-8B-R-GGUF | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v3.5-8B-R-GGUF), [Ollama](https://ollama.com/aisingapore/Llama-SEA-LION-v3.5-8B-R) |

<br>

**Llama-SEA-LION-v3.5-70B**

| Model                          | Download                                                                                                                                             |
| ------------------------------ | ---------------------------------------------------------------------------------------------------------------------------------------------------- |
| Llama-SEA-LION-v3.5-70B-R      | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v3.5-70B-R)                                                                          |
| Llama-SEA-LION-v3.5-70B-R-GGUF | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v3.5-70B-R-GGUF), [Ollama](https://ollama.com/aisingapore/Llama-SEA-LION-v3.5-70B-R) |

## SEA-LION v3

[Full HuggingFace SEA-LION v3 Collection](https://huggingface.co/collections/aisingapore/sea-lionv3-672589a39cdadd6a5b199581)

**Gemma-SEA-LION-v3-9B**

| Model                        | Download                                                                                                                                         |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ |
| Gemma-SEA-LION-v3-9B         | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-LION-v3-9B)                                                                           |
| Gemma-SEA-LION-v3-9B-IT      | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-LION-v3-9B-IT)                                                                        |
| Gemma-SEA-LION-v3-9B-IT-GGUF | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-LION-v3-9B-IT-GGUF), [Ollama](https://ollama.com/aisingapore/Gemma-SEA-LION-v3-9B-IT) |

<br>

**Llama-SEA-LION-v3-8B**

| Model                        | Download                                                                                                                                                     |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------ |
| Llama-SEA-LION-v3-8B         | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v3-8B), [Kaggle](https://www.kaggle.com/models/ai-singapore/llama3.1-8b-cpt-sea-lionv3-base) |
| Llama-SEA-LION-v3-8B-IT      | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v3-8B-IT)                                                                                    |
| Llama-SEA-LION-v3-8B-IT-GGUF | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v3-8B-IT-GGUF), [Ollama](https://ollama.com/aisingapore/Llama-SEA-LION-v3-8B-IT)             |

<br>

**Llama-SEA-LION-v3-70B**

| Model                         | Download                                                                                                                                           |
| ----------------------------- | -------------------------------------------------------------------------------------------------------------------------------------------------- |
| Llama-SEA-LION-v3-70B         | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v3-70B)                                                                            |
| Llama-SEA-LION-v3-70B-IT      | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v3-70B-IT)                                                                         |
| Llama-SEA-LION-v3-70B-IT-GGUF | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v3-70B-IT-GGUF), [Ollama](https://ollama.com/aisingapore/Llama-SEA-LION-v3-70B-IT) |

<br>

## SEA-LION v2

[Full HuggingFace SEA-LION v2 Collection](https://huggingface.co/collections/aisingapore/sea-lionv2-672589c4c7ea47e4174d3e7f)

| Model                        | Download                                                                                                                                         |
| ---------------------------- | ------------------------------------------------------------------------------------------------------------------------------------------------ |
| Llama-SEA-LION-v2-8B         | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v2-8B)                                                                           |
| Llama-SEA-LION-v2-8B-IT      | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v2-8B-IT)                                                                        |
| Llama-SEA-LION-v2-8B-IT-GGUF | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-LION-v2-8B-IT-GGUF), [Ollama](https://ollama.com/aisingapore/Llama-SEA-LION-v2-8B-IT) |

<br>

## SEA-LION v1

[Full HuggingFace SEA-LION v1 Collection](https://huggingface.co/collections/aisingapore/sea-lionv1-672589cd29a1781afa6be35e)

| Model                  | Download                                                                 |
| ---------------------- | ------------------------------------------------------------------------ |
| SEA-LION-v1-3B         | [HuggingFace](https://huggingface.co/aisingapore/SEA-LION-v1-3B)         |
| SEA-LION-v1-7B         | [HuggingFace](https://huggingface.co/aisingapore/SEA-LION-v1-7B)         |
| SEA-LION-v1-7B-IT      | [HuggingFace](https://huggingface.co/aisingapore/SEA-LION-v1-7B-IT)      |
| SEA-LION-v1-7B-IT-GGUF | [HuggingFace](https://huggingface.co/aisingapore/SEA-LION-v1-7B-IT-GGUF) |

<br>

## SEA-GUARD

[Full HuggingFace SEA-Guard Collection](https://huggingface.co/collections/aisingapore/sea-guard)

| Model               | Download                                                                     |
| ------------------- | ---------------------------------------------------------------------------- |
| Qwen-SEA-Guard-4B   | [HuggingFace](https://huggingface.co/aisingapore/Qwen-SEA-Guard-4B-040226)   |
| Qwen-SEA-Guard-8B   | [HuggingFace](https://huggingface.co/aisingapore/Qwen-SEA-Guard-8B-040226)   |
| Llama-SEA-Guard-8B  | [HuggingFace](https://huggingface.co/aisingapore/Llama-SEA-Guard-8B-040226)  |
| Gemma-SEA-Guard-12B | [HuggingFace](https://huggingface.co/aisingapore/Gemma-SEA-Guard-12B-040226) |


# SEA-HELM

SEA-HELM (Southeast Asian Holistic Evaluation of Language Models) is a comprehensive benchmark suite designed to evaluate the linguistic and cultural competencies of Large Language Models (LLMs) in Southeast Asian (SEA) languages.

Formerly known as BHASA, SEA-HELM has been expanded and integrated with HELM to provide a more rigorous and authentic evaluation. It addresses the critical need for multilingual and multicultural benchmarks in the rapidly evolving field of LLMs.

SEA-HELM serves as a valuable resource for developers looking to create and assess LLMs for linguistic accuracy and cultural sensitivity in the Southeast Asian context. It currently covers Filipino, Indonesian, Javanese, Sundanese, Tamil, Thai, and Vietnamese.

## Motivations

Several key factors motivated our development of SEA-HELM:

### Lack of Comprehensive SEA Language Benchmarks

Existing LLM benchmarks are capable of evaluating specific capabilities of LLMs in English and some mid- to low-resource languages, including those in the Southeast Asian (SEA) region. However, there has not been a comprehensive and authentic evaluation suite developed specifically for SEA languages, a problem exacerbated by the lack of both training and testing data on the internet.

### Need for Multilingual and Multicultural Evaluations

There's an increasing need for benchmarks that assess not only linguistic capabilities but also cultural representation and sensitivity. Evaluations of linguistic and cultural representation are essential for gauging the efficacy and fairness of language models.

SEA-HELM deliberately incorporates community participation by involving native speakers to ensure linguistic accuracy and cultural authenticity. For example, the evaluation suite includes a cultural evaluation dataset for Filipino developed in collaboration with community members from the Philippines.

## Core Pillars of SEA-HELM

SEA-HELM is organized into five core evaluation pillars, designed to comprehensively assess various competencies of LLMs.

Each evaluation task under SEA-HELM uses prompts in the target language to ensure that the LLM can interpret native instructions correctly. The datasets used are either originally written in the native language or carefully translated by native speakers to avoid translation errors. This ensures that the evaluation is authentic and relevant to the linguistic nuances of each language.

### 1. NLP Classics

This pillar focuses on evaluating the fundamental natural language processing abilities of LLMs in Southeast Asian languages, such as language understanding, generation, and reasoning.

#### Natural Language Understanding (NLU)

NLU tasks assesses the model's ability to comprehend text. The tasks included in this competency are question answering and sentiment analysis.

* **Question Answering (QA)** evaluates the model's ability to understand a question and extract the answer from a given text. SEA-HELM uses datasets like TyDi QA for Indonesian, XQuAD for Vietnamese and Thai, and IndicQA for Tamil. The models are prompted to answer questions by extracting the answer from a provided paragraph.
* **Sentiment Analysis:** determines sentiment expressed in a text. SEA-HELM uses datasets like NusaX for Indonesian, UIT-VSFC for Vietnamese, Wisesight Sentiment for Thai, and IndicSentiment for Tamil. The models are prompted to determine the sentiment of a given sentence and respond with a single word: Positive, Negative, or Neutral.

#### Natural Language Generation (NLG)

NLG tasks evaluates a model's ability to generate human-like text. The tasks included in this competency are machine translation and abstractive summarization.

* **Machine Translation** assesses the model's ability to translate text from one language to another. SEA-HELM evaluates translation between English and target SEA languages. The models are prompted to translate a given text into a specified language.
* **Abstractive Summarization** requires the model to read a document, identify the key points, and summarize them into a coherent and fluent text. SEA-HELM uses the XLSum dataset for all four target languages. The models are prompted to summarize an article in one or two sentences in the specified language

#### Natural Language Reasoning (NLR)

NLR tasks assesses the model's ability to reason and draw inferences from text. The tasks included in this competency are causal reasoning and Natural Language Inference (NLI).

NLI tasks involve determining whether a given premise entails or contradicts a hypothesis. SEA-HELM uses datasets like IndoNLI for Indonesian, XNLI for Vietnamese and Thai, and IndicXNLI for Tamil.

### 2. LLM-Specifics

This pillar focuses on evaluating capabilities unique to LLMs, such as instruction following and chat capabilities.

* **Instruction following** assesses the ability of LLMs to follow human instructions and adhere to specified formats using **SEA-IFEval**, a benchmark we manually created collaboratively with native speakers.
* **Chat capability** evaluates the ability of LLMs to engage in human-like conversations using **SEA-MTBench**, another manually translated and localised version of the popular MT-Bench dataset.

### 3. SEA Linguistics

This pillar focuses on systematically diagnosing language proficiency and grammatical understanding of language models in Southeast Asian languages by utilizing LINDSEA (LINguistic Diagnostics for Southeast Asian languages).

LINDSEA is a high-quality, manually-crafted linguistic dataset designed to provide a fine-grained evaluation of a model’s linguistic abilities, and is the first dataset of its kind created for SEA languages.

### 4. SEA Culture

This pillar is designed to assess cultural representation and sensitivity in language models. This pillar recognizes the importance of evaluating LLMs on their understanding and appropriate handling of cultural nuances, social norms, and values specific to Southeast Asian cultures.

SEA-HELM uses cultural diagnostics to probe for both cultural representation and sensitivity. The goal is to ensure that LLMs used in SEA contexts are not only linguistically accurate but also culturally aware and respectful.

SEA-HELM achieves authentic cultural representation through a strong participatory approach that includes native speaker communities. For example, SEA-HELM includes a cultural evaluation dataset for Filipino developed in collaboration with community members from the Philippines, which resulted in KALAHI.

### 5. Safety

This pillar ensures that language models (LLMs) do not produce harmful or unsafe outputs, especially in the context of Southeast Asian (SEA) languages and cultures. It recognizes that multilingual inputs, particularly in lower-resource languages common in the SEA region, can increase the likelihood of LLMs generating unsafe responses. The pillar aims to protect users interacting with these models in SEA languages from issues like hate speech.

Currently, the Safety pillar focuses on toxicity detection as its primary task, with coverage for Indonesian, Thai, Vietnamese, and Filipino. This involves identifying toxic content such as hate speech and abusive language in text, which is crucial for content moderation. SEA-HELM uses datasets like MLHSD for Indonesian, Thai Toxicity Tweet for Thai, ViHSD for Vietnamese, and PH Elections Toxicity for Filipino.

While passing these evaluations is a necessary step, it is not a guarantee of complete safety in real-world scenarios, as it is impossible to cover every type of unsafe response. SEA-HELM plans to expand this pillar to include a broader range of safety-related tasks in the future.

## SEA-HELM Leaderboard

To provide transparency, comparative insights and understanding of models' multilingual and multicultural performance, SEA-HELM maintains a [public leaderboard](https://leaderboard.sea-lion.ai).

The [leaderboard](https://leaderboard.sea-lion.ai) offers multiple views, including overall scores, language-specific scores, and detailed task scores, allowing users to delve deeper into the evaluation results.

## SEA-HELM Datasets and Resources

The full codebase for SEA-HELM is provided open-source on [Github](https://github.com/aisingapore/sea-helm).

Our SEA-HELM evaluation datasets are publicly available on [HuggingFace](https://huggingface.co/collections/aisingapore/sea-helm-evaluation-datasets-67593d0bb8c9f17f9f6b0fcb).

## Limitations and Future Work

While SEA-HELM aims for holistic evaluations, it is not yet exhaustive in its coverage of languages and tasks. Areas for improvement include:

* Expanding SEA language coverage beyond the current seven languages (e.g., Burmese, Khmer, Lao).
* Exploring automatic LLM evaluations for better benchmarking.
* Enhancing the safety pillar to include more real-world risk factors.
* Adding more cultural and linguistic assessments

## Resources

For more details and information on SEA-HELM, you can refer to the following materials:

1. [BHASA: A Holistic Southeast Asian Linguistic and Cultural Evaluation Suite for Large Language Models](https://arxiv.org/abs/2309.06085)
2. [SEA-HELM: Southeast Asian Holistic Evaluation of Language Models](https://arxiv.org/abs/2502.14301)
3. [Towards fair and comprehensive multilingual LLM benchmarking](https://cohere.com/blog/towards-fair-and-comprehensive-multilingual-and-multicultural-llm-benchmarking)


# Model Deployment & Inferencing

This section covers the various environments where SEA-LION models can be hosted, ranging from fully managed cloud services to self-hosted high-performance infrastructure. SEA-LION models are freely available for [download](/models/download_models). The following guides provide technical how-tos for setting up SEA-LION inference using different approaches:

1. [Using our provided SEA-LION API](/guides/inferencing/api)
2. Running SEA-LION on a local machine (coming soon)
3. Deploying SEA-LION on the cloud
   * [Create SEA-LION endpoint on Google Vertex AI](/guides/inferencing/vertex_ai)
   * [Importing and Using Llama-SEA-LION models in a Serverless On-Demand Environment with Amazon Bedrock](/guides/inferencing/amazon_bedrock)
   * [OpenAI-compatible APIs with Llama-SEA-LION models and Bedrock Access Gateway](/guides/inferencing/bedrock_access_gateway)
   * [Deploying Gemma-SEA-LION models using AWS Sagemaker AI](https://github.com/aisingapore/sealion/blob/main/guides/inferencing/Gemma-SEA-LION-v4-27B-Instruct.ipynb)
   * [Deploying SEA-LION using vLLM on Linux server](/guides/inferencing/vllm_linux)
4. Leveraging our Partner API Platforms
   * [Cloudflare Workers AI](/guides/inferencing/cloudflare)


# SEA-LION API

The SEA-LION API provides a quick and simple interface to our various SEA-LION models for text generation, translation, summarization, and more.

Usage of the SEA-LION API is subject to our [Terms of Use](https://sea-lion.ai/terms-of-use/) and [Privacy Policy](https://sea-lion.ai/privacy-policy/)

## Getting an API Key

To get started with SEA-LION API, you'll need to first create an API key via our [SEA-LION Playground](https://playground.sea-lion.ai/):

1. Sign in to SEA-LION Playground via your Google account
2. Navigate to our API Key Manager page by clicking on

* `API Key` on the side menu, or
* `Launch Key Manager` on the home dashboard

<figure><img src="/files/nmfrLpvdfwvxKQutYiOg" alt="" width="100%"><figcaption></figcaption></figure>

3. Click on the "Create New Trial API Key" button, and enter a name for your API key.

<figure><img src="/files/bqavH2kSyqtFlA9M53qd" alt="" width="100%"><figcaption></figcaption></figure>

An API key will be generated for you after you click "Create". **Make sure to copy or download the generated key** and keep it in a safe place since you won't be able to view it again.

<figure><img src="/files/bWHaK4hGnHApPhF2uqwt" alt="" width="400"><figcaption></figcaption></figure>

Only 1 API key is allowed to be created per user.

## How To Use Your API Key

### Step 1. Find the Available Models

To find the available SEA-LION models for your API key, use the following curl command.

```
curl 'https://api.sea-lion.ai/v1/models' \
  -H 'Authorization: Bearer YOUR_API_KEY'
```

Replace YOUR\_API\_KEY with your generated API key.

### Step 2: Call the API

SEA-LION's API endpoints for chat are compatible with OpenAI's API and libraries.

#### Calling our Instruct models

{% tabs %}
{% tab title="curl" %}

```
curl https://api.sea-lion.ai/v1/chat/completions \
  -H 'accept: text/plain' \
  -H 'Authorization: Bearer YOUR_API_KEY' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "aisingapore/Qwen-SEA-LION-v4.5-27B-IT",
    "messages": [
      {
        "role": "user",
        "content": "Tell me a Singlish joke!"
      }
    ]
  }'
```

{% endtab %}

{% tab title="python" %}

```python
from openai import OpenAI

client = OpenAI(
    api_key=YOUR_API_KEY,
    base_url="https://api.sea-lion.ai/v1"
)

completion = client.chat.completions.create(
    model="aisingapore/Qwen-SEA-LION-v4.5-27B-IT",
    messages=[
        {
            "role": "user",
            "content": "Tell me a Singlish joke!"
        }
    ]
)

print(completion.choices[0].message.content)
```

{% endtab %}
{% endtabs %}

#### Agentic Tool Use

The SEA-LION v4.5 model supports function calling, enabling agentic workflows where the model autonomously decides which tools to call, processes results, and continues until the task is complete.

{% tabs %}
{% tab title="python" %}

```python
from openai import OpenAI
import json

client = OpenAI(
    api_key=YOUR_API_KEY,
    base_url="https://api.sea-lion.ai/v1"
)

tools = [
    {
        "type": "function",
        "function": {
            "name": "get_weather",
            "description": "Get the current weather for a city",
            "parameters": {
                "type": "object",
                "properties": {
                    "city": {"type": "string", "description": "The city name"}
                },
                "required": ["city"]
            }
        }
    }
]

def get_weather(city):
    # Replace with a real weather API call
    mock_data = {"Singapore": "32°C, humid", "Jakarta": "30°C, cloudy"}
    return mock_data.get(city, "Weather data unavailable")

messages = [{"role": "user", "content": "What is the weather in Singapore and Jakarta?"}]

# Agentic loop: runs until the model stops calling tools
while True:
    response = client.chat.completions.create(
        model="aisingapore/Qwen-SEA-LION-v4.5-27B-IT",
        messages=messages,
        tools=tools,
        tool_choice="auto"
    )
    msg = response.choices[0].message
    messages.append(msg)

    if response.choices[0].finish_reason == "tool_calls":
        for tc in msg.tool_calls:
            args = json.loads(tc.function.arguments)
            result = get_weather(args["city"])
            messages.append({"role": "tool", "tool_call_id": tc.id, "content": result})
    else:
        print(msg.content)
        break
```

{% endtab %}
{% endtabs %}

#### Calling our Reasoning models

Our v3.5 models offers dynamic reasoning capabilities, and defaults to reasoning with `thinking_mode="on"` passed to the chat template. To use non-thinking mode ie. standard generations, pass `thinking_mode="off"` to the chat template instead.

{% tabs %}
{% tab title="curl" %}

```
curl https://api.sea-lion.ai/v1/chat/completions \
  -H 'accept: text/plain' \
  -H 'Authorization: Bearer YOUR_API_KEY' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "aisingapore/Llama-SEA-LION-v3.5-70B-R",
    "messages": [
      {
        "role": "user",
        "content": "Tell me a Singlish joke!"
      }
    ],
    "chat_template_kwargs": {
	    "thinking_mode": "off"
    }
  }'
```

{% endtab %}

{% tab title="python" %}

```python
from openai import OpenAI

client = OpenAI(
    api_key=YOUR_API_KEY,
    base_url="https://api.sea-lion.ai/v1"
)

completion = client.chat.completions.create(
    model="aisingapore/Llama-SEA-LION-v3.5-70B-R",
    messages=[
        {
            "role": "user",
            "content": "Tell me a Singlish joke!"
        }
    ],
    extra_body={
        "chat_template_kwargs": {
            "thinking_mode": "off"
        }
    },
)

print(completion.choices[0].message.content)
```

{% endtab %}
{% endtabs %}

{% hint style="info" %}
If you are not observing any changes in response when toggling `thinking_mode` on/off, your API responses might have been cached.

You can disable cache temporarily for your testing by setting the `no-cache` flag to `true`

{% tabs %}
{% tab title="curl" %}

```
curl https://api.sea-lion.ai/v1/chat/completions \
  -H 'accept: text/plain' \
  -H 'Authorization: Bearer YOUR_API_KEY' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "aisingapore/Llama-SEA-LION-v3.5-70B-R",
    "messages": [
      {
        "role": "user",
        "content": "Tell me a Singlish joke!"
      }
    ],
    "chat_template_kwargs": {
	    "thinking_mode": "off"
    },
    "cache": {
      "no-cache": true
    }
  }'
```

{% endtab %}

{% tab title="python" %}

```python
from openai import OpenAI

client = OpenAI(
    api_key=YOUR_API_KEY,
    base_url="https://api.sea-lion.ai/v1"
)

completion = client.chat.completions.create(
    model="aisingapore/Llama-SEA-LION-v3.5-70B-R",
    messages=[
        {
            "role": "user",
            "content": "Tell me a Singlish joke!"
        }
    ],
    extra_body={
        "chat_template_kwargs": {
            "thinking_mode": "off"
        },
        "cache": {
            "no-cache": True
        }
    },
)

print(completion.choices[0].message.content)
```

{% endtab %}
{% endtabs %}
{% endhint %}

#### Calling our Guard model

Our safety model, `aisingapore/SEA-Guard`, can be used to evaluate potentially harmful content. It returns a binary classification of `safe` and `unsafe`, and supports 3 flexible input modes, depending on your use case.

{% hint style="warning" %}
Note: The safety model **does not** support system prompts or multi-turn conversations.
{% endhint %}

**Mode 1: Prompt-only Classification**

{% tabs %}
{% tab title="curl" %}

```
curl https://api.sea-lion.ai/v1/chat/completions \
  -H 'accept: text/plain' \
  -H 'Authorization: Bearer YOUR_API_KEY' \
  -H 'Content-Type: application/json' \
  -d '{
  "messages": [
    {
      "role": "user",
      "content": {PROMPT}
    }
  ],
  "model": "aisingapore/SEA-Guard",
  "stream": false
}'
```

{% endtab %}

{% tab title="python" %}

```python
from openai import OpenAI

client = OpenAI(
    api_key=YOUR_API_KEY,
    base_url="https://api.sea-lion.ai/v1"
)

completion = client.chat.completions.create(
    model="aisingapore/SEA-Guard",
    messages=[
        {
            "role": "user",
            "content": {PROMPT}
        }
    ],
)

print(completion.choices[0].message.content)
```

{% endtab %}
{% endtabs %}

**Mode 2: Prompt + Response Classification**

{% tabs %}
{% tab title="curl" %}

```
curl https://api.sea-lion.ai/v1/chat/completions \
  -H 'accept: text/plain' \
  -H 'Authorization: Bearer YOUR_API_KEY' \
  -H 'Content-Type: application/json' \
  -d '{
  "messages": [
    {
      "role": "user",
      "content": "Human user:{PROMPT}\nAI assistant:{RESPONSE}."
    }
  ],
  "model": "aisingapore/SEA-Guard",
  "stream": false
}'
```

{% endtab %}

{% tab title="python" %}

```python
from openai import OpenAI

client = OpenAI(
    api_key=YOUR_API_KEY,
    base_url="https://api.sea-lion.ai/v1"
)

completion = client.chat.completions.create(
    model="aisingapore/SEA-Guard",
    messages=[
        {
            "role": "user",
            "content": "Human user:{PROMPT}\nAI assistant:{RESPONSE}."
        }
    ],
)

print(completion.choices[0].message.content)
```

{% endtab %}
{% endtabs %}

**Mode 3: Response-only Classification**

{% tabs %}
{% tab title="curl" %}

```
curl https://api.sea-lion.ai/v1/chat/completions \
  -H 'accept: text/plain' \
  -H 'Authorization: Bearer YOUR_API_KEY' \
  -H 'Content-Type: application/json' \
  -d '{
  "messages": [
    {
      "role": "user",
      "content": "Human user:\nAI assistant:{RESPONSE}."
    }
  ],
  "model": "aisingapore/SEA-Guard",
  "stream": false
}'
```

{% endtab %}

{% tab title="python" %}

```python
from openai import OpenAI

client = OpenAI(
    api_key=YOUR_API_KEY,
    base_url="https://api.sea-lion.ai/v1"
)

completion = client.chat.completions.create(
    model="aisingapore/SEA-Guard",
    messages=[
        {
            "role": "user",
            "content": "Human user:\nAI assistant:{RESPONSE}."
        }
    ],
)

print(completion.choices[0].message.content)
```

{% endtab %}
{% endtabs %}

#### Calling our Embedding models

[SEA-LION-ModernBERT-Embedding-600M](https://huggingface.co/aisingapore/SEA-LION-ModernBERT-Embedding-600M) produces 1024-dimensional embeddings and supports 11 Southeast Asian languages with an 8k token context window.

{% tabs %}
{% tab title="curl" %}

```
curl https://api.sea-lion.ai/v1/embeddings \
  -H 'Authorization: Bearer YOUR_API_KEY' \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "aisingapore/SEA-LION-ModernBERT-Embedding-600M",
    "input": [
      "Singapore is a tropical island city-state.",
      "The Lion City sits at the tip of the Malay Peninsula."
    ]
  }'
```

{% endtab %}

{% tab title="python" %}

```python
from openai import OpenAI

client = OpenAI(api_key="YOUR_API_KEY", base_url="https://api.sea-lion.ai/v1/embeddings")

result = client.embeddings.create(
    model="aisingapore/SEA-LION-ModernBERT-Embedding-600M",
    input=[
        "Singapore is a tropical island city-state.",
        "The Lion City sits at the tip of the Malay Peninsula.",
    ],
)

for item in result.data:
    print(f"index {item.index}: dim={len(item.embedding)}")
```

{% endtab %}
{% endtabs %}

#### Sentence Similarity

Embeddings can be used to measure semantic similarity between sentences using cosine similarity.

{% tabs %}
{% tab title="python" %}

```python
import math
import requests

def get_embeddings(texts, api_key):
    response = requests.post(
        "https://api.sea-lion.ai/v1/embeddings",
        headers={
            "Authorization": f"Bearer {api_key}",
            "Content-Type": "application/json"
        },
        json={"model": "aisingapore/SEA-LION-ModernBERT-Embedding-600M", "input": texts}
    )
    response.raise_for_status()
    return [item["embedding"] for item in response.json()["data"]]

def cosine_similarity(a, b):
    dot = sum(x * y for x, y in zip(a, b))
    norm_a = math.sqrt(sum(x ** 2 for x in a))
    norm_b = math.sqrt(sum(x ** 2 for x in b))
    return dot / (norm_a * norm_b)

api_key = YOUR_API_KEY

sentences = [
    "The food in Singapore is delicious.",
    "Singapore has amazing cuisine.",
    "The weather today is very hot.",
]

embeddings = get_embeddings(sentences, api_key)

for i in range(len(sentences)):
    for j in range(i + 1, len(sentences)):
        sim = cosine_similarity(embeddings[i], embeddings[j])
        print(f"[{i}] vs [{j}] similarity: {sim:.4f}")
        print(f"  '{sentences[i]}'")
        print(f"  '{sentences[j]}'")
```

{% endtab %}
{% endtabs %}

Sentences with similar meaning (e.g. food-related) score higher (≈0.95) than unrelated ones (≈0.63–0.65).

## Rate Limits

Limits help us mitigate misuse and manage API capacity and help ensure that everyone has fair access to the API.

SEA-LION API usage frequency will be subject to rate limits applied on requests per minute (RPM).

As of 04 Jun 2026, our rate limits is set to **10 requests per minute per user**.

If you have any questions or want to speak about getting a rate limit increase, reach out to <sealion@aisingapore.org>.


# Google Vertex AI

A step-by-step guide to deploying AI Singapore's SEA-LION models to run as an endpoint on Google's Vertex AI.

1. Go to Google's [Model Garden](https://console.cloud.google.com/vertex-ai/model-garden)
2. Scroll down (in the main page) to the "All Partners" section, and click on "Hugging Face".
3. Set the “Filter by Name” to “aisingapore”. A list of SEA-LION models will appear.

<figure><img src="/files/ETeVzmj0dsYpNbefJI2i" alt="" width="100%"><figcaption></figcaption></figure>

4. Click the model you wish to deploy, and a side panel will pop up. Here you can set the name of your endpoint, select the region and machine spec to deploy the model to. Then click on Deploy to complete the deployment.

   (Note: make sure you have enough hardware quota in your selected region)

<figure><img src="/files/VfZ0Z1WIkRitNmWFnWuA" alt="" width="50%"><figcaption></figcaption></figure>

5. The deployment takes some time (depending on the size of the model). When done, you will see the deployed endpoint in your [Endpoints page](https://console.cloud.google.com/vertex-ai/endpoints)

   Take note of your endpoint ID.

<figure><img src="/files/T54WfeXtbzmgdBwE1fXq" alt="" width="80%"><figcaption></figcaption></figure>

6. To test your deployed endpoint using a Python program,

   a. Get a Vertex AI service key from your administrator, or create one yourself if you have the right permissions on GCP.

   b. Install the [Google Cloud SDK](https://cloud.google.com/sdk/docs/install) for your system.

   c. Make sure your Google Cloud SDK is properly installed and authenticated by running the following command: `gcloud auth login`

   d. Create a new Python environment and install the following packages: `pip install google-cloud-aiplatform openai`
7. Create a Python program `vt-test.py` containing the following code:

```
import openai
import os
from google.auth import default
import google.auth.transport.requests

# Set the path to your service account key file
os.environ['GOOGLE_APPLICATION_CREDENTIALS'] = 'vt-svc-key.json'

PROJECT_ID = "<your GCP project ID>"
LOCATION = "<region of your deployed endpoint>"
ENDPOINT = "Your endpoint ID (from the above screenshot)"

# Get access token
# Note: the credential lives for 1 hour by default, and must be refreshed after expiration
# (https://cloud.google.com/docs/authentication/token-types#at-lifetime);

credentials, _ = default(scopes=["https://www.googleapis.com/auth/cloud-platform"])
credentials.refresh(google.auth.transport.requests.Request())

client = openai.OpenAI(
    base_url = f"https://{LOCATION}-aiplatform.googleapis.com/v1/projects/{PROJECT_ID}/locations/{LOCATION}/endpoints/{ENDPOINT}",
    api_key = credentials.token,
)

msg = "Write me a 5-line limerick"
max_tokens = 500
stream = False

resp = client.chat.completions.create(
    model = "",
    messages = [ {"role": "user" , "content": msg} ],
    max_tokens = max_tokens,
    stream = stream,
)
print(msg)
print(resp.choices[0].message.content)
```

8. Make sure your Vertex AI service key JSON file (named `vt-svc-key.json`) is in the same directory as your Python program.
9. Run your program `python vt-test.py` and see the result:

```
❯ python vt-test.py
Write me a 5-line limerick
There once was a fellow so bright,
Whose knowledge shone with great might.
In the city so grand,
He solved problems with hand,
And his name echoed day and night.
```

10. Below is a sample `bash` script `curl_test.sh` for calling the endpoint via CLI using `curl`:

```
#!/bin/bash

# 1. Set your variables
export PROJECT_ID="<your GCP project ID>"
export LOCATION="<region of your deployed endpoint>"
export ENDPOINT_ID="Your endpoint ID (from the above screenshot)"

# 2. Get a fresh access token
export ACCESS_TOKEN=$(gcloud auth print-access-token)

# 3. Define the request payload
read -r -d '' PAYLOAD << EOM
{
  "model": "",
  "messages": [
    {
      "role": "user",
      "content": "Write me a 5-line limerick"
    }
  ],
  "max_tokens": 500,
  "stream": false
}
EOM

# 4. Make the curl request
curl -X POST \
    -H "Authorization: Bearer ${ACCESS_TOKEN}" \
    -H "Content-Type: application/json" \
    "https://${LOCATION}-aiplatform.googleapis.com/v1/projects/${PROJECT_ID}/locations/${LOCATION}/endpoints/${ENDPOINT_ID}/chat/completions" \
    -d "${PAYLOAD}"
```

When run,

```
./curl_test.sh
{"id":"chatcmpl-da744c88eaa848929593366bec8f8900","object":"chat.completion","created":1760605976,
"model":"aisingapore/Gemma-SEA-LION-v4-27B-IT","choices":[{"index":0,"message":{"role":"assistant",
"content":"There once was a baker named Sue,\nWhose bread was a wonderful hue.\nWith berries so bright,\nA delectable sight,\nAnd a taste that was lovely and new!\n",
"refusal":null,"annotations":null,"audio":null,"function_call":null,"tool_calls":[],
"reasoning_content":null},"logprobs":null,"finish_reason":"stop","stop_reason":106}],
"service_tier":null,"system_fingerprint":null,
"usage":{"prompt_tokens":19,"total_tokens":59,"completion_tokens":40,"prompt_tokens_details":null},
"prompt_logprobs":null,"kv_transfer_params":null}
```


# Amazon Bedrock Custom Model Import

The [SEA-LION models](https://huggingface.co/aisingapore) are open source and freely available for research and commercial use. A common question we receive from developers is how to host and configure SEA-LION for model inference in their own environments. When it comes to deploying our models, organizations have several hosting options available to them. One such approach involves using the [Custom Model Import](https://aws.amazon.com/bedrock/custom-model-import/) feature in [Amazon Bedrock](https://aws.amazon.com/bedrock/), which allows for integration within Amazon’s cloud infrastructure.

At the time of writing, the [supported model architectures](https://docs.aws.amazon.com/bedrock/latest/userguide/model-customization-import-model.html#model-customization-import-model-architecture) include Llama 3 and Llama 3.1. The [Llama-SEA-LION-v3-8B-IT](https://huggingface.co/aisingapore/Llama-SEA-LION-v3-8B-IT) and [Llama-SEA-LION-v3-70B-IT](https://huggingface.co/aisingapore/Llama-SEA-LION-v3-70B-IT) models are supported. The SEA-LION models not built using the Llama architecture (such as Gemma2 and MPT) are not supported.

This article outlines the process of importing the SEA-LION model into Amazon Bedrock and includes a demo application that integrates with the imported model.

<figure><img src="/files/AUpA5U6ylP2idv7kJuN9" alt="" width="100%"><figcaption></figcaption></figure>

## Prerequisites

If you do not already have an Amazon Web Services (AWS) account, please [sign up](https://aws.amazon.com/resources/create-account/) first.

Note that in following this guide, the following AWS services are paid services and will incur costs:

* Amazon Bedrock
* Amazon S3

Please check that the following are installed on your development machine.

* [Git](https://git-scm.com/downloads)
* [Python](https://www.python.org/downloads/)

## Amazon Bedrock

[Amazon Bedrock](https://aws.amazon.com/bedrock/) is a fully managed service that simplifies the deployment and scaling of AI models. It provides access to high-performing foundation models, enabling a serverless experience for model deployment and integration. With **Custom Model Import**, users can upload their own models and use a unified platform for AI development.

### Pricing Model of Custom Model Import

There is no charge to import a custom model to Bedrock. Once you import a model, you will be able to access it **on-demand** without requiring to perform any control plane action.

You are only **charged for model inference**, based on the number of copies of your custom model required to service your inference volume and the duration each model copy is active, **billed in 5-minute windows**.

A model copy is a single instance of an imported model ready to serve inference requests. The price per model copy per minute depends on factors such as architecture, context length, AWS Region, compute unit version (hardware generation), and is tiered by model copy size.

A monthly storage cost per Custom Model Unit is applicable. The Custom Model Units needed to host a model depend on a variety of factors — notably the model architecture, model parameter count, and context length. The exact number of Custom Model Units needed will be determined at the time of import.

Please refer to the [Amazon Bedrock Pricing](https://aws.amazon.com/bedrock/pricing/) page, including the Custom Model Import pricing under the Pricing Details section, for the latest information.

## Import the SEA-LION Model

The model used in this guide is [Llama-SEA-LION-v3-8B-IT](https://huggingface.co/aisingapore/Llama-SEA-LION-v3-8B-IT). This section describes the steps to import the model.

Upload the contents of <https://huggingface.co/aisingapore/Llama-SEA-LION-v3-8B-IT> to an Amazon S3 bucket. Please take note of the [Amazon S3 Pricing](https://aws.amazon.com/s3/pricing/).

As the time of writing, Amazon Bedrock Custom Model Import is supported in the us-east-1 (N. Virginia) and us-west-2 (Oregon) regions. Please check the [Amazon Bedrock documentation](https://docs.aws.amazon.com/bedrock/latest/userguide/model-customization-import-model.html) for the latest information.

Select a supported region in the AWS Console. Navigate to Amazon Bedrock, **Imported Models**. Click the **Import Model** button.

Input the model name and the S3 location of the uploaded model. Update the other [settings](https://docs.aws.amazon.com/bedrock/latest/userguide/model-customization-import-model-job.html) accordingly. Click **Import Model** to start the import.

<figure><img src="/files/Czo97024v0TRR8qFBV1v" alt="" width="100%"><figcaption></figcaption></figure>

<figure><img src="/files/nI1VKtnWAaIM9TJleRpp" alt="" width="100%"><figcaption></figcaption></figure>

After the import is completed, locate the model in **Imported models**.

<figure><img src="/files/AUpA5U6ylP2idv7kJuN9" alt="" width="100%"><figcaption></figcaption></figure>

Click the copy button to copy the **ARN** (Amazon Resource Name).

<figure><img src="/files/3bnj9zOg52uuWlGCsVFX" alt="" width="100%"><figcaption></figcaption></figure>

## Demo

The demo application uses the [Amazon Bedrock Runtime](https://docs.aws.amazon.com/bedrock/latest/APIReference/API_Operations_Amazon_Bedrock_Runtime.html) to integrate with the imported model.

Clone the [repository](https://github.com/aisingapore/bedrock-access-gateway). The repository is a fork to include a functional demo and to support the imported models via the gateway.

```bash
git clone https://github.com/aisingapore/bedrock-access-gateway.git
```

Navigate to the demo directory.

```bash
cd bedrock-access-gateway/demo
```

Copy the environment file.

```bash
cp .env.example .env
```

Edit the value of `ENDPOINT_ARN` in the environment file and paste the **ARN** of the imported model.

Edit the value of `AWS_REGION` (e.g. us-east-1) in the environment file to match the region where the model was imported.

Before running the demo, it is a good practice to create a virtual environment to isolate the app. Please follow these steps to create a virtual environment, or feel free to use your preferred tool.

Initialise the virtual environment.

```bash
python -m venv venv
```

Activate the virtual environment.

```bash
source venv/bin/activate
```

Install the packages.

```bash
pip install -r requirements.txt
```

Set up the Boto3 credentials: <https://boto3.amazonaws.com/v1/documentation/api/latest/guide/credentials.html>

Run the demo.

```bash
python sealion_bedrock.py
```

<figure><img src="/files/wMP5E7H58gMaN1dREUx0" alt="" width="100%"><figcaption></figcaption></figure>

## Exception Handling

Amazon Bedrock Custom Model Import optimizes the hardware utilization by removing the models that are not active. The demo might throw an exception that indicates the model is not ready for inference. Please refer to <https://docs.aws.amazon.com/bedrock/latest/userguide/invoke-imported-model.html#handle-model-not-ready-exception> and customize your applications to handle it gracefully.

## Further Work

The demo uses the [AWS SDK](https://docs.aws.amazon.com/bedrock/latest/userguide/sdk-general-information-section.html) to integrate with the imported model. If you are looking for how to work with OpenAI-compatible APIs, given their popularity, please refer to the next [guide](/guides/inferencing/bedrock_access_gateway).

## Links

* SEA-LION models on Hugging Face: <https://huggingface.co/aisingapore>
* Import a customized model into Amazon Bedrock: <https://docs.aws.amazon.com/bedrock/latest/userguide/model-customization-import-model.html>


# Bedrock Access Gateway

In the [previous guide](/guides/inferencing/amazon_bedrock), we have shown how to leverage on Amazon Bedrock‘s Custom Model Import to deploy SEA-LION in the cloud. After importing the SEA-LION models, you can build applications with the AWS SDK. This guide describes an alternative method to build the applications with OpenAI-compatible APIs served by the **Bedrock Access Gateway**.

<figure><img src="/files/qIXR53llnPWBskD4P27k" alt="" width="100%"><figcaption></figcaption></figure>

## Prerequisites

The SEA-LION model is [imported](https://docs.aws.amazon.com/bedrock/latest/userguide/model-customization-import-model.html) and available on Amazon Bedrock. For imported models, you are charged for model inference. Please refer to the [Amazon Bedrock Pricing](https://aws.amazon.com/bedrock/pricing/) page for the latest information.

Please check that the following are installed on your development machine.

* [Git](https://git-scm.com/downloads)
* [Python](https://www.python.org/downloads/)

## Local Installation

If you have already cloned the <https://github.com/aisingapore/bedrock-access-gateway> repository from the previous [guide](/guides/inferencing/amazon_bedrock), skip to the next step. The repository is a fork to support the imported models. At the time of writing, the original repository supports [foundation models](https://docs.aws.amazon.com/bedrock/latest/userguide/models-supported.html) only.

```bash
git clone https://github.com/aisingapore/bedrock-access-gateway.git
```

Navigate to the src directory.

```bash
cd bedrock-access-gateway/src
```

Check the [access keys](https://docs.aws.amazon.com/sdkref/latest/guide/feature-static-credentials.html) and environment variables. Please ensure that `AWS_REGION` is set to the region where the model is imported.

* `AWS_ACCESS_KEY_ID`
* `AWS_SECRET_ACCESS_KEY`
* `AWS_SESSION_TOKEN`
* `AWS_REGION`

### Run with Uvicorn

Before running the gateway, it is a good practice to create a virtual environment to isolate the app. Please follow these steps to create a virtual environment, or feel free to use your preferred tool.

Initialise the virtual environment.

```bash
python -m venv venv
```

Activate the virtual environment.

```bash
source venv/bin/activate
```

Install the packages.

```bash
pip install -r requirements.txt
```

Start the gateway.

```bash
uvicorn api.app:app --host 0.0.0.0 --port 8000
```

### Run with Docker Compose

Alternatively, start the gateway with Docker Compose if it is installed.

```bash
docker compose up
```

## Test the APIs

List the models with the following curl script. The imported model is in the list with the format `arn:aws:bedrock:<AWS_REGION>:<ACCOUNT_ID>:imported-model/<MODEL_ID>`

```bash
curl http://localhost:8000/api/v1/models -H "Authorization: Bearer bedrock"
```

Test the chat completion API with the following curl script. Replace `<AWS_REGION>`, `<ACCOUNT_ID>` and `<MODEL_ID>` with the values from the imported model’s ARN (Amazon Resource Name) from the model list.

```bash
curl http://localhost:8000/api/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer bedrock" \
-d '{
  "model": "arn:aws:bedrock:<AWS_REGION>:<ACCOUNT_ID>:imported-model/<MODEL_ID>",
  "messages": [
    {
      "role": "user",
      "content": "Kalau orang sedang di GBK, biasanya sedang apa dia?"
    }
  ]
}'
```

## Demo

With the gateway, the demo from the [previous guide](/guides/inferencing/amazon_bedrock) can use the [Python library for OpenAI API](https://pypi.org/project/openai/) to integrate with the imported model. Follow the steps in the guide to set it up. Run the demo with the `--api` parameter.

```bash
python sealion_bedrock.py --api
```

<figure><img src="/files/WWqaCGGUekbzULcAkbuX" alt="" width="100%"><figcaption></figcaption></figure>

## Links

* SEA-LION Models on Hugging Face: <https://huggingface.co/aisingapore>
* Bedrock Access Gateway: <https://github.com/aisingapore/bedrock-access-gateway>
  * Production deployment: <https://github.com/aisingapore/bedrock-access-gateway?tab=readme-ov-file#deployment>


# vLLM on Linux

The following guide explains the procedures for deploying a SEA-LION model on a Linux server.

## Prerequisites

* OS: Linux
* Python: 3.9-3.12
* vLLM version: 0.10.1.1
* uv 0.7.x installed
* CUDA drivers version 12 installed

## Environment Setup

To get started, you will need to create an environment and install vLLM. The steps below outline how to install and use **uv** as the package manager, and then proceeding to install vLLM.

1. Navigate to the desired base directory. In this guide, the base folder is assumed to be `/home/sealion-user`.

```bash
cd /home/sealion-user
```

2. Create a new working directory and enter it

```bash
mkdir sealion_test && cd sealion_test
```

3. Initialize uv

```bash
uv init
```

4. Add the vllm dependency

```bash
uv add vllm
```

5. Clone vllm github repository and rename the top-level directory to `vllm_code`. The renaming is necessary to prevent Python from importing modules from the vllm repository directory instead of the packages installed in the environment.

```bash
git clone https://github.com/vllm-project/vllm.git && mv vllm vllm_code
```

## Model Deployment using vLLM

The steps below outline how to host a vLLM service using GPUs.

1. Navigate to working directory and activate environment:

```bash
cd /home/sealion-user/sealion_test && source/.venv/bin/activate
```

2. Set relevant values of environment variables. It is necessary to set VLLM\_CACHE\_ROOT to prevent errors arising from insufficient disk space as a result of using the default vLLM cache directory:

```bash
export CUDA_VISIBLE_DEVICES=0
export VLLM_CACHE_ROOT=/home/sealion-user/sealion_test/
```

3. Start the server with the desired model. The `aisingapore/Gemma-SEA-LION-v4-27B-IT` model is used in this example.

```bash
python -m vllm.entrypoints.openai.api_server --model aisingapore/Gemma-SEA-LION-v4-27B-IT
```

Alternatively, the server can be started with the below command:

```bash
vllm serve aisingapore/Gemma-SEA-LION-v4-27B-IT
```

4. Create a python script `main.py` similar to the following:

```bash
import requests

response = requests.post(
    "http://localhost:8000/v1/completions",
    headers={"Content-Type": "application/json"},
    json={
        "model": "aisingapore/Gemma-SEA-LION-v4-27B-IT",
        "prompt": "Prove or give a counter-example of the Birch and Swinnerton-Dyer conjecture.",
        "max_tokens": 3200,
        "temperature": 0.7,
    }
)
print(response.json())
```

5. Run the python script. The following assumes it is called `main.py`.

```bash
python main.py
```

You can also query the model with input prompts using curl method:

```bash
curl http://localhost:8000/v1/completions \
    -H "Content-Type: application/json" \
    -d '{
        "model": "aisingapore/Gemma-SEA-LION-v4-27B-IT",
        "prompt": "Prove or give a counter-example of the Birch and Swinnerton-Dyer conjecture.",
        "max_tokens": 3200,
        "temperature": 0.7
    }'
```

The output obtained should be similar to the following:

```bash
{'id': 'cmpl-7f41c8ff4bc9428e91505548b508747a', 'object': 'text_completion', 'created': 1755826519, 'model': 'aisingapore/Gemma-SEA-LION-v4-27B-IT', 'choices': [{'index': 0, 'text': '\n\nThe Birch and Swinnerton-Dyer (BSD) conjecture is one of the most important unsolved problems in mathematics. It relates the arithmetic of an elliptic curve to the analytic behavior of its L-function.\n\n**Statement of the Conjecture:**\n\nLet E be an elliptic curve defined over the rational numbers. Let L(E, s) be the L-function of E, which is a complex function defined for Re(s) > 1 by an Euler product.  The L-function can be analytically continued to the entire complex plane.\n\nThe conjecture states that:\n\n1. **The L-function L(E, s) has a pole at s = 1 if and only if E(Q) is finite.**  In other words, the L-function has a simple pole at s=1 if and only if the elliptic curve has finitely many rational points. If E(Q) is infinite, the L-function is analytic at s=1.\n\n2. **If the L-function has a pole at s = 1, the order of the pole is equal to the rank of the Mordell-Weil group E(Q).**  The rank of E(Q) is the dimension of the Mordell-Weil group, which measures the number of independent points of infinite order on the elliptic curve.\n\n3. **A precise formula relating the leading coefficient of the Taylor series of L(E, s) at s = 1 to several arithmetic invariants of E.** Specifically, if r is the rank of E(Q), then\n\n   lim_{s → 1} (s - 1)^r L(E, s) =  Ω_E * R_E *  ∏_{p | N} c_p *  |Sha(E)| / |E(Q)_{tor}|^2\n\n   where:\n    * Ω_E is the real period of E.\n    * R_E is the regulator of E.\n    * N is the conductor of E.\n    * c_p are the Tamagawa numbers at the primes p dividing N.\n    * Sha(E) is the Tate-Shafarevich group of E, which measures the failure of the Hasse principle.\n    * E(Q)_{tor} is the torsion subgroup of E(Q).\n\n**Status of the Conjecture:**\n\n* **Unproven:** The BSD conjecture remains unproven in general. It is considered one of the seven Millennium Prize Problems, with a $1 million reward for a correct proof.\n* **Partial Results:** Significant progress has been made:\n    * **Kolyvagin (1988):** Proved that if E has rank 0, then L(E, s) has a simple pole at s = 1.\n    * **Gross-Zagier (1986):** Proved the first half of BSD for elliptic curves with complex multiplication (CM curves). This result established the connection between the arithmetic invariants and the analytic behavior for a specific class of elliptic curves.\n    * **Rank 1 Curves:** Significant progress has been made towards proving BSD for rank 1 curves.\n    * **Taylor-Wiles-Katz:** Established modularity theorems, which are crucial for understanding the L-functions of elliptic curves.\n\n**Counter-Example?**\n\nThere are **no known counter-examples** to the BSD conjecture.  All the evidence so far supports its validity. However, the conjecture is incredibly difficult to prove, and the complexity of the terms involved makes it a formidable challenge.\n\n**Why can\'t I "prove or give a counter-example"?**\n\nThe difficulty lies in the following:\n\n* **Calculating L(E, s):** Computing the L-function to a high enough degree of accuracy to determine its behavior at s=1 is extremely challenging, even for relatively simple elliptic curves.\n* **Determining the Rank:** Finding the rank of an elliptic curve is also a difficult problem.\n* **Tate-Shafarevich Group:** The Tate-Shafarevich group Sha(E) is notoriously hard to compute. It is conjectured to be finite for all elliptic curves, but this remains unproven.\n* **Complex Analytic Continuation:** Understanding the analytic continuation of the L-function is also a complex problem.\n\n**In conclusion:**\n\nThe Birch and Swinnerton-Dyer conjecture is a profound statement about the deep connection between arithmetic and analysis.  Despite decades of research, it remains unproven.  There are no known counter-examples, and considerable evidence supports its truth.  A proof (or disproof) would be a major breakthrough in number theory.  Therefore, I cannot provide a proof or a counter-example. The best I can do is state the conjecture and summarize its current status.\n', 'logprobs': None, 'finish_reason': 'stop', 'stop_reason': 106, 'prompt_logprobs': None}], 'service_tier': None, 'system_fingerprint': None, 'usage': {'prompt_tokens': 20, 'total_tokens': 1020, 'completion_tokens': 1000, 'prompt_tokens_details': None}, 'kv_transfer_params': None}
```


# Cloudflare Workers AI

Our Gemma-SEA-LION-v4-27B-IT model is now supported on the Cloudflare Workers AI for on-demand inferencing, subject to Cloudflare's [Data Usage](https://developers.cloudflare.com/workers-ai/platform/data-usage/) and [Pricing Plans](https://developers.cloudflare.com/workers-ai/platform/pricing/).

This quickstart guide will instruct you through setting up your Workers AI account and REST API to easily experiment with Gemma-SEA-LION-v4-27B-IT. For more detailed information, you can refer directly to the latest guide on Cloudflare [here](https://developers.cloudflare.com/workers-ai/models/gemma-sea-lion-v4-27b-it/)

### Step 1. Create a Cloudflare Account

Sign up for a [Cloudflare account](https://dash.cloudflare.com/sign-up/workers-and-pages) if you have not already done so.

### Step 2. Get Cloudflare API token and Account ID

You need your API token and Account ID to use the REST API.

To get these values:

* In the Cloudflare dashboard, go to the `Workers AI` page and select `REST API`:

<figure><img src="/files/rOJe5f1Dixcc9GSYBPlo" alt="" width="100%"><figcaption></figcaption></figure>

* Get your API token:
  * Select Create a Workers AI API Token.
  * Review the prefilled information.
  * Select Create API Token.
  * Select Copy API Token.
  * Save that value for future use.
* For Get Account ID, copy the value for Account ID. Save that value for future use.

### 3. Run Gemma-SEA-LION-v4-27B-IT via Cloudflare Workers AI API

After creating your API token, authenticate and make requests to the Worker AI API using your API token in the request.

{% tabs %}
{% tab title="curl" %}

```bash
curl https://api.cloudflare.com/client/v4/accounts/{$CLOUDFLARE_ACCOUNT_ID}/ai/run/@cf/aisingapore/gemma-sea-lion-v4-27b-it \
  -X POST \
  -H 'Authorization: Bearer {$CLOUDFLARE_AUTH_TOKEN}' \
  -d '{
    "messages": [
        {
            "role": "system",
            "content": "You are a friendly assistant"
        },
        {
            "role": "user",
            "content": "Tell me a Singlish joke!"
        }
    ],
    "stream": false
    }'
```

{% endtab %}

{% tab title="python" %}

```python
import os
import requests

ACCOUNT_ID = os.environ.get("CLOUDFLARE_ACCOUNT_ID")
AUTH_TOKEN = os.environ.get("CLOUDFLARE_AUTH_TOKEN")

prompt = "Tell me a Singlish joke"
response = requests.post(
  f"https://api.cloudflare.com/client/v4/accounts/{ACCOUNT_ID}/ai/run/@cf/aisingapore/gemma-sea-lion-v4-27b-it",
    headers={"Authorization": f"Bearer {AUTH_TOKEN}"},
    json={
      "messages": [
        {"role": "system", "content": "You are a friendly assistant"},
        {"role": "user", "content": prompt}
      ]
    }
)
completion = response.json()
print(completion["result"]["choices"][0]["message"]["content"])
```

{% endtab %}
{% endtabs %}

Cloudflare also supports OpenAI compatible endpoints for text generation (/v1/chat/completions)

{% tabs %}
{% tab title="curl" %}

```bash
curl --request POST \
  --url https://api.cloudflare.com/client/v4/accounts/{$CLOUDFLARE_ACCOUNT_ID}/ai/v1/chat/completions \
  -H 'Authorization: Bearer {$CLOUDFLARE_AUTH_TOKEN}' \
  -H 'Content-Type: application/json' \
  -d '{
      "model": "@cf/aisingapore/gemma-sea-lion-v4-27b-it",
      "messages": [
        {
          "role": "user",
          "content": "Tell me a Singlish joke"
        }
      ]
    }'
```

{% endtab %}

{% tab title="python" %}

```python
from openai import OpenAI

ACCOUNT_ID = os.environ.get("CLOUDFLARE_ACCOUNT_ID")
AUTH_TOKEN = os.environ.get("CLOUDFLARE_AUTH_TOKEN")

client = OpenAI(
    api_key=AUTH_TOKEN,
    base_url=f"https://api.cloudflare.com/client/v4/accounts/{ACCOUNT_ID}/ai/v1" 
)

completion = client.chat.completions.create(
    model="@cf/aisingapore/gemma-sea-lion-v4-27b-it",
    messages=[
        {
            "role": "user",
            "content": "Tell me a Singlish joke!"
        }
    ]
)

print(completion.choices[0].message.content)
```

{% endtab %}
{% endtabs %}


# Fine-tuning


# Example Use Cases


# Capabilities & Tool-Use

## Introduction to Tool Calling

Tool calling is a powerful feature that enables Large Language Models (LLMs) to interact with external functions and APIs, extending their capabilities beyond text generation. SEA-LION models support tool calling with different implementations depending on the model version.

Tool calling allows models to:

* Access real-time information (weather, time, web search)
* Perform calculations and data processing
* Interact with external systems and APIs
* Execute specific functions based on user requests

This guide covers tool calling implementation for the SEA-LION model variants hosted on [SEA-LION API](/guides/inferencing/api), each with distinct behaviors and requirements. For demonstration purposes, the tools suggested in the [tool implementation page](/guides/tool_calling/tool_examples) will be used in the sample code snippets.

#### [Tool Implementation Example](/guides/tool_calling/tool_examples)

* [Tool Functions](/guides/tool_calling/tool_examples#tool-functions)
* [Tool Schema Definition](/guides/tool_calling/tool_examples#tool-schema-definition)
* [System Prompt](/guides/tool_calling/tool_examples#system-prompt-configuration)
* [Response Parsing](/guides/tool_calling/tool_examples#response-parsing)
* [Tool Execution](/guides/tool_calling/tool_examples#tool-execution-framework)
* [Customizing Tools](/guides/tool_calling/tool_examples#customizing-tools-for-your-application)

## Model-Specific Tool Calling Guides

### Gemma-SEA-LION-v4-27B-IT

**Key Characteristics:**

* Uses text-based tool calling format
* Requires parsing tool calls from response content
* Does not utilize standard `tool_calls` parameter
* Follows system prompt instructions for tool call formatting

Following the Gemma 3 chat template, Gemma-SEA-LION-v4-27B-IT does not parse the [`tools` parameter](https://platform.openai.com/docs/api-reference/chat/create#chat-create-tools), hence it is recommended to handle tool-calling via the parsing of the model's message response, similar to [this example](https://www.philschmid.de/gemma-function-calling) by Google DeepMind engineer Philipp Schmid.

When `tool_choice` is configured to enforce usage of a specific tool, the `tool_calls` parameter will be returned, but this removes flexibility from the LLM on determining whether tool call is required.

#### API Request Configuration

```python
# For Gemma-SEA-LION-v4-27B-IT, DO NOT include tools in request
request_data = {
    "model": "aisingapore/Gemma-SEA-LION-v4-27B-IT",
    "messages": messages,
    "temperature": 0,
    # Note: No tools or tool_choice parameters
}
```

If enforcing tool call:

```python
# For models that support native tool calling, include tools and enforce specific tool
request_data = {
    "model": "aisingapore/Gemma-SEA-LION-v4-27B-IT",
    "messages": messages,
    "temperature": 0,
    "tools": build_tool_schema(),
    "tool_choice": {
        "type": "function",
        "function": {"name": "get_current_weather"}  # Force use of specific tool
    }
}

# Alternative: Force any tool call (not a specific one)
request_data_any_tool = {
    "model": "aisingapore/Gemma-SEA-LION-v4-27B-IT", 
    "messages": messages,
    "temperature": 0,
    "tools": build_tool_schema(),
    "tool_choice": "required"  # Force model to use any available tool
}
```

#### Example Response (Tool-calling not enforced)

````json
{
  "choices": [{
    "message": {
      "content": "```tool_code\nget_time(timezone=\"Asia/Singapore\")\n```",
      "role": "assistant"
    }
  }]
}
````

### Llama-SEA-LION-v3-70B-IT

**Key Characteristics:**

* Supports standard OpenAI-style function calling
* Uses `tool_calls` parameter in responses
* Requires tools configuration in API request
* Works with `tool_choice: "auto"` setting

#### API Request Configuration

```python
# For Llama-SEA-LION-v3-70B-IT, include tools and tool_choice
request_data = {
    "model": "aisingapore/Llama-SEA-LION-v3-70B-IT",
    "messages": messages,
    "temperature": 0,
    "tools": build_tool_schema(),
    "tool_choice": "auto"
}
```

#### Response Handling

```python
def extract_tool_calls(data):
    """Extract tool calls from the response data."""
    choice = data.get("choices", [{}])[0] if data.get("choices") else {}
    return choice.get("message", {}).get("tool_calls")

# Usage
tool_calls = extract_tool_calls(response_data)
if tool_calls:
    # Execute tool calls directly
    tool_results = await execute_tool_calls(tool_calls, session)
```

#### Example Response

```json
{
  "choices": [{
    "message": {
      "content": null,
      "role": "assistant",
      "tool_calls": [{
        "function": {
          "arguments": "{\"timezone\": \"Asia/Singapore\"}",
          "name": "get_time"
        },
        "id": "chatcmpl-tool-920019c71dd14d96a262ec798b778ccd",
        "type": "function"
      }]
    }
  }]
}
```

### Llama-SEA-LION-v3.5-70B-R

**Key Characteristics:**

* Reasoning model without tool calling capability
* Tool-calling can be done via parsing from message response
* Similar to Gemma-SEA-LION-v4-27B-IT using tool-calling via message content
* Recommend not adding `tools`, `tool_choice` in API call

#### API Request Configuration

```python
# For reasoning models, do NOT include tools to avoid errors
def is_reasoning_model(model_name):
    return model_name.endswith('-R')

# Request configuration
if is_reasoning_model(api_config["model"]):
    request_data = {
        "model": "aisingapore/Llama-SEA-LION-v3.5-70B-R",
        "messages": messages,
        "temperature": 0,
        # No tools or tool_choice parameters
    }
else:
    request_data = {
        "model": api_config["model"],
        "messages": messages,
        "temperature": 0,
        "tools": tools,
        "tool_choice": "auto"
    }
```

## Implementation Example

Here's an examples that handles all three models, making use of the components provided in the [tool implementation page](/guides/tool_calling/tool_examples):

```python
async def process_user_message(user_message, messages, api_config, session):
    """Process a user message and handle tool calls for different model types."""
    messages.append({"role": "user", "content": user_message})
    tools = build_tool_schema()
    
    # Check model type
    is_reasoning_model = api_config["model"].endswith('-R')
    
    # Configure request based on model type
    request_data = {
        "model": api_config["model"],
        "messages": messages,
        "temperature": 0,
    }
    
    # Only add tools for non-reasoning models
    if not is_reasoning_model:
        request_data["tools"] = tools
        request_data["tool_choice"] = "auto"
    
    headers = {"Authorization": f"Bearer {api_config['api_key']}"}
    
    async with session.post(
        api_config["api_url"],
        json=request_data,
        headers=headers,
        timeout=30
    ) as response:
        data = await response.json()
    
    assistant_message = data.get("choices", [{}])[0].get("message")
    if not assistant_message:
        return
        
    messages.append(assistant_message)
    
    # Handle tool calls based on model type
    tool_calls = extract_tool_calls(data)
    if not tool_calls:
        # Tool call not found, parse tool call from message content
        message_content = assistant_message.get("content", "")
        if is_reasoning_model:
            # Check for tool call only in non-reasoning content to prevent excess calls
            message_content = message_content.split("</think>")[1].strip()
        tool_calls = parse_tool_calls_from_text(message_content)
    
    if tool_calls:
        # Execute tools and get final response
        tool_results = await execute_tool_calls(tool_calls, session)
        messages.extend(tool_results)
        
        # Get final response with tool results
        final_response = await session.post(
            api_config["api_url"],
            json={"model": api_config["model"], "messages": messages, "temperature": 0},
            headers=headers,
            timeout=30
        )
        final_data = await final_response.json()
        
        final_message = final_data.get("choices", [{}])[0].get("message")
        if final_message:
            print(final_message["content"])
            messages.append(final_message)
    else:
        # No tool calls, show direct response
        print(assistant_message.get("content", ""))
```

## Points to Take Note Of

### Model-Specific Considerations

1. **Gemma-SEA-LION-v4-27B-IT**:
   * Typically uses text parsing instead of standard tool calling
   * System prompt should explicitly define tool call format
   * `tools` parameter in API requests is only utilized when `tool_choice` is set to `"required"` or specific tool is enforced
   * Tool calls are wrapped in \`\`\`tool\_code blocks
   * Regex patterns needed for extraction
2. **Llama-SEA-LION-v3-70B-IT**:
   * Fully supports OpenAI-style tool calling
   * Uses `tools` and `tool_choice` in API requests
   * Returns structured `tool_calls` in response
   * Reliable for production tool calling applications
3. **Llama-SEA-LION-v3.5-70B-R**:
   * Reasoning model without tool calling capability
   * Tool-calling can be done via parsing from message response
   * Can reason about tool usage
   * Take note to parse from message content after reasoning segment, to prevent multiple redundant tool calls
   * Best used for complex reasoning tasks

### General Best Practices

1. **Error Handling**: Always implement proper error handling for tool execution failures and API timeouts.
2. **Model Detection**: Use model name suffixes to determine the appropriate tool calling approach:

   ```python
   is_reasoning_model = model_name.endswith('-R')
   ```
3. **Timeout Management**: Set appropriate timeouts for both LLM API calls and tool execution
4. **Response Validation**: Always validate tool call responses before processing:

   ```python
   if not tool_calls or not isinstance(tool_calls, list):
       # Handle no tool calls case
   ```
5. **Conversation Flow**: Maintain proper conversation history by adding all messages (user, assistant, tool results) to the messages array.
   * Gemma 3 chat template enforces the alternating of `user` and `assistant` in message history, hence in the example function `execute_tool_calls`, the role returned is `user` and not `tool` for the tool result
6. **Platform Considerations**: Some models may behave differently on different platforms (e.g., Ollama vs cloud APIs). Test your implementation on your target platform.
7. **Token Efficiency**: The text-based approach may use more tokens than standard function calling. Monitor usage accordingly.

### Security Considerations

* Validate all tool parameters before execution
* Implement rate limiting for external API calls
* Sanitize user inputs that will be passed to tools
* Consider implementing tool execution sandboxing for production environments

### Performance Optimization

* Cache tool results where appropriate (e.g., weather data for short periods)
* Implement parallel tool execution when multiple tools are called
* Use connection pooling for HTTP requests in tool implementations
* Consider implementing tool call batching for efficiency

### Relevant Links

* [Gemma 3 Function Calling Example](https://www.philschmid.de/gemma-function-calling)
* [OpenAI API Reference - `tools`](https://platform.openai.com/docs/api-reference/chat/create#chat-create-tools)
* [OpenAI Function Calling Guide](https://platform.openai.com/docs/guides/function-calling)


# Tool Implementation Example

This guide provides practical tool implementations for SEA-LION models, including working code, parsing logic, and execution frameworks. The implementations cover common patterns that you can adapt for your specific use cases.

## What's Included

This example includes the components typically needed for tool calling:

* **Tool Functions**: Working implementations for weather, time, and web search
* **Schema Definitions**: OpenAI-compatible tool schemas for model integration
* **System Prompt**: Sample prompt configured for model to execute tool-calling if it is not equipped to do so via chat template
* **Response Parsing**: Parsing logic for text-based tool calls
* **Execution Framework**: Tool execution system with error handling
* **Customization Guide**: Instructions for adapting these examples to your own tools

The three example tools (weather, time, web search) demonstrate different common patterns—external API integration, system functions, and data processing. You can use these directly or as templates when building your own tools.

### Tool Functions

#### Weather Tool Implementation

```python
async def get_weather(location, session):
    """Get weather information using Open-Meteo API."""
    if not location or not isinstance(location, str):
        raise ValueError("Invalid location")

    # Geocoding request to resolve location
    geocode_url = "https://geocoding-api.open-meteo.com/v1/search"
    geocode_params = {
        "name": location,
        "count": 1,
        "language": "en",
        "format": "json",
    }
    
    async with session.get(geocode_url, params=geocode_params, timeout=20) as response:
        geo_data = await response.json()
    
    results = geo_data.get("results", [])
    if not results:
        raise ValueError(f"No results found for location: {location}")
    
    first = results[0]
    latitude = first["latitude"]
    longitude = first["longitude"]
    resolved = {
        "name": first["name"],
        "country": first["country"],
        "admin1": first.get("admin1"),
        "latitude": latitude,
        "longitude": longitude,
        "timezone": first.get("timezone"),
    }

    # Weather forecast request
    forecast_url = "https://api.open-meteo.com/v1/forecast"
    forecast_params = {
        "latitude": latitude,
        "longitude": longitude,
        "current": "temperature_2m,apparent_temperature,relative_humidity_2m,weather_code,wind_speed_10m",
        "wind_speed_unit": "kmh",
        "timezone": "auto",
    }
    
    async with session.get(forecast_url, params=forecast_params, timeout=20) as response:
        forecast_data = await response.json()

    current = forecast_data.get("current")
    units = forecast_data.get("current_units")
    if not current:
        raise ValueError("Weather data unavailable")

    return {
        "locationQueried": location,
        "resolvedLocation": resolved,
        "current": current,
        "units": units,
        "source": "open-meteo.com",
    }
```

#### Time Tool Implementation

```python
from datetime import datetime
from zoneinfo import ZoneInfo

async def get_current_time(timezone):
    """Get current time in specified timezone."""
    if not timezone or not isinstance(timezone, str):
        raise ValueError("Invalid timezone")

    try:
        now = datetime.now()
        tz = ZoneInfo(timezone)
        time_in_timezone = now.astimezone(tz)
        
        formatted_time = time_in_timezone.strftime("%m/%d/%Y, %I:%M:%S %p %Z")
        
        return {
            "timezone": timezone,
            "currentTime": formatted_time,
            "utcTime": now.isoformat() + "Z",
            "timestamp": int(now.timestamp() * 1000)
        }
    except Exception as err:
        raise ValueError(f"Invalid timezone: {timezone}")
```

#### Web Search Tool Implementation

```python
import os

async def search_web(query, session):
    """Search web using SEARXNG (requires SEARXNG_URL environment variable)."""
    if not query or not isinstance(query, str):
        raise ValueError("Invalid search query")

    try:
        # You'll need to set up SEARXNG or use another search API
        search_url = f"{os.getenv('SEARXNG_URL')}/search"
        search_params = {
            "q": query,
            "format": "json",
            "engines": "google,duckduckgo,bing"
        }
        
        async with session.get(search_url, params=search_params, timeout=15) as response:
            search_data = await response.json()

        results = search_data.get("results", [])
        if not results:
            raise ValueError("No search results found")

        # Return top 5 results with cleaned data
        top_results = []
        for result in results[:5]:
            top_results.append({
                "title": result.get("title", ""),
                "url": result.get("url", ""),
                "content": result.get("content") or result.get("snippet", ""),
                "engine": result.get("engine", "")
            })

        return {
            "query": query,
            "resultsCount": len(results),
            "topResults": top_results,
            "source": "searxng"
        }
    except Exception as err:
        raise ValueError(f"Search failed: {str(err)}")
```

### Tool Schema Definition

```python
def build_tool_schema():
    return [
        {
            "type": "function",
            "function": {
                "name": "get_weather",
                "description": "Get the current weather information for a location.",
                "parameters": {
                    "type": "object",
                    "properties": {
                        "location": {
                            "type": "string",
                            "description": "The city or location to get weather for"
                        },
                    },
                    "required": ["location"],
                    "additionalProperties": False,
                },
            },
        },
        {
            "type": "function",
            "function": {
                "name": "get_time",
                "description": "Get the current time in a specified timezone.",
                "parameters": {
                    "type": "object",
                    "properties": {
                        "timezone": {
                            "type": "string",
                            "description": 'The timezone to get time for (e.g., "America/New_York", "Europe/London", "Asia/Singapore")'
                        },
                    },
                    "required": ["timezone"],
                    "additionalProperties": False,
                },
            },
        },
        {
            "type": "function",
            "function": {
                "name": "search_web",
                "description": "Search the web for information using SEARXNG search engine.",
                "parameters": {
                    "type": "object",
                    "properties": {
                        "query": {
                            "type": "string",
                            "description": "The search query to look for on the web"
                        },
                    },
                    "required": ["query"],
                    "additionalProperties": False,
                },
            },
        },
    ]
```

### System Prompt

````python
def build_system_message():
    return {
        "role": "system",
        "content": (
            "You are a helpful assistant with access to weather information, time checking, and web search capabilities. "
            "\n\nYou have three tools available: get_weather, get_time, and search_web. "
            "\n\nIMPORTANT: When you need to use a tool, you MUST wrap the function call in a tool_code code block using the exact format specified below:"
            "\n\nFOR WEATHER:\n```tool_code\nget_weather(location=\"City Name\")\n```"
            "\n\nFOR TIME:\n```tool_code\nget_time(timezone=\"Timezone\")\n```"
            "\n\nFOR WEB SEARCH:\n```tool_code\nsearch_web(query=\"search terms\")\n```"
            "\n\nCRITICAL REQUIREMENTS:"
            "\n- Always wrap tool calls in ```tool_code code blocks"
            "\n- Use the EXACT function call syntax shown above"
            "\n- Always use double quotes around parameter values"
            "\n- Do NOT use JSON format or any other format"
            "\n- When a user asks for weather, time, or current information, you MUST use the appropriate tool"
            "\n- Do NOT try to answer from your knowledge alone for these requests"
            "\n\nAfter receiving tool results, provide a helpful response based on the actual data returned."
        )
    }
````

### Response Parsing

````python
def parse_tool_calls_from_text(content):
    """Parse tool calls from text content using regex patterns."""
    if not content or not isinstance(content, str):
        return None
    
    # Extract content from tool_code blocks
    tool_code_pattern = r'```tool_code\s*([\s\S]*?)\s*```'
    tool_code_matches = re.findall(tool_code_pattern, content)
    
    search_content = '\n'.join(tool_code_matches) if tool_code_matches else content
    
    # Define regex patterns for different tool calls
    patterns = [
        (r'get_weather\s*\(\s*location\s*=\s*[\'"]([^\'"]+)[\'"]\s*\)', "get_weather", "location"),
        (r'get_time\s*\(\s*timezone\s*=\s*[\'"]([^\'"]+)[\'"]\s*\)', "get_time", "timezone"),
        (r'search_web\s*\(\s*query\s*=\s*[\'"]([^\'"]+)[\'"]\s*\)', "search_web", "query"),
    ]
    
    results = []
    for pattern, func_name, param_name in patterns:
        matches = re.findall(pattern, search_content)
        for i, match in enumerate(matches):
            results.append({
                "id": f"text_parsed_{func_name}_{i}",
                "function": {
                    "name": func_name,
                    "arguments": json.dumps({param_name: match})
                }
            })
    
    return results if results else None
````

### Tool Execution Framework

```python
async def execute_tool_calls(tool_calls, session):
    """Execute the tool calls and return results."""
    results = []
    
    for call in tool_calls:
        name = call.get("function", {}).get("name")
        args = call.get("function", {}).get("arguments")
        try:
            args = json.loads(args) if isinstance(args, str) else args
        except json.JSONDecodeError:
            args = {}

        try:
            if name == "get_weather":
                location = args.get("location")
                result = await get_weather(location, session)
            elif name == "get_time":
                timezone = args.get("timezone")
                result = await get_current_time(timezone)
            elif name == "search_web":
                query = args.get("query")
                result = await search_web(query, session)
            else:
                result = {"error": f"Unknown tool: {name}"}

            results.append({
                "tool_call_id": call.get("id"),
                "role": "user",
                "name": name,
                "content": json.dumps(result)
            })
        except Exception as err:
            error_result = {"name": name, "error": str(err)}
            results.append({
                "tool_call_id": call.get("id"),
                "role": "user", 
                "name": name,
                "content": json.dumps(error_result)
            })
    
    return results
```

### Customizing Tools for Your Application

The example tools above demonstrate different types of functionality:

* **get\_weather**: External API integration with error handling
* **get\_time**: System/library function usage with timezone handling
* **search\_web**: Complex data processing and result formatting

To create your own tools:

1. **Define your function**: Create an async function that takes parameters and returns structured data
2. **Add error handling**: Always validate inputs and handle exceptions gracefully
3. **Update the tool schema**: Define the function signature for the model
4. **Modify the execution framework**: Add your function to the `execute_tool_calls` function
5. **Update system prompts**: Include your tool in the system message

Example of a custom tool:

```python
# Custom database query tool
async def query_database(table, filters):
    """Query database with filters."""
    if not table or not isinstance(table, str):
        raise ValueError("Invalid table name")
    
    # Your database logic here
    # This is just an example structure
    try:
        # connection = await get_db_connection()
        # results = await connection.fetch(f"SELECT * FROM {table} WHERE {filters}")
        results = [{"id": 1, "name": "example"}]  # placeholder
        
        return {
            "table": table,
            "filters": filters,
            "results": results,
            "count": len(results)
        }
    except Exception as err:
        raise ValueError(f"Database query failed: {str(err)}")

# Add to tool schema
{
    "type": "function",
    "function": {
        "name": "query_database",
        "description": "Query a database table with optional filters.",
        "parameters": {
            "type": "object",
            "properties": {
                "table": {"type": "string", "description": "The table name to query"},
                "filters": {"type": "string", "description": "SQL WHERE clause filters"}
            },
            "required": ["table"],
            "additionalProperties": False,
        },
    },
}
```


# SEA-LION MCP Server

SEA-LION Model Context Protocol (MCP) Server provides a simple way to access SEA-LION’s multilingual models and Southeast Asian language tools from MCP-compatible clients.

The Model Context Protocol (MCP) is an open standard for connecting AI models to external tools, data sources, and services. It defines a common interface so that clients (such as IDEs, chat assistants, or AI agents) can discover and use capabilities provided by servers without custom integrations for each tool. Powered by the SEA-LION API, the MCP server exposes capabilities including translation and localisation, and safety classification.

Usage of the SEA-LION API is subject to our [Terms of Use](https://sea-lion.ai/terms-of-use/) and [Privacy Policy](https://sea-lion.ai/privacy-policy/)

## **Connect via Claude Desktop**

Claude Desktop handles the OAuth flow for you automatically. It's the most straightforward way to connect.

1. Open Claude Desktop and go to **Settings → Connectors → Customize**
2. Click Add Custom Connector and select Web
3. Enter a name (e.g. SEA-LION MCP) and the URL: <https://mcp.sea-lion.ai>
4. Click Add and Connect

<figure><img src="/files/jbqrvbZk2nq27UQcPbtl" alt="Claude Desktop Add custom connector dialog with the SEA-LION MCP name and URL filled in" width="70%"><figcaption></figcaption></figure>

5. Click on **Allow Access** and sign in with your Google account

<figure><img src="/files/e8z14DvqjVPRf3PPStSl" alt="SEA-LION Application Access Request screen with the Allow Access button" width="70%"><figcaption></figcaption></figure>

## Connect via Claude Code CLI

To add the MCP server to Claude Code CLI:

```bash
claude mcp add --transport http sealion https://mcp.sea-lion.ai/
```

You'll then be prompted to sign in: your browser opens, and you authorise access with your Google account.

Verify the connection with:

```bash
claude mcp list
```

You should see `sealion` listed as ✓ Connected.

## Using Claude

Once integrated, SEA-LION MCP works invisibly in the background. Users can simply talk to the main LLM such as Claude naturally.

**Example Prompt:**

Can you translate this article into Malay for our Singapore audience?

<figure><img src="/files/NkB1uVMDI385tg5xQW2r" alt="Translation prompt with Claude requesting permission to use the SEA-LION Translate localize tool" width="100%"><figcaption></figcaption></figure>

<figure><img src="/files/3cHEVhze3PvWc6heB3BW" alt="The article translated into Malay and localized for a Singaporean audience" width="100%"><figcaption></figcaption></figure>

Claude automatically calls SEA-LION's translation engine, localising the content for the right tone and context.

Beyond translation, SEA-LION can also detect Southeast Asian language variants and moderate content with cultural context specific to the region.

## Rate Limits

Limits help us mitigate misuse and manage API capacity and help ensure that everyone has fair access to the API.

SEA-LION MCP usage frequency will be subject to rate limits applied on requests per minute (RPM).

As of 04 Jun 2026, our rate limits is set to **10 requests per minute per user**.

If you have any questions or want to speak about getting a rate limit increase, reach out to <sealion@aisingapore.org>.


# Agentic Frameworks

SEA-LION models can be used as intelligent agents capable of tool-calling, multi-step reasoning, and complex task execution. The SEA-LION v3 family, such as `aisingapore/Llama-SEA-LION-v3-70B-IT`, can be utilized for agent-based applications with tool-calling capabilities.

Agents built with SEA-LION can:

* Execute complex, multi-step workflows
* Integrate with external tools and APIs
* Perform real-time information retrieval
* Handle specialized tasks like translation, calculation, and web search
* Chain multiple operations together intelligently

## Available Frameworks

The following guides provide technical how-tos for building SEA-LION agents using different frameworks:

### [Strands Agents SDK](/guides/agents/strands_sdk)

Learn how to build powerful agents using AWS's Strands Agents SDK, including:

* Basic agent setup using in-built tools
* Custom tool development
* Using agents as tools
* Multi-tool agent example

### [Google Agent Development Kit](/guides/agents/google_adk)

Develop and deploy agents with Google ADK:

* Basic agent setup with custom tools and agent-as-a-tool
* Custom tool development
* Using agents as tools
* Running multi-tool agent via in-built web interface/CLI

## Recommended Models

For optimal agent performance, we recommend using the following SEA-LION models:

* **`aisingapore/Llama-SEA-LION-v3-70B-IT`/`aisingapore/Llama-SEA-LION-v3-8B-IT`** - Suitable for tool-calling, long context tasks or creative response.
* **`aisingapore/Gemma-SEA-LION-v3-9B-IT`** - Suitable for smaller tasks with succinct response.
* **`aisingapore/Llama-SEA-LION-v3.5-70B-R`/`aisingapore/Llama-SEA-LION-v3.5-8B-R`** - Display reasoning process and long context tasks.

## Key Features

**Tool Integration**: SEA-LION agents can seamlessly integrate with both built-in tools (calculations, time queries) and custom tools (web search, translation, APIs).

**Multi-Agent Systems**: Agents can be composed together in multiple ways. Specialized agents serving as tools, running in sequential workflow, or passing to each other in a loop.

**Streaming Responses**: Built-in support for real-time streaming responses for better user experience.

**Flexible Configuration**: Easy environment-based configuration for different deployment scenarios.


# Strands Agents SDK

This guide demonstrates how to use SEA-LION models as intelligent agents using AWS's Strands Agent SDK. The examples show how to build both general-purpose agents with multiple tools and specialized single-purpose agents.

## Prerequisites

* Python 3.10+
* AWS Strands Agents SDK installed ([Quickstart](https://strandsagents.com/0.1.x/user-guide/quickstart/))
* [SEA-LION API](https://docs.sea-lion.ai/guides/inferencing/api) or [Bedrock](https://docs.sea-lion.ai/guides/inferencing/amazon_bedrock) access
* Environment variables configured

The sample code in this guide will be configuring the model via SEA-LION API, which follows [an OpenAI-compatible format](https://strandsagents.com/0.1.x/user-guide/concepts/model-providers/openai/). Strands Agent ADK is also compatible with other model providers. For configuring SEA-LION in Strands SDK via Bedrock, you may refer to [this page](https://strandsagents.com/0.1.x/user-guide/concepts/model-providers/amazon-bedrock/). For setting up of SEA-LION in Bedrock, refer to [this link](https://docs.sea-lion.ai/guides/inferencing/amazon_bedrock).

## Environment Setup

Create a `.env` file with your configuration:

```env
API_KEY=your-api-key-here
API_BASE_URL=https://api.sea-lion.ai/v1
MODEL=aisingapore/Llama-SEA-LION-v3-70B-IT

# Environment variable for custom tool
SEARXNG_URL=https://your-searxng-instance-url-here
```

## Basic Agent Setup

Here's how to create a basic SEA-LION agent with tool-calling capabilities:

**`agent.py`**

```python
import os
from dotenv import load_dotenv
from strands import Agent
from strands.models.openai import OpenAIModel
from strands_tools import calculator, current_time

load_dotenv()

# Configure the SEA-LION model
model = OpenAIModel(
    client_args={
        "api_key": os.environ.get("API_KEY"),
        "base_url": os.environ.get("API_BASE_URL"),
    },
    model_id=os.environ.get("MODEL", "aisingapore/Llama-SEA-LION-v3-70B-IT"),
    params={
        "max_tokens": 1000,
        "stream": True,
    },
)

# Initialize the agent
agent = Agent(
    model=model,
    system_prompt="You are a helpful assistant who answers questions in a concise and informative manner. You can use tools to assist with calculations and current time queries.",
    tools=[current_time, calculator]
)

# Use the agent
demo_prompts = [
    "What is the current time in Singapore?",
    "What is sixty five + 34?"
]

for prompt in demo_prompts:
    agent(prompt)
```

Refer to [the documentation](https://strandsagents.com/0.1.x/user-guide/concepts/tools/example-tools-package/) for more examples of in-built tools.

## Custom Tool Development

You can create custom tools for your agents.\
Here's an example of custom web search tool using [a locally-deployed SearXNG search engine](https://docs.searxng.org/admin/installation-docker.html):

**`tools.py`**

```python
import os
import requests

from dotenv import load_dotenv
from strands import tool
from typing import Optional, List, Dict, Any
from urllib.parse import urljoin

load_dotenv()

@tool
def searxng_search(
    query: str,
    categories: Optional[List[str]] = None,
    engines: Optional[List[str]] = None,
    language: Optional[str] = None,
    pageno: int = 1,
    time_range: Optional[str] = None,
    format: str = "json",
    safesearch: int = 1,
    timeout: int = 10
) -> Dict[str, Any]:
    """
    Send a search query to a SearXNG search engine.
    
    Args:
        query: The search query string
        categories: List of search categories (e.g., ["general", "images"])
        engines: List of specific engines to use (e.g., ["google", "bing"])
        language: Language code (e.g., "en", "de")
        pageno: Page number for pagination (default: 1)
        time_range: Time range filter ("day", "month", "year")
        format: Output format ("json", "csv", "rss")
        safesearch: Safe search level (0=off, 1=moderate, 2=strict)
        timeout: Request timeout in seconds
    
    Returns:
        Dictionary containing search results
    """
    
    search_url = urljoin(os.getenv("SEARXNG_URL", "http://localhost:8080"), '/search')
    
    params = {
        'q': query,
        'pageno': pageno,
        'format': format,
        'safesearch': safesearch
    }
    
    # Add optional parameters
    if categories:
        params['categories'] = ','.join(categories)
    if engines:
        params['engines'] = ','.join(engines)
    if language:
        params['language'] = language
    if time_range:
        params['time_range'] = time_range
    
    try:
        response = requests.get(search_url, params=params, timeout=timeout)
        response.raise_for_status()
        
        if format == 'json':
            return response.json()
        else:
            return {'raw_response': response.text, 'status_code': response.status_code}
            
    except requests.RequestException as e:
        raise requests.RequestException(f"SearXNG search failed: {str(e)}")
```

## Agent as a Tool

Create a specialized agent that can be used as a tool by other agents:

**`translate_agent.py`**

```python
import os
import sys

from dotenv import load_dotenv
from strands import Agent, tool
from strands.models.openai import OpenAIModel

load_dotenv()

# Use a smaller model for specialized tasks
translation_model = OpenAIModel(
    client_args={
        "api_key": os.environ.get("API_KEY"),
        "base_url": os.environ.get("API_BASE_URL"),
    },
    model_id="aisingapore/Gemma-SEA-LION-v3-9B-IT",
    params={
        "max_tokens": 1000,
        "stream": True,
    },
)

@tool
def translator(input_text: str, target_language: str) -> str:
    """
    Translates the input text to the target language.
    
    Args:
        input_text (str): The text to translate.
        target_language (str): The language to translate to (e.g. "English").
        
    Returns:
        str: The translated text.
    """
    TRANSLATOR_PROMPT = f"""
    You are a translation engine. Your sole purpose is to translate text.

    - If the input is not in {target_language}, translate it to {target_language}.
    - If language is not specified or unclear, default to English.

    You must only output the translated text. Do not include any additional explanations, notes, formatting, or any other text besides the translated content.
    """
    
    try:
        agent = Agent(
            model=translation_model,
            system_prompt=TRANSLATOR_PROMPT,
            tools=[]  # No tools needed for translation
        )
        response = agent(input_text)
        return response
    except Exception as e:
        return f"Translation failed: {str(e)}"

# Usage example
if __name__ == "__main__":
    try:
        print("🤖 Strands Translator Agent CLI")
        language = input("Enter the target language (e.g. English): ").strip()

        if not language:
            language = "English"
        print(f"Target language set to: {language}")
        print("Type your text below. Type 'quit', 'exit', or press Ctrl+C to exit.")

        # Initialize the agent
        
        while True:
            try:
                
                # Get user input
                user_input = input("\nUser: ").strip()
                
                # Check for exit commands
                if user_input.lower() in ['quit', 'exit', 'q']:
                    print("Goodbye! 👋")
                    break
                
                # Skip empty inputs
                if not user_input:
                    continue
                
                # Get agent response
                print("Translation: ", end="", flush=True)
                translator(user_input,language)
                
            except KeyboardInterrupt:
                print("\n\nGoodbye! 👋")
                break
            except EOFError:
                print("\nGoodbye! 👋")
                break

    except Exception as e:
        print(f"An error occurred initializing the agent: {e}")
        sys.exit(1)
```

## Multi-Tool Agent Example

Combine multiple tools for a comprehensive agent, making use of Strands tools, custom tools and agents as tools:

**`agent_multi.py`**

```python
import os

from dotenv import load_dotenv
from strands import Agent
from strands.models.openai import OpenAIModel
from strands_tools import calculator, current_time

# Import your custom tools from your tools.py script, or place in this script with @tools decorator
from .tools import searxng_search

# Import your specialised agent from your agent's script, or place in this script with @tools decorator
from .translate_agent import translator

load_dotenv()
# Configure model
model = OpenAIModel(
    client_args={
        "api_key": os.environ.get("API_KEY"),
        "base_url": os.environ.get("API_BASE_URL"),
    },
    model_id=os.environ.get("MODEL", "aisingapore/Llama-SEA-LION-v3-70B-IT"),
    params={
        "max_tokens": 1000,
        "stream": True,
    },
)

# Create comprehensive agent
agent = Agent(
    model=model,
    system_prompt="""You are a helpful assistant who answers questions in a concise and informative manner. 
    You can use tools to assist with calculations, current time queries, web search, and translation. 
    For current_time, you can specify a timezone in the format 'Asia/Singapore'. 
    For searxng_search, make sure to provide the links to user as well. 
    If you need more information, ask the user for clarification.""",
    tools=[current_time, calculator, searxng_search, translator]
)

# Interactive CLI
def main():
    print("🤖 SEA-LION Agent CLI")
    print("Type your questions below. Type 'quit' to exit.")
    
    while True:
        try:
            user_input = input("\nUser: ").strip()
            
            if user_input.lower() in ['quit', 'exit', 'q']:
                print("Goodbye! 👋")
                break
            
            if not user_input:
                continue
            
            print("Agent: ", end="", flush=True)
            agent(user_input)
            
        except KeyboardInterrupt:
            print("\n\nGoodbye! 👋")
            break

if __name__ == "__main__":
    main()
```

## Best Practices

### Model Selection

* Use `aisingapore/Llama-SEA-LION-v3-70B-IT` for complex reasoning and tool-calling
* Use `aisingapore/Gemma-SEA-LION-v3-9B-IT` for specialized, single-purpose agents

### Agent Configuration

* Use clear, descriptive names for your agents
* Provide specific instructions about tool usage guidelines
* Include examples when tools have specific parameter formats
* Clearly define the agent's role and capabilities

### Session Management

* Agent conversation history can be accessed via the `agent.messages` property
* To handle long-context conversations, `SlidingWindowConversationManager` can be configured
* More information and examples on conversation/state/session management available [here](https://strandsagents.com/0.1.x/user-guide/concepts/agents/sessions-state/)

### Tool Documentation

Always provide clear docstrings for custom tools, including:

* Purpose and functionality
* Parameter descriptions with types
* Return value format
* Usage examples
* Exception handling

## Troubleshooting

**Common Issues:**

1. **Tool not being called**: Ensure your system prompt mentions the tool and its use cases
2. **API connection errors**: Verify your API\_KEY and API\_BASE\_URL in the `.env` file
3. **Timeout issues**: Adjust the timeout parameters in your tools
4. **Memory usage**: Use streaming responses for better performance with large outputs

**Performance Tips:**

* Use appropriate model sizes for your use case
* Implement proper error handling and fallbacks
* Monitor token usage to optimize costs

**Strands Agents SDK Specific**

* Configuring specialised agents as tools allows for them to do focused singular tasks. For using agents in a workflow or loop, consider other multi-agents systems in Strands SDK like [Workflow](https://strandsagents.com/0.1.x/user-guide/concepts/multi-agent/workflow/) or [Graph](https://strandsagents.com/0.1.x/user-guide/concepts/multi-agent/graph/)


# Google ADK

This guide demonstrates how to use SEA-LION models as intelligent agents using Google's Agent Development Kit (ADK). The examples show how to build both general-purpose agents with multiple tools and specialized single-purpose agents.

## Prerequisites

* Python 3.9+
* Google ADK installed ([Quickstart](https://google.github.io/adk-docs/get-started/quickstart/))
* [SEA-LION API](https://docs.sea-lion.ai/guides/inferencing/api) or [Google Vertex AI](https://docs.sea-lion.ai/guides/inferencing/vertex_ai) access
* Environment variables configured

The sample code in this guide will be configuring the model via SEA-LION API through LiteLLM, which follows [an OpenAI-compatible format](https://google.github.io/adk-docs/agents/models/#using-openai-provider). Google ADK is also compatible with other model providers including Google's own Vertex AI. For configuring SEA-LION in Google ADK, you may refer to [this page](https://google.github.io/adk-docs/agents/models/#vertex-ai). [Click here for instructions on deploying SEA-LION in Vertex AI.](https://docs.sea-lion.ai/guides/inferencing/vertex_ai)

## Environment Setup

Create a `.env` file with your configuration:

```env
## For using SEA-LION API, uncomment if needed
OPENAI_API_KEY=your-sea-lion-api-key-here
OPENAI_API_BASE=https://api.sea-lion.ai/v1
MODEL=aisingapore/Llama-SEA-LION-v3-70B-IT

## For using Google Vertex AI, uncomment if needed
# GOOGLE_CLOUD_PROJECT="YOUR_PROJECT_ID"
# GOOGLE_CLOUD_LOCATION="YOUR_VERTEX_AI_LOCATION" # e.g., us-central1
GOOGLE_GENAI_USE_VERTEXAI=TRUE # Set to TRUE if NOT using Google AI Studio, due to function declaration issue: https://github.com/google/adk-python/issues/26#issuecomment-2911749485

# Environment variable for custom tool
SEARXNG_URL=https://your-searxng-instance-url-here
```

## Project Structure

```
parent_folder/
    adk_agent/
        __init__.py
        agent.py
        .env
        tools.py
        translator.py
        chat-cli.py
```

## Basic Agent Setup

Here's how to create a basic SEA-LION agent with tool-calling capabilities:

**`agent.py`**

```python
import os
import sys

from dotenv import load_dotenv
from google.adk.agents import Agent
from google.adk.models.lite_llm import LiteLlm
from google.adk.tools.agent_tool import AgentTool

# Import your custom tools from your tools.py script, or place in this script
from .tools import searxng_search

# Import your specialised agent from your agent's script, or place in this script
from .translator import root_agent as translator

load_dotenv()

try:
   root_agent = Agent(
      name="root_agent",
      model=LiteLlm(
         model=f"openai/{os.getenv('MODEL', 'aisingapore/Llama-SEA-LION-v3-70B-IT')}"
      ),
      description="Agent to answer questions using search and translation tools.",
      instruction="You are a helpful assistant who answers questions in a concise and informative manner. " \
      "You can use tools to assist with searches (`searxng_search`) and translations (`translator`). "\
      "If using `searxng_search`, make sure to provide the links you use at the end of your response, if not already provided",
      tools=[searxng_search, AgentTool(agent=translator)],
   )
except Exception as e:
   print(f"An error occurred while creating the root agent: {e}")
   sys.exit(1)
```

**`__init__.py`**

```python
from . import agent
```

## Custom Tool Development

You can create custom tools for your agents.\
Here's an example of custom web search tool using [a locally-deployed SearXNG search engine](https://docs.searxng.org/admin/installation-docker.html):

**`tools.py`**

```python
import requests
from dotenv import load_dotenv
import os
from typing import Optional, List, Dict, Any
from urllib.parse import urljoin

load_dotenv()
SEARXNG_URL = os.getenv("SEARXNG_URL", "http://localhost:8080")

def searxng_search(
    query: str,
    categories: Optional[List[str]] = None,
    engines: Optional[List[str]] = None,
    language: Optional[str] = None,
    pageno: int = 1,
    time_range: Optional[str] = None,
    format: str = "json",
    safesearch: int = 1,
    timeout: int = 10
) -> Dict[str, Any]:
    """
    Send a search query to a SearXNG search engine.
    Quick local set up for SearXNG via Docker can be found at
    https://docs.searxng.org/admin/installation-docker.html
    
    Args:
        query: The search query string
        categories: List of search categories (e.g., ["general", "images"])
        engines: List of specific engines to use (e.g., ["google", "bing"])
        language: Language code (e.g., "en", "de")
        pageno: Page number for pagination (default: 1)
        time_range: Time range filter ("day", "month", "year")
        format: Output format ("json", "csv", "rss")
        safesearch: Safe search level (0=off, 1=moderate, 2=strict)
        timeout: Request timeout in seconds
    
    Returns:
        Dictionary containing search results
    
    Raises:
        requests.RequestException: If the request fails
        ValueError: If the response is not valid JSON (when format="json")
    """
    
    # Prepare the search endpoint URL
    search_url = urljoin(SEARXNG_URL, '/search')
    
    # Build parameters
    params = {
        'q': query,
        'pageno': pageno,
        'format': format,
        'safesearch': safesearch
    }
    
    # Add optional parameters
    if categories:
        params['categories'] = ','.join(categories)
    
    if engines:
        params['engines'] = ','.join(engines)
    
    if language:
        params['language'] = language
    
    if time_range:
        params['time_range'] = time_range
    
    # Make the request
    try:
        response = requests.get(search_url, params=params, timeout=timeout)
        response.raise_for_status()
        
        if format == 'json':
            return response.json()
        else:
            return {'raw_response': response.text, 'status_code': response.status_code}
            
    except requests.RequestException as e:
        raise requests.RequestException(f"SearXNG search failed: {str(e)}")
    except ValueError as e:
        if format == 'json':
            raise ValueError(f"Invalid JSON response from SearXNG: {str(e)}")
        raise
```

## Agent as a Tool

Create a specialized agent that can be used as a tool by other agents:

**`translator.py`**

```python
import os

from dotenv import load_dotenv
from google.adk.agents import LlmAgent
from google.adk.models.lite_llm import LiteLlm

load_dotenv()

TRANSLATOR_PROMPT = """
You are a translation engine. Your sole purpose is to translate text. You do not have any tools or additional capabilities.

- Translate the input text provided to the target language instructed to you.
- If the target language is not specified or unclear, default to English.

You must only output the translated text. Do not include any additional explanations, notes, formatting, or any other text besides the translated content.
"""

root_agent = LlmAgent(
   name="translator_agent",
   model=LiteLlm(
        model=f"openai/{os.getenv('MODEL', 'aisingapore/Gemma-SEA-LION-v3-9B-IT')}"
    ),
   description="Agent that translates text to a specified language.",
   instruction=TRANSLATOR_PROMPT,
   tools=[]
)
```

## Running the Agent

### Web Interface (Recommended)

Google ADK provides a built-in web interface. Simply run:

```bash
adk web
```

This will start a web server accessible at `http://localhost:8000` using your `agent.py` configuration.

### CLI Interface (Optional)

For command-line interaction, you can use the optional CLI script:

**`chat-cli.py`**

```python
import asyncio
import datetime
import getpass
import os
import sys

from dotenv import load_dotenv
from google.adk.agents import Agent
from google.adk.models.lite_llm import LiteLlm
from google.adk.sessions import InMemorySessionService
from google.adk.tools.agent_tool import AgentTool
from google.adk.runners import Runner
from google.genai.types import Content, Part

# Import your custom tools from your tools.py script, or place in this script
from .tools import searxng_search

# Import your specialised agent from your agent's script, or place in this script
from .translator import root_agent as translator

load_dotenv()

# Set up session service and identifiers
session_service = InMemorySessionService()
APP_NAME = "google-adk-agent"
SESSION_ID = f"session-{datetime.datetime.now().strftime('%Y%m%d%H%M%S')}"
USER_ID = f"user-{getpass.getuser()}"

async def setup_session():
   try:
      session = await session_service.create_session(
         app_name=APP_NAME,
         user_id=USER_ID,
         session_id=SESSION_ID,
         state={}
      )
      print(f"Session created with ID: {SESSION_ID} for user: {USER_ID}")
      return session
   except Exception as e:
      print(f"An error occurred while creating the session: {e}")
      sys.exit(1)

try:
   root_agent = Agent(
      name="root_agent",
      model=LiteLlm(
         model=f"openai/{os.getenv('MODEL', 'aisingapore/Llama-SEA-LION-v3-70B-IT')}"
      ),
      description="Agent to answer questions using search and translation tools.",
      instruction="You are a helpful assistant who answers questions in a concise and informative manner. " \
      "If you deem necessary, use tools to assist with searches (`searxng_search`) and translations (`translator`). " \
      "If using `searxng_search`, make sure to provide the links at the end of your response, if not already provided",
      tools=[searxng_search, AgentTool(agent=translator)],
   )
except Exception as e:
   print(f"An error occurred while creating the root agent: {e}")
   sys.exit(1)

async def main():
   try:
      await setup_session()
      print("🤖 Google ADK Agent CLI")
      print("Type your questions below. Type 'quit', 'exit', or press Ctrl+C to exit.")

      # Initialize the agent runner
      runner = Runner(
         app_name=APP_NAME,
         agent=root_agent,
         session_service=session_service,
         )

      while True:
         try:
            # Get user input
            user_input = input("\nUser: ").strip()

            # Check for exit commands
            if user_input.lower() in ['quit', 'exit', 'q']:
               print("Goodbye! 👋")
               break

            # Run the agent with the user input
            user_message = Content(role='user', parts=[Part(text=user_input)])

            agent_response_text = ""
            print("Bot: ", end="", flush=True)

            async for event in runner.run_async(
                  user_id=USER_ID,
                  session_id=SESSION_ID,
                  new_message=user_message
            ):
                  if event.content and event.content.parts:
                     text_chunk = event.content.parts[0].text
                     if text_chunk:
                        print(text_chunk, end="", flush=True)
                        if event.is_final_response():
                              agent_response_text += text_chunk

                  if event.is_final_response():
                     if not agent_response_text and event.content and event.content.parts and event.content.parts[0].text:
                        agent_response_text = event.content.parts[0].text
                     print()
                     break

            if not agent_response_text:
                  if event and event.error_message:
                     print(f"\nError: {event.error_message}")
                  else:
                     print("\n(No text response from bot)")

         except KeyboardInterrupt:
            print("\nExiting... Goodbye! 👋")
            break

   except Exception as e:
      print(f"An error occurred: {e}")
   finally:
      print("\nExiting the agent CLI. Goodbye! 👋")

if __name__ == "__main__":
   asyncio.run(main())
```

To run the CLI:

```bash
python -m adk_agent.chat-cli
```

## Best Practices

### Model Selection

* Use `aisingapore/Llama-SEA-LION-v3-70B-IT` for complex reasoning and tool-calling
* Use `aisingapore/Gemma-SEA-LION-v3-9B-IT` for specialized, single-purpose agents

### Agent Configuration

* Use clear, descriptive names for your agents
* Provide specific instructions about tool usage guidelines
* Include examples when tools have specific parameter formats
* Clearly define the agent's role and capabilities

### Session Management

* Google ADK uses session-based interactions for maintaining context
* Use `InMemorySessionService` for development and testing
* Consider persistent session storage for production applications
* More information and examples on session/memory management available [here](https://google.github.io/adk-docs/sessions/)

### Tool Documentation

Always provide clear docstrings for custom tools, including:

* Purpose and functionality
* Parameter descriptions with types
* Return value format
* Usage examples
* Exception handling

## Troubleshooting

**Common Issues:**

1. **Tool not being called**: Ensure your instruction mentions the tool and its use cases
2. **LiteLLM connection errors**: Verify your API keys and base URLs in the `.env` file
3. **Function declaration issues**: Set `GOOGLE_GENAI_USE_VERTEXAI=TRUE` if not using Google AI Studio
4. **Session creation failures**: Check your session service configuration and permissions
5. **Async/await issues**: Ensure proper async handling in CLI implementations

**Performance Tips:**

* Use appropriate model sizes for your use case
* Implement proper error handling and fallbacks
* Monitor token usage to optimize costs

**Google ADK Specific:**

* Web interface (`adk web`) is the recommended way to interact with agents
* CLI implementation requires proper session management
* Agent composition using `AgentTool` allows for specialised agents to do focused singular tasks. For using agents in a workflow or loop, consider [the difference between sub-agents and agent tools](https://google.github.io/adk-docs/tools/function-tools/#key-difference-from-sub-agents). If it suits your use-case, you can [explore workflow agents and sub-agents](https://google.github.io/adk-docs/agents/workflow-agents/)


# Publications

List of SEA-LION research papers published:

* [SEA-GUARD: Culturally Grounded Multilingual Safeguard for Southeast Asia](https://arxiv.org/pdf/2602.01618)
* [SEA-LION: Southeast Asian Languages in One Network](http://arxiv.org/abs/2504.05747)
* [SEA-HELM: Southeast Asian Holistic Evaluation of Language Models](https://arxiv.org/abs/2502.14301)
* [BHASA: A Holistic Southeast Asian Linguistic and Cultural Evaluation Suite for Large Language Models](https://arxiv.org/abs/2309.06085v2)


