SMALL LANGUAGE MODELS

The rapid development of Artificial Intelligence has been characterised by the emergence of increasingly capable language models capable of understanding, generating and reasoning across human language with unprecedented sophistication. While much public attention has focused upon very large foundation models containing hundreds of billions of parameters, a parallel and increasingly significant movement has emerged towards the development of Small Language Models. These models seek to provide many of the practical capabilities associated with large-scale Artificial Intelligence whilst substantially reducing computational complexity, infrastructure requirements, operational cost and energy consumption. Rather than viewing scale as the sole determinant of capability, Small Language Models demonstrate that carefully designed architectures, high-quality training data and task-specific optimisation may produce systems capable of delivering highly effective performance within clearly defined operational domains.

The growing importance of Small Language Models reflects a broader transition within Artificial Intelligence from experimental capability towards practical deployment. Organisations increasingly require systems that are explainable, economically sustainable, secure, responsive and capable of operating within constrained computing environments. Many enterprise applications do not require the extensive general knowledge possessed by very large models. Instead, they require reliable expertise within specific organisational contexts, rapid response times, enhanced privacy and deployment on local infrastructure or edge computing platforms. Small Language Models satisfy these requirements by sacrificing unnecessary breadth in favour of efficiency, precision and operational control.

Consequently, Small Language Models are becoming an increasingly important component of contemporary Artificial Intelligence strategy. Rather than competing directly with very large foundation models, they complement them by providing organisations with specialised, adaptable and resource-efficient capabilities suitable for deployment across industry, government, healthcare, finance, manufacturing, defence and scientific research. Their future significance lies not simply in their reduced size but in their ability to democratise Artificial Intelligence by making advanced language technologies accessible beyond the largest technology organisations.

Efficient Artificial Intelligence Beyond Unrestricted Scale

Artificial Intelligence has undergone extraordinary transformation during the past decade. Progress in deep learning, neural network architecture, computational hardware and the availability of large-scale datasets has enabled language models to achieve levels of performance previously regarded as unattainable. Modern systems demonstrate impressive capabilities in language understanding, question answering, summarisation, software development, translation, reasoning and content generation, fundamentally changing expectations regarding the practical capabilities of Artificial Intelligence.

Much of this progress has been driven by continual increases in model size. Researchers frequently associated improved performance with increasing numbers of parameters, larger training datasets and greater computational expenditure. This scaling paradigm produced foundation models capable of performing an exceptionally broad range of tasks with comparatively little task-specific training. The remarkable achievements of these systems established a new generation of Artificial Intelligence capable of general linguistic competence across numerous domains.

Nevertheless, increasing scale has introduced significant practical limitations. Extremely large models require substantial computational infrastructure, considerable electrical power, specialist hardware and extensive financial investment. Their deployment frequently depends upon cloud computing services operated by comparatively small numbers of organisations possessing the necessary resources. Response latency, environmental impact, operational cost and data governance have therefore become increasingly significant considerations alongside model capability.

These practical constraints have stimulated renewed interest in computational efficiency. Researchers have recognised that many real-world applications require neither unrestricted generality nor encyclopaedic knowledge. Instead, organisations frequently require reliable performance within well-defined operational environments supported by comparatively limited but highly relevant knowledge. This observation has encouraged the development of Small Language Models capable of delivering high-quality performance whilst operating within substantially reduced computational budgets.

The emergence of Small Language Models therefore represents more than a technical refinement. It reflects an important philosophical shift within Artificial Intelligence. Rather than assuming that larger models necessarily provide superior practical solutions, researchers increasingly recognise that effectiveness depends upon the relationship between capability, efficiency, context and organisational requirements. Artificial Intelligence should therefore be evaluated according to its fitness for purpose rather than according to computational scale alone.

Defining Purpose-Built and Resource-Efficient Language Models

Small Language Models are Artificial Intelligence models specifically designed to provide advanced language understanding, reasoning and generation using substantially fewer computational resources than large-scale foundation models. Their defining characteristic is not merely a reduced number of parameters but an architectural emphasis upon efficiency, optimisation and purposeful deployment. Rather than attempting to represent every conceivable domain of human knowledge, Small Language Models are typically trained or adapted to perform exceptionally well within clearly defined areas of application whilst maintaining comparatively modest computational requirements.

The significance of Small Language Models lies in the recognition that language capability exists upon a continuum rather than as a simple function of size. Large foundation models derive considerable flexibility from extensive parameter counts and broad training corpora. Small Language Models achieve effectiveness through selective optimisation, efficient representation, carefully curated data and domain-specific adaptation. Their objective is therefore not to replicate every capability of larger systems but to maximise useful performance relative to available computational resources.

This distinction fundamentally alters the relationship between Artificial Intelligence and organisational deployment. Instead of requiring substantial remote computing infrastructure, Small Language Models may often operate upon local servers, embedded devices, industrial equipment, mobile platforms or edge computing environments. Their reduced memory requirements permit lower latency, improved privacy, decreased operational expenditure and greater organisational control over sensitive information. Consequently, the concept of intelligence embodied by Small Language Models is one of practical sufficiency rather than unrestricted generality. They demonstrate that effective Artificial Intelligence need not always be enormous, provided that its design remains closely aligned with the environment in which it operates.

From Early Statistical Models to Efficient Transformers

The intellectual origins of Small Language Models may be traced to the earliest stages of natural language processing, when computational resources imposed severe limitations upon model complexity. Statistical language models, hidden Markov models and recurrent neural networks were necessarily modest in scale, yet they established many of the theoretical principles that continue to influence contemporary language modelling. As computational capability increased, progressively larger neural architectures became feasible, culminating in the introduction of transformer architectures that fundamentally altered the trajectory of Artificial Intelligence research.

The publication of the transformer architecture in 2017 represented a decisive turning point. Attention mechanisms enabled substantially improved handling of long-range linguistic dependencies whilst facilitating efficient parallel computation. Subsequent years witnessed an extraordinary expansion in model scale as researchers demonstrated that increasing parameters, training data and computational expenditure frequently produced corresponding improvements in language capability. This period became characterised by an implicit assumption that larger models represented the inevitable future of Artificial Intelligence.

As deployment expanded beyond research laboratories, however, practical realities became increasingly apparent. Organisations encountered substantial financial costs associated with training and operating very large models. Infrastructure demands limited accessibility, while latency and energy consumption constrained deployment within mobile, embedded and industrial environments. Questions concerning privacy, sovereignty, governance and environmental sustainability further encouraged investigation into more efficient approaches.

These concerns stimulated renewed research into model compression, knowledge distillation, parameter sharing, quantisation, sparse architectures and efficient transformer variants. Rather than viewing efficiency as a secondary objective, researchers increasingly recognised it as a central design principle. Small Language Models emerged from this broader movement, demonstrating that architectural innovation could recover much of the capability associated with larger systems whilst substantially reducing computational requirements.

Recent developments have accelerated this trend. Advances in synthetic data generation, instruction tuning, reinforcement learning, retrieval augmentation and domain-specific fine-tuning have enabled comparatively compact models to perform increasingly sophisticated reasoning tasks. Improvements in hardware acceleration and optimisation algorithms have further strengthened their practical viability. Consequently, Small Language Models have evolved from computational compromises into strategically important technologies in their own right.

Parameter Efficiency, Compression and Curated Training Data

Although Small Language Models frequently employ transformer architectures similar to those used within larger systems, their design philosophy differs in several important respects. Every architectural decision seeks to maximise useful capability whilst minimising unnecessary computational expenditure. Efficiency therefore becomes an organising principle governing model structure, parameter allocation, memory management and inference behaviour.

Parameter efficiency represents one of the most important design considerations. Rather than increasing parameter counts indiscriminately, Small Language Models seek to maximise the amount of useful information represented by each parameter. Improved training procedures, higher-quality datasets and more effective optimisation strategies frequently permit substantial reductions in model size without proportionate reductions in performance.

Attention mechanisms are similarly refined to reduce computational complexity. Various efficient attention methods reduce memory consumption and processing time while preserving the contextual understanding necessary for effective language reasoning. Alternative architectural innovations may further reduce computational requirements through sparse computation, shared parameters or hierarchical representations.

Compression techniques contribute significantly to the practical success of Small Language Models. Knowledge distillation enables compact models to learn behavioural characteristics from substantially larger teacher models. Quantisation reduces numerical precision whilst maintaining acceptable performance. Pruning removes redundant parameters that contribute comparatively little to model accuracy. Collectively, these techniques illustrate that much of the computational complexity present within very large models may be unnecessary for many practical applications.

Training methodology has likewise evolved. Rather than relying exclusively upon enormous collections of heterogeneous internet data, many Small Language Models benefit from carefully curated datasets specifically aligned with intended operational domains. High-quality specialised information frequently contributes more effectively to practical capability than substantially larger quantities of poorly filtered material. Consequently, dataset quality increasingly complements architectural efficiency as a determinant of model performance.

Tokenisation, Embeddings, Attention and Language Generation

Every Small Language Model combines several fundamental components that collectively determine its behaviour. Tokenisation transforms natural language into computational representations suitable for numerical processing. Embedding layers convert these representations into high-dimensional mathematical spaces that preserve semantic relationships. Transformer layers process contextual information through attention mechanisms, enabling each token to influence the interpretation of neighbouring tokens according to linguistic context.

Feed-forward neural networks perform nonlinear transformations that enable increasingly abstract representations of language to emerge throughout successive computational layers. Positional encoding preserves information regarding word order, ensuring that grammatical structure remains meaningful during processing. Output layers convert internal representations into probability distributions from which responses may be generated according to the objectives established during training.

Despite their reduced scale, Small Language Models retain these essential architectural principles. Their distinction lies not in the absence of capability but in more efficient implementation. Each component is optimised to maximise useful performance while limiting unnecessary computational complexity. This emphasis upon efficiency allows relatively compact models to maintain impressive linguistic competence across many practical applications.

Fitness for Purpose, Explainability and Operational Control

The philosophy underlying Small Language Models differs fundamentally from that associated with unrestricted scaling. Their objective is not to become universal repositories of human knowledge but to provide dependable intelligence within clearly defined operational boundaries. Capability is therefore measured by relevance rather than comprehensiveness.

This philosophy aligns closely with organisational requirements. Enterprises rarely require unlimited conversational breadth when performing specialised tasks such as legal document analysis, financial reporting, engineering documentation or clinical record management. Instead, they require consistency, accuracy, security and efficient deployment. Small Language Models therefore embody a more disciplined conception of Artificial Intelligence in which computational capability is deliberately aligned with operational necessity.

This approach also improves explainability. Smaller architectures frequently permit more systematic evaluation, testing and refinement than extremely large models. Their behaviour may be more predictable within specialised domains because training objectives remain comparatively focused. Although explainability remains an active area of research across all neural architectures, the comparatively constrained nature of Small Language Models often simplifies governance, validation and regulatory compliance.

Their design philosophy consequently reflects maturity within Artificial Intelligence engineering. Progress is no longer measured exclusively through scale but through the intelligent balance between capability, efficiency, transparency and practical value.

Curated Training, Instruction Tuning and Domain Adaptation

The effectiveness of Small Language Models depends not only upon their architecture but also upon the methods through which they acquire, refine and apply knowledge. Training generally proceeds through several distinct stages, each contributing different forms of capability. Initial pre-training exposes the model to extensive collections of textual information, enabling it to learn the statistical structure of language, semantic relationships, grammatical organisation and broad patterns of human communication. Rather than memorising individual documents, the model develops distributed internal representations that allow it to predict, interpret and generate language across many different contexts.

Unlike the largest foundation models, however, Small Language Models frequently derive greater benefit from the careful selection and curation of training data than from simply increasing its volume. Because parameter capacity is more limited, every example contributes proportionately more to the knowledge acquired during training. Considerable attention is therefore given to eliminating duplicated material, improving linguistic quality, reducing noise and ensuring that the resulting corpus accurately reflects the intended operational domain. High-quality domain-specific information often produces more reliable practical performance than substantially larger collections of heterogeneous internet content.

Following pre-training, models are commonly refined through supervised instruction tuning, during which they learn to respond appropriately to carefully constructed examples of human interaction. This stage improves conversational behaviour, factual presentation, reasoning style and adherence to user instructions. Additional refinement may employ preference optimisation or reinforcement learning based upon human evaluation, enabling the model to produce responses that better reflect human expectations regarding usefulness, clarity, accuracy and safety.

Domain adaptation represents one of the greatest strengths of Small Language Models. Organisations may fine-tune a general model using proprietary documents, technical manuals, legal guidance, engineering specifications, scientific literature or internal knowledge repositories without requiring the computational resources necessary to retrain an entire foundation model. This enables enterprises to create highly specialised Artificial Intelligence systems possessing expertise aligned closely with their own operational environment. The resulting model becomes an organisational asset whose knowledge reflects the language, terminology, procedures and objectives of the institution in which it operates.

Retrieval-Augmented Knowledge and Authoritative Information

An increasingly important development concerns retrieval-augmented generation, through which Small Language Models supplement their internal knowledge by accessing trusted external information during inference. Rather than attempting to memorise every possible fact during training, the model retrieves relevant documents from authorised repositories before generating its response. This approach reduces hallucination, improves factual accuracy, supports continual updating and enables organisations to maintain authoritative knowledge without repeatedly retraining the underlying model. The distinction between learned knowledge and retrieved knowledge therefore becomes an important architectural principle within modern enterprise Artificial Intelligence.

Efficiency, Low Latency, Privacy and Deployment Flexibility

Several characteristics distinguish Small Language Models from larger foundation models whilst simultaneously explaining their growing strategic importance. The most immediately apparent is computational efficiency. Smaller architectures require substantially less memory, processing power and electrical energy, permitting deployment upon local servers, desktop computers, mobile devices, industrial controllers and edge computing platforms where large-scale models would be impractical.

Reduced inference latency constitutes another significant advantage. Many industrial, medical, financial and engineering applications require responses within fractions of a second. Delays introduced by remote cloud infrastructure or exceptionally large computational workloads may be unacceptable where timely decision-making is essential. Small Language Models frequently provide considerably faster response times while maintaining sufficient linguistic capability for specialised operational tasks.

Privacy and data sovereignty have become increasingly important characteristics. Organisations operating within regulated industries frequently possess information that cannot be transmitted to external cloud services without creating legal, commercial or ethical concerns. Because Small Language Models may operate entirely within private infrastructure, they enable sensitive information to remain under organisational control. This capability is particularly valuable within healthcare, defence, financial services, legal practice, scientific research and public administration, where confidentiality represents a fundamental operational requirement.

Energy efficiency provides a further strategic advantage. The environmental consequences of training and operating extremely large Artificial Intelligence models have become an increasingly important area of public discussion. Although all computational systems consume energy, smaller models generally require substantially fewer computational resources throughout both training and operational deployment. As organisations pursue sustainability objectives alongside technological innovation, energy-efficient Artificial Intelligence is likely to become increasingly attractive.

Finally, Small Language Models exhibit considerable flexibility. Their comparatively modest computational requirements allow organisations to develop multiple specialised models rather than relying upon a single universal system. Individual models may therefore be optimised for finance, engineering, healthcare, legal analysis, scientific research or customer service whilst remaining computationally practical to maintain and update independently.

Operational Advantages and Capability Trade-Offs

The advantages of Small Language Models arise principally from their efficiency, adaptability and practicality. Their lower infrastructure requirements reduce operational expenditure, allowing smaller organisations to adopt advanced Artificial Intelligence without substantial investment in specialist hardware or cloud services. Their ability to operate within constrained environments extends Artificial Intelligence into applications previously inaccessible to very large models, including embedded systems, autonomous equipment, mobile devices and industrial automation.

Domain-specific optimisation frequently produces superior performance within narrowly defined tasks. A carefully trained Small Language Model specialising in engineering documentation or pharmaceutical regulation may outperform substantially larger general-purpose models when evaluated exclusively within those domains because its internal representations have been optimised for precisely those forms of knowledge. This illustrates an important principle of Artificial Intelligence engineering: capability should always be evaluated relative to intended purpose rather than through general benchmarks alone.

Operational governance also benefits from reduced complexity. Smaller models are generally easier to validate, monitor and update. Organisations can retrain or fine-tune them more frequently, respond more rapidly to changing regulations and maintain greater oversight of the knowledge incorporated into operational systems. Their comparatively limited scale also facilitates deployment within highly regulated environments requiring detailed evidence concerning model behaviour.

Capacity, Generalisation and Engineering Limitations

These advantages should not obscure important limitations. Small Language Models necessarily possess less representational capacity than the largest foundation models. Their breadth of general knowledge may therefore be more restricted and they may perform less effectively when confronted with unfamiliar subjects requiring extensive multidisciplinary reasoning. Tasks demanding exceptionally broad contextual understanding, multilingual competence across many languages or sophisticated creative synthesis may exceed the capabilities of comparatively compact architectures.

Their performance also depends more heavily upon careful training and high-quality domain adaptation. Poorly selected datasets or inadequate fine-tuning may reduce effectiveness substantially because limited parameter capacity leaves less opportunity to compensate for deficiencies within the training process. Successful deployment therefore depends upon disciplined engineering rather than simply selecting a smaller architecture.

These limitations should be interpreted within the broader context of organisational requirements. Artificial Intelligence should not be judged according to theoretical maximum capability alone but according to its ability to perform required tasks efficiently, reliably and responsibly. In many practical settings, Small Language Models represent the more appropriate technological choice precisely because their capabilities align closely with operational objectives.

Distillation, Sparse Computation, Multimodality and Reasoning

Contemporary research into Small Language Models reflects a broader movement towards efficient Artificial Intelligence. Rather than pursuing unrestricted growth in parameter count, researchers increasingly investigate methods for improving capability through architectural innovation, optimisation and more effective learning strategies. Considerable attention is devoted to improving parameter efficiency, enabling compact models to represent increasingly complex linguistic relationships without proportional increases in computational cost.

Knowledge distillation remains an active area of investigation. Large teacher models transfer behavioural characteristics to substantially smaller student models, allowing much of the reasoning capability of the larger system to be preserved within a more efficient architecture. Researchers continue to refine these techniques, seeking improved preservation of reasoning, factual accuracy and contextual understanding.

Sparse computation represents another important direction. Instead of activating every parameter during every inference, models increasingly employ selective computation, allowing only the most relevant components to participate in generating a response. Closely related developments include Mixture of Experts architectures, adaptive computation and conditional parameter activation, all of which seek to improve efficiency without sacrificing capability.

Research into multimodal Small Language Models is expanding rapidly. Compact architectures are increasingly being developed that integrate textual, visual, audio and structured information whilst remaining suitable for deployment upon local hardware. Such systems may eventually support sophisticated document understanding, industrial inspection, autonomous robotics and healthcare diagnostics without requiring remote computational infrastructure.

Reasoning represents another major area of investigation. Early generations of Small Language Models were frequently criticised for comparatively limited logical reasoning when compared with very large foundation models. Recent work increasingly demonstrates that improved training methodologies, synthetic reasoning datasets and structured problem-solving techniques may substantially enhance reasoning capability without dramatic increases in model size.

Applications Across Industry, Healthcare, Finance and Government

The practical significance of Small Language Models becomes most apparent when examining their deployment across contemporary organisations. Many enterprise applications require dependable linguistic competence rather than unrestricted conversational ability. Technical documentation, regulatory compliance, engineering knowledge management, financial reporting and customer support each involve specialised vocabularies and comparatively stable knowledge domains well suited to compact language models.

Within manufacturing, Small Language Models support maintenance procedures, equipment diagnostics, quality assurance and technical documentation. Their ability to operate locally enables intelligent assistance even where continuous internet connectivity cannot be guaranteed. Healthcare organisations employ similar approaches for clinical documentation, administrative support and secure knowledge retrieval whilst maintaining appropriate control over sensitive patient information.

Financial institutions benefit from efficient language models capable of analysing contracts, summarising reports, monitoring regulatory developments and supporting internal decision-making. Legal organisations similarly employ domain-adapted models to assist document review, legislative analysis and case preparation. Scientific institutions increasingly use specialised language models trained upon research literature, enabling more efficient exploration of complex technical information.

Government agencies represent another important area of adoption. Public sector organisations frequently require Artificial Intelligence systems capable of operating within secure environments under strict governance arrangements. Small Language Models provide opportunities for local deployment whilst supporting document management, citizen services, administrative efficiency and knowledge preservation.

Education likewise presents considerable opportunity. Institutions may develop locally hosted educational assistants trained upon approved curricular material, enabling students to receive guidance consistent with institutional standards whilst maintaining appropriate privacy and pedagogical oversight. The flexibility of Small Language Models therefore extends beyond technical efficiency into broader questions of institutional autonomy and responsible knowledge management.

Private Deployment, Governance and Risk-Based Assurance

Security and governance occupy an increasingly central position within the deployment of Artificial Intelligence. Small Language Models offer several important advantages in this regard because their reduced computational requirements permit operation within private organisational infrastructure. Sensitive information need not leave institutional boundaries, reducing exposure to external threats and simplifying compliance with regulatory requirements concerning personal information, commercial confidentiality and national security.

Governance extends beyond data protection. Organisations must understand the provenance of training information, the limitations of model capability and the circumstances under which human oversight remains essential. Small Language Models generally permit more frequent retraining and validation than extremely large architectures, enabling organisations to maintain closer alignment between operational knowledge and changing regulatory or scientific developments.

Explainability also assumes particular importance. Although all neural networks remain mathematically complex, domain-specific Small Language Models often produce behaviour that is easier to evaluate because their operational scope is comparatively constrained. Performance may be assessed against clearly defined organisational objectives rather than broad and ambiguous measures of general intelligence. This facilitates auditing, monitoring and continual improvement throughout the operational life of the system.

Future governance frameworks are likely to distinguish increasingly between general-purpose foundation models and specialised enterprise models. Regulatory expectations concerning transparency, accountability, testing and operational assurance may therefore become progressively aligned with the specific risks associated with individual deployments rather than with model size alone.

Edge Intelligence, Hybrid Architectures and Emerging Trajectories

The future evolution of Small Language Models is unlikely to involve simple increases in parameter count. Instead, progress will probably arise through more efficient architectures, improved reasoning, better integration with external knowledge and closer alignment between Artificial Intelligence capability and organisational purpose. Future models are expected to become increasingly specialised whilst simultaneously remaining capable of collaboration with larger foundation models through hybrid computational architectures.

Edge computing will represent an important area of expansion. As computational hardware becomes progressively more capable, intelligent language systems will increasingly operate directly upon industrial equipment, medical devices, autonomous vehicles, scientific instruments and personal computing platforms. Artificial Intelligence will therefore become progressively decentralised, moving away from exclusive dependence upon large cloud infrastructures towards distributed intelligent systems operating closer to their users.

Specialised Agents and Coordinated Enterprise Workflows

Another important trajectory concerns collaboration between Small Language Models and intelligent agents. Rather than functioning solely as conversational systems, compact models will increasingly support autonomous workflows capable of retrieving information, coordinating software, analysing documents and executing complex organisational procedures. Their efficiency makes them particularly suitable for continuous operation within enterprise environments where numerous specialised agents may cooperate simultaneously.

Continued advances in synthetic data generation, automated evaluation, self-improving training methodologies and efficient reasoning algorithms are likely to narrow further the performance difference between compact models and substantially larger systems. The distinction between small and large may therefore become progressively less important than the suitability of each architecture for its intended operational environment.

Small Language Models as Foundations for Intelligent Enterprise Systems

Small Language Models represent one of the most significant developments within contemporary Artificial Intelligence because they challenge the assumption that increasing computational scale inevitably represents the most desirable path towards practical intelligence. Their emergence reflects the growing maturity of Artificial Intelligence engineering, in which efficiency, governance, privacy, sustainability and organisational relevance increasingly complement raw computational capability as measures of technological success.

Rather than competing directly with the largest foundation models, Small Language Models occupy a distinct and increasingly valuable position within the Artificial Intelligence ecosystem. They provide organisations with efficient, adaptable and secure language technologies capable of operating within constrained environments whilst maintaining high standards of performance within specialised domains. Their comparatively modest computational requirements democratise access to advanced Artificial Intelligence by enabling deployment across institutions that lack the extensive infrastructure necessary to support very large models.

The future of Artificial Intelligence is unlikely to be characterised by a single universal architecture. Instead, it will consist of complementary systems operating at different scales according to their intended purpose. Foundation models will continue to provide broad general capabilities, whilst Small Language Models will deliver efficient domain-specific expertise wherever responsiveness, privacy, operational control and sustainability are paramount.

In this evolving landscape, the importance of Small Language Models will continue to grow. They represent not merely smaller versions of larger systems but a distinct philosophy of Artificial Intelligence engineering in which capability is measured by appropriateness, efficiency and practical value. As organisations increasingly seek trustworthy, economically sustainable and strategically aligned Artificial Intelligence, Small Language Models are likely to become one of the principal technological foundations upon which the next generation of intelligent enterprise systems will be built.

Bibliography

  • Bommasani, R. and others, On the Opportunities and Risks of Foundation Models (Stanford: Stanford University, 2021).
  • Brown, T. B. and others, ‘Language Models are Few-Shot Learners’, Advances in Neural Information Processing Systems, 33 (2020), 1877-1901.
  • Dettmers, T. and others, ‘QLoRA: Efficient Finetuning of Quantised Large Language Models’, Advances in Neural Information Processing Systems (2023).
  • Devlin, J., Ming-Wei Chang, Kenton Lee and Kristina Toutanova, ‘BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding’, Proceedings of NAACL-HLT (2019).
  • Hinton, G., Oriol Vinyals and Jeffrey Dean, ‘Distilling the Knowledge in a Neural Network’, Neural Information Processing Systems Deep Learning Workshop (2015).
  • Hoffmann, J. and others, ‘Training Compute-Optimal Large Language Models’, Proceedings of Advances in Neural Information Processing Systems, 35 (2022).
  • Kaplan, J. and others, ‘Scaling Laws for Neural Language Models’, arXiv (2020).
  • Liu, P. and others, ‘Pre-train, Prompt and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing’, ACM Computing Surveys, 55.9 (2023).
  • Raffel, C. and others, ‘Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer’, Journal of Machine Learning Research, 21 (2020), 1-67.
  • Touvron, H. and others, ‘Llama 2: Open Foundation and Fine-Tuned Chat Models’, arXiv (2023).
  • Vaswani, A. and others, ‘Attention is All You Need’, Advances in Neural Information Processing Systems, 30 (2017).
  • Wolf, T. and others, ‘Transformers: State-of-the-Art Natural Language Processing’, Proceedings of EMNLP(2020).
  • Xu, C. and others, ‘A Survey of Efficient Large Language Models’, ACM Computing Surveys (2024).
  • Zhang, S. and others, ‘Small Language Models: Survey, Measurements and Future Directions’, arXiv (2024).

X is a registered trade mark of GENERAL INTELLIGENCE PLC.
It was registered in 1896 with company number: SC003234