Researchers are figuring out how large language models work

TO MOST PEOPLE, the inner workings of a car engine or a computer are a mystery. It might as well be a black box: never mind what goes on inside, as long as it works. Besides, the people who design and build such complex systems know how they work in great detail, and can diagnose and fix them when they go wrong. But that is not the case for large language models (LLMs), such as GPT-4, Claude and Gemini, which are at the forefront of the boom in artificial intelligence (AI).

LLMs are built using a technique called deep learning, in which a network of billions of neurons, simulated in software and modelled on the structure of the human brain, is exposed to trillions of examples of something to discover inherent patterns. Trained on text strings, LLMs can hold conversations, generate text in a variety of styles, write software code, translate between languages and more besides.

  • Related Posts

    Dogs trained to detect spotted lanternfly eggs

    BLACKSBURG, Virgina – Researchers at Virginia Tech say man’s best friend may also be one of nature’s best defenses against an invasive pest. For the first time, a study shows…

    Interstellar object could be on ‘reconnaissance mission,’ expert warns

    Gabbard on aliens: “continuing to look for the truth” Director of National Intelligence Tulsi Gabbard plays coy about government knowledge of UFO, but says she will continue to “look for…

    Leave a Reply

    Your email address will not be published. Required fields are marked *