老司机午夜视频,日韩一道本,成人网站在线视频

一 LLama.cpp

LLama.cpp 支持x86，arm，gpu的編譯。

github下載llama.cpp

https://github.com/ggerganov/llama.cpp.git

2. gem5支持arm架構(gòu)比較好，所以我們使用編譯LLama.cpp。

以下是我對(duì)Makefile的修改

開始編譯：

make UNAME_M=aarch64

編譯會(huì)使用到aarch64-linux-gnu-gcc-10，編譯成功可以生成一個(gè)main 文件，這里我把main重命名成main_arm_backup了。

可以使用file main查看一下文件：

3. 下載一個(gè)大模型的model到llama.cpp/models的目錄下，這里我下載了llama-2-7b-chat.Q2_K.gguf。

這個(gè)模型2bit量化，跑起來不到3G的內(nèi)存。

GGML_TYPE_Q2_K - "type-1" 2-bit quantization in super-blocks containing 16 blocks, each block having 16 weight. Block scales and mins are quantized with 4 bits. This ends up effectively using 2.5625 bits per weight (bpw)

4.此時(shí)我們可以本地運(yùn)行以下main和模型，我的prompt是How are you

./main -m ./models/llama-2-7b-chat.Q2_K.gguf -p "How are you"-n 16

下圖最下面一行就是模型自動(dòng)生成的

二 gem5

gem5下載編譯好后，我們可以使用gem5.fast運(yùn)行模型了。

build/ARM/gem5.fast

--outdir=./m5out/llm_9

./configs/example/se.py -c

$LLAMA_path/llama.cpp/main-arm

'--options=-m $LLAMA_path/llama-2-7b-chat.Q2_K.gguf -p Hi -n 16'

--cpu-type=ArmAtomicSimpleCPU --mem-size=8GB -n 8

此時(shí)我的prompt是Hi，預(yù)期是n=8，跑8核。

上圖是gem5運(yùn)行大模型時(shí)生成的simout，我增加了AtomicCPU 運(yùn)行指令數(shù)量的打印，這是在gem5的改動(dòng)。

如果你下載的是gem5的源碼，那么現(xiàn)在運(yùn)行起來應(yīng)該只是最前面大模型的輸出。

模型的回答是Hi，I'm a30-year-old male, and 但是我預(yù)期的是8核，實(shí)際上運(yùn)行起來：

可以看出來，實(shí)際上只跑起來4核，定位后發(fā)現(xiàn)，模型默認(rèn)是4核，需要增加-t 8選項(xiàng)，即threadnumber設(shè)置成8，下面的紅色標(biāo)注的command.

build/ARM/gem5.fast

--outdir=./m5out/llm_9

./configs/example/se.py -c

$LLAMA_path/llama.cpp/main-arm

'--options=-m$LLAMA_path/llama-2-7b-chat.Q2_K.gguf -p Hi -n 16 -t 8'

--cpu-type=ArmAtomicSimpleCPU --mem-size=8GB -n8

如上圖所示，8核都跑起來了，處理到Hi這個(gè)token的時(shí)候，CPU0執(zhí)行了2.9 Billion指令，相對(duì)于4核時(shí)的5.4 Billion約減少了一半。

審核編輯：劉清

聲明：本文內(nèi)容及配圖由入駐作者撰寫或者入駐合作網(wǎng)站授權(quán)轉(zhuǎn)載。文章觀點(diǎn)僅代表作者本人，不代表電子發(fā)燒友網(wǎng)立場(chǎng)。文章及其配圖僅供工程師學(xué)習(xí)之用，如有內(nèi)容侵權(quán)或者其他違規(guī)問題，請(qǐng)聯(lián)系本站處理。舉報(bào)投訴

ARM

ARM

+關(guān)注

關(guān)注
134

文章
9312

瀏覽量
375168
gpu

gpu

+關(guān)注

關(guān)注
28

文章
4912

瀏覽量
130681
Linux系統(tǒng)

Linux系統(tǒng)

+關(guān)注

關(guān)注
4

文章
603

瀏覽量
28321
大模型

大模型

+關(guān)注

關(guān)注
2

文章
3033

瀏覽量
3839

原文標(biāo)題：大模型筆記【3】 gem5 運(yùn)行模型框架LLama

文章出處：【微信號(hào)：處理器與AI芯片，微信公眾號(hào)：處理器與AI芯片】歡迎添加關(guān)注！文章轉(zhuǎn)載請(qǐng)注明出處。

女人自慰AV免费观看内涵网,日韩国产剧情在线观看网址,神马电影网特片网,最新一级电影欧美,在线观看亚洲欧美日韩,黄色视频在线播放免费观看,ABO涨奶期羡澄,第一导航fulione,美女主播操b

搜索歷史

大模型筆記之gem5運(yùn)行模型框架LLama介紹

評(píng)論