Optimize tp broadcast #2889

grimoire · 2024-12-12T11:05:04Z

requirement:

Refactor VLM modules #2810

* qwen2-vl * internvl * qwen2

…M#2772) * qwen2-vl * internvl * qwen2 * get image_tokens_per_patch for internvl2 * deepseek-vl * cogvlm * glm4v * update internvl * internvl_llava * llava * glm4v * upate internvl * cogvlm * deepseek * llava_hf * rollback llava, internvl-llava

* qwen2-vl * internvl * qwen2 * get image_tokens_per_patch for internvl2 * deepseek-vl * cogvlm * glm4v * update internvl * internvl_llava * llava * glm4v * upate internvl * cogvlm * deepseek * llava_hf * rollback llava, internvl-llava * refactor qwen * update internvl * update llava_hf * update qwen2-vl * llava_next * update llava_next * update llava * update llava * update llava

* qwen2-vl * internvl * qwen2 * get image_tokens_per_patch for internvl2 * deepseek-vl * cogvlm * glm4v * update internvl * internvl_llava * llava * glm4v * upate internvl * cogvlm * deepseek * llava_hf * rollback llava, internvl-llava * refactor qwen * update internvl * update llava_hf * update qwen2-vl * llava_next * update llava_next * update llava * update llava * update llava * qwen2

* qwen2-vl * internvl * qwen2 * get image_tokens_per_patch for internvl2 * deepseek-vl * cogvlm * glm4v * update internvl * internvl_llava * llava * glm4v * upate internvl * cogvlm * deepseek * llava_hf * rollback llava, internvl-llava * refactor qwen * update internvl * update llava_hf * update qwen2-vl * llava_next * update llava_next * update llava * update llava * update llava * qwen2 * fix internvl

* qwen2-vl * internvl * qwen2 * get image_tokens_per_patch for internvl2 * deepseek-vl * cogvlm * glm4v * update internvl * internvl_llava * llava * glm4v * upate internvl * cogvlm * deepseek * llava_hf * rollback llava, internvl-llava * refactor qwen * update internvl * update llava_hf * update qwen2-vl * llava_next * update llava_next * update llava * update llava * update llava * qwen2 * fix internvl * phi3-vision

* qwen2-vl * internvl * qwen2 * get image_tokens_per_patch for internvl2 * deepseek-vl * cogvlm * glm4v * update internvl * internvl_llava * llava * glm4v * upate internvl * cogvlm * deepseek * llava_hf * rollback llava, internvl-llava * refactor qwen * update internvl * update llava_hf * update qwen2-vl * llava_next * update llava_next * update llava * update llava * update llava * qwen2 * fix internvl * phi3-vision * refactor yi-vl * refactor mllama

* internvl2 v2 * cogvlm * deepseek-vl * glm-4v * llava-hf * llava-next * llava * internvl-llava * mllama * phi3-vision * qwen * qwen2 * yi-vl * xcomposer * minicpm * molmo * update * update

* feature: support qwen2.5 fuction_call (InternLM#2737) * feat: support qwen2.5 tools_call * fix: npe bug * fix: 模版不一致 * fix: adopting review suggestions * fix: adopting review suggestions * fix: adopting review suggestions * fix: adopting review suggestions * feat: Support multi tools calling * feat: Support multi tools calling * fix: Add '\n' between each tool * fix: Add ensure_ascii=False * bugfix: rfind * bugfix: tools_call -> tool_calls * bugfix: add toolName in tool_response * fix: some '\n' error * fix: remove toolname * fix: replace '\n' to self.separator * feat: add doc with multiple tool calling * fix：update doc * feat: add qwen2.5 prompt template test * feat: add qwen2.5 no tool call prompt test --------- Co-authored-by: gaozixiang <[email protected]> * Update supported models & Ascend doc (InternLM#2765) * update ascend supported model list * fix markdown * fix markdown * fix lint * Update get_started.md * Update get_started.md * [CI] Split vl testcases into turbomind and pytorch backend (InternLM#2751) * updaet * update * update * update * update * update * update * update * update * update * update * update * update * update * update * update * update * [Feature] support minicpm-v_2_6 for pytorch engine. (InternLM#2767) * support minicpmv_2_6. * update supported_models. * update supported_models. * Support qwen2-vl AWQ quantization (InternLM#2787) * Support qwen2-vl AWQ quantization * Update config.yaml --------- Co-authored-by: zhulinJulia24 <[email protected]> * [dlinfer] Fix qwenvl rope error for dlinfer backend (InternLM#2795) * Optimize update_step_ctx on Ascend (InternLM#2804) * opt update_ctx for ascend * fix lint --------- Co-authored-by: 逝夜长歌 <[email protected]> Co-authored-by: gaozixiang <[email protected]> Co-authored-by: jinminxi104 <[email protected]> Co-authored-by: zhulinJulia24 <[email protected]> Co-authored-by: zhoushenglong <[email protected]> Co-authored-by: AllentDan <[email protected]> Co-authored-by: Wei Tao <[email protected]>

Resolve conflicts

* WIP * support mrope * support long context * support causal=false * fix mask * flash attn bound * optimize * Moskau, Moskau, wirf die Gläser an die Wand * YMCA * optimize mllama * update processor * support cogvlm * all work and no play make jack a dull boy * upgrade triton * support qwen2vl * support internvl * phi3-v WIP * glm4v WIP * support chatglm and cogvlm * use image tokens * support llava * support internvl-mono * phi3v, mllama * add llavanext * use img token ids * support multiimage chatglm cogvlm * fix ut

…tor-vl

Fix VLM batch inference error

…tor-vl

lmdeploy/pytorch/model_inputs.py

AllentDan

Tested OK with Qwen2-VL

RunningLeon

LGTM

lvhan028 · 2024-12-17T13:41:06Z

Will it affect backends supported by dlinfer?

lvhan028 and others added 30 commits November 18, 2024 21:28

refactor VL modules for internvl and qwen2-vl (InternLM#2764)

464d451

* qwen2-vl * internvl * qwen2

Refactor VLM module for minicpm and molmo (InternLM#2794)

fff7b91

Refactor VLM modules for xcomposer series (InternLM#2796)

45cf22d

Refactor VLM modules for internvl-llava (InternLM#2797)

38eec0d

Refactor VLM modules v2 (InternLM#2806)

c4638d5

* internvl2 v2 * cogvlm * deepseek-vl * glm-4v * llava-hf * llava-next * llava * internvl-llava * mllama * phi3-vision * qwen * qwen2 * yi-vl * xcomposer * minicpm * molmo * update * update

Remove vl template (InternLM#2809)

b07211b

resolve conflicts

70875eb

Merge pull request InternLM#2812 from lvhan028/resolve-conflicts

787f765

Resolve conflicts

minor-fix

36a15e3

minor-fix (InternLM#2813)

4e6760e

fix

f7c167e

fix

e977361

fix mono

5dc967d

fix docs

ae7015a

read norm_type

c746fd3

Merge branch 'main' into refactor-vl

dc9757c

super().collect_images->self.collect_images

d577acb

Merge branch 'refactor-vl' of github.com:InternLM/lmdeploy into refac…

30ea075

…tor-vl

add note in supported models

bbcf9a5

define the parameters clearly

b3a2887

better streaming

8f7a56f

merge main

c2a5b44

grimoire and others added 22 commits December 10, 2024 11:51

fix llava

92b09d0

fix minicpm 2.6

db367f4

fix callback

4e2f1f8

fix minicpm v2.5

715fbb3

fix minicpm v2.6

1a6d88f

Merge branch 'refactor-vl' into fix-refactor-vl

8ee759f

update llava_next.py

f0a4422

remove hardcode from xcomposer2.py

ee022ad

Merge pull request InternLM#2879 from lvhan028/fix-refactor-vl

02a25eb

Fix VLM batch inference error

rollback supported_models

a21abe3

change to staticmethod

6a9342e

optimize tp

798298b

solve conflict

1b6ea24

Merge branch 'refactor-vl' into optimize-tp-broadcast

a9aacda

Merge branch 'refactor-vl' of github.com:InternLM/lmdeploy into refac…

d005bc8

…tor-vl

Merge branch 'refactor-vl' into optimize-tp-broadcast

e9517e1

fix vlm quantization

4107d4f

Merge branch 'refactor-vl' of github.com:InternLM/lmdeploy into refac…

fdaa601

…tor-vl

update doc

e5a5085

update

bc93e73

Merge branch 'refactor-vl' into optimize-tp-broadcast

67188eb

solve conflict

6cdc76a

lvhan028 requested review from AllentDan and RunningLeon December 16, 2024 04:20

AllentDan reviewed Dec 17, 2024

View reviewed changes

lmdeploy/pytorch/model_inputs.py Show resolved Hide resolved

AllentDan approved these changes Dec 17, 2024

View reviewed changes

lvhan028 added the improvement label Dec 17, 2024

RunningLeon approved these changes Dec 17, 2024

View reviewed changes

lvhan028 merged commit 8afb84c into InternLM:main Dec 17, 2024
5 checks passed

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Optimize tp broadcast #2889

Optimize tp broadcast #2889

grimoire commented Dec 12, 2024 •

edited

Loading

AllentDan left a comment

RunningLeon left a comment

lvhan028 commented Dec 17, 2024

Optimize tp broadcast #2889

Optimize tp broadcast #2889

Conversation

grimoire commented Dec 12, 2024 • edited Loading

AllentDan left a comment

Choose a reason for hiding this comment

RunningLeon left a comment

Choose a reason for hiding this comment

lvhan028 commented Dec 17, 2024

grimoire commented Dec 12, 2024 •

edited

Loading