fix --use-cpu-initialization error when expert is not tensor-parallel #413

taozhiwei · 2024-07-03T09:42:14Z

Use --use-cpu-initialization will be failed When non-expert is tensor-parallel and expert is not tensor-parallel.
because per_partition_size is equal to master_weight.shape[partition_dim], so my_weight_list len is 0 except in 0 rank , so we can not use torch.cat, We should use assign.
Please help review @GuanhuaWang @tjruwase, thanks.

Signed-off-by: taozhiwei <[email protected]>

GuanhuaWang · 2024-07-08T23:47:13Z

Hi @taozhiwei ,

just curious, what is the case that expert not TP but other layers are TP? given experts usually have more aggregated weights compared with other parts.

taozhiwei · 2024-07-12T03:57:52Z

Hi @taozhiwei ,

just curious, what is the case that expert not TP but other layers are TP? given experts usually have more aggregated weights compared with other parts.

when not set the parameter --enable-expert-tensor-parallelism, that expert is not TP. for example, ds_pretrain_gpt_125M_MoE64.sh is not set the parameter, When adding the parameter --use-cpu-initialization directly, an error will be reported.When I need to compare whether the convergence curves are completely consistent, I will add --use-cpu-initialization @GuanhuaWang

GuanhuaWang · 2024-07-12T17:35:01Z

Hi @taozhiwei ,
just curious, what is the case that expert not TP but other layers are TP? given experts usually have more aggregated weights compared with other parts.

when not set the parameter --enable-expert-tensor-parallelism, that expert is not TP. for example, ds_pretrain_gpt_125M_MoE64.sh is not set the parameter, When adding the parameter --use-cpu-initialization directly, an error will be reported.When I need to compare whether the convergence curves are completely consistent, I will add --use-cpu-initialization @GuanhuaWang

Hi @taozhiwei , I think I should rephrase my question since I am not asking configurations: What are the application scenarios for expert not using TP but rest using TP (i.e. expert not TP but non-expert TP)? To me, there is no such application given experts usually much larger than non-expert, thus if TP is enabled, it will always apply on expert first.

To me, for TP enabled cases, there are only two:

expert TP, non-expert not TP
both expert and non-expert TP

fix --use-cpu-initialization error when expert is not tensor-parallel

d92c9de

Signed-off-by: taozhiwei <[email protected]>

taozhiwei requested review from GuanhuaWang, arashb, awan-10, duli2012, eltonzheng, tjruwase and xiaoxiawu-microsoft as code owners July 3, 2024 09:42

tjruwase removed request for arashb, awan-10, duli2012, eltonzheng and xiaoxiawu-microsoft July 8, 2024 19:19

iamdeepakgit approved these changes Jul 13, 2024

View reviewed changes

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Uh oh!

fix --use-cpu-initialization error when expert is not tensor-parallel #413

fix --use-cpu-initialization error when expert is not tensor-parallel #413

Uh oh!

taozhiwei commented Jul 3, 2024 •

edited

Loading

Uh oh!

GuanhuaWang commented Jul 8, 2024

Uh oh!

taozhiwei commented Jul 12, 2024 •

edited

Loading

Uh oh!

GuanhuaWang commented Jul 12, 2024 •

edited

Loading

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants

fix --use-cpu-initialization error when expert is not tensor-parallel #413

Are you sure you want to change the base?

fix --use-cpu-initialization error when expert is not tensor-parallel #413

Uh oh!

Conversation

taozhiwei commented Jul 3, 2024 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

GuanhuaWang commented Jul 8, 2024

Uh oh!

taozhiwei commented Jul 12, 2024 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

GuanhuaWang commented Jul 12, 2024 • edited Loading Uh oh! There was an error while loading. Please reload this page.

Uh oh!

Uh oh!

Reviewers

Assignees

Labels

Projects

Milestone

Development

Uh oh!

3 participants

taozhiwei commented Jul 3, 2024 •

edited

Loading

taozhiwei commented Jul 12, 2024 •

edited

Loading

GuanhuaWang commented Jul 12, 2024 •

edited

Loading