Large and small model combined cross-language news dialogue abstract generation method
Through the method of combining size and model, the big model is used to classify news dialogue types and extract keywords, and the initial summary is generated by combining small models. The content is optimized through the big model. Finally, the small model is judged for evaluation and optimization, which solves the dialogue understanding, information screening and cross-language migration problems in the generation of cross-language news dialogue summary, achieving high-quality and efficient summary generation.
Patent Information
- Application Number
- CN202510309972.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-06
AI Technical Summary
The prior art has limitations in dialogue understanding, information screening and cross-language migration in the generation of cross-language news dialogue summary, resulting in unstable summary quality and information omission.
Using a combination of size and model, we first classify news conversations and extract keywords based on the big model, and then train the multilingual dialogue summary small model to generate initialization summary, and optimize the summary content through the big model collaborative technology, and finally use the judgment small model for formal evaluation and optimization.
It realizes more precise and smooth cross-language news conversation summary generation, improves the quality and consistency of summary, and solves the problems of information omissions and poor language migration.
Smart Images

Figure CN120104790A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of natural language processing, specifically to the field of cross-language summary generation, and in particular to a cross-language news dialogue summary generation method combining large and small models. Background Art
[0002] As an important research direction in the field of natural language processing, cross-language news dialogue summary generation aims to understand and compress the text of news dialogues of multiple parties and generate high-quality summaries in different languages. This task has the characteristics of complex content and diverse languages. It not only involves information interaction in the dialogue, but also requires extracting the core information of the news event, while achieving cross-language conversion and adapting to multiple language phenomena. Therefore, improving the efficiency of the model while ensuring the quality of the summary is the core issue of research in this field.
[0003] Unlike traditional single-article news summaries, news conversations often involve multi-party exchanges, including interviews, comments, etc. The text is long, the content structure is complex, and different speakers may express different positions. In addition, issues such as reference resolution, viewpoint integration, and topic coherence in news conversations make summary generation more difficult than general document summaries. In cross-language scenarios, English-to-Chinese summary generation not only requires information compression, but also involves issues such as language transfer, syntactic structure conversion, and cultural adaptation, which further increases the challenge of the task.
[0004] At present, existing methods usually rely on a single large model or a small model to handle this task. For the large model, with its powerful text understanding ability, it can effectively capture the structure and information context of the conversation, and can optimize the generation effect through the Few-Shot method; while the small model builds an efficient cross-language summary generation model through a special pre-training task or framework. However, using these two models alone will have different limitations in the following aspects: 1) Conversation understanding: The large model has a strong context modeling ability, but when processing long texts, information loss, misunderstanding of conversation logic or hallucinations may occur; and the small model has limited parameters and is difficult to capture cross-sentence relationships and semantic dependencies between speakers, which may lead to a lack of coherence in the summary or omission of key information; 2) Information screening: Conversation texts usually contain a lot of redundant information, such as greetings, repeated expressions, emotional words, etc. Although the large model can understand the content from a global perspective, it is easy to generate lengthy summaries and cannot effectively remove irrelevant information; the small model can screen out key information after targeted training, but it is more sensitive to noise, which can easily lead to unstable summary quality; 3) Cross-language transfer: The large model has certain translation capabilities, but in the process of generating English to Chinese summaries, it may be difficult to take into account both information compression and language style adjustment at the same time, resulting in the target language summary being stiff or lacking in news style; and the small model may have difficulty in accurately modeling the language transfer process due to limited training data, resulting in missing summary information or unsmooth translation.
[0005] In order to solve the above problems, the present invention designs a cross-language news dialogue summary generation method that combines a large model and a small model. This method gives full play to the powerful text understanding ability of the large model and the efficient reasoning ability of the small model, realizes the complementary advantages of the two, and thus obtains a more accurate and fluent summary result. Summary of the invention
[0006] The purpose of the present invention is to provide a cross-language news dialogue summary generation method combining large and small models, taking into account the characteristics of news dialogues and making full use of the capabilities of large and small models to enhance the generation of high-quality target language summaries in the field of multilingual news dialogues.
[0007] In order to achieve the above object, the present invention provides the following scheme:
[0008] S1: Classify news conversation types based on the big model;
[0009] S2: Train a small multilingual dialogue summary model to generate an initial target language summary;
[0010] S3: Based on the initial target language summary and news dialogue type, the content of the summary is optimized using large model collaboration technology;
[0011] S4: Use the small decision model to formally evaluate the content-optimized summary, and use the large model to optimize based on the evaluation to obtain the final summary.
[0012] Further, step S1 includes:
[0013] S11: Pre-definition of news dialogue types: define three types of dialogue: news report dialogue, commentary dialogue and interview dialogue: News report dialogue is a dialogue centered around news events, with the purpose of delivering news information, reporting social events or public affairs, usually consisting of interactions between reporters and witnesses, etc., and the content of the discussion involves the background, process, impact, etc. of the news events; Commentary dialogue is a discussion centered around news events, social phenomena or public issues, usually led by experts or critics, with the purpose of conducting in-depth analysis, commentary and evaluation of news events or social issues; Interview dialogue focuses on a special interview with an important person, focusing on the person's views, experiences and opinions;
[0014] S12: Use the big model to classify news dialogue types: First, design a source language prompt template for news dialogue classification, which details the definition of each news dialogue type. Then, combine the template with the source language news dialogue and call the big model to obtain the corresponding news dialogue type in the target language. The dialogue type will guide the big model in the direction of summary content iteration and the angle of summary form optimization.
[0015] Further, step S2 includes:
[0016] S21: Secondary training based on the multilingual pre-trained model mBART model to obtain a multilingual dialogue summary model: Four secondary pre-training tasks are designed according to the characteristics of the dialogue. Among them, the action filling task is similar to the masking task, which randomly masks the content of the action triple "who-doing-what" to help the model effectively capture the structural features in the dialogue, so as to better understand the actions and role information in the dialogue; the discourse replacement task disrupts the order of the sentences in the dialogue and requires the model to restore the original order of the sentences in the dialogue, enhancing the model's understanding of the fluidity of the dialogue and the relationship between different parts; the single language dialogue summary task helps the model learn the ability to generate summaries; the machine translation task model can learn the conversion between the source language and the target language, improving the model's cross-language ability;
[0017] S22: Fine-tune the multilingual dialogue summary model to obtain the target language summary of the multilingual news dialogue: Use fine-tuning technology to adapt the multilingual dialogue summary model to the news dialogue summary task from a specific source language to a target language to improve the model's performance in cross-language news dialogue summarization. The target language summary finally generated will serve as the initial version of the subsequent large model summary iteration.
[0018] Further, step S3 includes:
[0019] S31: Keyword extraction from news dialogues: First, a source language prompt template is designed for the keyword extraction task. Then, the prompt template and the source language news dialogue are input into the big model to extract keywords from the news dialogues, providing additional knowledge support for subsequent summary content optimization.
[0020] Optionally, the news conversation keyword extraction method in step S31 may specifically include:
[0021] The news dialogue can be divided into blocks, extracted and then merged, or directly extracted from the entire news dialogue. The selection of the method depends on the length of the news dialogue and the granularity of the required keyword extraction.
[0022] Optionally, the design method of the news dialogue keyword prompt template in step S31 may specifically include:
[0023] General prompt template and prompt template customized according to the conversation type. Among them, the prompt template customized according to the conversation type needs to use the conversation type information obtained in step S1 to ensure the accurate extraction of keywords. The specific design is as follows: news report conversations need to extract key entities such as events, time, and place; commentary conversations need to extract key entities such as people and topics; interview conversations need to extract key entities such as people, identity information, company name, and topics.
[0024] S32: Design a summary evaluation template based on the target language for use by the evaluation model and a summary generation prompt template for use by the summary generation model;
[0025] S33: combining the initialization target language summary and summary evaluation prompt template generated in step S2, calling the evaluation model to obtain the summary score and target language modification suggestions;
[0026] S34: The summary generation model optimizes the summary according to the optimization suggestions of the evaluation model and generates an improved summary;
[0027] S35: Use the optimized summary as a new evaluation object and continue to perform evaluation and feedback optimization;
[0028] S36: Repeat the evaluation and optimization process until the evaluation model determines that the summary has no obvious content omissions or factual errors, or reaches the maximum number of iterations, and outputs the final summary optimization result as the result.
[0029] Further, step S4 includes:
[0030] S41: Design a feedback template for suggestions on different indicators;
[0031] Optionally, summary indicators may include three categories: fluency, consistency indicators, and attribute relevance indicators. Fluency mainly measures the semantic relevance between sentences in the generated summary; consistency indicators mainly measure the relevance between news dialogues and summaries; and attribute relevance mainly measures whether the generated summary contains the core features that a specific news dialogue summary type should have.
[0032] S42: Use the summary result after step S3 as the initial summary;
[0033] S43: using the small judgment model to score each indicator of the summary based on the CTRLEval algorithm, if each indicator reaches a predetermined threshold, the summary is output as an optimized result;
[0034] S44: if a certain indicator does not reach a threshold, the summary generation model is called to modify the summary using a prompt generated by a suggestion feedback template corresponding to the indicator;
[0035] S45: The revised summary is used as a new evaluation object, and the indicator calculation, indicator judgment and feedback optimization steps are iteratively performed until each indicator of the summary reaches the corresponding threshold, and the summary is output as the final result. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0037] Figure 1 A flow chart of a cross-language news dialogue summary generation method combining large and small models provided by the present invention;
[0038] Figure 2 A data flow diagram of the cross-language news dialogue summary generation method combining large and small models of the present invention in practical application;
[0039] Figure 3 A prompt template for the news dialogue type classification step in the English-Chinese news dialogue summary generation task applied by the present invention;
[0040] Figure 4 A prompt template for the news dialogue keyword extraction step in the English-Chinese news dialogue summary generation task applied by the present invention;
[0041] Figure 5 The invention is applied to the English-Chinese news dialogue summary generation task, and the summary generation large model is used to perform summary optimization prompt template;
[0042] Figure 6 This is a prompt template for summarization evaluation in a large evaluation model used in the present invention for summarization generation task of English-Chinese news dialogues. DETAILED DESCRIPTION
[0043] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0044] The purpose of the present invention is to provide a cross-language news dialogue summary generation method combining large and small models, taking into account the characteristics of news dialogues and making full use of the capabilities of large and small models to enhance the generation of high-quality target language summaries in the field of multilingual news dialogues.
[0045] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0046] like Figure 1 and Figure 2 As shown, the cross-language news dialogue summary generation method combining large and small models provided by the present invention includes:
[0047] S1: Classify news conversations based on the big model:
[0048] S11: Pre-definition of news dialogue types: Three types of news dialogues are defined: news reporting dialogues, commentary dialogues, and interview dialogues. Among them, news reporting dialogues are dialogues centered around news events, with the purpose of delivering news information, reporting social events or public affairs, and are usually composed of interactions between reporters and witnesses, etc. The content of the discussion involves the background, occurrence process, impact, etc. of the news events; commentary dialogues are discussion dialogues centered around news events, social phenomena or public issues, and are usually led by experts or critics, with the purpose of conducting in-depth analysis, commentary and evaluation of news events or social issues; interview dialogues focus on special interviews with a certain important person, focusing on discussing the person's views, experiences and opinions.
[0049] S12: Use the big model to classify news dialogue types: First, design a source language prompt template for news dialogue classification, which details the definition of each news dialogue type; then, combine the template with the source language news dialogue, and call the big model to obtain the corresponding target language news dialogue type. The dialogue type will guide the big model to iterate the summary content and optimize the summary form.
[0050] When applied to the English-Chinese news dialogue summary generation task, this step will update Figure 3 The {src_news_dialog} placeholder part in the prompt template shown is input into the large model for news dialogue type recognition. Similarly, if it is applied to cross-language news dialogue generation tasks in other language pairs, the prompt template can be converted into the corresponding target language version to adapt to different language environments. In addition, the news dialogue classification model can be a general large language model, such as ChatGPT, Qwen, Llama, etc.
[0051] S2: Train a small multilingual dialogue summary model to generate an initial target language summary:
[0052] S21: Secondary training based on the multilingual pre-trained model mBART model to obtain a multilingual dialogue summary model: Four secondary pre-training tasks are designed according to the characteristics of the dialogue. Among them, the action filling task is similar to the masking task, which randomly masks the content in the action triple "who-doing-what" to help the model effectively capture the structural features in the dialogue, so as to better understand the actions and role information in the dialogue; the discourse replacement task disrupts the order of the sentences in the dialogue, requiring the model to restore the original order of the sentences in the dialogue, and enhance the model's understanding of the fluidity of the dialogue and the relationship between different parts; the single language dialogue summary task helps the model learn the ability to generate summaries; the machine translation task model can learn the conversion between the source language and the target language, and improve the model's cross-language capabilities;
[0053] S22: Fine-tune the multilingual dialogue summary model to obtain the target language summary of the multilingual news dialogue: Use fine-tuning technology to adapt the multilingual dialogue summary model to the news dialogue summary task from a specific source language to a target language to improve the model's performance in cross-language news dialogue summarization. The target language summary finally generated will serve as the initial version of the subsequent large model summary iteration.
[0054] In the application of English-to-Chinese news dialogue summary generation tasks, it is considered that existing multilingual pre-trained models such as mBART and mT5 mainly learn the underlying language modeling capabilities, but have deficiencies in cross-language conversion and dialogue understanding. In order to improve its performance in the cross-language dialogue summary task, the present invention designs a secondary pre-training task based on the characteristics of the dialogue based on mBART, so that the model can master the dialogue structure and cross-language conversion capabilities at the same time. In the fine-tuning stage, the parallel corpus from English to Chinese in the XMediaSum dataset is used for fine-tuning to further enhance the model's adaptability to cross-language news dialogue summaries. Finally, the small model generates an initial target language summary as input for subsequent iterative optimization of the large model, laying the foundation for high-quality summary generation.
[0055] S3: Based on the initial target language summary and news dialogue type, the content of the summary is optimized using large model collaboration technology:
[0056] S31: News dialogue keyword extraction: First, design a source language prompt template for the keyword extraction task. Then, input the prompt template and the source language news dialogue into the big model to extract keywords from the news dialogue and provide additional knowledge support for subsequent summary content optimization.
[0057] Optionally, the keyword extraction method in step S31 may specifically include:
[0058] The news dialogue can be divided into blocks, extracted and then merged, or directly extracted from the entire news dialogue. The selection of the method depends on the length of the news dialogue and the granularity of the required keyword extraction.
[0059] Optionally, the design method of the keyword extraction prompt template in step S31 may specifically include:
[0060] General prompt template and prompt template customized according to the conversation type. Among them, the prompt template customized according to the conversation type needs to use the conversation type information obtained in step S1 to ensure the accurate extraction of keywords. The specific design is that news report conversations need to extract key entities such as events, time, and place, commentary conversations need to extract key entities such as people and topics, and interview conversations need to extract key entities such as people, identity information, company name, and topics.
[0061] When applied to the task of summarizing English-Chinese news dialogues, such as Figure 4 The task template shown is in English, i.e. the source language. If it is applied to the cross-language news dialogue generation task of other language pairs, the prompt template can be converted into the corresponding source language version to adapt to different language environments. At the same time, in the selection of prompt word templates, prompt templates based on the dialogue type address are used. It is necessary to dynamically select prompt words for the task refinement part based on the news dialogue classification results in S1 to ensure efficient and accurate extraction of key entities of the appropriate category, providing clearer guidance for evaluating large models.
[0062] S32: Design a summary evaluation prompt template based on the target language for use by the evaluation large model and a summary generation prompt template for use by the summary generation large model;
[0063] When applied to the summary generation task of English-Chinese news dialogue, the language of the summary generation prompt template and the summary evaluation prompt template are both Chinese, that is, the target language, such as Figure 5 and Figure 6 Similarly, if it is applied to the cross-language news dialogue generation task of other language pairs, the prompt template can be converted into the corresponding target language version to adapt to different language environments.
[0064] At the same time, the keyword extraction results provided by step S31 are added to the evaluation prompt template to help the evaluation model focus on key information and provide more targeted modification suggestions. Through the quantitative evaluation of key information, the quality of the summary is improved, and the risk of hallucinations generated by the large model is alleviated to a certain extent. In addition, the requirements of the task refinement part of the prompt template specify specific requirements according to different types of conversations. It is necessary to select the requirements of the corresponding category based on the news conversation classification results in S1 to ensure that each type of summary can better meet the specific content focus. Through this classification refinement, the model can generate summaries that meet the requirements in a more targeted manner and improve the overall quality.
[0065] S33: combining the initialization target language summary and summary evaluation prompt template generated in step S2, calling the evaluation model to obtain the summary score and target language modification suggestions;
[0066] When applied to the English-to-Chinese news dialogue summary generation task, this step fills the source language news dialogue into the {src_news_dialog} placeholder in the prompt template, and updates the target language summary generated in step S2 to the {tgt_news_sum} placeholder in the prompt template. Then, the updated prompt template is input into the evaluation model. The evaluation model analyzes and scores the summary, and generates modification suggestions based on the analysis results, providing guiding feedback to help optimize the quality of the summary. This process ensures that the evaluation model focuses on the key information in the summary and improves the accuracy of the feedback.
[0067] S34: The summary generation model optimizes the summary according to the optimization suggestions of the evaluation model and generates an improved summary;
[0068] When applied to the summary generation task of English-to-Chinese news dialogue, this step updates the {tgt_suggest} placeholder in the summary generation model prompt template according to the optimization suggestions provided by the evaluation model. At this time, the {tgt_news_sum} placeholder content is the target language summary generated in step S2. The summary generation model will adjust and optimize the summary content according to the evaluation suggestions, enhance the accuracy and information completeness of the summary, and ensure that the summary is more in line with the main information of the source language dialogue.
[0069] S35: Use the optimized summary as a new evaluation object and continue to perform evaluation and feedback optimization;
[0070] When applied to the summary generation task of English-to-Chinese news dialogues, this step updates the {tgt_news_sum} placeholder in the evaluation model prompt template to the optimized summary of the summary generation model in the previous step, and continues to enter the next round of evaluation and feedback optimization cycle. The core of this step is to maintain the high quality of the summary, while gradually eliminating the shortcomings in the summary in each round of optimization to ensure the accuracy and conciseness of the summary.
[0071] S36: Repeat the evaluation and optimization process until the evaluation model determines that the summary has no obvious content omissions or factual errors, or reaches the maximum number of iterations, and outputs the final summary optimization result as the result.
[0072] When applied to the summary generation task of English-to-Chinese news dialogues, in each iteration, {tgt_news_sum} in the evaluation prompt word is updated to the summary optimized by the summary generation model. The evaluation model will check whether the summary meets the conditions for terminating the iteration. If the summary has not yet reached the standard, {tgt_news_sum} in the summary generation model is updated to the summary after the previous round of optimization, and {tgt_suggest} is the modification suggestion given by the evaluation model in this round, and the summary is optimized again. This process will continue to cycle until the summary meets the evaluation criteria or reaches the maximum iteration round, and finally output the optimization result of the summary generation model in the last iteration.
[0073] In addition, the summary generation model, evaluation model and keyword extraction can all be general large-scale language models, such as ChatGPT, Qwen, Llama, etc.
[0074] S4: Use the small decision model to formally evaluate the content-optimized summary, and use the large model to optimize based on the evaluation to obtain the final summary:
[0075] S41: Design a feedback template for suggestions on different indicators;
[0076] Optionally, summary indicators may include three categories: fluency, consistency indicators, and attribute relevance indicators. Fluency mainly measures the semantic relevance between sentences in the generated summary; consistency indicators mainly measure the relevance between news dialogues and summaries; and attribute relevance mainly measures whether the generated summary contains the core features that a specific news dialogue summary type should have.
[0077] When applied to the English-to-Chinese news dialogue summary generation task, the feedback template for fluency is "The coherence score of the summary is {coh_result}, which is less than the coherence threshold, which means that the coherence and fluency between sentences in the summary are low. Please improve the summary to improve its coherence."; The feedback template for fluency is "The consistency score of the summary is {cons_result}, which is less than the consistency threshold, which means that the consistency between the summary and the dialogue is low. Please refine the summary to improve its consistency.";
[0078] The feedback template for attribute relevance is "The attribute relevance score of the summary is {ar_result}, which is less than the attribute relevance threshold, which means that the summary does not meet the requirements of {news_type} summary. Please add key information to improve its attribute relevance." Among them, {coh_result}, {cons_result}, {ar_result}, and {news_type} are placeholders, and the corresponding data will be filled in in sequence when used later.
[0079] S42: Use the summary result after step S3 as the initial summary;
[0080] When applied to the English-to-Chinese news dialogue summary generation task, the target language summary obtained from step S3 is used as the initial version and as the starting point for summary optimization in subsequent steps.
[0081] S43: using the small judgment model to score each indicator of the summary based on the CTRLEval algorithm, if each indicator reaches a predetermined threshold, the summary is output as an optimized result;
[0082] When applied to the English-to-Chinese news dialogue summary generation task, based on the initial summary in step S42, the Pegasus summary model is used as a small judgment model to evaluate the summary and calculate the score of each indicator. If all indicators reach the predetermined threshold, the summary is considered to be sufficiently optimized and is output as the final summary result.
[0083] S44: if a certain indicator does not reach a threshold, the summary generation model is called to modify the summary using a prompt generated by a suggestion feedback template corresponding to the indicator;
[0084] When applied to the summary generation task of English-to-Chinese news dialogues, when the evaluation result of a certain indicator by the judgment model based on the Pegasus summary model is lower than the set threshold, the feedback template of the indicator is updated and used as the actual content of the {tgt_suggest} placeholder of the summary generation model. Then, the summary generation model is called to make corresponding corrections to the summary to optimize the indicator score.
[0085] S45: The revised summary is used as a new evaluation object, and the indicator calculation, indicator judgment and feedback optimization steps are iteratively performed until each indicator of the summary reaches the corresponding threshold, and the summary is output as the final result.
[0086] When applied to the English-to-Chinese news dialogue summary generation task, the revised summary in step S44 is used as a new evaluation object, and the indicator calculation, judgment and feedback optimization are performed again. This process is repeated until each indicator reaches a predetermined threshold, and finally an optimized summary is output.
[0087] In the above, the specific embodiments of the present invention are described with reference to the accompanying drawings. However, those skilled in the art will appreciate that various changes and substitutions may be made to the specific embodiments of the present invention without departing from the spirit and scope of the present invention. These changes and substitutions are all within the scope defined by the claims of the present invention.
Claims
1. A cross-language news dialogue summary generation method combining large and small models, characterized in that The following steps are included: S1: Classify news conversation types based on the big model; S2: Train a small multilingual dialogue summary model to generate an initial target language summary; S3: Based on the initial target language summary and news dialogue type, the content of the summary is optimized using large model collaboration technology; S4: Use the small decision model to formally evaluate the content-optimized summary, and use the large model to optimize based on the evaluation to obtain the final summary.
2. A cross-language news dialogue summarization method combining large and small models according to claim 1, characterized in that: The step S1 specifically includes: S11: Pre-definition of news dialogue types: define three types of dialogue: news report dialogue, commentary dialogue and interview dialogue. Among them, news report dialogue is a dialogue centered around news events, with the purpose of delivering news information, reporting social events or public affairs. It usually consists of interactions between reporters and witnesses, and the content of the discussion involves the background, process, impact, etc. of the news events. Commentary dialogue is a discussion centered around news events, social phenomena or public issues, usually led by experts or critics, with the purpose of conducting in-depth analysis, commentary and evaluation of news events or social issues. Interview dialogue focuses on a special interview with an important person, focusing on the person's views, experiences and opinions. S12: Use the big model to classify news dialogue types: First, design a source language prompt template for news dialogue classification, which details the definition of each news dialogue type. Then, combine the template with the source language news dialogue and call the big model to obtain the corresponding news dialogue type in the target language. The dialogue type will guide the big model in the direction of summary content iteration and the angle of summary form optimization.
3. A cross-language news dialogue summarization method combining large and small models according to claim 1, characterized in that: The step S2 specifically includes: S21: Secondary training based on the multilingual pre-trained model mBART to obtain a multilingual dialogue summary model: Four secondary pre-training tasks are designed according to the characteristics of the dialogue. Among them, the action filling task is similar to the masking task. It randomly masks the content in the action triple "who-doing-what" to help the model effectively capture the structural features in the dialogue so as to better understand the actions and role information in the dialogue. The discourse replacement task disrupts the order of the sentences in the dialogue and requires the model to restore the original order of the sentences in the dialogue to enhance the model's understanding of the fluidity of the dialogue and the relationship between different parts. The single-language dialogue summary task helps the model learn the ability to generate summaries. The machine translation task model can learn the conversion between the source language and the target language, improving the model's cross-language ability. S22: Fine-tune the multilingual dialogue summary model to obtain the target language summary of the multilingual news dialogue: Use fine-tuning technology to adapt the multilingual dialogue summary model to the news dialogue summary task from a specific source language to a target language to improve the model's performance in cross-language news dialogue summarization. The target language summary finally generated will serve as the initial version of the subsequent large model summary iteration.
4. A cross-language news dialogue summarization method combining large and small models according to claim 1, characterized in that: The step S3 specifically includes: S31: Keyword extraction from news dialogues: First, a source language prompt template is designed for the keyword extraction task. Then, the prompt template and the source language news dialogue are input into the big model to extract keywords from the news dialogues, providing additional knowledge support for subsequent summary content optimization. S32: Design a summary evaluation prompt template based on the target language for use by the evaluation large model and a summary generation prompt template for use by the summary generation large model; S33: combining the initialization target language summary and summary evaluation prompt template generated in step S2, calling the evaluation model to obtain the summary score and target language modification suggestions; S34: The summary generation model optimizes the summary according to the optimization suggestions of the evaluation model and generates an improved summary; S35: Use the optimized summary as a new evaluation object and continue to perform evaluation and feedback optimization; S36: Repeat the evaluation and optimization process until the evaluation model determines that the summary has no obvious content omissions or factual errors, or reaches the maximum number of iterations, and outputs the final summary optimization result as the result.
5. A cross-language news dialogue summarization method combining large and small models according to claim 1, characterized in that: The step S4 specifically includes: S41: Design a feedback template for suggestions on different indicators; S42: Use the summary result after step S3 as the initial summary; S43: using the small judgment model to score each indicator of the summary based on the CTRLEval algorithm, if each indicator reaches a predetermined threshold, the summary is output as an optimized result; S44: if a certain indicator does not reach a threshold, the summary generation model is called to modify the summary using a prompt generated by a suggestion feedback template corresponding to the indicator; S45: The revised summary is used as a new evaluation object, and the indicator calculation, indicator judgment and feedback optimization steps are iteratively performed until each indicator of the summary reaches the corresponding threshold, and the summary is output as the final result.
6. The cross-language news dialogue summary generation method combining large and small models according to claim 4 is characterized in that: The news conversation keyword extraction method in step S31 may be: The news dialogue can be divided into blocks, extracted and then merged, or the entire news dialogue can be directly extracted. The choice of method depends on the length of the news dialogue and the granularity of the required keyword extraction.
7. The cross-language news dialogue summary generation method combining large and small models according to claim 4 is characterized in that: The design method of the news dialogue keyword extraction prompt template in step S31 may be as follows: There are general prompt templates and prompt templates customized according to the conversation type. Among them, the prompt template customized according to the conversation type needs to use the conversation type information obtained in step S1 to ensure the accurate extraction of keywords. The specific design is that news report-type conversations need to extract key entities such as events, time, and place; commentary-type conversations need to extract key entities such as people and topics; interview-type conversations need to extract key entities such as people, identity information, company names, and topics.
8. The cross-language news dialogue summary generation method combining large and small models according to claim 5 is characterized in that: The indicator selection in step S41 may specifically include: Fluency, consistency index and attribute relevance index. Among them, fluency mainly measures the semantic relevance between sentences in the generated summary; consistency index mainly measures the relevance between news dialogue and summary; attribute relevance mainly measures whether the generated summary contains the core features that a specific news dialogue summary type should have.