News generation system based on language large model
By designing a news generation system based on language big model, the accuracy problem caused by training data deviations when generating news is solved. The accuracy of news components is judged through feature extraction and comparison analysis, and the model is improved based on the type of inaccurate information, so that high-quality and reliable news content can be generated.
Patent Information
- Application Number
- CN202510107045.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When generating news, the existing language models may have one-sided or inaccurate problems due to the deviation and incompleteness of the training data when generating news, which affects its accuracy and reliability.
Design a news generation system based on language big models, obtain news components generated by multiple language big models, perform topic recognition, feature extraction and comparison analysis, judge the accuracy of news components, and determine the improvement direction of language big models based on the type of inaccurate information.
By specifically calculating and analyzing the accuracy of the generated news structure, counting the impact proportional coefficients of inaccurate information feedback from different types, determining the improvement direction of the language model, and continuously improving the model to ensure the high quality and reliability of the news content it generates.
Smart Images

Figure CN120030254A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a news generation system based on a language large model. Background Art
[0002] The language big model is an artificial intelligence technology with powerful language understanding and generation capabilities. It is trained on massive text data, which can come from various sources such as the Internet, books, news, and academic papers. Through large-scale data training and deep neural network architecture, it can handle various natural language processing tasks. With the continuous development and improvement of the language big model, the news generation system can quickly and efficiently generate various types of news content. These systems can understand the structure, semantics, and context of language by learning and analyzing large amounts of text data, thereby generating high-quality news reports.
[0003] At present, large language models are usually generated based on existing data and are unable to conduct in-depth investigations and verify facts like professional journalists. They often produce one-sided or inaccurate reports due to deviations in training data. Statistical data analysis tools show that accuracy issues occur frequently and to a significant degree. The reason is that the model lacks different knowledge. When faced with text processing tasks in different fields, the generated results are erroneous. The problem may be that the model's training data is not comprehensive enough. The model may not be able to accurately determine its exact meaning, thus affecting the overall processing effect. Inappropriate content generation also occurs from time to time, which reduces the accuracy and reliability of large language models when generating news.
[0004] In view of the above technical defects, a solution is now proposed. Summary of the invention
[0005] The purpose of the present invention is to generate different types of news components through a language macro model, and obtain the accuracy of the news topic and news content based on the news topic and news content of the generated news components, and obtain the type of inaccurate information generated based on the accuracy, thereby determining the improvement direction of the language macro model.
[0006] In order to achieve the above-mentioned purpose, the present invention adopts the following technical scheme: a news generation system based on a large language model, comprising a language model acquisition module, a model operation analysis module, an operation result analysis module, and a judgment generation module;
[0007] The language model acquisition module is used to acquire news components generated by multiple language models, and set standard data sources and news components to be generated according to the language models. The news components include news topics and news content. When generating news components, the news components include four modes: news article generation, news picture generation, news video material generation, and news data generation.
[0008] The model operation analysis module includes a topic identification unit and an operation analysis unit. The topic identification unit is used to obtain the topic content of the news component in real time, predict the news category to which the news component belongs when generating news according to the topic content of the news component, and transmit it to the operation result analysis module;
[0009] The operation analysis module is used to obtain the news content of the generated news component, perform different feature extraction restrictions on the news content based on the news component generation method, perform feature extraction on the news content of the generated news component, perform feature dimension processing on the extracted features, and send them to the operation result analysis module;
[0010] The operation result analysis module is used to obtain the news category and extracted features of the news component, compare the extracted features with the standard data source based on the language big model, analyze the subject content of the news component and the accuracy of the news content of the news component, obtain the comprehensive accuracy based on the combination of the two, and judge the conformity status of the generated news component through the comprehensive accuracy. The conformity status includes the standard status, qualified status and unqualified status. If it is judged to be unqualified, feedback is given to record the location and type of inaccurate information;
[0011] The judgment generation module is used to generate a judgment analysis report according to the compliance status of the news component, obtain the impact ratio coefficient of each generation method of the news component according to the judgment analysis report, determine the improvement direction of the language model according to the size of the impact ratio coefficient, and send it to the language model data control terminal.
[0012] Furthermore, the standard data source and the news components to be generated are set according to the language model:
[0013] S1: Collect the news components to be generated, obtain the standard news components generated by the news components to be generated, and set the standard news components as the comparison subject and news content of the language model as the standard data source;
[0014] S2: extracting news topics and news contents from standard news components, analyzing the news categories to which the standard news components belong, and obtaining a preset news category set;
[0015] S3: When generating news components, there are four methods: news article generation, news picture generation, news video material generation, and news data generation.
[0016] Furthermore, the specific process of predicting the news category to which the news component belongs when generating news is as follows:
[0017] Model building: Based on the news components that need to be generated, a large amount of news component data is collected from news media websites, news databases, and social media, covering different fields and types of news reports, to generate a classification model;
[0018] Feature extraction: Extract features from the subject content of the news component and analyze the extracted features. Determine the strength of the relationship between the extracted features and each category based on the size of the correlation. Determine the possibility of the category based on the strength of the relationship and make a classification result of the possibility of the category.
[0019] Output results: The classification results are output through the classification model, and the category label of the news component topic is generated based on the classification results.
[0020] Furthermore, the specific process of judging whether the generated news component meets the status is as follows:
[0021] S01: obtaining news content of the generated news component, and performing pre-processing on the news content of the news component and comparing and analyzing it with the standard data source, and analyzing the accuracy of the news content of the news component and the news theme;
[0022] S02: pre-setting the comparative analysis accuracy of the news component during analysis, setting limiting conditions for the compliance status according to the accuracy, and dividing the compliance status into three types;
[0023] S03: Obtain the accuracy of the news content and news theme of the news component, combine the accuracy to obtain the comprehensive accuracy, and obtain the news component compliance status analysis result under the limited conditions according to the comprehensive accuracy;
[0024] S04: According to the analysis result of the news component conforming to the state, a state instruction is generated according to the type of the conforming state.
[0025] Furthermore, the specific process of combining the accuracies to obtain the comprehensive accuracy is as follows:
[0026] Step 1: The accuracy of the news topic is preset to L, and the accuracy of the news content of the news component is preset to S, including the generation of news articles A, the generation of news illustrations B, the generation of news video materials C, and the generation of news data D;
[0027] Step 2: Based on the accuracy of the news topic preset as L and the accuracy of the news content of the news component as S, obtain the combined accuracy of the two:
[0028] Step 3: Preliminary calculation of the accuracy rate L of the news topic according to the following formula: L = w 1 X 1 +w 2 X 2 +w 3 X 3 , w 1 +w 2 +w 3 =1, X is the target factor to be evaluated, and W is the weight coefficient of each target factor;
[0029] Then calculate the accuracy of the news content: S = α × A + β × B + γ × C + ε × D, α + β + γ + ε = 1, α, β, γ, ε are all assigned weight coefficients;
[0030] The comprehensive accuracy of the combination of the two is: P=L+S.
[0031] Furthermore, the specific process of classifying the compliance status into three types in the above S02 is as follows:
[0032] The combined accuracy of the two is preset to be y1, y2, y3,
[0033] The first one is: if y1>P>y2, it is the standard state and no feedback is given;
[0034] The second type is: if y2>P>y3, it is a qualified state; no feedback is given, but error information is recorded;
[0035] The third type is: y3>P, which is an unqualified state, and feedback is given to record the location and type of inaccurate information.
[0036] Furthermore, the specific process of generating a judgment analysis report according to the news component compliance status is as follows:
[0037] Step 1: Analyze the accuracy of news components, obtain the overall news components, and check the key features and information sources in the news components;
[0038] Step 2: Based on the location and type of inaccurate information in the acquired news component, determine whether it is an error in the main body and news content, an error in the accompanying image, an error in the video, or an error in the data;
[0039] Step 3: Based on the inaccurate information, the language model is analyzed in the news generation process, and the influence ratio coefficient of each generation method of the generated news component is
[0040] Step 4: Determine the direction of improvement of the language model based on the size of the impact coefficient.
[0041] Furthermore, the judgment generation module also includes an impact improvement module, which is used to obtain an impact ratio coefficient data value, determine the impact ratio of inaccurate information according to the impact ratio coefficient, and optimize and improve the large language model being trained.
[0042] In summary, due to the adoption of the above technical solution, the beneficial effects of the present invention are:
[0043] The news generation system based on the language big model constructs the news body and news content of the news components by using the language big model, constructs different types of news bodies and news content, obtains the errors and inaccurate information generated during news generation from different types of news bodies and news content, generates the accuracy of each news body and news content construction through specific calculation and analysis, and then uses data analysis technology to count the impact ratio coefficient of each generation method of inaccurate information of different types of feedback, determines the improvement direction of the language big model, deeply analyzes the feedback content, explores the reasons behind it, continuously improves the language big model, and takes effective measures to ensure its accuracy, so as to provide readers with high-quality and reliable news services. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] Figure 1 A schematic diagram of the system flow structure of the present invention is shown. DETAILED DESCRIPTION
[0045] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0046] Embodiment 1:
[0047] like Figure 1 As shown, a news generation system based on a large language model includes a language model acquisition module, a model operation analysis module, an operation result analysis module, and a judgment generation module;
[0048] The language model acquisition module is used to acquire news components generated by multiple language models, and set standard data sources and news components to be generated according to the language models. The news components include news topics and news content. When generating news components, the news components include four modes: news article generation, news picture generation, news video material generation, and news data generation.
[0049] Establish data screening standards during the data collection process, strictly review the collected data to ensure its accuracy, completeness and timeliness, formulate regular data update plans, such as conducting a comprehensive check and update of data every week or month, use automated tools to monitor news events and data changes, and promptly incorporate new data into model training. Establish a data quality assessment indicator system, conduct quantitative assessments of data accuracy, completeness, consistency, etc., regularly clean and repair data, and remove erroneous or duplicate data.
[0050] The model operation analysis module includes a topic identification unit and an operation analysis unit. The topic identification unit is used to obtain the topic content of the news component in real time, predict the news category to which the news component belongs when generating news according to the topic content of the news component, and transmit it to the operation result analysis module;
[0051] The operation analysis module is used to obtain the news content of the generated news component, perform different feature extraction restrictions on the news content based on the news component generation method, perform feature extraction on the news content of the generated news component, perform feature dimension processing on the extracted features, and send them to the operation result analysis module;
[0052] The operation result analysis module is used to obtain the news category and extracted features of the news component, compare the extracted features with the standard data source based on the language big model, analyze the subject content of the news component and the accuracy of the news content of the news component, obtain the comprehensive accuracy based on the combination of the two, and judge the conformity status of the generated news component through the comprehensive accuracy. The conformity status includes the standard status, qualified status and unqualified status. If it is judged to be unqualified, feedback is given to record the location and type of inaccurate information;
[0053] When extracting features, the news subject and news content are regarded as a set of words. The order and grammatical relationship of the words are not considered. The frequency of each word in the text is counted to construct the feature vectors of the subject and news content.
[0054] The judgment generation module is used to generate a judgment analysis report according to the compliance status of the news component, obtain the impact ratio coefficient of each generation method of the news component according to the judgment analysis report, determine the improvement direction of the language model according to the size of the impact ratio coefficient, and send it to the language model data control terminal.
[0055] Set standard data sources and news components to be generated based on the language model:
[0056] S1: Collect the news components to be generated, obtain the standard news components generated by the news components to be generated, and set the standard news components as the comparison subject and news content of the language model as the standard data source;
[0057] S2: extracting news topics and news contents from standard news components, analyzing the news categories to which the standard news components belong, and obtaining a preset news category set;
[0058] S3: When generating news components, there are four methods: news article generation, news picture generation, news video material generation, and news data generation.
[0059] The specific process of predicting the news category to which the news component belongs when generating news is as follows:
[0060] Model building: Based on the news components that need to be generated, a large amount of news component data is collected from news media websites, news databases, and social media, covering different fields and types of news reports, to generate a classification model;
[0061] Feature extraction: Extract features from the subject content of the news component and analyze the extracted features. Determine the strength of the relationship between the extracted features and each category based on the size of the correlation. Determine the possibility of the category based on the strength of the relationship and make a classification result of the possibility of the category.
[0062] Output results: The classification results are output through the classification model, and the category label of the news component topic is generated based on the classification results.
[0063] The specific process of judging whether the generated news component meets the status is as follows:
[0064] S01: obtaining news content of the generated news component, and performing pre-processing on the news content of the news component and comparing and analyzing it with the standard data source, and analyzing the accuracy of the news content of the news component and the news theme;
[0065] S02: pre-setting the comparative analysis accuracy of the news component during analysis, setting limiting conditions for the compliance status according to the accuracy, and dividing the compliance status into three types;
[0066] S03: Obtain the accuracy of the news content and news theme of the news component, combine the accuracy to obtain the comprehensive accuracy, and obtain the news component compliance status analysis result under the limited conditions according to the comprehensive accuracy;
[0067] S04: According to the analysis result of the news component conforming to the state, a state instruction is generated according to the type of the conforming state.
[0068] The specific process of combining the accuracies to obtain the comprehensive accuracy is as follows:
[0069] Step 1: The accuracy of the news topic is preset to L, and the accuracy of the news content of the news component is preset to S, including the generation of news articles A, the generation of news illustrations B, the generation of news video materials C, and the generation of news data D;
[0070] News release production can be evaluated in terms of grammatical accuracy, content completeness, logical coherence, etc.
[0071] The generation of news illustrations can consider image clarity, relevance to news content, and visual appeal;
[0072] News video material generation can be evaluated from the aspects of video clarity, editing rationality, and content richness;
[0073] News data generation can be considered from the aspects of data accuracy, timeliness and authority.
[0074] Step 2: Based on the accuracy of the news topic preset as L and the accuracy of the news content of the news component as S, obtain the combined accuracy of the two:
[0075] Step 3: Preliminary calculation of the accuracy rate L of the news topic according to the following formula: L = w 1 X 1 +w 2 X 2 +w 3 X 3 , w 1 +w 2 +w 3 =1, X is the target factor to be evaluated, and W is the weight coefficient of each target factor;
[0076] Then calculate the accuracy of the news content: S = α × A + β × B + γ × C + ε × D, α + β + γ + ε = 1, α, β, γ, ε are all assigned weight coefficients;
[0077] The comprehensive accuracy of the combination of the two is: P=L+S.
[0078] The specific process of classifying the compliance status into three types in the above S02 is as follows:
[0079] The combined accuracy of the two is preset to be y1, y2, y3,
[0080] The first one is: if y1>P>y2, it is the standard state and no feedback is given;
[0081] The second type is: if y2>P>y3, it is a qualified state; no feedback is given, but error information is recorded;
[0082] The third type is: y3>P, which is an unqualified state, and feedback is given to record the location and type of inaccurate information.
[0083] The specific process of generating a judgment analysis report based on the news component compliance status is as follows:
[0084] Step 1: Analyze the accuracy of news components, obtain the overall news components, and check the key features and information sources in the news components;
[0085] Step 2: Based on the location and type of inaccurate information in the acquired news component, determine whether it is an error in the main body and news content, an error in the accompanying image, an error in the video, or an error in the data;
[0086] Step 3: Based on the inaccurate information, the language model is analyzed in the news generation process, and the influence ratio coefficient of each generation method of the generated news component is
[0087] Step 4: Determine the direction of improvement of the language model based on the size of the impact coefficient.
[0088] When searching for problems with the language big model, check whether the language big model is trained based on inaccurate or incomplete data, resulting in factual errors in the generated news; analyze the reliability of the data source and how the big model processes the data; determine whether the big model lacks verification steps for information when generating news; consider whether it is necessary to introduce more external verification mechanisms to improve the accuracy of the news; analyze whether the big model has misunderstandings about input instructions or questions, resulting in the generation of erroneous news content; check for problems in language understanding and semantic analysis; observe whether there are over-generalizations or exaggerations in the news, which may be the tendency of the big model when generating subjects and news content; consider how to guide the big model to avoid over-generalization and exaggeration, maintain objectivity and accuracy, confirm whether the big model can obtain the latest information in a timely manner and reflect it in the news; analyze the shortcomings of the big model in processing timeliness; think about the lack of human judgment and common sense in the big model when generating news, resulting in inaccurate content; explore how to combine human expertise and judgment to improve news quality.
[0089] The judgment generation module also includes an impact improvement module, which is used to obtain an impact ratio coefficient data value, determine the impact ratio of inaccurate information according to the impact ratio coefficient, and optimize and improve the language model being trained.
[0090] The present invention constructs news bodies and news contents of news components by using a language macro model, constructs different types of news bodies and news contents, obtains error points and inaccurate information generated when generating news from different types of news bodies and news contents, generates the accuracy of each news body and news content when constructed through specific calculation and analysis, and then uses data analysis technology to count the influence proportion coefficient of each generation method of inaccurate information of different types of feedback, determines the improvement direction of the language macro model, deeply analyzes the feedback content, explores the reasons behind it, continuously improves the language macro model, and takes effective measures to ensure its accuracy, so as to provide readers with high-quality and reliable news services.
[0091] The size of the interval and threshold is set to facilitate comparison. The size of the threshold depends on the amount of sample data and the number of bases set by technical personnel in this field for each group of sample data; as long as it does not affect the proportional relationship between the parameter and the quantized value.
[0092] The above formulas are all dimensionless and numerical calculations. The formula is a formula obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formula are set by technicians in this field according to actual conditions.
[0093] In the two embodiments provided in the present application, it should be understood that the disclosed system can be implemented in other ways; for example, the device embodiments described above are only schematic, for example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed; another point, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, the indirect coupling or communication connection of the modules can be electrical, mechanical or other forms;
[0094] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A news generation system based on a language big model, characterized in that: It includes a language model acquisition module, a model operation analysis module, an operation result analysis module, and a judgment generation module; The language model acquisition module is used to acquire news components generated by multiple language models, and set standard data sources and news components to be generated according to the language models. The news components include news topics and news content. When generating news components, the news components include four modes: news article generation, news picture generation, news video material generation, and news data generation. The model operation analysis module includes a topic identification unit and an operation analysis unit. The topic identification unit is used to obtain the topic content of the news component in real time, predict the news category to which the news component belongs when generating news according to the topic content of the news component, and transmit it to the operation result analysis module; The operation analysis module is used to obtain the news content of the generated news component, perform different feature extraction restrictions on the news content based on the news component generation method, perform feature extraction on the news content of the generated news component, perform feature dimension processing on the extracted features, and send them to the operation result analysis module; The operation result analysis module is used to obtain the news category and extracted features of the news component, compare the extracted features with the standard data source based on the language big model, analyze the subject content of the news component and the accuracy of the news content of the news component, obtain the comprehensive accuracy based on the combination of the two, and judge the conformity status of the generated news component through the comprehensive accuracy. The conformity status includes the standard status, qualified status and unqualified status. If it is judged to be unqualified, feedback is given to record the location and type of inaccurate information; The judgment generation module is used to generate a judgment analysis report according to the compliance status of the news component, obtain the impact ratio coefficient of each generation method of the news component according to the judgment analysis report, determine the improvement direction of the language model according to the size of the impact ratio coefficient, and send it to the language model data control terminal.
2. The news generation system based on the language big model according to claim 1 is characterized in that: Set standard data sources and news components to be generated based on the language model: S1: Collect the news components to be generated, obtain the standard news components generated by the news components to be generated, and set the standard news components as the comparison subject and news content of the language model as the standard data source; S2: extracting news topics and news contents from standard news components, analyzing the news categories to which the standard news components belong, and obtaining a preset news category set; S3: When generating news components, there are four methods: news article generation, news picture generation, news video material generation, and news data generation.
3. The news generation system based on the language big model according to claim 2 is characterized in that: The specific process of predicting the news category to which the news component belongs when generating news is as follows: Model building: Based on the news components that need to be generated, a large amount of news component data is collected from news media websites, news databases, and social media, covering different fields and types of news reports, to generate a classification model; Feature extraction: Extract features from the subject content of the news component and analyze the extracted features. Determine the strength of the relationship between the extracted features and each category based on the size of the correlation. Determine the possibility of the category based on the strength of the relationship and make a classification result of the possibility of the category. Output results: The classification results are output through the classification model, and the category label of the news component topic is generated based on the classification results.
4. The news generation system based on the language big model according to claim 1 is characterized in that: The specific process of judging whether the generated news component meets the status is as follows: S01: obtaining news content of the generated news component, and performing pre-processing on the news content of the news component and comparing and analyzing it with the standard data source, and analyzing the accuracy of the news content of the news component and the news theme; S02: pre-setting the comparative analysis accuracy of the news component during analysis, setting limiting conditions for the compliance status according to the accuracy, and dividing the compliance status into three types; S03: Obtain the accuracy of the news content and news theme of the news component, combine the accuracy to obtain the comprehensive accuracy, and obtain the news component compliance status analysis result under the limited conditions according to the comprehensive accuracy; S04: According to the analysis result of the news component conforming to the state, a state instruction is generated according to the type of the conforming state.
5. The news generation system based on the language big model according to claim 1 is characterized in that: The specific process of combining the accuracies to obtain the comprehensive accuracy is as follows: Step 1: The accuracy of the news topic is preset to L, and the accuracy of the news content of the news component is preset to S, including the generation of news articles A, the generation of news illustrations B, the generation of news video materials C, and the generation of news data D; Step 2: Based on the accuracy of the news topic preset as L and the accuracy of the news content of the news component as S, obtain the combined accuracy of the two: Step 3: Preliminarily calculate the accuracy rate L of the news topic according to the following formula: L = w1X1 + w2X2 + w3X3, w1 + w2 + w3 = 1, X is the target factor to be evaluated, and W is the weight coefficient of each target factor; Then calculate the accuracy of the news content: S = α × A + β × B + γ × C + ε × D, α + β + γ + ε = 1, α, β, γ, ε are all assigned weight coefficients; The comprehensive accuracy of the combination of the two is: P=L+S.
6. The news generation system based on the language big model according to claim 5 is characterized in that: The specific process of classifying the compliance status into three types in the above S02 is as follows: The combined accuracy of the two is preset to be y1, y2, y3, The first one is: if y1>P>y2, it is the standard state and no feedback is given; The second type is: if y2>P>y3, it is a qualified state; no feedback is given, but error information is recorded; The third type is: y3>P, which is an unqualified state, and feedback is given to record the location and type of inaccurate information.
7. The news generation system based on the language big model according to claim 6 is characterized in that: The specific process of generating a judgment analysis report based on the news component compliance status is as follows: Step 1: Analyze the accuracy of news components, obtain the overall news components, and check the key features and information sources in the news components; Step 2: Based on the location and type of inaccurate information in the acquired news component, determine whether it is an error in the main body and news content, an error in the accompanying image, an error in the video, or an error in the data; Step 3: Based on the inaccurate information, the language model is analyzed in the news generation process, and the influence ratio coefficient of each generation method of the generated news component is Step 4: Determine the direction of improvement of the language model based on the size of the impact coefficient.
8. The news generation system based on the language big model according to claim 1 is characterized in that: The judgment generation module also includes an impact improvement module, which is used to obtain an impact ratio coefficient data value, determine the impact ratio of inaccurate information according to the impact ratio coefficient, and optimize and improve the language model being trained.