Data processing method and system based on large language model

By performing word segmentation, part-of-speech annotation and weight assignment on the power Q&A dataset, and pre-training with large language models, the problems of large language model understanding deviation and inaccurate generation in the power Q&A scenario are solved, and higher answer accuracy and fewer answers are achieved.

CN120104758APending Publication Date: 2025-06-06SHANDONG HAILIANXUN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510482833.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

In the power Q&A scenario, the diversity and complexity of the content of the user's questions lead to bias or misunderstanding of the large language model when comprehension, resulting in inaccurate or inappropriate answers.

Method used

A data processing method based on a large language model is provided. By performing word segmentation and part-of-speech annotation of text data in the question and answer data set, preset logical word list and target word list, weight assignment for each word, combining TF-IDF value, logical weight factor and numeric weight factor, the final weight of each word is calculated, and pre-trained on the question and answer data set to generate a more accurate question and answer model.

Benefits of technology

By enhancing the semantic understanding of the text, reducing logical errors and numerical errors, improving the model's attention to the target words, significantly improving the accuracy of the generation of the question-and-answer model and reducing the phenomenon of answering non-questions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120104758A_ABST
    Figure CN120104758A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing, in particular to a data processing method and system based on a large language model.The method comprises the steps that word segmentation and part-of-speech tagging are conducted on text data in a question and answer data set; presetting a logic word list to obtain a logic weight factor of each word; obtaining a digital weight factor of each word; obtaining a TF-IDF value of each word to obtain a first weight of each word; presetting a target word list to obtain a target expression capability value of each word; obtaining a target word weight factor of each word, and obtaining a second weight of each word; confirming the final weight of each word; and pre-training the question and answer data set, and combining with the final weight to obtain a question and answer model. The invention aims to improve the answering accuracy of the question and answer model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of natural language processing, and in particular to a data processing method and system based on a large language model. Background Art

[0002] With the development of natural language processing technology, large language models are applied to a variety of task scenarios. Large language models can be used to achieve tasks such as intelligent question answering, text generation, and text classification. By introducing natural language processing technology into the field of intelligent question answering, the question answering model can have more powerful cognitive and decision-making capabilities.

[0003] With the continuous development of the Internet and the increasing demand for power customer service, manual customer service can no longer safely meet the needs of large-scale power users. With the help of a large language model, the intelligent question-and-answer assistant can quickly understand and answer users' questions, and provide personalized answers and solutions, improving the efficiency of power services, and becoming an emerging service method to meet user needs. Since the data sources in the power question-and-answer scenario are wide and complex, they often contain a lot of noise and errors. If the large language model is trained directly using this data, it may result in inaccurate or inappropriate answers. In addition, the content of user questions is diverse and complex, which can easily cause the model to have deviations or misunderstandings when understanding, resulting in incorrect answers. Summary of the invention

[0004] In view of the above, it is necessary to provide a data processing method and system based on a large language model to solve the above problems.

[0005] The first aspect of the present application provides a data processing method based on a large language model, the method comprising: Perform word segmentation and part-of-speech tagging on the text data in the question-answering dataset; Preset a logical word list, judge each word in the text data, and obtain the logical weight factor of each word; obtain the digital weight factor of each word according to the part-of-speech tagging results of each word and adjacent words in the text data; obtain the TF-IDF value of each word, combine the logical weight factor and the digital weight factor, and obtain the first weight of each word; Based on the preset target word list, each word in the text data and its adjacent words are subjected to similarity analysis to obtain the target expression capability value of each word; for each word, the interval between it and the words in the sentence whose target expression capability value is greater than the preset threshold is analyzed to obtain the target word weight factor of each word, and the second weight of each word is obtained by combining the TF-IDF value and the target expression capability value of each word; Based on the first weight of each word and the magnitude relationship of the TF-IDF value, determining the final weight of each word in the first weight and the second weight; Pre-train the question-answering dataset and combine it with the final weights to obtain the question-answering model.

[0006] Preferably, the logical weight factor of each word is obtained as follows: When each word is an element in a preset logical word list, the first preset value is used as the logical weight factor of each word; otherwise, the second preset value is used as the logical weight factor of each word; wherein the first preset value is greater than the second preset value, and the second preset value is greater than or equal to 1.

[0007] Preferably, the digital weight factor of each word is obtained as follows: When the part-of-speech tagging results of each word and its previous word are not numbers, the third preset value is used as the digital weight factor of each word; when the part-of-speech tagging result of each word is not a number, and the part-of-speech tagging result of its previous word is a number, the fourth preset value is used as the digital weight factor of each word; otherwise, the fifth preset value is used as the digital weight factor of each word; wherein the fifth preset value is greater than the fourth preset value, the fourth preset value is greater than the third preset value, and the third preset value is greater than or equal to 1.

[0008] Preferably, the first weight of each word is specifically a result of forward fusion of the TF-IDF value, the logical weight factor and the digital weight factor of each word.

[0009] Preferably, the purpose expression capability value of each word is obtained as follows: Record each word and the combination of its preceding and following words as the phrase list of each word; for each word in the phrase list, calculate the maximum text editing distance between each word in the phrase list and each word in the preset target word list as the target expression ability value of each word.

[0010] Preferably, the target word weight factor of each word is obtained as follows: The minimum interval between each word and the words in the sentence where the word is located and whose target expression ability is greater than a preset threshold is obtained; when the minimum interval is less than the preset interval threshold, the negative correlation mapping result of the minimum interval is used as the target word weight factor of each word.

[0011] Preferably, the second weight of each word is specifically the product of the TF-IDF value of each word, the target word weight factor and the target expression ability value.

[0012] Preferably, obtaining the final weight of each word includes: when the first weight of each word is greater than the TF-IDF value of the corresponding word, using the first weight as the final weight of the corresponding word; otherwise, using the second weight as the final weight of each word.

[0013] Preferably, the process of pre-training the question-answering dataset and obtaining the question-answering model in combination with the final weight is specifically as follows: Load the BERT pre-trained model from the transformers library, use the question-answering dataset and the pre-trained model as the input of the BERT model, concatenate the final weight of each word with the word embedding vector output by the BERT model in the embedding layer to obtain a new word embedding vector, and send it to the Transformer encoder to obtain the question-answering model.

[0014] In a second aspect, an embodiment of the present application also provides a data processing system based on a large language model, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor implements the steps of any one of the above methods when executing the computer program.

[0015] The above scheme aims to solve the problem of inaccurate generation results of the electric power question-answering model, and proposes a data processing method and system based on a large language model. First, a list of logical words is preset, and each word in the text data is judged to obtain the logical weight factor of each word, which helps to enhance the semantic understanding of the text, provide a data basis for subsequent logical error judgment, and help the model capture key relationships in the text, thereby optimizing the task execution effect; according to the part-of-speech tagging results of each word and adjacent words in the text data, a numerical weight factor for each word is obtained to provide a data basis for subsequent judgment of numbers and unit names; the TF-IDF value of each word is obtained, and the first weight of each word is obtained by combining the logical weight factor and the numerical weight factor. By analyzing the phenomenon of logical errors and numerical errors in the answer sentences, the words representing logic and numbers in the data set are given higher weights, so that the large language model can pay more attention to these words and reduce the response time. The logical errors and numerical errors in the answer sentences improve the generation accuracy of the model; then, based on the preset target word list, a similarity analysis is performed on each word and its adjacent words in the text data to obtain the target expression ability value of each word, effectively quantifying the ability of each word in expressing the target concept; for each word, the interval between it and the words in the sentence whose target expression ability value is greater than the preset threshold is analyzed to obtain the target word weight factor of each word, and the second weight of each word is obtained by combining the TF-IDF value and the target expression ability value of each word; by analyzing the phenomenon of irrelevant answers in the answer sentences, the words representing the purpose in the dataset are given a higher weight, so that the large language model pays more attention to the target words, reducing the phenomenon of irrelevant answers in the generated results and improving the model's answer accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A flowchart of a data processing method based on a large language model provided in one embodiment of the present application; Figure 2 A flowchart for obtaining the final weight is provided for one embodiment of the present application. DETAILED DESCRIPTION

[0017] In the description of the embodiments of the present application, words such as "exemplary", "or", "for example" and the like are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of words such as "exemplary", "or", "for example" and the like is intended to present related concepts in a concrete manner.

[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art in the present application. The terms used in the specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application.

[0019] It should also be noted that the terms "first" and "second" in this application and the accompanying drawings are used to distinguish similar objects, rather than to describe a specific order or sequence. The method disclosed in the embodiments of the present application or the method shown in the flow chart includes one or more steps for implementing the method. Without departing from the scope of protection of the present application, the execution order of multiple steps can be interchanged with each other, and some steps can also be deleted.

[0020] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs.

[0021] The specific scheme of the data processing method and system based on the large language model provided by the present application is described in detail below with reference to the accompanying drawings.

[0022] See also Figure 1 , which shows a flowchart of a data processing method based on a large language model provided by an embodiment of the present application, the method comprising the following steps: The first step is to segment and tag the text data in the question-answering dataset.

[0023] We obtain conversations between users and customer service from the power customer service system and construct a power question-and-answer dataset. The constructed question-and-answer dataset contains a total of 500,000 conversations. These conversations are in Chinese format and cover a variety of power scenarios and topics.

[0024] In the power data set, there are noises such as typos, spaces, and line breaks. The TextBlob library is used to clean and denoise the text, and the Jieba toolkit is used to segment and tag the text data. The text data is divided into a training set and a test set according to a preset ratio. In this embodiment, the preset ratio is 8:2. The implementer can adjust the ratio according to the actual situation, and this application does not limit this.

[0025] The second step: preset a logical word list, judge each word in the text data, and obtain the logical weight factor of each word; obtain the digital weight factor of each word based on the part-of-speech tagging results of each word and adjacent words in the text data; obtain the TF-IDF value of each word, combine the logical weight factor and the digital weight factor to obtain the first weight of each word.

[0026] In the field of power question answering, large language models generate answers based on the information provided by users through a series of retrieval, matching, and reasoning processes. However, since large language models do not have true understanding capabilities, they may generate content that appears reasonable on the surface but is actually inaccurate or wrong.

[0027] In the question-and-answer dataset of this application, each question and answer consists of two parts: a question and an answer. For example: Question 1: "Question: If the electricity consumption of a data center is expected to exceed 750 billion kWh in 2030, accounting for 5%-10% of the total electricity consumption in society, what trend does this reflect? Answer: This reflects the rapid growth of electricity consumption in data centers and the urgent need for digital transformation of power systems. One of the existing methods to measure the importance of a word is the term frequency-inverse document frequency (TF-IDF) method. This method uses TF to represent the term frequency and IDF to represent the rarity of the term in all documents. The more frequently a word appears in a document and the rarer it is in all documents, the more important the word is. However, this method does not take into account the importance of some special words to the text, such as logical words, which may appear frequently in a specific document or have a high frequency of appearance in the entire document set, so their TF-IDF values ​​will be lower. But in some fields, such as power documents, logical words are crucial to the understanding of the text. Once these words are wrong, it may cause significant deviations in the semantics of the text.

[0028] When training the power question-answering system, in order to make the model answer more accurately, it is necessary to improve the common inaccurate phenomena of the power question-answering system. Common inaccurate phenomena of the generated results include opposite description logic, numerical errors, and irrelevant answers. Among them, opposite description logic means that the answer sentence contains descriptions that are contrary to the actual situation or principle, such as answering "rapid growth" as "rapid reduction"; numerical errors mean that the numbers in the answer sentence do not match the actual situation, such as answering "exceeding 750 billion kWh in 2030" as "exceeding 750 billion kWh in 2020"; irrelevant answers mean that the content of the answer does not directly respond to the core of the question, such as "Question: What parameters need to be considered when selecting a high-voltage circuit breaker? Answer: High-voltage circuit breakers play an important control and protection role in the power system."

[0029] For phenomena that describe opposite logics, the words that represent logic in the data set can be given a higher weight. A preset logic word list A, wherein the elements in the logic word list A include but are not limited to lower than, higher than, greater than, less than, better than, stronger, weaker, larger, smaller, higher, lower, rising, and falling. For each word, when each word is an element in the preset logic word list, the first preset value is used as the logic weight factor of each word; otherwise, the second preset value is used as the logic weight factor of each word. In this embodiment, the first preset value is greater than the second preset value, and the second preset value is greater than or equal to 1, wherein the first preset value is 2 and the second preset value is 1.

[0030] For the phenomenon of digital errors, the numbers in the data set can be given a higher weight. In the process of part-of-speech tagging by the jieba toolkit, the number is marked as "m", and the part-of-speech tagging is used to determine whether a word is a number; in addition, the number is usually followed by the "unit name", and the "unit name" also plays a vital role in expressing the meaning of the number, so when a word is a "unit name", it should also be given a higher weight. Specifically, when the part-of-speech tagging results of each word and its previous word are not numbers, the third preset value is used as the digital weight factor of each word; when the part-of-speech tagging result of each word is not a number, and the part-of-speech tagging result of the previous word is a number, the fourth preset value is used as the digital weight factor of each word; otherwise, the fifth preset value is used as the digital weight factor of each word; wherein the fifth preset value is greater than the fourth preset value, the fourth preset value is greater than the third preset value, and the third preset value is greater than or equal to 1. In this embodiment, the fifth preset value is 2, the fourth preset value is 1.5, and the third preset value is 1; the implementer can adjust the size of the preset value according to the actual situation.

[0031] The result of forward fusion of the TF-IDF value, the logical weight factor and the digital weight factor of each word is used as the first weight of each word. In this embodiment, the forward fusion of multiple variables adopts a multiplication calculation method.

[0032] It should be understood that when one of the words in the data set is an element in the preset logical word list, its importance is greater and a larger weight should be assigned; when a word is not an element in the preset logical word list, its corresponding weight should not be affected. Furthermore, when one of the words in the data set and its previous word are judged to be numbers, it means that both words are very important. Let their digital weight factor take a value greater than 1, so that the word can obtain a larger first weight; when the previous word of the word is judged to be a number, it means that the word may be a "unit name", so let the digital weight factor take a value greater than 1, so that each word can obtain a larger first weight. It should be noted here that when there is no word in front of a word, only the word is judged; when each word is neither a number nor a unit, let Equal to 1, so that the first weight of each word is not affected.

[0033] The third step: Based on the preset target word list, perform similarity analysis on each word and its adjacent words in the text data to obtain the target expression ability value of each word; for each word, analyze the interval between it and the words in the sentence whose target expression ability value is greater than the preset threshold, and obtain the target word weight factor of each word, and combine the TF-IDF value and the target expression ability value of each word to obtain the second weight of each word.

[0034] For the situation where the answer is irrelevant to the question, it is because the model does not have enough understanding of the question statement, and does not pay enough attention to certain keywords in the question statement that represent the purpose. For example, when the user wants to ask the "cause" of a certain power failure, the answer obtained is "solution". Therefore, this application considers enhancing the words representing the purpose in the question statement. Common question statements such as "How is the frequency of the power system adjusted?", "What are the main functions of the smart grid?", "What is the rated power of the transformer?", etc. Therefore, a purpose word list B is preset, wherein the purpose word list B includes but is not limited to "how", "how", "how much", "is it", and "what". When segmenting the data set, some purpose words will be divided into two words, such as "is it" will be divided into "yes" and "no", so when judging whether a word is a purpose word, it is also necessary to judge whether the combination of the word and its adjacent words is a purpose word.

[0035] When the target word appears in the question sentence, in order to make the target word get more attention, it should be given a higher weight. In addition, in addition to the target word, other words that are close to the target word will also make a greater contribution to the expression purpose of the sentence, such as "What is the principle of this" and "What is the fault of this device", both sentences contain "what", but the former focuses on "principle" and the latter focuses on "fault", so the words that are close to the target word should also be given a greater weight.

[0036] In this application, the i-th word in the data set is taken as an example for analysis, and the combination of the i-th word and its adjacent words before and after it is recorded as the phrase list of the i-th word , its formula form is: ,in, , , Respectively represent the i-1th, i-th, and i+1th words in the data set. Indicated by and The phrases that make up Indicated by and Next, determine the semantic similarity between each word in list R and each word in target word list B: for each word in the phrase list, calculate the maximum text editing distance between each word in the phrase list and each word in the preset target word list as the target expression ability value of each word.

[0037] It should be understood that the larger the value of the purpose expression ability value is, the more similar each word is to the words in the purpose word list, and the stronger the purpose expression ability value of each word is.

[0038] For each word, analyze the interval between it and the words in the sentence where the target expression ability value is greater than the preset threshold, obtain the target word weight factor of each word, and combine the TF-IDF value and the target expression ability value of each word to obtain the second weight of each word: obtain the minimum interval between each word and the words in the sentence where the target expression ability is greater than the preset threshold; when the minimum interval is less than the preset interval threshold, use the negative correlation mapping result of the minimum interval as the target word weight factor of each word; use the positive fusion result of each word's TF-IDF value, target word weight factor and target expression ability value as the second weight of each word.

[0039] In this embodiment, the target word weight factor of the i-th word is recorded as , its formula form is: ;in, is the target word weight factor of the i-th word; g is the minimum interval between the i-th word and the word whose target expression ability value in the current sentence is greater than the preset threshold n, g is an integer and , n represents the preset threshold for separating the target word from the common word, and in this embodiment, the value is 1; for example, when the i-th word itself is a word with a target expression ability value greater than n, g=0; when the word adjacent to the i-th word is a word with a target expression ability value greater than n, g=1; t is the interval threshold. When the value of g exceeds the interval threshold, it means that the i-th word is far away from the target word and cannot express the purpose of the sentence. The value range of t is an integer in [2,4], and in this embodiment, the value is 2.

[0040] Furthermore, the product of the TF-IDF value of each word, the target word weight factor and the target expression ability value is used as the second weight of each word.

[0041] It should be understood that when the minimum interval between the i-th word and the word with a target expression ability value greater than n is smaller, the i-th word is more able to express the purpose of the sentence, and the word should be given a greater weight. When the minimum interval between the i-th word and the target word is greater than or equal to the interval threshold, it means that the i-th word is far away from the word with a target expression ability value greater than n, and usually cannot express the purpose of the sentence, so in this case, , so that the second weight of the i-th word no longer increases.

[0042] The fourth step: based on the first weight of each word and the size relationship of the TF-IDF value, confirm the final weight of each word in the first weight and the second weight.

[0043] For the i-th word, first calculate the first weight of the word. When the first weight is greater than the TF-IDF value, it means that the i-th word is one of the "logical words, numbers, units", then the word can no longer be the target word, so the first weight is used as the final weight of the i-th word; when the first weight is equal to the TF-IDF value, it means that the i-th word is neither a logical word, nor a number or a unit, then it is necessary to further judge the i-th word, calculate the second weight of the i-th word, and use the second weight as the final weight of the i-th word; finally, the final weight of the i-th word is obtained.

[0044] Traverse all the words in the data set in turn, calculate the final weight of each word, and obtain the final weight of each word. The final weight acquisition flow chart is as follows Figure 2 shown. Figure 2 In , G1i and G2i represent the first weight and second weight of the i-th word respectively; Ci represents the TF-IDF value of the i-th word.

[0045] The fifth step: pre-train the question-answering dataset and combine it with the final weights to obtain the question-answering model.

[0046] Load the BERT pre-trained model from the transformers library, use the question-answering dataset and the pre-trained model as the input of the BERT model, and concatenate the final weight of each word with the word embedding vector output by the BERT model in the embedding layer to obtain a new word embedding vector. Send the new word embedding vector to the Transformer encoder to complete the model training task. During the training process, the initial learning rate of the BERT model is set to 0.0001, the batch size is set to 32, the number of training rounds is 1000, and the optimizer uses Adam. The training of the model is a well-known technology, and the specific process will not be repeated here. By concatenating the final weight of each word with the word embedding vector of each word, the model pays more attention to some words (logical words, numbers, units, and purpose words), which improves the accuracy of the question-answering model.

[0047] After the training is completed, a question-answering model in the power field is obtained. When a question statement is input into the question-answering model, the question-answering model can output the corresponding answer statement.

[0048] Based on the same inventive concept as the above method, an embodiment of the present application also provides a data processing system based on a large language model, including a memory, a processor, and a computer program stored in the memory and running on the processor, and when the processor executes the computer program, the steps of any one of the above-mentioned data processing methods based on a large language model are implemented.

[0049] In summary, this application proposes a data processing method and system based on a large language model to address the problem of inaccurate generation results of the electric power question-answering model. First, a list of logical words is preset, and each word in the text data is judged to obtain the logical weight factor of each word, which helps to enhance the semantic understanding of the text, provide a data basis for subsequent logical error judgments, and help the model capture key relationships in the text, thereby optimizing the task execution effect; based on the part-of-speech tagging results of each word and adjacent words in the text data, a digital weight factor for each word is obtained to provide a data basis for subsequent judgments of numbers and unit names; the TF-IDF value of each word is obtained, and the first weight of each word is obtained by combining the logical weight factor and the digital weight factor. By analyzing the phenomenon of logical errors and numerical errors in the answer sentences, the words representing logic and numbers in the data set are given higher weights, so that the large language model can pay more attention to these words and reduce the response time. The logical errors and numerical errors in the answer sentences improve the generation accuracy of the model; then, based on the preset target word list, a similarity analysis is performed on each word and its adjacent words in the text data to obtain the target expression ability value of each word, effectively quantifying the ability of each word in expressing the target concept; for each word, the interval between it and the words in the sentence whose target expression ability value is greater than the preset threshold is analyzed to obtain the target word weight factor of each word, and the second weight of each word is obtained by combining the TF-IDF value and the target expression ability value of each word; by analyzing the phenomenon of irrelevant answers in the answer sentences, the words representing the purpose in the dataset are given a higher weight, so that the large language model pays more attention to the target words, reducing the phenomenon of irrelevant answers in the generated results and improving the model's answer accuracy.

[0050] The flowchart and block diagram in the accompanying drawings show the possible architecture, function and operation of the system, method and computer program product according to the embodiment of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the function marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two continuous boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. In the description corresponding to the flowchart and the block diagram in the accompanying drawings, the operations or steps corresponding to different boxes can also occur in a different order from the order disclosed in the description, and sometimes there is no specific order between different operations or steps. For example, two continuous operations or steps can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, which can depend on the functions involved. Each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented by a dedicated hardware-based system that performs the specified functions or actions, or may be implemented by a combination of dedicated hardware and computer instructions.

[0051] For those skilled in the art, it is obvious that the present application is not limited to the details of the above exemplary embodiments, and the present application can be implemented in other specific forms without departing from the basic features of the present application. Therefore, no matter from which point of view, the above embodiments of the present application should be regarded as exemplary and non-restrictive; the technical solutions recorded in the above embodiments are modified, or some of the technical features are replaced by equivalents, which does not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. A data processing method based on a large language model, characterized in that: The method comprises the following steps: Perform word segmentation and part-of-speech tagging on the text data in the question-answering dataset; Preset a logical word list, judge each word in the text data, and obtain the logical weight factor of each word; obtain the digital weight factor of each word according to the part-of-speech tagging results of each word and adjacent words in the text data; obtain the TF-IDF value of each word, combine the logical weight factor and the digital weight factor, and obtain the first weight of each word; Based on the preset target word list, each word in the text data and its adjacent words are subjected to similarity analysis to obtain the target expression capability value of each word; for each word, the interval between it and the words in the sentence whose target expression capability value is greater than the preset threshold is analyzed to obtain the target word weight factor of each word, and the second weight of each word is obtained by combining the TF-IDF value and the target expression capability value of each word; Based on the first weight of each word and the magnitude relationship of the TF-IDF value, determining the final weight of each word in the first weight and the second weight; Pre-train the question-answering dataset and combine it with the final weights to obtain the question-answering model.

2. The data processing method based on a large language model according to claim 1, characterized in that: The logical weight factor of each word is obtained as follows: When each word is an element in a preset logical word list, the first preset value is used as the logical weight factor of each word; otherwise, the second preset value is used as the logical weight factor of each word; wherein the first preset value is greater than the second preset value, and the second preset value is greater than or equal to 1.

3. The data processing method based on a large language model according to claim 1, characterized in that: The digital weight factor of each word is obtained as follows: When the part-of-speech tagging results of each word and its previous word are not numbers, the third preset value is used as the digital weight factor of each word; when the part-of-speech tagging result of each word is not a number, and the part-of-speech tagging result of its previous word is a number, the fourth preset value is used as the digital weight factor of each word; Otherwise, the fifth preset value is used as the digital weight factor of each word; wherein the fifth preset value is greater than the fourth preset value, the fourth preset value is greater than the third preset value, and the third preset value is greater than or equal to 1.

4. The data processing method based on a large language model according to claim 1, characterized in that: The first weight of each word is specifically the result of forward fusion of the TF-IDF value, the logical weight factor and the digital weight factor of each word.

5. The data processing method based on a large language model according to claim 1, characterized in that: The purpose expression capability value of each word is obtained as follows: Record each word and the combination of its preceding and following words as the phrase list of each word; for each word in the phrase list, calculate the maximum text editing distance between each word in the phrase list and each word in the preset target word list as the target expression ability value of each word.

6. The data processing method based on a large language model according to claim 1, characterized in that: The target word weight factor of each word is obtained as follows: The minimum interval between each word and the words in the sentence where the word is located and whose target expression ability is greater than a preset threshold is obtained; when the minimum interval is less than the preset interval threshold, the negative correlation mapping result of the minimum interval is used as the target word weight factor of each word.

7. The data processing method based on a large language model according to claim 1, characterized in that: The second weight of each word is specifically the product of the TF-IDF value of each word, the target word weight factor and the target expression ability value.

8. The data processing method based on a large language model according to claim 1, characterized in that: The obtaining of the final weight of each word includes: when the first weight of each word is greater than the TF-IDF value of the corresponding word, taking the first weight as the final weight of the corresponding word; otherwise, taking the second weight as the final weight of each word.

9. The data processing method based on a large language model according to claim 1, characterized in that: The process of pre-training the question-answering dataset and obtaining the question-answering model by combining the final weights is as follows: Load the BERT pre-trained model from the transformers library, use the question-answering dataset and the pre-trained model as the input of the BERT model, concatenate the final weight of each word with the word embedding vector output by the BERT model in the embedding layer to obtain a new word embedding vector, and send it to the Transformer encoder to obtain the question-answering model.

10. A data processing system based on a large language model, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 9 are implemented.