Data processing method and terminal
By determining the model word element information of the input text in the big model, formulating processing strategies, and using global relative position characteristics to control the self-attention mechanism, the information loss problem when the big model is processed in a long text is solved, and the coherence and accuracy of the generated results are achieved.
Patent Information
- Application Number
- CN202411775315.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-04
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-12-04
AI Technical Summary
When large models process longer text, they may cause information loss or processing errors because longer text is beyond the processing scope of large models.
By determining the model word element information of the input text, a processing strategy is formulated, and when the processing strategy is the first processing strategy, the unit unit information value of the vocabulary unit is used to perform streamlined compression processing to generate the target input text. At the same time, the self-attention mechanism of the large model is controlled by using global relative position features to capture long-distance dependencies in the text.
It effectively avoids the loss of information caused by input text beyond the processing scope of the large model, ensures the consistency, consistency and accuracy of the generated results, and can accurately reason and restore deleted information.
Smart Images

Figure CN119250072B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a data processing method and terminal. Background Art
[0002] In recent years, with the rapid development of artificial intelligence, large models have been increasingly used in natural language processing. When large models process longer texts, the longer texts may exceed the processing range of the large models, resulting in information loss or errors in the model processing of the large models. Summary of the invention
[0003] The embodiments of the present application provide a data processing method and terminal, which can solve the technical problem that when a large model processes a longer text, the longer text may exceed the processing range of the large model, thereby causing information loss or errors in the model processing process of the large model.
[0004] In a first aspect, an embodiment of the present application provides a data processing method, the method comprising:
[0005] Determine an input text for a target large model and model word unit number information of the input text, and determine a processing strategy for the input text based on the model word unit number information;
[0006] When the processing strategy is the first processing strategy, determining the vocabulary units of the input text and the unit self-information values of the vocabulary units, determining the target vocabulary units from the vocabulary units by the unit self-information values, and generating the target input text corresponding to the target vocabulary units;
[0007] The global relative position feature of the target input text is determined, and the target large model is controlled to process the target input text using the global relative position feature to obtain a first model output result.
[0008] Optionally, the determining of the vocabulary units of the input text and the unit self-information values of the vocabulary units comprises:
[0009] Performing vocabulary unit division processing on the input text to obtain vocabulary units of the input text;
[0010] Determine a word-gram sequence corresponding to the vocabulary unit, and calculate the word-gram self-information value of each word-gram in the word-gram sequence; wherein the word-gram self-information value is used to characterize the information content of the word-gram;
[0011] Based on the word-gram self-information value of each word-gram in the word-gram sequence, the unit self-information value of the vocabulary unit corresponding to the word-gram sequence is determined.
[0012] Optionally, calculating the word-gram self-information value of each word-gram in the word-gram sequence includes:
[0013] Calculating the word-unit self-information value of each word-unit in the word-unit sequence using the first calculation formula;
[0014] The first calculation formula satisfies the following formula:
[0015] I(T i ) = -log2P(T i |C(T0~T N ,T i ));
[0016] Among them, T i is the word with sequence number i in the word sequence, N≥i≥0 and i is an integer, and N+1 is the total number of words in the word sequence; T0~T N is the word sequence from T0 to T N All words of C(T0~T N ,T i ) is T0 to T N Remove T from all the words i The word obtained after i ) is T i The word self-information value of P(T i |C(T0~T N ,T i )) is T0 to T N Remove T from all the words i Under the condition that the word generation event corresponding to the remaining word occurs, T i The occurrence probability value of the corresponding word generation event;
[0017] The step of determining the unit self-information value of the vocabulary unit corresponding to the word-gram sequence based on the word-gram self-information value of each word-gram in the word-gram sequence includes:
[0018] The word unit self-information value of each word unit in the word unit sequence is accumulated by using the second calculation formula to obtain the unit self-information value of the vocabulary unit corresponding to the word unit sequence;
[0019] The second calculation formula satisfies the following formula:
[0020] I(u) = ;
[0021] Wherein, u is the vocabulary unit, and I(u) is the unit self-information value of u.
[0022] Optionally, determining a target vocabulary unit from the vocabulary units by using the unit self-information value comprises:
[0023] Sorting the unit self-information values of the vocabulary units in the input text to obtain a unit self-information value ranking;
[0024] Obtaining a preset screening parameter and the total number of vocabulary units of the input text, determining a reference screening rank by using the preset screening parameter and the total number of vocabulary units, and searching for a reference unit self-information value corresponding to the reference screening rank from the unit self-information value ranking; wherein the preset screening parameter is greater than 0 and less than 1;
[0025] The unit self-information value in the vocabulary unit that is greater than or equal to the reference unit self-information value is taken as the target unit self-information value, and the target vocabulary unit corresponding to the target unit self-information value is determined from the vocabulary unit.
[0026] Optionally, determining the global relative position feature of the target input text includes:
[0027] Inputting the target input text into the target macro model, and using the target macro model to determine a model word-unit sequence corresponding to the target input text;
[0028] The relative position feature of each target word in the model word sequence is determined, and the global relative position feature of the target input text is obtained based on the relative position feature of each target word.
[0029] Optionally, the using the global relative position feature to control the target large model to process the target input text to obtain a first model output result includes:
[0030] Determine the word embedding features corresponding to each target word in the model word sequence using the target large model;
[0031] Integrate the relative position feature corresponding to the target word into the word embedding feature corresponding to the target word to update the word embedding feature;
[0032] Based on the target large model, a model attention score corresponding to the word embedding feature is determined, and a first model output result is generated based on the model attention score and the word embedding feature.
[0033] Optionally, the determining of the input text for the target large model and model word unit number information of the input text, and determining a processing strategy for the input text based on the model word unit number information, includes:
[0034] Acquire an input text for a target large model, determine the number of input word units and the number of predicted output word units of the input text for the target large model, and obtain model word unit number information based on the number of input word units and the number of predicted output word units;
[0035] Obtaining an input word unit number threshold of the target large model, determining a proportional parameter corresponding to the input word unit number threshold, and obtaining a length parameter value based on a product of the input word unit number threshold and the proportional parameter, wherein the proportional parameter is greater than 0 and less than 1;
[0036] When the number of input word units and the number of predicted output word units are both greater than the length parameter value, it is determined that the processing strategy for the input text is the first processing strategy.
[0037] Optionally, the method further comprises:
[0038] When the number of input word units and the number of predicted output word units are both less than or equal to the length parameter value, and the sum of the number of input word units and the number of predicted output word units is greater than or equal to the input word unit threshold, and the number of input word units is greater than the predicted output word unit number, determining that the processing strategy for the input text is the second processing strategy;
[0039] When the processing strategy is the second processing strategy, determining the vocabulary unit of the input text and the unit self-information value of the vocabulary unit, determining the target vocabulary unit from the vocabulary unit by the unit self-information value, and determining the target input text corresponding to the target vocabulary unit;
[0040] The target input text is input into the target macro model, and the target input text is processed based on the target macro model to obtain a second model output result.
[0041] Optionally, the method further comprises:
[0042] When the number of input word units and the number of predicted output word units are both less than or equal to the length parameter value, and the sum of the number of input word units and the number of predicted output word units is greater than or equal to the input word unit threshold, and the number of input word units is less than or equal to the predicted output word unit number, determining that the processing strategy for the input text is a third processing strategy;
[0043] When the processing strategy is the third processing strategy, the original global relative position features of the input text are determined, and the target large model is controlled to process the input text using the original global relative position features to obtain a third model output result.
[0044] In a second aspect, an embodiment of the present application provides a terminal, the terminal comprising:
[0045] Processor; and
[0046] A memory arranged to store computer executable instructions which, when executed, cause the processor to perform any of the methods described above.
[0047] The beneficial effects brought about by the technical solutions provided by some embodiments of the present application include at least:
[0048] By determining the input text for the target large model and the model word number information of the input text, the processing strategy for the input text is determined based on the model word number information, so as to execute the corresponding processing strategy based on the model word number information of the input text, so as to realize that various types of input texts can be processed based on this method. When the processing strategy is the first processing strategy, in order to avoid the input text exceeding the processing range of the target large model and causing the large-scale loss of the context information of the input text, the input text can be simplified and compressed based on the distribution of the unit self-information value of the vocabulary unit, and then the target input text corresponding to the target vocabulary unit is generated, so as to avoid the input text exceeding the processing range of the target large model and causing the large-scale loss of the context information of the input text.
[0049] At the same time, the global relative position feature is used to control the self-attention mechanism of the target large model to effectively capture the long-distance dependencies in the target input text, thereby strengthening the target large model to deepen the learning and understanding of the target input text based on contextual information, so that the target large model can ensure the coherence, consistency and accuracy of the generated first model output results during the processing. In addition, since the target input text is a text that is streamlined and compressed based on the distribution of information in the input text, controlling the self-attention mechanism in the target large model to effectively capture the long-distance dependencies in the text can also enable the target large model to accurately infer and restore the deleted information, and convert the deleted information into known information of the target input text, thereby further strengthening the target large model's learning and understanding of the target input text. Finally, the global relative position features of the target input text are determined, and the global relative position features are used to control the target large model to process the target input text to obtain the first model output result, thereby using the global relative position features to control the self-attention mechanism of the target large model to effectively capture the long-distance dependencies in the target input text, thereby strengthening the target large model to deepen the learning and understanding of the target input text based on contextual information, so that the target large model can ensure the coherence, consistency and accuracy of the generated first model output results during the processing process, thereby solving the technical problem that longer texts exceed the processing range of the large model, resulting in information loss or errors in the model processing of the large model. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.
[0051] Figure 1 An exemplary system architecture diagram of a data processing method provided in an embodiment of the present application;
[0052] Figure 2 A flowchart of a data processing method provided in an embodiment of the present application;
[0053] Figure 3 A schematic diagram of a process for determining a unit self-information value of a vocabulary unit provided in an embodiment of the present application;
[0054] Figure 4 A schematic diagram of a process for determining a target vocabulary unit provided in an embodiment of the present application;
[0055] Figure 5 A schematic diagram of a process for determining a global relative position feature provided in an embodiment of the present application;
[0056] Figure 6 A schematic diagram of a process for determining an output result of a first model provided in an embodiment of the present application;
[0057] Figure 7 A structural intention of a data processing device provided in an embodiment of the present application;
[0058] Figure 8 A schematic diagram of the structure of a terminal provided in an embodiment of the present application. DETAILED DESCRIPTION
[0059] In order to make the features and advantages of the embodiments of the present application more obvious and easy to understand, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of them. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the embodiments of the present application.
[0060] In recent years, with the rapid development of artificial intelligence, data generation and data processing for office documents have begun to rely on collaborative processing by intelligent agents. An intelligent agent refers to a software, system or entity that can perceive the environment, make decisions and take actions. Some intelligent agents will generate corresponding codes in the process of collaborative data generation and data processing of office documents. As a part of the intelligent agent, the large model will face longer prompt word texts when generating codes for targeted data generation and data processing of office documents. When the large model processes longer texts, the longer texts may exceed the maximum input length of the large model, that is, the context length. Due to the limited memory of the model, the longer texts may exceed the processing range of the large model, resulting in information loss or errors in the processing of the model during the large model processing.
[0061] In order to solve the existing technical problems, the embodiment of the present application provides a data processing method, which includes: determining the input text for the target large model and the model word number information of the input text, and determining the processing strategy for the input text based on the model word number information; when the processing strategy is the first processing strategy, determining the vocabulary unit of the input text and the unit self-information value of the vocabulary unit, determining the target vocabulary unit from the vocabulary unit through the unit self-information value, and generating the target input text corresponding to the target vocabulary unit; determining the global relative position feature of the target input text, and using the global relative position feature to control the target large model to process the target input text, and obtain the first model output result. Thereby solving the technical problem that the longer text exceeds the processing range of the large model, resulting in information loss or errors in the processing of the model processing of the large model.
[0062] See also Figure 1 , Figure 1 An exemplary system architecture diagram of a data processing method provided in an embodiment of the present application.
[0063] like Figure 1 As shown, the system architecture may include a terminal 101, a network 102, and a server 103. The network 102 is used to provide a medium for a communication link between the terminal 101 and the server 103. The network 102 may include various types of wired communication links or wireless communication links, for example, the wired communication link includes an optical fiber, a twisted pair, or a coaxial cable, and the wireless communication link includes a Bluetooth communication link, a Wi-Fi (Wireless-Fidelity) communication link, or a microwave communication link.
[0064] The terminal 101 can interact with the server 103 through the network 102 to receive messages from the server 103 or send messages to the server 103, or the terminal 101 can interact with the server 103 through the network 102 to receive messages or data sent by other users to the server 103. The terminal 101 can be hardware or software. When the terminal 101 is hardware, it can be various terminals, including but not limited to smart watches, smart phones, tablet computers, laptop portable computers and desktop computers. When the terminal 101 is software, it can be installed in the terminals listed above, which can be implemented as multiple software or software modules (for example: used to provide distributed services), or it can be implemented as a single software or software module, which is not specifically limited here.
[0065] The server 103 may be a business server that provides various services. It should be noted that the server 103 may be hardware or software. When the server 103 is hardware, it may be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server 103 is software, it may be implemented as multiple software or software modules (for example, for providing distributed services), or as a single software or software module, which is not specifically limited here.
[0066] In an embodiment of the present application, terminal 101 can determine the input text for the target large model and the model word number information of the input text, and determine the processing strategy for the input text based on the model word number information; when the processing strategy is the first processing strategy, determine the vocabulary units of the input text and the unit self-information values of the vocabulary units, determine the target vocabulary units from the vocabulary units through the unit self-information values, and generate the target input text corresponding to the target vocabulary units; determine the global relative position features of the target input text, and use the global relative position features to control the target large model to process the target input text to obtain the first model output result.
[0067] It should be understood that the number of the above terminals, networks and servers is only illustrative, and any number of terminals, networks and servers may be required according to implementation requirements. Of course, in an embodiment provided by the present application, the system architecture may not include a server.
[0068] See also Figure 2 , Figure 2 A flowchart of a data processing method provided for an embodiment of the present application. The execution subject of the embodiment of the present application can be a terminal that executes the data processing method, or a processor in the terminal that executes the data processing method, or a data processing service in the terminal that executes the data processing method. For the convenience of description, the specific execution process of the data processing method is introduced below by taking the execution subject as the processor in the terminal as an example.
[0069] The data processing method includes:
[0070] S202: Determine the input text for the target large model and the model word number information of the input text, and determine the processing strategy for the input text based on the model word number information.
[0071] Among them, before the target large model receives the input text, the input text is obtained and the processing strategy is matched for the input text. Generally, the processing type of the target large model for the input text can be divided into four categories. The first category is to input the long text into the target large model, and the target large model outputs the corresponding long text after processing. In this case, it is a long text input and long text output type; the second category is to input the long text into the target large model, and the target large model outputs the corresponding short text after processing. In this case, it is a long text input and short text output type; the third category is to input the short text into the target large model, and the target large model outputs the corresponding long text after processing. In this case, it is a short text input and long text output type; the fourth category is to input the short text into the target large model, and the target large model outputs the corresponding short text after processing. In this case, it is a short text input and short text output type.
[0072] Among them, in order to determine the processing type of the target large model for the input text, the word segmentation tool corresponding to the target large model (such as the tokenizer tool) can be used to determine the input taken number corresponding to the input text, that is, the input word number, and obtain the input taken number threshold of the target large model, that is, the input word number threshold. The input word number threshold is the maximum context length corresponding to the target large model. At the same time, the word segmentation tool corresponding to the target large model is used to calculate the number of takens that the target large model needs to generate for the input text and the task corresponding to the input text, that is, the predicted output word number of the target large model, and determine a scale parameter greater than 0 and less than 1. The length parameter value is obtained based on the product of the input word number threshold and the scale parameter. For example, the scale parameter can be 0.7. Of course, the scale parameter page can be other values set based on actual needs, and there is no restriction here.
[0073] When the number of input words and the number of predicted output words are both greater than the length parameter value, the target large model uses the long text input and long text output type for the text; when the number of input words and the predicted output words are both less than or equal to the length parameter value, and the sum of the number of input words and the predicted output words is greater than or equal to the input word number threshold, and the number of input words is greater than the predicted output word number, the target large model uses the long text input and short text output type for the text; when the number of input words and the predicted output words are both less than or equal to the length parameter value, and the sum of the number of input words and the predicted output words is greater than or equal to the input word number threshold, and the number of input words is less than or equal to the predicted output word number, the target large model uses the short text input and long text output type for the text; when the number of input words and the predicted output words are both less than or equal to the length parameter value, and the sum of the number of input words and the predicted output words is less than the input word number threshold, the target large model uses the short text input and short text output type for the text.
[0074] When the processing type of the target large model for the input text is short text input and short text output type, the target large model performs model processing on the input text normally. When the processing type of the target large model for the input text is long text input and long text output type, in order to avoid the input text exceeding the processing range of the target large model and causing large-scale loss of context information of the input text, the input text can be streamlined and compressed based on the distribution of information in the input text, thereby avoiding the input text exceeding the processing range of the target large model and causing large-scale loss of context information of the input text.
[0075] At the same time, since the target large model outputs long texts, in order to achieve high-quality output of long texts, the model usually needs to have a certain degree of language generation ability and context understanding ability. However, the model's language generation ability and context understanding ability often rely on high-quality training data to train the target large model. In order to avoid secondary model training for the target large model, so that the target large model can also achieve high-quality output of long texts, the global relative position features corresponding to the text received by the target large model can be used to control the self-attention mechanism in the target large model to effectively capture the long-distance dependencies in the text, thereby strengthening the target large model's learning and understanding of the text it receives, so that the target large model can ensure the coherence, consistency and accuracy of the generated long text during the processing process.
[0076] Moreover, since the target large model receives text that is simplified and compressed based on the distribution of information in the input text, the self-attention mechanism in the target large model is controlled to effectively capture long-distance dependencies in the text, and the target large model can accurately infer and restore the deleted information, and use the deleted information as the known information of the received text, thereby further strengthening the target large model's learning and understanding of the received text. It should be understood that large models usually have a certain degree of reasoning ability, which does not require additional training. For example, the input text is "1+1=2", and the input text is simplified and compressed to "1+1=". At this time, the basic reasoning ability of the large model can accurately infer and restore the deleted information, that is, "2", and use 2 as the information carried by the received text.
[0077] When the processing type of the target large model for the input text is long text input and short text output type, in order to avoid the input text exceeding the processing range of the target large model and causing large-scale loss of context information of the input text, the input text can be streamlined and compressed based on the distribution of information in the input text, thereby avoiding the input text exceeding the processing range of the target large model and causing large-scale loss of context information of the input text.
[0078] When the target large model processes short text input and long text output for input text, in order to enable the target large model to achieve high-quality output of long text, the target large model can control the self-attention mechanism in the target large model based on the global relative position features corresponding to the input text to effectively capture the long-distance dependencies in the text, thereby strengthening the target large model's learning and understanding of the input text, so that the target large model can ensure the coherence, consistency and accuracy of the generated long text during the processing process.
[0079] The processing strategies for input text include a first processing strategy, a second processing strategy and a third processing strategy. The first processing strategy corresponds to the input text processing type of long text input and long text output type, the second processing strategy corresponds to the input text processing type of long text input and short text output type, and the third processing strategy corresponds to the input text processing type of short text input and long text output type.
[0080] Exemplarily, determining the input text for the target large model and the model word number information of the input text, and determining the processing strategy for the input text based on the model word number information, further includes:
[0081] Obtain input text for the target large model, determine the input word number and predicted output word number of the input text for the target large model, and obtain model word number information based on the input word number and the predicted output word number; obtain the input word number threshold of the target large model, determine the proportional parameter corresponding to the input word number threshold, and obtain the length parameter value based on the product of the input word number threshold and the proportional parameter, where the proportional parameter is greater than 0 and less than 1;
[0082] When the number of input words and the predicted output words are both greater than the length parameter value, the processing strategy for the input text is determined to be the first processing strategy; when the number of input words and the predicted output words are both less than or equal to the length parameter value, and the sum of the number of input words and the predicted output words is greater than or equal to the input word number threshold, and the number of input words is greater than the predicted output word number, the processing strategy for the input text is determined to be the second processing strategy; when the number of input words and the predicted output words are both less than or equal to the length parameter value, and the sum of the number of input words and the predicted output words is greater than or equal to the input word number threshold, and the number of input words is less than or equal to the predicted output word number, the processing strategy for the input text is determined to be the third processing strategy.
[0083] S204: When the processing strategy is the first processing strategy, determine the vocabulary units of the input text and the unit self-information values of the vocabulary units, determine the target vocabulary units from the vocabulary units through the unit self-information values, and generate the target input text corresponding to the target vocabulary units.
[0084] When the processing strategy for the input text is determined to be the first processing strategy based on the model word number information, the input text may be divided into vocabulary units, and the vocabulary units may be phrases or sentences segmented from the input text.
[0085] Specifically, you can use the NLTK (Natural Language Toolkit) sentence tagger to obtain sentence-level vocabulary units and tag the parts of speech of words in the sentence, and then use spaCy to merge the tags into noun phrases. spaCy is a natural language processing library. It is easy to understand that the length of noun phrases is usually shorter than that of verb phrases, so noun phrase segmentation is performed here to segment as many vocabulary units as possible. For example, you can also segment the input text based on punctuation (period, question mark, exclamation mark, etc.), or use NLP (Natural Language Processing) libraries (such as NLTK, spaCy, etc.) to accurately segment the input text. Of course, you can also use machine learning or deep learning models to train a phrase divider to accurately segment the input text.
[0086] The input text has multiple vocabulary units. After the vocabulary units of the input text are determined, the unit self-information value corresponding to the vocabulary unit, that is, the amount of information carried by the vocabulary unit, is calculated. The unit self-information value corresponding to the vocabulary unit can be calculated based on the vocabulary unit information model.
[0087] Exemplarily, the vocabulary unit information model can be obtained by training the basic large model with sample vocabulary units carrying sample unit self-information values. Specifically, the sample vocabulary units carrying sample unit self-information values can be input into the basic large model, and the basic large model performs model processing on the sample vocabulary units to calculate the reference unit self-information values of the sample vocabulary units, and constructs the model loss function of the basic large model based on the parameters corresponding to the reference unit self-information values and the parameters corresponding to the sample unit self-information values, and then inputs the reference unit self-information values and the sample unit self-information values into the model loss function to determine the model loss value, and then updates the model parameters of the basic large model based on the model loss value until the basic large model completes the training to obtain the vocabulary unit information model.
[0088] After determining the unit self-information value of each vocabulary unit in the input text, the vocabulary units with low unit self-information values can be filtered out from the vocabulary units to obtain the target vocabulary units, and then the input text is simplified and compressed based on the information distribution of the input text. Of course, the input text can also be simplified and compressed based on the input word number threshold of the target large model and the information distribution of the input text, so as to control the input word number of the target input text to be less than or equal to the input word number threshold.
[0089] After determining the target vocabulary unit, the text corresponding to the target vocabulary unit in the input text is retained, and the remaining text is deleted, while maintaining the text order of the input text, and the deleted text is automatically replaced by the subsequent text. That is, the relative order relationship between the target vocabulary units is determined based on the input text, and then the target vocabulary units are sorted using the relative order relationship to generate the target input text corresponding to the target vocabulary unit.
[0090] S206: Determine the global relative position feature of the target input text, and use the global relative position feature to control the target large model to process the target input text to obtain a first model output result.
[0091] After determining the target input text, the relative position representation between each character in the target input text and the remaining characters, that is, the relative position feature of each character, is determined, and then the global relative position feature of the target input text is obtained based on the relative position feature of each character.
[0092] After obtaining the global relative position feature of the target input text, the global relative position feature is used to control the self-attention mechanism of the target large model to effectively capture the long-distance dependencies in the target input text, thereby strengthening the target large model to deepen the learning and understanding of the target input text based on contextual information, so that the target large model can ensure the coherence, consistency and accuracy of the generated first model output results during the processing. In addition, since the target input text is a text that is simplified and compressed based on the distribution of information in the input text, controlling the self-attention mechanism in the target large model to effectively capture the long-distance dependencies in the text can also enable the target large model to accurately infer and restore the deleted information, and convert the deleted information into known information of the target input text, thereby further strengthening the target large model's learning and understanding of the target input text.
[0093] In other embodiments provided by the present application, when the processing strategy is the second processing strategy, the vocabulary units of the input text and the unit self-information values of the vocabulary units are determined, the target vocabulary units are determined from the vocabulary units by the unit self-information values, and the target input text corresponding to the target vocabulary units is determined; the target input text is input into the target large model, and the target input text is processed based on the target large model to obtain the second model output result. The specific description here can refer to the relevant description of S202, which will not be repeated here.
[0094] When the processing strategy is the third processing strategy, the original global relative position feature of the input text is determined, and the original global relative position feature is used to control the target large model to process the input text to obtain the third model output result. For a specific description here, please refer to the relevant description of S202, which will not be repeated here.
[0095] The present application provides a data processing method, by determining the input text for the target large model and the model word number information of the input text, and determining the processing strategy for the input text based on the model word number information, so as to execute the corresponding processing strategy based on the model word number information of the input text, so as to realize that the method can process various types of input texts. When the processing strategy is the first processing strategy, in order to avoid the input text exceeding the processing range of the target large model and causing the large-scale loss of the context information of the input text, the input text can be streamlined and compressed based on the distribution of the unit self-information value of the vocabulary unit, and then the target input text corresponding to the target vocabulary unit is generated, so as to avoid the input text exceeding the processing range of the target large model and causing the large-scale loss of the context information of the input text.
[0096] At the same time, the global relative position features are used to control the self-attention mechanism of the target large model to effectively capture the long-distance dependencies in the target input text, thereby strengthening the target large model to deepen the learning and understanding of the target input text based on contextual information, so that the target large model can ensure the coherence, consistency and accuracy of the generated first model output results during the processing process. Moreover, since the target input text is a text that is simplified and compressed based on the distribution of information in the input text, the self-attention mechanism in the target large model is controlled to effectively capture the long-distance dependencies in the text, and the target large model can also accurately infer and restore the deleted information, and the deleted information is the known information of the target input text, thereby further strengthening the target large model's learning and understanding of the target input text. Finally, the global relative position feature of the target input text is determined, and the global relative position feature is used to control the target large model to process the target input text to obtain the first model output result, thereby using the global relative position feature to control the self-attention mechanism of the target large model to effectively capture the long-distance dependencies in the target input text, and then strengthen the target large model to deepen the learning and understanding of the target input text based on context information, so that the target large model can ensure the coherence, consistency and accuracy of the generated first model output result during the processing. This solves the technical problem that a longer text exceeds the processing range of the large model, resulting in information loss or errors in the model processing of the large model.
[0097] See also Figure 3 , Figure 3 A schematic diagram of a flow chart of determining a unit self-information value of a vocabulary unit provided in an embodiment of the present application. Determining the vocabulary unit of the input text and the unit self-information value of the vocabulary unit in S204 includes:
[0098] S302: Perform vocabulary unit division processing on the input text to obtain vocabulary units of the input text.
[0099] When the processing strategy for the input text is determined to be the first processing strategy based on the model word number information, the input text may be divided into vocabulary units, and the vocabulary units may be phrases or sentences segmented from the input text.
[0100] Specifically, you can use machine learning or deep learning models to train a phrase splitter to accurately segment the input text. Of course, you can also use the NLTK sentence tagger to obtain sentence-level vocabulary units and tag the parts of speech of words in the sentence, and then use spaCy, a natural language processing library, to merge the tags into noun phrases. For example, you can also segment the input text based on punctuation, or use an NLP library to accurately segment the input text.
[0101] S304: Determine a word-gram sequence corresponding to the vocabulary unit, and calculate the word-gram self-information value of each word-gram in the word-gram sequence; wherein the word-gram self-information value is used to characterize the information content of the word-gram.
[0102] Among them, after determining the vocabulary unit, the target large model can be used to determine the word-gram sequence corresponding to the vocabulary unit, that is, the taken sequence corresponding to the vocabulary unit. Then the word-gram self-information value of each word-gram in the word-gram sequence is calculated. It should be understood that in a word-gram sequence, the higher the probability that the information carried by the word-gram is inferred by other words in the word-gram sequence, the smaller the word-gram self-information value of the word-gram; the lower the probability that the information carried by the word-gram is inferred by other words in the word-gram sequence, the larger the word-gram self-information value of the word-gram. Therefore, the word-gram self-information value can characterize the amount of information in the word-gram.
[0103] S306: Based on the word-gram self-information value of each word-gram in the word-gram sequence, determine the unit self-information value of the vocabulary unit corresponding to the word-gram sequence.
[0104] Among them, after determining the word-gram self-information value of each word-gram in the word-gram sequence, the word-gram self-information value of each word-gram in the word-gram sequence is accumulated to obtain the unit self-information value of the vocabulary unit corresponding to the word-gram sequence. The unit self-information value of the vocabulary unit can represent the amount of information carried by the vocabulary unit. By determining the unit self-information value of each vocabulary unit corresponding to the input text, the distribution of the amount of information in the input text can be determined.
[0105] In the embodiment provided in the present application, the vocabulary units of the input text are first determined, and then the word-gram self-information value of each word-gram in the word-gram sequence is calculated, and the word-gram self-information value can represent the amount of information of the word-gram. Then, the word-gram self-information value of each word-gram in the word-gram sequence is accumulated to obtain the unit self-information value of the vocabulary unit corresponding to the word-gram sequence, and the unit self-information value of the vocabulary unit can represent the amount of information carried by the vocabulary unit. By determining the unit self-information value of each vocabulary unit corresponding to the input text, the distribution of the amount of information in the input text can be determined.
[0106] Exemplarily, in the embodiment provided in the present application, calculating the word-gram self-information value of each word-gram in the word-gram sequence in S304 includes:
[0107] The first calculation formula is used to calculate the word-unit self-information value of each word-unit in the word-unit sequence;
[0108] The first calculation formula satisfies the following formula:
[0109] I(T i ) = -log2P(T i |C(T0~T N ,T i ));
[0110] Among them, T i is the word with sequence number i in the word sequence, N ≥ i ≥ 0 and i is an integer, N + 1 is the total number of words in the word sequence; T0~T N is the word sequence from T0 to T N All words of C(T0~T N ,T i ) is T0 to T N Remove T from all the words i The word obtained after i ) is T i The word self-information value of P(T i |C(T0~T N ,T i )) is T0 to T N Remove T from all the words i Under the condition that the word generation event corresponding to the remaining word occurs, T i The occurrence probability value of the corresponding word generation event;
[0111] In S306, based on the word-gram self-information value of each word-gram in the word-gram sequence, determining the unit self-information value of the vocabulary unit corresponding to the word-gram sequence includes:
[0112] The second calculation formula is used to accumulate the word-gram self-information value of each word-gram in the word-gram sequence to obtain the unit self-information value of the vocabulary unit corresponding to the word-gram sequence;
[0113] The second calculation formula satisfies the following formula:
[0114] I(u) = ;
[0115] Where u is a vocabulary unit and I(u) is the unit self-information value of u.
[0116] Specifically, when N is 5 and i is 3, then C(T0~T N ,T i ) can be expressed as T0, T1, T2, T4 and T5, P (T i |C(T0~T N ,T i )) can be expressed as P(T3|T0, T1, T2, T4, T5), that is, the probability value of the word generation event corresponding to T3 under the condition that the word generation event corresponding to the remaining words after T3 is removed from all the words from T0 to T5 occurs. It is easy to understand that the larger the probability value of P(T3|T0, T1, T2, T4, T5), the smaller the word self-information value of T3; the smaller the probability value of P(T3|T0, T1, T2, T4, T5), the larger the word self-information value of T3.
[0117] In the embodiment provided by the present application, when the probability that the information carried by a word is inferred by other words in the word sequence is higher, the word-gram self-information value of the word-gram is smaller; when the probability that the information carried by a word-gram is inferred by other words in the word sequence is lower, the word-gram self-information value of the word-gram is larger. Therefore, the word-gram self-information value of each word in the word-gram sequence can be determined based on the probability that the information carried by the word-gram is inferred by other words in the word sequence, and then the word-gram self-information value of each word in the word-gram sequence is accumulated to obtain the unit self-information value of the vocabulary unit corresponding to the word-gram sequence.
[0118] See also Figure 4 , Figure 4 A schematic diagram of a process for determining a target vocabulary unit provided in an embodiment of the present application. Figure 4 As shown, in S204, determining the target vocabulary unit from the vocabulary units by using the unit self-information value includes:
[0119] S402: Sort the unit self-information values of the vocabulary units in the input text by size to obtain a unit self-information value ranking.
[0120] The total number of vocabulary units in the input text may be M, and the size sorting process may include descending order and ascending order. Here, taking ascending order as an example, after the unit self-information values of the vocabulary units in the input text are sorted in ascending order, the unit self-information value sorting is obtained. At this time, the vocabulary unit sorting corresponding to the unit self-information value sorting may be u1,...,u M .
[0121] S404: Obtain preset filtering parameters and the total number of vocabulary units in the input text, determine the reference filtering rank based on the preset filtering parameters and the total number of vocabulary units, and search for the reference unit self-information value corresponding to the reference filtering rank from the unit self-information value ranking; wherein the preset filtering parameter is greater than 0 and less than 1.
[0122] Among them, the preset screening parameters can be predetermined, or the maximum number of taken that should be retained in the input text is determined based on the input word number threshold of the target large model, and then the preset screening parameters are determined based on the number of taken corresponding to each vocabulary unit in the input text and the maximum number of taken that should be retained in the input text, so that the preset screening parameters can effectively control the number of taken of the corresponding proportion of vocabulary units to be less than or equal to the maximum number of taken that should be retained in the input text. It is easy to understand that the preset screening parameter is greater than 0 and less than 1.
[0123] After determining the preset screening parameters, the reference screening rank is determined by multiplying the preset screening parameters by the total number of vocabulary units. When arranging in ascending order, the integer can be rounded up after obtaining the product of the preset screening parameters and the total number of vocabulary units; when arranging in descending order, the integer can be rounded down after obtaining the product of the preset screening parameters and the total number of vocabulary units. Afterwards, the reference unit self-information value corresponding to the reference screening rank is searched from the unit self-information value ranking. Specifically, the reference unit self-information value I corresponding to the reference screening rank P is searched from the unit self-information value ranking. P It can be expressed as np.percentile([I(u1),...,I(u M )],P).
[0124] S406: Taking the unit self-information value in the vocabulary unit that is greater than or equal to the reference unit self-information value as the target unit self-information value, and determining the target vocabulary unit corresponding to the target unit self-information value from the vocabulary unit.
[0125] Among them, the unit self-information value of the vocabulary unit that is greater than the reference unit self-information value can be expressed as I(u j )≥I P , M ≥ j ≥ 1 and j is an integer. At this time, the target vocabulary unit can be expressed as Q = u j |I(u j )≥I P , M ≥ j ≥ 1 and j is an integer. 1
[0126] In the embodiment provided in the present application, the unit self-information value in the vocabulary unit that is greater than or equal to the reference unit self-information value is used as the target unit self-information value, so as to determine the target vocabulary unit corresponding to the target unit self-information value from the vocabulary unit, and then screen out the vocabulary units carrying higher information, thereby eliminating the vocabulary units carrying less information, thereby reducing the prompt word length corresponding to the input of the target large model.
[0127] See also Figure 5 , Figure 5 A schematic diagram of a process for determining global relative position features provided in an embodiment of the present application. Figure 5 As shown, determining the global relative position feature of the target input text in S206 includes:
[0128] S502: Input the target input text into the target macro model, and use the target macro model to determine the model word sequence corresponding to the target input text.
[0129] Among them, after determining the target vocabulary unit, the target vocabulary unit is used to generate the target input text, and then the target vocabulary unit is input into the target large model. The input layer of the target large model can be used to determine the model word sequence corresponding to the target input text, that is, the taken sequence of the target input text.
[0130] S504: Determine the relative position feature of each target word in the model word sequence, and obtain the global relative position feature of the target input text based on the relative position feature of each target word.
[0131] Among them, the relative position features of each target word in the model word sequence can be determined based on the rotation position coding, and the rotation position coding generates a position coding by applying a rotation matrix to each element in the sequence, so that the model can take into account the relative position relationship of the elements in the sequence. After determining the relative position features of each target word in the model word sequence, the global relative position features of the target input text are obtained based on the set of relative position features of each target word.
[0132] In the embodiment provided in the present application, a model word sequence corresponding to a target input text is determined by a target large model, and then the relative position feature of each target word in the model word sequence is determined, and then the global relative position feature of the target input text is determined based on the relative position feature of each target word.
[0133] See also Figure 6 , Figure 6 A schematic diagram of a process for determining the output result of the first model provided in an embodiment of the present application. Figure 6 As shown, in S206, the target large model is controlled by using the global relative position feature to process the target input text to obtain a first model output result, including:
[0134] S602: Use the target large model to determine the word embedding features corresponding to each target word in the model word sequence.
[0135] Determine the word embedding features corresponding to each target word in the model word sequence, so that each target word in the model word sequence can be converted into a model feature that can be understood by the target large model, thereby facilitating subsequent model processing.
[0136] S604: Integrate the relative position feature corresponding to the target word-gram into the word embedding feature corresponding to the target word-gram to update the word embedding feature.
[0137] Among them, each target word has its corresponding relative position feature, and the relative position feature corresponding to the target word is integrated into the word embedding feature corresponding to the target word, so that the updated word embedding feature carries the corresponding relative position information, so that the self-attention mechanism of the target large model can be controlled to effectively capture the long-distance dependencies in the target input text, thereby strengthening the target large model to deepen the learning and understanding of the target input text based on contextual information.
[0138] S606: Determine the model attention score corresponding to the word embedding feature based on the target large model, and generate a first model output result based on the model attention score and the word embedding feature.
[0139] Among them, the model attention score of the target large model is calculated through word embedding features. Since the updated word embedding features carry the corresponding relative position information, the model attention scores corresponding to the updated word embedding features also effectively represent the relative position features, so that the target large model can effectively determine the importance weights assigned to different parts when processing data based on the model attention scores.
[0140] At the same time, after obtaining the model attention score, the model attention score can be corrected by the temperature coefficient t. The obtained model attention score can be divided by t, where t is greater than 0 and less than 1. At this time, the calculation formula of the model attention score can be , q is the query in the attention mechanism, k is the key in the attention mechanism. At the same time, t satisfies At this time, the correction effect on the model attention score is better, making the model attention distribution smoother. L is the original input word number threshold of the target large model, and L' is the original equivalent input word number threshold of the target large model, that is, the target input word number corresponding to the maximum length input text that can be processed by the target large model after applying this data processing method.
[0141] Specifically, the embodiment of the present application can introduce the YaRN algorithm to perform model extrapolation. YaRN is essentially a combination of NTK-by-parts Interpolation and attention distribution correction strategy, which reduces the rotation arc of the low-frequency part when rotating the position encoding.
[0142] In this embodiment, the global relative position feature is used to control the self-attention mechanism of the target large model to effectively capture the long-distance dependencies in the target input text, thereby strengthening the target large model to deepen the learning and understanding of the target input text based on contextual information, so that the target large model can ensure the coherence, consistency and accuracy of the generated first model output results during the processing.
[0143] See also Figure 7 , Figure 7 The present invention provides a data processing device. Figure 7 As shown, the data processing device 700 includes:
[0144] A determination module 710, adapted to determine input text for a target large model and model word number information of the input text, and determine a processing strategy for the input text based on the model word number information;
[0145] A generating module 720, adapted to, when the processing strategy is the first processing strategy, determine the vocabulary units of the input text and the unit self-information values of the vocabulary units, determine the target vocabulary units from the vocabulary units by the unit self-information values, and generate the target input text corresponding to the target vocabulary units;
[0146] The processing module 730 is adapted to determine the global relative position feature of the target input text, and use the global relative position feature to control the target large model to process the target input text to obtain a first model output result.
[0147] Optionally, the generating module 720 includes:
[0148] A division unit, adapted to perform vocabulary unit division processing on the input text to obtain vocabulary units of the input text;
[0149] A calculation unit, adapted to determine a word-gram sequence corresponding to a vocabulary unit, and calculate a word-gram self-information value of each word-gram in the word-gram sequence; wherein the word-gram self-information value is used to characterize the information content of the word-gram;
[0150] The determination unit is adapted to determine the unit self-information value of the vocabulary unit corresponding to the word-gram sequence based on the word-gram self-information value of each word-gram in the word-gram sequence.
[0151] Optionally, the calculation unit is adapted to calculate the word-unit self-information value of each word-unit in the word-unit sequence by using a first calculation formula;
[0152] The first calculation formula satisfies the following formula:
[0153] I(T i ) = -log2P(T i |C(T0~T N ,T i ));
[0154] Among them, T i is the word with sequence number i in the word sequence, N ≥ i ≥ 0 and i is an integer, N + 1 is the total number of words in the word sequence; T0~T N is the word sequence from T0 to T N All words of C(T0~T N ,T i ) is T0 to T N Remove T from all the words i The word obtained after i ) is T i The word self-information value of P(T i |C(T0~T N ,T i )) is T0 to T N Remove T from all the words iUnder the condition that the word generation event corresponding to the remaining word occurs, T i The occurrence probability value of the corresponding word generation event;
[0155] The determining unit is further adapted to use the second calculation formula to perform cumulative processing on the word-unit self-information value of each word-unit in the word-unit sequence to obtain the unit self-information value of the vocabulary unit corresponding to the word-unit sequence;
[0156] The second calculation formula satisfies the following formula:
[0157] I(u) = ;
[0158] Where u is a vocabulary unit, and I(u) is the unit self-information value of u.
[0159] Optionally, the generating module 720 includes:
[0160] A sorting unit, adapted to sort the unit self-information values of the vocabulary units in the input text to obtain a unit self-information value sorting;
[0161] an acquisition unit adapted to acquire a preset screening parameter and the total number of vocabulary units of the input text, determine a reference screening rank by the preset screening parameter and the total number of vocabulary units, and search for a reference unit self-information value corresponding to the reference screening rank from the unit self-information value ranking; wherein the preset screening parameter is greater than 0 and less than 1;
[0162] The target vocabulary unit determination unit is adapted to take the unit self-information value in the vocabulary unit that is greater than or equal to the reference unit self-information value as the target unit self-information value, and determine the target vocabulary unit corresponding to the target unit self-information value from the vocabulary unit.
[0163] Optionally, the processing module 730 includes:
[0164] An input unit, adapted to input a target input text into a target macro model, and determine a model word-unit sequence corresponding to the target input text using the target macro model;
[0165] The global relative position feature determination unit is adapted to determine the relative position feature of each target word in the model word sequence, and obtain the global relative position feature of the target input text based on the relative position feature of each target word.
[0166] Optionally, the processing module 730 includes:
[0167] A word embedding feature determination unit, adapted to determine the word embedding features corresponding to each target word in the model word-unit sequence using the target large model;
[0168] An updating unit, adapted to integrate the relative position feature corresponding to the target word unit into the word embedding feature corresponding to the target word unit, so as to update the word embedding feature;
[0169] A generation unit is adapted to determine a model attention score corresponding to a word embedding feature based on a target large model, and to generate a first model output result based on the model attention score and the word embedding feature.
[0170] Optionally, the determination module 710 includes:
[0171] A model word unit number information determination unit, adapted to obtain an input text for a target large model, determine an input word unit number and a predicted output word unit number of the input text for the target large model, and obtain model word unit number information based on the input word unit number and the predicted output word unit number;
[0172] a length parameter value determining unit, adapted to obtain an input word number threshold of a target large model, determine a proportional parameter corresponding to the input word number threshold, and obtain a length parameter value based on a product of the input word number threshold and the proportional parameter, wherein the proportional parameter is greater than 0 and less than 1;
[0173] The first processing strategy determining unit is adapted to determine that the processing strategy for the input text is the first processing strategy when the number of input word units and the number of predicted output word units are both greater than the length parameter value.
[0174] Optionally, the determining module 710 further includes:
[0175] a second processing strategy determining unit adapted to determine that the processing strategy for the input text is the second processing strategy when the number of input word units and the number of predicted output word units are both less than or equal to the length parameter value, the sum of the number of input word units and the number of predicted output word units is greater than or equal to the input word unit threshold, and the number of input word units is greater than the number of predicted output word units;
[0176] a target input text determination unit adapted to, when the processing strategy is the second processing strategy, determine the vocabulary units of the input text and the unit self-information values of the vocabulary units, determine the target vocabulary units from the vocabulary units by the unit self-information values, and determine the target input text corresponding to the target vocabulary units;
[0177] The second model output result determination unit is adapted to input the target input text into the target macro model, process the target input text based on the target macro model, and obtain the second model output result.
[0178] Optionally, the determining module 710 further includes:
[0179] a third processing strategy determining unit, adapted to determine that the processing strategy for the input text is the third processing strategy when the number of input word units and the number of predicted output word units are both less than or equal to the length parameter value, the sum of the number of input word units and the number of predicted output word units is greater than or equal to the input word unit threshold, and the number of input word units is less than or equal to the predicted output word unit number;
[0180] The third model output result determination unit is suitable for determining the original global relative position features of the input text when the processing strategy is the third processing strategy, and using the original global relative position features to control the target large model to process the input text to obtain the third model output result.
[0181] See also Figure 8 , Figure 8 A schematic diagram of the structure of a terminal provided in an embodiment of the present application. Figure 8 As shown, the terminal 800 may include: at least one processor 801 , at least one network interface 804 , a user interface 803 , a memory 805 , and at least one communication bus 802 .
[0182] The communication bus 802 is used to realize the connection and communication between these components.
[0183] The user interface 803 may include a display screen (Display) and a camera (Camera). The optional user interface 803 may also include a standard wired interface and a wireless interface.
[0184] The network interface 804 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).
[0185] Among them, the processor 801 may include one or more processing cores. The processor 801 uses various interfaces and lines to connect various parts in the entire terminal 800, and executes various functions and processes data of the terminal 800 by running or executing instructions, programs, code sets or instruction sets stored in the memory 805, and calling data stored in the memory 805. Optionally, the processor 801 can be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). The processor 801 can integrate one or a combination of CPU (Central Processing Unit), GPU (Graphics Processing Unit) and modem. Among them, the CPU mainly processes the operating system, user interface and application program, etc.; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; the modem is used to process wireless communication. It can be understood that the above-mentioned modem may not be integrated into the processor 801, but implemented by a single chip.
[0186] Among them, the memory 805 may include RAM (Random Access Memory) and may also include ROM (Read-Only Memory). Optionally, the memory 805 includes a non-transitory computer-readable storage medium. The memory 805 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 805 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area may store data involved in the above-mentioned method embodiments, etc. The memory 805 may optionally be at least one storage device located away from the aforementioned processor 801. As Figure 8 As shown, the memory 805 as a computer storage medium may include an operating system, a network communication module, a user interface module, and a data processing program.
[0187] exist Figure 8 In the terminal 800 shown, the user interface 803 is mainly used to provide an input interface for the user and obtain data input by the user; and the processor 801 can be used to call the data processing program stored in the memory 805 and specifically perform the following operations:
[0188] Determine input text for the target large model and model word number information of the input text, and determine a processing strategy for the input text based on the model word number information;
[0189] When the processing strategy is the first processing strategy, the vocabulary units of the input text and the unit self-information values of the vocabulary units are determined, the target vocabulary units are determined from the vocabulary units by the unit self-information values, and the target input text corresponding to the target vocabulary units is generated;
[0190] The global relative position feature of the target input text is determined, and the global relative position feature is used to control the target large model to process the target input text to obtain a first model output result.
[0191] Optionally, when the processor 801 determines the vocabulary units of the input text and the unit self-information values of the vocabulary units, it specifically performs:
[0192] Performing vocabulary unit division processing on the input text to obtain vocabulary units of the input text;
[0193] Determine a word unit sequence corresponding to the vocabulary unit, and calculate the word unit self-information value of each word unit in the word unit sequence; wherein the word unit self-information value is used to represent the information content of the word unit;
[0194] Based on the word-gram self-information value of each word-gram in the word-gram sequence, the unit self-information value of the vocabulary unit corresponding to the word-gram sequence is determined.
[0195] Optionally, when the processor 801 calculates the word-gram self-information value of each word-gram in the word-gram sequence, it specifically executes:
[0196] The first calculation formula is used to calculate the word-unit self-information value of each word-unit in the word-unit sequence;
[0197] The first calculation formula satisfies the following formula:
[0198] I(T i ) = -log2P(T i |C(T0~T N ,T i ));
[0199] Among them, T i is the word with sequence number i in the word sequence, N ≥ i ≥ 0 and i is an integer, N + 1 is the total number of words in the word sequence; T0~T N is the word sequence from T0 to T N All words of C(T0~T N ,T i ) is T0 to T N Remove T from all the words i The word obtained after i ) is T i The word self-information value of P(T i |C(T0~T N ,T i )) is T0 to T N Remove T from all the words i Under the condition that the word generation event corresponding to the remaining word occurs, T i The occurrence probability value of the corresponding word generation event;
[0200] When the processor 801 determines the unit self-information value of the vocabulary unit corresponding to the word-gram sequence based on the word-gram self-information value of each word-gram in the word-gram sequence, it specifically executes:
[0201] The second calculation formula is used to accumulate the word-gram self-information value of each word-gram in the word-gram sequence to obtain the unit self-information value of the vocabulary unit corresponding to the word-gram sequence;
[0202] The second calculation formula satisfies the following formula:
[0203] I(u) = ;
[0204] Where u is a vocabulary unit and I(u) is the unit self-information value of u.
[0205] Optionally, when the processor 801 determines the target vocabulary unit from the vocabulary units by using the unit self-information value, it specifically executes:
[0206] The unit self-information values of the vocabulary units in the input text are sorted by size to obtain the unit self-information value sorting;
[0207] Obtaining a preset screening parameter and the total number of vocabulary units of the input text, determining a reference screening ranking by using the preset screening parameter and the total number of vocabulary units, and searching for a reference unit self-information value corresponding to the reference screening ranking from the unit self-information value ranking; wherein the preset screening parameter is greater than 0 and less than 1;
[0208] The unit self-information value in the vocabulary unit that is greater than or equal to the reference unit self-information value is taken as the target unit self-information value, and the target vocabulary unit corresponding to the target unit self-information value is determined from the vocabulary unit.
[0209] Optionally, when the processor 801 determines the global relative position feature of the target input text, it specifically performs:
[0210] Input the target input text into the target macro model, and use the target macro model to determine the model word sequence corresponding to the target input text;
[0211] The relative position feature of each target word in the model word sequence is determined, and the global relative position feature of the target input text is obtained based on the relative position feature of each target word.
[0212] Optionally, when the processor 801 uses the global relative position feature to control the target large model to process the target input text and obtains the first model output result, specifically executes:
[0213] Using the target large model, determine the word embedding features corresponding to each target word in the model word sequence;
[0214] Integrate the relative position feature corresponding to the target word into the word embedding feature corresponding to the target word to update the word embedding feature;
[0215] A model attention score corresponding to the word embedding feature is determined based on the target large model, and a first model output result is generated based on the model attention score and the word embedding feature.
[0216] Optionally, the processor 801 determines the input text for the target large model and the model word number information of the input text, and determines the processing strategy for the input text based on the model word number information, specifically performing:
[0217] Obtaining input text for a target large model, determining the number of input word units and the number of predicted output word units of the input text for the target large model, and obtaining model word unit number information based on the number of input word units and the number of predicted output word units;
[0218] Obtain an input word number threshold of the target large model, determine a proportional parameter corresponding to the input word number threshold, and obtain a length parameter value based on the product of the input word number threshold and the proportional parameter, where the proportional parameter is greater than 0 and less than 1;
[0219] When the number of input word units and the number of predicted output word units are both greater than the length parameter value, the processing strategy for the input text is determined to be the first processing strategy.
[0220] Optionally, the processor 801 is further adapted to execute:
[0221] When the number of input word units and the number of predicted output word units are both less than or equal to the length parameter value, and the sum of the number of input word units and the number of predicted output word units is greater than or equal to the input word unit threshold, and the number of input word units is greater than the predicted output word unit number, determining that the processing strategy for the input text is the second processing strategy;
[0222] When the processing strategy is the second processing strategy, determining the vocabulary units of the input text and the unit self-information values of the vocabulary units, determining the target vocabulary units from the vocabulary units by the unit self-information values, and determining the target input text corresponding to the target vocabulary units;
[0223] The target input text is input into the target macro model, and the target input text is processed based on the target macro model to obtain a second model output result.
[0224] Optionally, the processor 801 is further adapted to execute:
[0225] When the number of input word units and the number of predicted output word units are both less than or equal to the length parameter value, and the sum of the number of input word units and the number of predicted output word units is greater than or equal to the input word unit threshold, and the number of input word units is less than or equal to the predicted output word unit number, determining that the processing strategy for the input text is the third processing strategy;
[0226] When the processing strategy is the third processing strategy, the original global relative position features of the input text are determined, and the original global relative position features are used to control the target large model to process the input text to obtain the third model output result.
[0227] In the several embodiments provided in the embodiments of the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.
[0228] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0229] In addition, each functional module in each embodiment of the present application can be integrated into a processing module, or each module can exist physically separately, or two or more modules can be integrated into one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.
[0230] If the integrated module is implemented in the form of a software function module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.
[0231] It should be noted that, for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present application are not limited by the described order of actions, because according to the embodiments of the present application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of the present application.
[0232] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0233] The above is a description of a data processing method and terminal provided in an embodiment of the present application. For technicians in this field, according to the ideas of the embodiments of the present application, there may be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the embodiments of the present application.
Claims
1. A data processing method, wherein: The method comprises: Determine an input text for a target large model and model word unit number information of the input text, and determine a processing strategy for the input text based on the model word unit number information; When the processing strategy is the first processing strategy, determining the vocabulary units of the input text and the unit self-information values of the vocabulary units, determining the target vocabulary units from the vocabulary units by the unit self-information values, and generating the target input text corresponding to the target vocabulary units; Determine the global relative position feature of the target input text, and use the global relative position feature to control the target large model to process the target input text to obtain a first model output result; The step of determining the vocabulary unit of the input text and the unit self-information value of the vocabulary unit includes: Performing vocabulary unit division processing on the input text to obtain vocabulary units of the input text; Determine a word-gram sequence corresponding to the vocabulary unit, and calculate the word-gram self-information value of each word-gram in the word-gram sequence; wherein the word-gram self-information value is used to characterize the information content of the word-gram; Determining a unit self-information value of a vocabulary unit corresponding to the word-gram sequence based on the word-gram self-information value of each word-gram in the word-gram sequence; The step of calculating the word-gram self-information value of each word-gram in the word-gram sequence includes: Calculating the word-unit self-information value of each word-unit in the word-unit sequence using the first calculation formula; The first calculation formula satisfies the following formula: I(T i )=-log2P(T i |C(T0~T N ,T i )); Among them, T i is the word with sequence number i in the word sequence, N≥i≥0 and i is an integer, and N+1 is the total number of words in the word sequence; T0~T N is the word sequence from T0 to T N All words of C(T0~T N ,T i ) is T0 to T N Remove T from all the words i The word obtained after i ) is T i The word self-information value of P(T i |C(T0~T N ,T i )) is T0 to T N Remove T from all the words i Under the condition that the word generation event corresponding to the remaining word occurs, T i The occurrence probability value of the corresponding word generation event; The step of determining the unit self-information value of the vocabulary unit corresponding to the word-gram sequence based on the word-gram self-information value of each word-gram in the word-gram sequence includes: The word unit self-information value of each word unit in the word unit sequence is accumulated by using the second calculation formula to obtain the unit self-information value of the vocabulary unit corresponding to the word unit sequence; The second calculation formula satisfies the following formula: I (in)= ; Wherein, u is the vocabulary unit, and I(u) is the unit self-information value of u.
2. The method according to claim 1, wherein: The step of determining a target vocabulary unit from the vocabulary units by using the unit self-information value comprises: Sorting the unit self-information values of the vocabulary units in the input text to obtain a unit self-information value ranking; Obtaining a preset screening parameter and the total number of vocabulary units of the input text, determining a reference screening rank by using the preset screening parameter and the total number of vocabulary units, and searching for a reference unit self-information value corresponding to the reference screening rank from the unit self-information value ranking; wherein the preset screening parameter is greater than 0 and less than 1; The unit self-information value in the vocabulary unit that is greater than or equal to the reference unit self-information value is taken as the target unit self-information value, and the target vocabulary unit corresponding to the target unit self-information value is determined from the vocabulary unit.
3. The method according to claim 1 or 2, wherein: The determining of the global relative position feature of the target input text comprises: Inputting the target input text into the target macro model, and using the target macro model to determine a model word-unit sequence corresponding to the target input text; The relative position feature of each target word in the model word sequence is determined, and the global relative position feature of the target input text is obtained based on the relative position feature of each target word.
4. The method according to claim 3, wherein: The using of the global relative position feature to control the target large model to process the target input text to obtain a first model output result includes: Determine the word embedding features corresponding to each target word in the model word sequence using the target large model; Integrate the relative position feature corresponding to the target word into the word embedding feature corresponding to the target word to update the word embedding feature; Based on the target large model, a model attention score corresponding to the word embedding feature is determined, and a first model output result is generated based on the model attention score and the word embedding feature.
5. The method according to claim 1, wherein: The step of determining an input text for a target large model and model word unit number information of the input text, and determining a processing strategy for the input text based on the model word unit number information, includes: Acquire an input text for a target large model, determine the number of input word units and the number of predicted output word units of the input text for the target large model, and obtain model word unit number information based on the number of input word units and the number of predicted output word units; Obtaining an input word unit number threshold of the target large model, determining a proportional parameter corresponding to the input word unit number threshold, and obtaining a length parameter value based on a product of the input word unit number threshold and the proportional parameter, wherein the proportional parameter is greater than 0 and less than 1; When the number of input word units and the number of predicted output word units are both greater than the length parameter value, it is determined that the processing strategy for the input text is the first processing strategy.
6. The method according to claim 5, wherein: The method further comprises: When the number of input word units and the number of predicted output word units are both less than or equal to the length parameter value, and the sum of the number of input word units and the number of predicted output word units is greater than or equal to the input word unit threshold, and the number of input word units is greater than the predicted output word unit number, determining that the processing strategy for the input text is the second processing strategy; When the processing strategy is the second processing strategy, determining the vocabulary unit of the input text and the unit self-information value of the vocabulary unit, determining the target vocabulary unit from the vocabulary unit by the unit self-information value, and determining the target input text corresponding to the target vocabulary unit; The target input text is input into the target macro model, and the target input text is processed based on the target macro model to obtain a second model output result.
7. The method according to claim 6, wherein: The method further comprises: When the number of input word units and the number of predicted output word units are both less than or equal to the length parameter value, and the sum of the number of input word units and the number of predicted output word units is greater than or equal to the input word unit threshold, and the number of input word units is less than or equal to the predicted output word unit number, determining that the processing strategy for the input text is a third processing strategy; When the processing strategy is the third processing strategy, the original global relative position features of the input text are determined, and the target large model is controlled to process the input text using the original global relative position features to obtain a third model output result.
8. A terminal, wherein: The terminal includes: Processor; and A memory arranged to store computer executable instructions which, when executed, cause the processor to perform a method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Intelligent question answering method based on pre-training model, computer equipment and storage medium
CN118332095A
Text information processing method and device based on artificial intelligence and electronic equipment
CN118568204A