Text processing method and device based on large model, equipment and storage medium
Through a text processing method based on large-models, combined with historical correlation data and semantic information, the goal impact results of professional text are predicted, and the problem of insufficient analysis is solved and more efficient text analysis is achieved.
Patent Information
- Application Number
- CN202311873273.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-08
AI Technical Summary
The analysis of professional texts is insufficiently accurate and inefficient, and lacks sufficient professional information processing capabilities.
Through the text processing method based on the big model, the text information to be processed is obtained, the correlation influencing factors are determined based on its historical correlation data information, the semantic information is analyzed to predict the target impact results, and the analysis accuracy is improved using the forgetting gate module and the full connection layer.
It improves the analysis accuracy and efficiency of professional texts and can more accurately predict the future impact results of the target industry.
Smart Images

Figure CN120277214A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of text analysis, and particularly to a text processing method, apparatus, device, and storage medium based on a large model. Background Art
[0002] Professional texts are often lengthy and highly specialized. Without sufficient professional information capabilities and the ability to analyze and process professional texts, the analysis accuracy of professional texts is usually insufficient, and the analysis efficiency is relatively low. Summary of the Invention
[0003] Embodiments of the present invention provide a text processing method, apparatus, device, and storage medium based on a large model, aiming to improve the accuracy of analyzing professional texts.
[0004] In a first aspect, embodiments of the present invention provide a text processing method, which includes:
[0005] Obtain the text information to be processed;
[0006] Determine the associated influencing factor information according to the historical associated data information corresponding to the text information to be processed;
[0007] Determine the target influencing result information according to the associated influencing factor information and the semantic information of the text information to be processed.
[0008] Optionally, the determining the target influencing result information according to the associated influencing factor information and the semantic information of the text information to be processed includes:
[0009] Perform calculation processing on the associated influencing factor information and the semantic information to obtain calculation result information;
[0010] Obtain first feature information according to the calculation result information and the associated influencing factor information;
[0011] Output the target influencing result information according to the first feature information.
[0012] Optionally, the outputting the target influencing result information according to the first feature information includes:
[0013] Input the first feature information into a forgetting gate module to obtain second feature information;
[0014] Process the second feature information based on a fully connected layer and output the target influencing result information.
[0015] Optionally, before the determining the associated influencing factor information according to the historical associated data information corresponding to the text information to be processed, the method further includes:
[0016] The word embedding layer based on the information extraction model performs feature extraction processing on the to-be-processed text information to obtain text feature information;
[0017] If the to-be-processed text information contains keyword information, update the text feature information according to the weight of the keyword information;
[0018] The information recognition module based on the information extraction model processes the updated text feature information to obtain the target industry information of the to-be-processed text information;
[0019] Obtain the historical associated data information corresponding to the to-be-processed text information according to the target industry information.
[0020] Optionally, the information recognition module based on the information extraction model processes the updated text feature information to obtain the target industry information of the to-be-processed text information, including:
[0021] According to the text structure division rule, obtain the first text feature information corresponding to the title of the to-be-processed text information and the second text feature information corresponding to the content from the text feature information;
[0022] Extract the first local feature information of the first text feature information based on the title channel of the information recognition module, and extract the second local feature information of the second text feature information based on the content channel of the information recognition module;
[0023] Determine the target industry information according to the first local feature information and the second local feature information.
[0024] Optionally, the historical associated data information includes multiple associated object data within a target time period;
[0025] The determining the associated influence factor information according to the historical associated data information corresponding to the to-be-processed text information includes:
[0026] Perform de-centralization processing on the associated object data to obtain target associated object data;
[0027] Extract independent object change data from the target associated object data as the associated influence factor information.
[0028] Optionally, after determining the target influence result information according to the associated influence factor information and the semantic information of the to-be-processed text information, the method further includes:
[0029] Input the target industry information, the to-be-processed text information, and the target influence result information into a text processing model to obtain a text sequence;
[0030] Obtain target object information and calculate target text content in the text sequence whose relevance to the target object information is greater than a relevance threshold;
[0031] Perform marking processing on the target text content and output a text analysis report corresponding to the text information to be processed.
[0032] In a second aspect, an embodiment of the present invention provides a text processing device, and the text processing device includes:
[0033] An acquisition unit for acquiring text information to be processed;
[0034] A first determination unit for determining associated influencing factor information according to historical associated data information corresponding to the text information to be processed;
[0035] A second determination unit for determining target influence result information according to the associated influencing factor information and semantic information of the text information to be processed;
[0036] Preferably, the second determination unit determines target influence result information according to the associated influencing factor information and semantic information of the text information to be processed, including:
[0037] Perform calculation processing on the associated influencing factor information and the semantic information to obtain calculation result information;
[0038] Obtain first feature information according to the calculation result information and the associated influencing factor information;
[0039] Output the target influence result information according to the first feature information;
[0040] Preferably, the second determination unit outputs the target influence result information according to the first feature information, including:
[0041] Input the first feature information into a forgetting gate module to obtain second feature information;
[0042] Process the second feature information based on a fully connected layer and output the target influence result information;
[0043] Preferably, before the first determination unit determines the associated influencing factor information according to the historical associated data information corresponding to the text information to be processed, the method further includes:
[0044] Perform feature extraction processing on the text information to be processed based on a word embedding layer of an information extraction model to obtain text feature information;
[0045] If the text information to be processed contains keyword information, update the text feature information according to the weight of the keyword information;
[0046] Process the updated text feature information through the information recognition module of the information extraction model to obtain the target industry information of the text information to be processed;
[0047] Obtain the historical associated data information corresponding to the text information to be processed according to the target industry information;
[0048] Preferably, the first determination unit processes the updated text feature information through the information recognition module of the information extraction model to obtain the target industry information of the text information to be processed, including:
[0049] Obtain the first text feature information corresponding to the title of the text information to be processed and the second text feature information corresponding to the content from the text feature information according to the text structure division rule;
[0050] Extract the first local feature information of the first text feature information based on the title channel of the information recognition module, and extract the second local feature information of the second text feature information based on the content channel of the information recognition module;
[0051] Determine the target industry information according to the first local feature information and the second local feature information;
[0052] Preferably, the historical associated data information includes multiple associated object data within a target time period, and the first determination unit determines the associated influence factor information according to the historical associated data information corresponding to the text information to be processed, including:
[0053] Perform a decentralization process on the associated object data to obtain target associated object data;
[0054] Extract independent object change data from the target associated object data as the associated influence factor information;
[0055] Preferably, after the second determination unit determines the target influence result information according to the associated influence factor information and the semantic information of the text information to be processed, the method further includes:
[0056] Input the target industry information, the text information to be processed, and the target influence result information into a text processing model to obtain a text sequence;
[0057] Obtain target object information, and calculate the target text content in the text sequence whose relevance to the target object information is greater than the relevance threshold;
[0058] Perform a tagging process on the target text content and output a text analysis report corresponding to the text information to be processed.
[0059] In a third aspect, an embodiment of the present invention further provides a text processing device, including a memory storing a computer program, and when the computer program is executed by a processor, the processor is caused to execute the steps of any one of the text processing methods based on a large model provided by the embodiments of the present invention.
[0060] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, which includes a computer program, and when the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of any one of the text processing methods based on a large model provided by the embodiments of the present invention.
[0061] The present invention first obtains text information to be processed; determines associated influencing factor information according to historical associated data information corresponding to the text information to be processed; and determines target influencing result information according to the associated influencing factor information and semantic information of the text information to be processed. When analyzing the text information to be processed, determining the associated influencing factor information according to the historical associated data information of the target industry and analyzing the target influencing result information of the target industry in the future in combination with the semantic information of the text information to be processed can improve the accuracy of analyzing professional texts. Description of the Drawings
[0062] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for description in the embodiments. Obviously, the following drawings are only some embodiments of the present invention, and those skilled in the art can obtain other drawings without creative efforts based on these drawings.
[0063] Figure 1 is a flowchart of an embodiment of the text processing method based on a large model provided in an embodiment of the present invention;
[0064] Figure 2 is a flowchart of another embodiment of the text processing method based on a large model provided in an embodiment of the present invention;
[0065] Figure 3 is a flowchart of yet another embodiment of the text processing method based on a large model provided in an embodiment of the present invention;
[0066] Figure 4 is a schematic diagram of the architecture of the influencing result prediction model provided in an embodiment of the present invention;
[0067] Figure 5It is a schematic diagram of the architecture of the information extraction model provided in the embodiments of the present invention;
[0068] Figure 6 It is a flowchart of the application of the text processing method based on a large model provided in the embodiments of the present invention;
[0069] Figure 7 It is a schematic diagram of the structure of the text processing device provided in the embodiments of the present invention;
[0070] Figure 8 It is a schematic diagram of the structure of the text processing device provided in the embodiments of the present invention. Specific embodiments
[0071] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention. At the same time, in the description of the embodiments of the present invention, terms such as "first" and "second" are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present invention, the meaning of "a plurality" is two or more unless otherwise specifically defined.
[0072] The embodiments of the present invention provide a text processing method, device, device and storage medium based on a large model.
[0073] Specifically, this embodiment will be described from the perspective of the text processing device, which can be specifically integrated in the text processing device, that is, the text processing method in the embodiments of the present invention can be executed by the text processing device.
[0074] The following will be described in detail with reference to the accompanying drawings. In this embodiment, the execution entity is a text processing device as an example. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments. Although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than that shown in the drawings.
[0075] According to the description of the background technology of the present invention, in the related technology, the accuracy of manual analysis of professional documents is too low.
[0076] To solve the above problems, the present invention discloses a text processing method based on a large model. Please refer to Figure 1 , the specific process of this text processing method can be as follows in steps S10 to S30, where:
[0077] Step S10: Obtain the text information to be processed;
[0078] In this embodiment, the entity executing the text processing method may be a text processing device, which may be a terminal device such as a mobile phone, a tablet, a computer, etc., or a server. It can obtain the latest released information of the specified relevant department, etc., so as to obtain professional documents, and can use the text content of the professional documents as the text information to be processed, or obtain the file download address input or selected by the user through the text processing interface, and download the professional documents based on the file download address, or receive the file uploaded by the user, and in response to the analysis instruction, use the text content in the uploaded file as the text information to be processed for recognition. Professional documents can be management documents issued by relevant departments such as decision-making documents or guiding documents, which put forward some decisions or guiding opinions, and these decisions or guiding opinions can often change the current situation or future trends of some industries, thus causing changes in the market conditions of these industries.
[0079] Step S20: Determine the associated influencing factor information according to the historical associated data information corresponding to the text information to be processed;
[0080] The target industry is the industry involved or affected in the text information to be processed. It can be the industry where the target object of the text information to be processed is located, or can also include other industries affected due to the changes in the industry where the target object is located. For example, the text content in the professional document issued for the aquaculture industry may affect the food processing industry. Industries can be classified with reference to GICS (Global Industry Classification Standard).
[0081] For the industry where the target object of the text information to be processed is located, it can be determined through the acquisition channel of the text information to be processed. For example, for a website, since relevant departments are in charge of different industries, when putting forward professional documents for a certain industry, they may give the classification information of the professional documents. According to the classification information of the text information to be processed such as professional documents in the acquisition channel where the text information to be processed is located, the industry where the target object is located can be determined and used as the target industry. In another implementation, the text information to be processed can be analyzed to obtain the target industry involved in the text information to be processed, which can be achieved through the target information extraction model. Filter out the noise mentioned in the text information to be processed but not belonging to the industry where the management object is located. Based on the target industry model, the text information to be processed can also be analyzed and recognized to obtain the industry where the actual target object of the text information to be processed is located. In this way, even if the text information to be processed and its related information do not directly mention the industry where the target object is located, the industry where the management object of the text information to be processed can be obtained by combining the text content, and the industry where the target object is located can be further used as the target industry.
[0082] Further, it is also possible to determine other industries that will be affected by the text information to be processed based on the industry where the target object of the text information to be processed is located. Preset the association relationships between various industries and other industries. According to this association relationship, the other industries associated with the industry where the management object is located can be determined as the other industries that will be affected by the text information to be processed, and they can also be included within the target industry. Further still, there may be more than two industries in the industry where the target object is located and the other industries that will be affected by the text information to be processed. Some of these industries can be selected as the target industry. Optionally, the stock holding information of the user can be obtained to determine the industries of the stocks purchased by the user. Then, among the industry where the target object is located and the other industries that will be affected by the text information to be processed, the industries of the stocks purchased by the user are selected as the target industry.
[0083] After determining the target industry, it is necessary to analyze the target impact result information of the target industry in combination with the text information to be processed. The target impact result information is the change situation of the target industry in the future industry under the influence of the text information to be processed. Analyzing the target impact result information requires understanding the rules of the target industry's changes. Such rules can be analyzed from the historical association data information. Obtain the historical association data information of the target industry within a preset period in the past. The association data information is associated with the industry changes. The association data information can be information affected by the changes of the target industry or information that can reflect the industry changes. The association data information can be the associated object data of the target object in the industry, and this target object is associated with the industry changes. For an industry, the prices of consumer goods, securities, etc. under the industry are associated with the industry changes. Therefore, the associated object data can be the associated price data, taking the price as the target object of the industry. And the changes of the target industry can be reflected in the stock market and information such as the stock prices in the stock market. Therefore, the association data information can be stock data, and the historical association data information can be the historical stock data within a preset period in the past. Specifically, it can be the historical stock data of the target industry in the previous quarter before the release of a professional document, for example, before the release of a professional document. The historical association data information obtains the past associated price change situation of the target industry. Furthermore, based on the historical association data information, the rules affecting the associated price changes of the industry can be summarized, so as to learn the influence of the professional document on the associated price and extract the associated influence factor information.
[0084] Step S30: Determine the target impact result information according to the associated influence factor information and the semantic information of the text information to be processed.
[0085] In this embodiment, to extract semantic information from the text information to be processed, the text information of the text information to be processed can be input into the pre-trained Bert encoding. After inputting the Bert encoding, a semantic information represented by a vector is formed. The semantic information is merged with the associated influencing factor information. After the merger, the data is input into the impact result prediction model, and the target impact result information that will occur in the target industry after being affected by the text information to be processed is output.
[0086] Optionally, to facilitate the impact result prediction model to distinguish the semantic information and the associated influencing factor information data of two pieces of text information to be processed, after merging the two, the semantic information of the text information to be processed and the stock impact data can be distinguished by marking. For example, use " ” to indicate the start of the semantic information of the text information to be processed, with " " to represent the end of the semantic information of the text information to be processed. Then, the merged semantic information and associated influencing factor information are input into the impact result prediction model. The impact result prediction model can predict the probabilities of the respective impact result categories corresponding to them based on the input information, and output the impact result category with the highest probability as the target impact result information. Based on the impact result prediction model to analyze the complex and abstract text information to be processed, the target impact result information in the future of the target industry involved in the text information to be processed can be accurately predicted, which can improve the accuracy of text analysis.
[0087] In the technical solution disclosed in this embodiment, the text information to be processed is obtained; according to the historical associated data information corresponding to the text information to be processed, the associated influencing factor information is determined; according to the associated influencing factor information and the semantic information of the text information to be processed, the target impact result information is determined. When analyzing the text information to be processed, the associated influencing factor information is extracted from the historical associated data information of the target industry, and the target impact result information in the future of the target industry is analyzed by combining the semantic information of the text information to be processed, which can improve the accuracy of analyzing professional texts.
[0088] Optionally, referring to Figure 2 , based on any of the above embodiments, in another embodiment of the text processing method based on a large model of the present invention, the step S30 includes:
[0089] S31. Perform calculation processing on the associated influencing factor information and the semantic information to obtain calculation result information;
[0090] In this embodiment, calculation processing is performed on the associated influencing factor information and the semantic information to obtain calculation result information. The calculation result information can be the attention score between the associated influencing factor information and the semantic information. The impact result prediction model used to analyze the text information to be processed can be an improved GPT model. To make it more suitable for performing the task of analyzing the text information to be processed, the attention mechanism in the impact result prediction model can be improved. Referring toFigure 4 After the influence result prediction model receives the merged sequence X of the semantic information of the associated influence factor information and the text information to be processed, it strengthens the connection between data through the self-attention mechanism (self-attention layer), and then inputs the sequence X processed based on the self-attention mechanism into the dynamic cross-attention layer (Dynamic cross attention layer) of the influence result prediction model. The dynamic cross-attention layer divides the sequence X into two parts: the data feature D corresponding to the associated influence factor information and the data feature P corresponding to the semantic information of the text information to be processed. By using the data feature D and the data feature P to calculate the cross-attention score between the semantic information of the text information to be processed and the associated influence factor information, as the calculation result information obtained after calculating and processing the associated influence factor information and the semantic information, the i-th attention score is:
[0091]
[0092] S32. Obtain first feature information according to the calculation result information and the associated influence factor information;
[0093] In this embodiment, if the calculation result information is the attention score, the attention feature can be obtained according to the attention score and the associated influence factor information, as the feature information output by the attention layer of the influence result prediction model. The first feature information can be the attention feature, which contains more information carried in the associated influence factor information and the semantic information.
[0094] Further, obtaining first feature information according to the calculation result information and the associated influence factor information includes:
[0095] Set the dynamic weight corresponding to the cross-attention score according to the semantic information;
[0096] Obtain first feature information according to the cross-attention score, that is, the calculation result information, the dynamic weight, and the associated influence factor information;
[0097] In this embodiment, the semantic information is the feature extracted based on the text information to be processed, which contains real-time features that can affect the industry. The dynamic weight needs to be set through these real-time features that can be learned in the semantic information. These real-time features can be obtained through the preset learning parameters and the semantic information. Obtain the preset learning parameters, which are determined when training the influence result prediction model. After multiplying the preset learning parameters and the semantic information element by element and performing normalization processing, the dynamic weight is obtained.
[0098] Specifically, the dynamic weight is calculated based on the following formula:
[0099] weight1(X)=Softmax(W i ⊙P)
[0100] Among them, weight i (X) represents the dynamic weight of the i-th cross attention score, W i refers to the learnable parameters, A i It is the cross attention score between the associated influencing factor information and the semantic information of the text information to be processed.
[0101] After determining the dynamic weight of the i-th cross-attention score, it can be calculated with the i-th cross-attention score by linear combination to obtain the cross-attention score after adding the dynamic weight.
[0102]
[0103] After determining the cross attention score after adding the dynamic weight, the first feature information is obtained by combining the associated influencing factor information as the output feature of the dynamic cross attention layer. Specifically, the first feature information V i (D) indicates the associated influencing factor information.
[0104] S33, outputting the target impact result information according to the first feature information;
[0105] Combining the associated influencing factor information can maintain the original data structure. The dynamic weight is set based on the semantic information and can be integrated into the associated influencing factor information to obtain the first feature information that pays more attention to the impact of the text information to be processed on the industry. Based on the first feature information, more accurate target impact result information can be obtained.
[0106] The impact result prediction model also includes a classification module (classification layer), which includes a fully connected layer and an output layer. The fully connected layer can use a softmax function, which divides the future target impact result information into multiple impact result categories, such as an obvious upward trend (denoted as 1, the increase is more than 10%), a stable fluctuation (denoted as 0, the fluctuation range is between -10% and 10%), and an obvious downward trend (denoted as -1, the decline exceeds 10%). The fully connected layer processes the first feature information and maps the attention layer to different impact result categories. The probability that the first feature information belongs to different impact result categories can be obtained. The output layer selects the impact result category with a probability greater than the probability threshold or the largest probability as the target impact result information corresponding to the text information to be processed.
[0107] In the technical solution disclosed in this embodiment, the cross-attention score between the associated influencing factor information and the semantic information of the text information to be processed is calculated. Using extended attention, the relationship between the associated influencing factor information and the text information to be processed can be analyzed in more detail; dynamic weights are dynamically set through the semantic information of the text information to be processed, and learnable preset learning parameters are introduced to adjust the cross-attention scores between different associated influencing factor information and the semantic information of the text information to be processed, so as to pay more attention to the impact of the text information to be processed on the industry, and more accurately analyze the target impact result information of the text information to be processed on the target industry in the future.
[0108] Further, step S34 includes:
[0109] Inputting the first feature information into a forget gate module to obtain second feature information;
[0110] The second feature information is processed based on the fully connected layer, and the target impact result information is output. In this embodiment, the impact result prediction model may also include a forget gate module, which is set before the fully connected layer. The first feature information must first be input into the preset forget gate module (Forget layer) to obtain the updated first feature information as the second feature information, and the second feature information is input into the fully connected layer to obtain the target impact result information. The preset forget gate module can use a sigmoid activation function:
[0111] F = sigmoid(W f *Z+b f )
[0112] Among them, W f is the weight of the forget gate module, b f is and b f Bias, in the process of training the impact result prediction model, it can be learned and optimized through back propagation and gradient descent. Z is the output from the previous step, that is, the first feature information from the dynamic cross attention layer output that has not been updated. In this way, adding a forget gate after the attention mechanism can help the impact result prediction model forget unnecessary information, so that the model can focus on important information in the process of learning the relationship between the associated impact factor information and the text information to be processed, thereby improving the accuracy of the model prediction and more accurately analyzing the target impact result information of the target industry in the future.
[0113] Optionally, refer to Figure 4, after the forgetting gate module, pooling processing can also be performed to further update the first feature information, making the obtained second feature information more prominent. The pooling processing can use a max pooling module (MaxPooling layer) to focus on the important peaks in the attention feature. In addition, in the impact result prediction model, a residual normalization module (add&Norm) and a Feed-Forward layer can also be set after the forgetting gate module to enhance the expressive ability of the model, so as to more accurately analyze the target impact result information of the target industry in the future.
[0114] Optionally, referring to Figure 3 , based on any of the above embodiments, in another embodiment of the text processing method based on a large model of the present invention, before step S10, the method further includes:
[0115] Step S40: Perform feature extraction processing on the to-be-processed text information based on the word embedding layer of the information extraction model to obtain text feature information;
[0116] In this embodiment, the target industry can be analyzed from the to-be-processed text information. The to-be-processed text information is input into the information extraction model, and the information extraction model includes a word embedding layer and an information recognition module. The information extraction model can be an industry extraction model. Referring to Figure 5 , the industry extraction model can include a word embedding layer and an industry recognition module. After obtaining the to-be-processed text information, after performing operations such as sentence splitting, word segmentation, and stop word removal on the text information of the to-be-processed text, the processed text is input into the preset word embedding layer, and the word embedding representations of each word can be obtained, which can form the text feature information corresponding to the to-be-processed text information. The word embedding layer can include an ElMo embedding model, which can obtain the word embedding representations related to the context of each word.
[0117] Step S50: If the to-be-processed text information contains keyword information, update the text feature information according to the weight of the keyword information;
[0118] In this embodiment, the keyword information can include multiple industry keywords. Detect the words in the to-be-processed text information that can match the preset industry keywords. If so, the to-be-processed text information contains the preset industry keyword. If the to-be-processed text information contains the preset industry keyword, obtain the weight associated with the preset industry keyword, and update the word embedding representation corresponding to the preset industry keyword in the to-be-processed text information according to the weight, so as to update the text feature information corresponding to the to-be-processed text information and achieve keyword embedding.
[0119] Specifically, the word embedding representation can be updated by multiplication according to the weight. The formula is as follows:
[0120]
[0121] Among them, keyword_embedding i represents the word embedding representation of the i-th keyword, represents the word embedding generated by the E word embedding model (such as the ELMo model), ⊙ represents element-wise multiplication, and α i represents the weight corresponding to the i-th keyword, and has the same dimension as The size of the updated word embedding is C*B, where C is the maximum length of the text and B is the dimension of the word embedding representation.
[0122] Specifically, the weight corresponding to the keyword can be obtained through the training of the industry extraction model. Further, prepare a dataset with labels, including multiple management texts and related industry labels. Adopt a CNN model to learn the word weights and output a vector with the same number of words as in the management text, that is, α mentioned above. The loss function used in the training can be the mean squared error.
[0123] Step S60: Process the updated text feature information through the information recognition module of the information extraction model to obtain the target industry information of the to-be-processed text information.
[0124] In this embodiment, after obtaining the text feature information with the keyword information weight updated, input it into the industry recognition module. The industry recognition module can map to the industry involved in the to-be-processed text information according to the updated text feature information, so as to determine the target industry information.
[0125] Step S70: Obtain the historical associated data information corresponding to the to-be-processed text information according to the target industry information.
[0126] In this embodiment, after determining the target industry information, obtain the historical associated data information of the target industry corresponding to the to-be-processed text information in the past period according to the target industry information.
[0127] In the technical solution disclosed in this embodiment, based on the preset word embedding layer, obtain the text feature information of the to-be-processed text information, and update the text feature information through the weight corresponding to the preset keyword information, so as to pay more attention to the keywords helpful for the output related industries. Based on the information recognition module to process the updated text feature information, it can efficiently and accurately identify the more accurate target industry information, so as to obtain the corresponding historical associated data information, and thus analyze the to-be-processed text information more accurately.
[0128] Further, the processing of the updated text feature information by the information recognition module based on the information extraction model to obtain the target industry includes:
[0129] According to the text structure division rules, obtain the first text feature information corresponding to the title of the text information to be processed and the second text feature information corresponding to the content from the text feature information;
[0130] Extract the first local feature information of the first text feature information based on the title channel of the information recognition module, and extract the second local feature information of the second text feature information based on the content channel of the information recognition module;
[0131] Determine the target industry information according to the first local feature information and the second local feature information.
[0132] In this embodiment, when relevant departments issue professional documents, they need to have a certain degree of authority. Although the text content may be long and obscure, the titles of professional documents are highly condensed statements, and some will clearly indicate the industries involved. For example, in the professional document "Notice on Several Issues Concerning the Development of the Real Estate Industry", the real estate industry is clearly indicated, which is very crucial for the industry extraction task. Since there are significant differences between the titles and contents of professional documents, resulting in different expression effects, title channels and content channels can be involved in the industry extraction module for professional document titles and contents.
[0133] According to the preset text structure division rules, which can refer to some requirements such as word count and typesetting, determine the title and text of the text information to be processed. Thus, the first text feature information corresponding to the title of the text information to be processed and the second text feature information corresponding to the content of the text information to be processed can be obtained from the text feature information of the text information to be processed according to the text order or position of the title and text of the text information to be processed. The first text feature information includes the word embedding representation corresponding to the title, and the second text feature information includes the word embedding representation corresponding to the content.
[0134] Input the first text feature information corresponding to the title of the text information to be processed into the title channel of the information recognition module, and input the second text feature information corresponding to the content of the text information to be processed into the content channel of the information recognition module. In the title channel, the first local feature information in the first text feature information can be captured through convolution operations. The size of the convolution kernel can be 3. Taking the i-th word as the center, perform sliding convolution on the first text feature information to obtain the feature vectors corresponding to each word embedding representation, and form the first local feature information. The specific formula is as follows:
[0135]
[0136] Among them, represents the feature vector of the i-th word embedding in the first text feature information, conv represents the convolution operation, Ti represents the word embedding of the i-th word in the first text feature information, and "[]" represents the vector concatenation operation, W title , b title represent the convolutional kernel weights and bias terms of the title channel.
[0137] In the content channel, the second local feature information in the second text feature information can be captured through a convolution operation. The convolutional kernel size can be 5. Centered on the i-th word, a sliding convolution is performed. The calculation formula is similar to that of the title channel, and the feature vectors represented by each word embedding in the second text feature information are obtained to form the second local feature information.
[0138] The second local feature information and the second local feature information are concatenated to obtain the merged feature vector representation g = [g title , g text . The merged feature vector representation g is input into the transformer layer to understand the deeper semantics in the text information to be processed, and then input into a fully connected layer. The softmax function is used to give the probability corresponding to each industry involved in the text information to be processed. The industry with the maximum probability or a probability greater than the probability threshold is used as the industry that the text information to be processed may affect, thereby determining the target industry.
[0139] Furthermore, before using the information recognition module, the information recognition module needs to be trained first. The loss function used for training the information recognition module selects the cross-entropy loss function, and the accuracy rate is used as the evaluation index to evaluate the training effect of the model. The accuracy rate is the ratio of the number of correctly predicted samples to all samples to be predicted. When the accuracy rate is greater than the target accuracy rate, the information recognition module training is completed.
[0140] In this way, the first text feature information corresponding to the title of the text information to be processed and the second text feature information corresponding to the content are respectively recognized and processed through the title channel and the content channel, and special attention is paid to the semantic information covered in the title, so as to obtain the target industry more accurately and efficiently, and improve the accuracy and efficiency of analyzing the text information to be processed.
[0141] Optionally, based on any of the above embodiments, in another embodiment of the text processing method based on a large model of the present invention, the step S20 includes:
[0142] Step S21, perform decentralization processing on the associated object data to obtain target associated object data;
[0143] Step S22, extract independent object change data from the target associated object data as the associated influence factor information.
[0144] In this embodiment, the historical associated data information may include multiple associated object data within a preset time period. Taking the historical associated data information as stock data, the associated object data may be associated price data. Taking the stock price data as an example, multiple stock price data within a preset time period of the target industry can be downloaded from a stock trading platform. The stock price data may include opening price, closing price, highest price, lowest price, rise and fall (compared with the previous day) data, etc. The preset time may be the time interval where the time of releasing the text information to be processed is located, including the release time. For example, the week (including) where the release time is located and the previous quarter. The stock price data can be expressed as:
[0145]
[0146] If there is data missing in the multiple associated object data within the preset time period, interpolation method can be used to fill the data, then de - centralize the data, and then whiten the data to make it satisfy E(ZZ T ) = 1 (Z is the whitened matrix). Obtain the preset independent component data, and use the independent component analysis algorithm (ICA algorithm) to extract the independent components from these associated object data. These independent components are independent object variation data. The data generated by the associated object data under the influence of various factors is the manifestation of various influencing factors. The object variation data obtained after comparing multiple associated object data essentially contains the potential information of various influencing factors including professional documents. Through the independent component algorithm, the redundant components in the associated object data can be removed, and the remaining independent object variation data can be used as the associated influencing factor information.
[0147] The extracted associated influencing factor information and the semantic information of the text information to be processed need to be input into the influence result prediction model. The influence result prediction model needs to be trained first. Different types of target influence result information labels need to be made. The production process is as follows: Obtain multiple professional documents and the associated object data within a preset time period after the release of the professional documents. Determine the target influence result information corresponding to the professional documents according to the associated object data within the preset time period after the release, and associate it with the professional documents as labels. Based on multiple professional documents and their target influence result information labels, train the influence result prediction model. Still taking the associated object data as stock price data as an example, the closing price data of the stock price in the next quarter after the release of the professional document can be downloaded, and the fluctuation range of the closing price is calculated. If it is within 10% above and below the current closing price, it is considered that the stock price fluctuates steadily, denoted as 0. If the increase amplitude is greater than 10%, it is considered that there is an obvious increase, denoted as 1. If the decrease amplitude exceeds 10%, it is considered that there is an obvious decrease, denoted as - 1. After the training is completed, the cross - entropy loss function is used to calculate the classification loss of the influence result prediction model, and the mean square error loss function is used to optimize the parameters of the forgetting gate until the influence result prediction model reaches the training purpose.
[0148] In the technical solution disclosed in this embodiment, a plurality of associated object data within an objective preset time period are obtained, and the independent component analysis algorithm is used to process the plurality of associated object data, removing redundancy and noise in the data, capturing the characteristics of the associated object data more efficiently, and obtaining independent object change data as associated influence factor information. On the one hand, the plurality of associated object data within the preset time period are relatively objective, and based on this, the associated influence factor information can be used to objectively analyze the target influence result information, improving the accuracy of the analysis. On the other hand, the associated object data are relatively easy to obtain. The independent component analysis method can remove redundant data and capture the characteristics of the associated object data more efficiently, thereby efficiently extracting independent associated object change data as associated influence factor information and improving the analysis efficiency.
[0149] In some other embodiments, the historical associated data information may include historical professional documents of one or more target industries, and target influence result information within a preset time period after the release of the historical professional documents. Based on the target influence result information, market influence factors are extracted from the historical professional documents as associated influence factor information. The market influence factor is a factor that will affect market changes. Determine whether the market influence factors are included in the to-be-processed text information based on the semantic information of the to-be-processed text information, so as to determine the target influence result information associated with the market influence factor as the target influence result information of the to-be-processed text information.
[0150] Further, based on any of the above embodiments, after step S30, the method further includes:
[0151] Input the target industry information, the to-be-processed text information, and the target influence result information into a text processing model to obtain a text sequence;
[0152] Obtain target object information, and calculate the target text content in the text sequence whose correlation with the target object information is greater than the correlation threshold;
[0153] Perform marking processing on the target text content, and output a text analysis report corresponding to the to-be-processed text information.
[0154] In this embodiment, the text processing flow can be as Figure 6As shown in the figure, after the user selects the text information to be processed through the analysis instruction, the text information to be processed is obtained and sent to the industry extraction model. The industry extraction model consists of a word embedding layer, a CNN network, and a transformer network, and can extract the target industry involved in the target associated file from the text information to be processed. After obtaining the historical associated data information of the target industry and extracting the associated influencing factor information, it forms a sequence with the semantic information of the text information to be processed and is input into the influence result prediction model. The influence result prediction model can be based on an improved GPT model architecture to obtain the target influence result information of the target industry in the future. Since it is user-oriented, the analysis result of the text information to be processed also needs to be displayed to the user. The target industry of the target associated file and its target influence result information are integrated into a new text sequence to generate a text analysis report and feedback it to the user.
[0155] The target industry information, the text information to be processed, and the target influence result information are input into a text processing model, such as the GPT model, and the model generates a text sequence that is easy for users to understand. The target object information can also be obtained. The target object information can be the relevant information of the preset text analysis report output object. The target object information can be the asset holding information of the output object (such as a user or a device). If the historical associated data information is stock data and the associated object data is stock price data, then the target object information can be the stock holding information of the target user who needs to query the text analysis report, determine data such as the industry where the stocks are held and the holding quantity, calculate the target text content in the text sequence whose relevance to the target object information is greater than the preset relevance, and perform marking processing on the target text content to distinguish it from other content in the display. After the marking processing, the text analysis report corresponding to the text information to be processed is output, which is convenient for reading. For example, when there are multiple target industries, if the target industry matches or is associated with the industry where the target user holds stocks, the corresponding content in the text content for this target industry and its target influence result information is used as the target file content and is marked, such as being bolded or highlighted. In this way, when the target user views the text analysis report corresponding to the text information to be processed, they can quickly and conveniently view the relevant analysis content they need.
[0156] This embodiment also provides a text processing device, which can be specifically integrated in a text processing device. For example, as Figure 7 shown, the text processing device may include:
[0157] An acquisition unit 1001, configured to acquire the text information to be processed;
[0158] A first determination unit 1002, configured to determine the associated influence factor information according to the historical associated data information corresponding to the text information to be processed;
[0159] A second determination unit 1003, configured to determine target influence result information according to the semantic information of the associated influence factor information and the text information to be processed;
[0160] Preferably, the second determination unit 1003 determines target influence result information according to the semantic information of the associated influence factor information and the text information to be processed, including:
[0161] Performing calculation processing on the associated influence factor information and the semantic information to obtain calculation result information;
[0162] Obtaining first feature information according to the calculation result information and the associated influence factor information;
[0163] Outputting the target influence result information according to the first feature information;
[0164] Preferably, the second determination unit 1003 outputs the target influence result information according to the first feature information, including:
[0165] Inputting the first feature information into a forgetting gate module to obtain second feature information;
[0166] Processing the second feature information based on a fully connected layer and outputting the target influence result information;
[0167] Preferably, before the first determination unit 1002 determines the associated influence factor information according to the historical associated data information corresponding to the text information to be processed, the method further includes:
[0168] Performing feature extraction processing on the text information to be processed based on the word embedding layer of the information extraction model to obtain text feature information;
[0169] If the text information to be processed contains keyword information, updating the text feature information according to the weight of the keyword information;
[0170] Processing the updated text feature information based on the information recognition module of the information extraction model to obtain the target industry information of the text information to be processed;
[0171] Obtaining the historical associated data information corresponding to the text information to be processed according to the target industry information;
[0172] Preferably, the first determination unit 1002 processes the updated text feature information based on the information recognition module of the information extraction model to obtain the target industry information of the text information to be processed, including:
[0173] Obtain the first text feature information corresponding to the title of the text information to be processed and the second text feature information corresponding to the content from the text feature information according to the text structure division rule;
[0174] Extract the first local feature information of the first text feature information based on the title channel of the information recognition module, and extract the second local feature information of the second text feature information based on the content channel of the information recognition module;
[0175] Determine the target industry information according to the first local feature information and the second local feature information;
[0176] Preferably, the historical association data information includes multiple association object data within a target time period, and the first determination unit 1002 determines the association influence factor information according to the historical association data information corresponding to the text information to be processed, including:
[0177] Perform a decentralization process on the association object data to obtain target association object data;
[0178] Extract independent object change data from the target association object data as the association influence factor information;
[0179] Preferably, after the second determination unit 1003 determines the target influence result information according to the association influence factor information and the semantic information of the text information to be processed, the method further includes:
[0180] Input the target industry information, the text information to be processed, and the target influence result information into a text processing model to obtain a text sequence;
[0181] Obtain target object information, and calculate the target text content in the text sequence whose relevance to the target object information is greater than the relevance threshold;
[0182] Perform a marking process on the target text content, and output a text analysis report corresponding to the text information to be processed.
[0183] In this embodiment, obtain the text information to be processed; determine the association influence factor information according to the historical association data information corresponding to the text information to be processed; determine the target influence result information according to the association influence factor information and the semantic information of the text information to be processed. When analyzing the text information to be processed, determining the association influence factor information for the historical association data information of the target industry and analyzing the target influence result information of the target industry in the future in combination with the semantic information of the text information to be processed can improve the accuracy of analyzing professional texts.
[0184] As Figure 8 shownFigure 8 Schematic structural diagram of a text processing device provided by an embodiment of the present invention. The text processing device 1100 includes a processor 1101 having one or more processing cores, a memory 1102 having one or more computer-readable storage media, and a computer program stored on the memory 1102 and executable on the processor. Among them, the processor 1101 is electrically connected to the memory 1102. Those skilled in the art can understand that the structural diagram of the text processing device shown in the figure does not constitute a limitation on the text processing device, and may include more or fewer components than shown, or combine some components, or have different component arrangements.
[0185] The processor 1101 is the control center of the text processing device 1100, connecting various parts of the entire text processing device 1100 through various interfaces and lines. By running or loading software programs and / or units stored in the memory 1102, and calling data stored in the memory 1102, it executes various functions of the text processing device 1100 and processes data, thereby monitoring the text processing device 1100 as a whole. The processor 1101 can be a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc., and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention.
[0186] In an embodiment of the present invention, the processor 1101 in the text processing device 1100 will load the instructions corresponding to the processes of one or more application programs into the memory 1102 according to the following steps, and the processor 1101 will run the application programs stored in the memory 1102 to implement various functions, such as:
[0187] Obtain text information to be processed;
[0188] Determine associated influencing factor information according to the historical associated data information corresponding to the text information to be processed;
[0189] Determine target influencing result information according to the associated influencing factor information and the semantic information of the text information to be processed.
[0190] For the specific implementation of each of the above operations, reference can be made to the previous embodiments, and details are not described herein again.
[0191] Optionally, as Figure 8As shown in the figure, the text processing device 1100 further includes: a touch display screen 1103, a radio frequency circuit 1104, an audio circuit 1105, an input unit 1106, and a power supply 1107. Among them, the processor 1101 is electrically connected to the touch display screen 1103, the radio frequency circuit 1104, the audio circuit 1105, the input unit 1106, and the power supply 1107 respectively. Those skilled in the art can understand that Figure 8 the structure of the text processing device shown in the figure does not constitute a limitation on the text processing device, and it may include more or fewer components than shown in the figure, or combine certain components, or have different component arrangements.
[0192] The touch display screen 1103 can be used to display a graphical user interface and receive operation instructions generated by the user acting on the graphical user interface. The touch display screen 1103 may include a display panel and a touch panel. Among them, the display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the text processing device. These graphical user interfaces can be composed of graphics, text, icons, videos, and any combination thereof. Optionally, the display panel can be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. The touch panel can be used to collect touch operations of the user on or near it (such as the user using a finger, a stylus, or any suitable object or accessory to operate on the touch panel or near the touch panel), and generate corresponding operation instructions, and the operation instructions execute the corresponding program. Optionally, the touch panel can include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the touch position of the user and detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into touch point coordinates, and then sends it to the processor 1101, and can receive and execute the command sent by the processor 1101. The touch panel can cover the display panel. After the touch panel detects a touch operation on or near it, it is transmitted to the processor 1101 to determine the type of touch event. Subsequently, the processor 1101 provides a corresponding visual output on the display panel according to the type of touch event. In the embodiments of the present invention, the touch panel and the display panel can be integrated into the touch display screen 1103 to realize input and output functions. However, in some embodiments, the touch panel and the touch panel can be implemented as two independent components to realize input and output functions. That is, the touch display screen 1103 can also be used as a part of the input unit 1106 to realize the input function.
[0193] The radio frequency circuit 1104 can be used to transmit and receive radio frequency signals to establish wireless communication with a network device or other text processing devices through wireless communication, and transmit and receive signals with the network device or other text processing devices.
[0194] The audio circuit 1105 can be used to provide an audio interface between the user and the text processing device through a speaker and a microphone. The audio circuit 1105 can transmit the electrical signal converted from the received audio data to the speaker, which is then converted into a sound signal for output. On the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 1105 and then converted into audio data. After the audio data is output to the processor 1101 for processing, it is sent through the radio frequency circuit 1104 to, for example, another text processing device, or the audio data is output to the memory 1102 for further processing. The audio circuit 1105 may also include an earphone jack to provide communication between the peripheral earphone and the text processing device.
[0195] The input unit 1106 can be used to receive input digital, character information or user feature information (such as fingerprint, iris, facial information, etc.), and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0196] The power supply 1107 is used to supply power to each component of the text processing device 1100. Optionally, the power supply 1107 can be logically connected to the processor 1101 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 1107 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0197] Although Figure 8 not shown in the figure, the text processing device 1100 may also include a camera, a sensor, a Wi-Fi module, a Bluetooth module, etc., which will not be elaborated here.
[0198] In the above embodiments, the descriptions of each embodiment have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0199] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0200] Therefore, an embodiment of the present invention provides a computer-readable storage medium, in which multiple computer programs are stored. The computer programs can be loaded by a processor to execute any one of the text processing methods provided by the embodiments of the present invention. The computer programs can execute the steps of the following text processing method:
[0201] Obtain the text information to be processed;
[0202] Determine the associated influencing factor information according to the historical associated data information corresponding to the text information to be processed;
[0203] Determine the target influence result information according to the associated influencing factor information and the semantic information of the text information to be processed.
[0204] For the specific implementation of each of the above operations, reference can be made to the previous embodiments, which will not be elaborated here.
[0205] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.
[0206] Since the computer program stored in the computer-readable storage medium can execute any one of the text processing methods based on the large model provided by the embodiments of the present invention, the beneficial effects that can be achieved by any one of the text processing methods based on the large model provided by the embodiments of the present invention can be realized. For details, refer to the previous embodiments, which will not be elaborated here.
[0207] In the above embodiments of the text processing device, computer-readable storage medium, text processing device, and computer program product, the descriptions of each embodiment have their own focuses. For the parts not elaborated in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes and the beneficial effects that can be brought by the above-described text processing device, computer-readable storage medium, computer program product, text processing device, and their corresponding units can refer to the description of the text processing method based on the large model in the above embodiments, which will not be elaborated here specifically.
[0208] The above has introduced in detail a text processing method, text processing device, text processing device, computer-readable storage medium, and computer program product based on the large model provided by the embodiments of the present invention. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those skilled in the art, according to the idea of the present invention, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present invention.
Claims
1. A text processing method based on a large model, characterized in that, The method includes: Obtaining the text information to be processed; Determining the associated influencing factor information according to the historical associated data information corresponding to the text information to be processed; Determining the target influencing result information according to the associated influencing factor information and the semantic information of the text information to be processed.
2. The text processing method based on a large model according to claim 1, wherein The determining the target influencing result information according to the associated influencing factor information and the semantic information of the text information to be processed includes: Performing a calculation process on the associated influencing factor information and the semantic information to obtain calculation result information; Obtaining first feature information according to the calculation result information and the associated influencing factor information; Outputting the target influencing result information according to the first feature information.
3. The text processing method based on a large model according to claim 2, wherein The outputting the target influencing result information according to the first feature information includes: Inputting the first feature information into a forgetting gate module to obtain second feature information; Processing the second feature information based on a fully connected layer and outputting the target influencing result information.
4. The text processing method based on a large model according to claim 1, wherein Before the determining the associated influencing factor information according to the historical associated data information corresponding to the text information to be processed, the method further includes: Performing feature extraction processing on the text information to be processed based on the word embedding layer of an information extraction model to obtain text feature information; If the text information to be processed contains keyword information, updating the text feature information according to the weight of the keyword information; Processing the updated text feature information based on the information recognition module of the information extraction model to obtain the target industry information of the text information to be processed; Obtaining the historical associated data information corresponding to the text information to be processed according to the target industry information.
5. The text processing method based on a large model according to claim 4, characterized in that, The processing the updated text feature information based on the information recognition module of the information extraction model to obtain the target industry information of the text information to be processed includes: Obtaining first text feature information corresponding to the title of the text information to be processed and second text feature information corresponding to the content from the text feature information according to text structure division rules; Extracting first local feature information of the first text feature information based on the title channel of the information recognition module, and extracting second local feature information of the second text feature information based on the content channel of the information recognition module; Determining the target industry information according to the first local feature information and the second local feature information.
6. The text processing method based on a large model according to claim 1, wherein The historical associated data information includes multiple associated object data within a target time period; The determining the associated influencing factor information according to the historical associated data information corresponding to the text information to be processed includes: Performing a decentralization process on the associated object data to obtain target associated object data; Extracting independent object change data from the target associated object data as the associated influencing factor information.
7. The text processing method based on a large model according to claim 6, wherein, After the determining the target influencing result information according to the associated influencing factor information and the semantic information of the text information to be processed, the method further includes: Inputting the target industry information, the text information to be processed, and the target influencing result information into a text processing model to obtain a text sequence; Obtain the target object information and calculate the target text content in the text sequence whose relevance to the target object information is greater than the relevance threshold; Perform marking processing on the target text content and output a text analysis report corresponding to the text information to be processed.
8. A text processing device, characterized in that, The text processing device includes: An acquisition unit for acquiring text information to be processed; A first determination unit for determining associated influencing factor information according to the historical associated data information corresponding to the text information to be processed; A second determination unit for determining target influence result information according to the associated influencing factor information and the semantic information of the text information to be processed; Preferably, the second determination unit determines the target influence result information according to the associated influencing factor information and the semantic information of the text information to be processed, including: Perform calculation processing on the associated influencing factor information and the semantic information to obtain calculation result information; Obtain first feature information according to the calculation result information and the associated influencing factor information; Output the target influence result information according to the first feature information; Preferably, the second determination unit outputs the target influence result information according to the first feature information, including: Input the first feature information into a forgetting gate module to obtain second feature information; Process the second feature information based on a fully connected layer and output the target influence result information; Preferably, before the first determination unit determines the associated influencing factor information according to the historical associated data information corresponding to the text information to be processed, the method further includes: Perform feature extraction processing on the text information to be processed based on the word embedding layer of the information extraction model to obtain text feature information; If the text information to be processed contains keyword information, update the text feature information according to the weight of the keyword information; Process the updated text feature information based on the information recognition module of the information extraction model to obtain the target industry information of the text information to be processed; Obtain the historical associated data information corresponding to the text information to be processed according to the target industry information; Preferably, the first determination unit processes the updated text feature information based on the information recognition module of the information extraction model to obtain the target industry information of the text information to be processed, including: Obtain the first text feature information corresponding to the title of the text information to be processed and the second text feature information corresponding to the content from the text feature information according to the text structure division rule; Extract the first local feature information of the first text feature information based on the title channel of the information recognition module, and extract the second local feature information of the second text feature information based on the content channel of the information recognition module; Determine the target industry information according to the first local feature information and the second local feature information; Preferably, the historical associated data information includes multiple associated object data within a target time period, and the first determination unit determines the associated influencing factor information according to the historical associated data information corresponding to the text information to be processed, including: Decentralize the associated object data to obtain target associated object data; Extract independent object change data from the target associated object data as the associated influence factor information; Preferably, after the second determination unit determines the target influence result information according to the semantic information of the associated influence factor information and the text information to be processed, the method further includes: Input the target industry information, the text information to be processed, and the target influence result information into a text processing model to obtain a text sequence; Obtain target object information and calculate target text content in the text sequence whose relevance to the target object information is greater than a relevance threshold; Perform marking processing on the target text content and output a text analysis report corresponding to the text information to be processed.
9. A text processing device, characterized in that, It includes a processor and a memory. The memory stores a computer program. When the computer program is executed by the processor, the processor executes the steps of the text processing method according to any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a computer program. When the computer program runs on an electronic device, the computer program is used to cause the electronic device to execute the steps of the text processing method according to any one of claims 1-7.