Enterprise text data multi-dimensional processing method based on large language model and electronic equipment
Through the multi-dimensional processing method of enterprise text data based on large language model, the problems of insufficient recognition accuracy of abnormal text data and insufficient integration of multi-source data in the prior art are solved, and high-precision, real-time and multi-dimensional abnormal text data recognition and monitoring are achieved.
Patent Information
- Application Number
- CN202510203852.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-06
AI Technical Summary
The prior art has problems in the identification of abnormal text data, such as insufficient recognition accuracy, limitation of generalization capabilities, insufficient integration of multi-source data, insufficient real-time and dynamics, and lack of multi-dimensional analysis capabilities.
The multi-dimensional processing method of enterprise text data based on the large language model is adopted. By obtaining the keywords and weights of the text data to be processed, reordering and assigning new weights, obtaining the corresponding context content, and using multiple text classification models to predict text fragments, combining the large language model, BERT model and keyword judgment rules for multi-dimensional classification.
It improves the recognition accuracy of abnormal text data, enhances the generalization ability and robustness of the text classification model, ensures real-time monitoring and dynamic analysis of abnormal text data, and can comprehensively capture the characteristics of abnormal text data from multiple dimensions.
Smart Images

Figure CN119940360A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to a method and electronic device for multi-dimensional processing of enterprise text data based on a large language model. Background Art
[0002] With the popularization of the Internet, data publishers, such as enterprises, can publish various types of text data through different data publishing platforms. For some reasons, data publishers may arbitrarily publish some abnormal text data that does not meet the publishing requirements. These abnormal text data may have some adverse effects on reading users. Therefore, it is necessary to identify these abnormal text data in real time and accurately. At present, natural language processing technology has been widely used in the field of text data processing. For example, text classification technology uses machine learning algorithms and deep learning technologies to classify a large amount of text data, and can classify text data such as corporate public information, news reports, and social media content into normal text data and abnormal text data. In addition, keyword extraction technology uses methods such as TF-IDF and TextRank to extract iconic keywords from large-scale text data, which lays a solid foundation for the identification of abnormal text data. In text data such as corporate public information, words with higher frequency of occurrence can be used as potential abnormal text data for further analysis. Big data analysis technology collects data from multiple sources including news reports, social media dynamics, corporate announcements, etc. on the Internet through data collection and integration, and integrates them into a unified analysis platform. Time series analysis technology identifies abnormal behaviors within a specific period of time by analyzing the time series of data. Machine learning and deep learning technologies also play a key role in the discovery of clues to illegal fundraising. Through supervised learning algorithms, such as support vector machines (SVM), random forests, neural networks, etc., a large amount of labeled historical data can be used to train models to automatically identify potential abnormal text data. Recurrent neural networks (RNN) and long short-term memory networks (LSTM) perform well in processing sequence data (such as text and time series), and can better capture the deep semantics and time correlation in text. In addition, unsupervised learning algorithms such as cluster analysis and principal component analysis can discover hidden patterns and anomalies in data, providing preliminary clues for the discovery of clues to abnormal text data. Financial public opinion monitoring technology can timely discover and predict abnormal corporate behavior by monitoring financial public opinion information on platforms such as social media and news websites. For example, by analyzing discussions about specific financial products or companies on platforms such as Weibo and forums, abnormally active discussion topics and emotional tendencies can be identified. Combined with the above technologies, the real-time early warning system processes and analyzes massive amounts of information in real time, and issues early warning signals in a timely manner to assist regulatory authorities in quickly responding to and processing abnormal text data. Although existing technologies have made some progress in identifying abnormal text data, there are still some obvious shortcomings and deficiencies. First, traditional keyword extraction technologies such as TF-IDF and TextRank methods are still insufficient in recognition accuracy when faced with complex and changing text content. These algorithms have limited understanding of context and deep semantics, and are prone to missing key illegal fundraising information or misjudging irrelevant content.Especially in abnormal text data, keywords may be cleverly hidden, and traditional methods are difficult to effectively capture these hidden signals. Secondly, existing text classification models have limitations in generalization ability. Although machine learning and deep learning technologies have performed well in text classification, their effectiveness is highly dependent on the breadth and quality of training data. If the training data cannot cover new and hidden abnormal text data content, the detection ability of the classification model will be limited, and it is easy to make misjudgments and missed judgments. In addition, when dealing with diverse data sources, existing technologies also have the problem of insufficient integration and unified analysis. Abnormal text data comes from a wide range of sources, including news reports, social media dynamics, corporate announcements, etc. Existing technologies often focus on the analysis of a single data source and lack effective integration of multi-source data, resulting in incomplete and inaccurate judgments. Furthermore, many current technical solutions have the problem of insufficient real-time and dynamic performance when dealing with abnormal text data. Abnormal text data often changes very quickly and frequently. If the monitoring system cannot be updated and dynamically analyzed in real time, it is easy to miss the best time for discovery and early warning. Finally, existing methods usually lack the ability to analyze from multiple dimensions, which is crucial for efficiently identifying abnormal text data features from multiple angles. Single-dimensional information analysis often fails to fully capture the complexity and diversity of abnormal text data, resulting in a significant reduction in the effectiveness of monitoring and early warning. Summary of the invention
[0003] In view of the above technical problems, the technical solution adopted by the present invention is:
[0004] According to a first aspect of the present invention, a method for multi-dimensional processing of enterprise text data based on a large language model is provided, the method comprising the following steps:
[0005] S100, obtaining enterprise text data that currently needs to be processed as text data to be processed.
[0006] S200, obtaining keywords and corresponding weights in the text data to be processed, and reordering the obtained keywords in descending order of weight to obtain ordered keywords.
[0007] S300, assigning new weights to the sorted keywords as the final weights of the keywords of the text data to be processed;
[0008] S400, based on the final weight corresponding to each keyword, the corresponding context content is obtained from the text data to be processed as the text segment corresponding to the keyword, thereby obtaining the text segment corresponding to the text data to be processed.
[0009] S500, using n text classification models to predict the category label of each text segment, n≥3, wherein the category label includes a first label representing that the text data is normal text data and a second label representing that the text data is abnormal text data, and the text classification type includes a large language model.
[0010] According to a second aspect of the present invention, there is provided an electronic device comprising a processor and a memory; the processor is used to execute the steps of the method described in the first aspect of the present invention by calling a program or instruction stored in the memory.
[0011] According to a second aspect of the present invention, there is provided a computer-readable storage medium storing a program or instructions, wherein the program or instructions enable a computer to execute the steps of the method according to the first aspect of the present invention.
[0012] The present invention has at least the following beneficial effects:
[0013] The multi-dimensional processing method of enterprise text data based on a large language model provided by an embodiment of the present invention includes: obtaining keywords and corresponding weights in the text data to be processed, and reordering the obtained keywords in order of weight from large to small to obtain the ordered keywords; assigning new weights to the ordered keywords as the final weights of the keywords; based on the final weights corresponding to each keyword, obtaining the corresponding context content from the text data to be processed as the corresponding text fragment; using multiple text classification models to predict the category label of each text fragment, the category label includes a first label that characterizes the text data as normal text data and a second label that characterizes the text data as abnormal text data. The present invention performs data processing based on keyword extraction and reordering, and integrates multiple methods for multi-dimensional classification, which can improve the accuracy of information recognition and enhance generalization ability and robustness.
[0014] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0016] Figure 1 A flowchart of a method for multi-dimensional processing of enterprise text data based on a large language model provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0017] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0018] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used herein in the specification of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The term "and / or" used herein includes any and all combinations of one or more related listed items.
[0019] It should be noted that some exemplary embodiments are described as processes or methods depicted as flow charts. Although the flow charts describe the steps as sequential processes, many of the steps can be implemented in parallel, concurrently or simultaneously. In addition, the order of the steps can be rearranged. The process can be terminated when its operation is completed, but it can also have additional steps not included in the accompanying drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0020] The embodiment of the present invention provides a method for multi-dimensional processing of enterprise text data based on a large language model. Figure 1 As shown, the method may include the following steps:
[0021] S100, obtaining enterprise text data that currently needs to be processed as text data to be processed.
[0022] In an embodiment of the present invention, the enterprise text data may be enterprise text data released by one or more enterprises. The text data may be obtained from multiple data sources. In an illustrative embodiment, the data source may include news websites, social media platforms, enterprise bulletin boards, court judgments, regulatory announcements, forums, etc.
[0023] In the embodiment of the present invention, the text data can be obtained by the following steps:
[0024] (1) Use the Scrapy crawler framework to write crawler scripts to crawl data from news websites and corporate bulletin boards. The crawlers are run regularly to ensure continuous data updates.
[0025] (2) Social Media API: Obtain social media updates through Weibo API and collect discussions about corporate investment and financial management.
[0026] (3) Data cleaning: Clean the collected data, including deduplication, removal of invalid information and incomplete data entries. Use regular expressions to remove HTML tags and special characters and convert the data into a standard format.
[0027] (4) Data storage: The cleaned data is stored in the distributed database Elasticsearch for subsequent rapid retrieval and analysis.
[0028] Further, in S100, if the amount of text data to be processed in the current database is greater than or equal to a set amount threshold or if the current monitoring time is greater than or equal to a set monitoring time threshold, the text data to be processed in the current database is used as text data to be processed.
[0029] In an embodiment of the present invention, the text data that needs to be processed in the current database is text data that has not been classified, that is, newly acquired text data. The quantity threshold can be set based on actual needs, as long as it can ensure that the data can be processed as promptly as possible. The monitoring time threshold can be set based on actual needs, as long as it can ensure that the data can be processed as promptly as possible, for example, it can be set to a few hours or a few days. S200, obtain the keywords and corresponding weights in the text data to be processed, and re-sort the obtained keywords in order of weight from large to small to obtain sorted keywords.
[0030] Furthermore, S200 may specifically include:
[0031] S210, respectively using m keyword extraction methods to obtain keywords and weights corresponding to the keywords in the text data to be processed, to obtain m keyword sets and m keyword weight sets corresponding to the text data to be processed; m≥2.
[0032] In an exemplary embodiment of the present invention, the keyword extraction method may include a TF-IDF method and a TextRank method.
[0033] The TF-IDF (Term Frequency-Inverse Document Frequency) method is an algorithm widely used in the field of text mining and information retrieval to evaluate the importance of a term in a collection of documents. The TF-IDF calculation combines two parts: term frequency (TF) and inverse document frequency (IDF). First, term frequency is used to calculate how often a specific term appears in a document. It measures the number of times a term appears in a document, compared to the total number of terms in the document. A higher term frequency means that the term is more important in the document. Second, inverse document frequency is used to measure the discriminating power of a term in the entire collection of documents. If a term is widely present in the entire collection of documents, its discriminating power is low and its inverse document frequency value is small. Conversely, if a term only appears in a few documents, its discriminating power is high and its inverse document frequency value is large. TF-IDF multiplies the above two to get the weight of each term in a specific document. The higher the weight value, the more important the term is in the document and the stronger its discriminating power is in the entire collection of documents. In this way, TF-IDF can effectively extract the keywords of a document and provide key data for subsequent text processing and analysis.
[0034] The TextRank method is an algorithm based on graph theory, which is used to extract important keywords or sentences from text. Its principle is inspired by Google's PageRank algorithm. It ranks the importance of words by building a relationship network between words (i.e., word graph). In the TextRank method, the text is first converted into a graph structure, in which each node represents a word, and the edges connecting these nodes represent the co-occurrence relationship between words. Specifically, if two words co-occur within a pre-set window size, there will be an edge between them, representing their relationship. Next, the TextRank method iteratively calculates this word graph. In each iteration, the importance (i.e., weight) of the word node is updated by the "vote" of its neighboring nodes. Initially, all word nodes have the same weight, and as the iteration proceeds, the weight gradually stabilizes. In the end, the word with the highest weight is the most important keyword in the text. The advantage of this algorithm is that it does not require supervised learning or a predefined vocabulary. It only needs to calculate the importance of each word through the graph structure based on the co-occurrence relationship of the words, so it is extremely suitable for keyword extraction from text. The keywords extracted by the TextRank method can effectively represent the core content of the text and lay a solid foundation for subsequent text recognition and analysis.
[0035] Those skilled in the art know that any method for obtaining keywords and weights corresponding to the keywords in the text data to be processed based on the TF-IDF method and the TextRank method falls within the protection scope of the present invention.
[0036] S220, taking the union of the m keyword sets as the intermediate keyword set of the text data to be processed, and fusing the m keyword weight sets corresponding to the text data to be processed, to obtain the first fusion weight of each keyword in the intermediate keyword set, and then to obtain the first keyword fusion weight set corresponding to the intermediate keyword set.
[0037] In S220, the first fusion weight W1 of any keyword i in the intermediate keyword set is f (i) The following conditions are met: W1 f (i)=∑ m j=1 k j × j (i), where w j (i) is the weight of keyword i obtained by the j-th keyword extraction method, k j is the weight corresponding to the j-th keyword extraction method, i ranges from 1 to q, q is the number of keywords in the intermediate keyword set, and j ranges from 1 to m.
[0038] In an embodiment of the present invention, the weight corresponding to each keyword extraction method can be determined based on actual conditions. For example, the keywords and weights obtained by each keyword extraction method are compared with the keywords and weights extracted by manual observation, and the weight of each keyword extraction method is determined based on the comparison result. For example, if the keywords and weights obtained by a certain keyword extraction method are more similar to the keywords and weights extracted by manual observation, the weight of the keyword extraction method is greater. In an illustrative embodiment, when the keyword extraction methods are the TF-IDF method and the TextRank method, the weights of the TF-IDF method and the TextRank method may be the same, that is, both are 0.5, or the weights of the TF-IDF method and the TextRank method may be different, for example, the weight corresponding to the TF-IDF method is less than the weight corresponding to the TextRank method, for example, the weight corresponding to the TF-IDF method is 0.4, and the weight of the TextRank method is 0.6, etc.
[0039] Furthermore, in the embodiment of the present invention, the intermediate keyword set is a keyword set obtained after preprocessing the keywords in the union of the m keyword sets, and the preprocessing includes removing stop words, standardizing words (such as converting all words to lowercase), and removing noise and irrelevant punctuation marks, etc. Through preprocessing, the uniformity and purity of the keywords can be ensured, and redundant information can be prevented from interfering with the subsequent re-ranking process.
[0040] In the embodiment of the present invention, by fusing the results of multiple keyword extraction methods, the importance of a single word in a specific document can be evaluated.
[0041] S230, constructing a first semantic network relationship graph based on the co-occurrence matrix corresponding to the intermediate keyword set, and constructing a second semantic network relationship graph based on the adjacency matrix corresponding to the intermediate keyword set.
[0042] In an embodiment of the present invention, the co-occurrence matrix is used to record the number of times two keywords appear in the same context paragraph, and the adjacency matrix reflects the co-occurrence intensity of the two keywords, that is, the distance between the keywords in the text data. The nodes in the first semantic network relationship diagram are the keywords in the intermediate keyword set, and the two keywords with a co-occurrence relationship are connected by a connecting line. The weight of the connecting line is the number of times the corresponding two keywords appear in the same context paragraph. The more times they appear, the greater the weight. The nodes in the second semantic network relationship diagram are the keywords in the intermediate keyword set, and the two keywords with a co-occurrence relationship are connected by a connecting line. The weight of the connecting line is the distance between the corresponding two keywords in the text data. The greater the distance, the greater the weight. By constructing this semantic network, the relationship between the keywords in the text data can be further understood, providing a basis for subsequent weight adjustment.
[0043] S240, based on the first semantic network relationship graph, obtain the first weight of each keyword in the intermediate keyword set, and then obtain the first keyword weight set corresponding to the intermediate keyword set, and based on the second semantic network relationship graph, obtain the second weight of each keyword in the intermediate keyword set, and then obtain the second keyword weight set corresponding to the intermediate keyword set; fuse the first keyword weight set and the second keyword weight set to obtain the second fused weight of each keyword in the intermediate keyword set, and then obtain the second keyword fused weight set corresponding to the intermediate keyword set.
[0044] In an embodiment of the present invention, the importance of keywords is refined and weighted by the relationship in the semantic network relationship graph, and the importance of keywords is recalculated using the node centrality measurement in graph theory. In an illustrative embodiment, the importance of keywords is re-ranked by using the PageRank method, and the calculation is iterated continuously to finally obtain a new weight value of the keyword. The PageRank method can well reflect the relative importance of keywords in the semantic network, thereby making the weight of keywords more accurate and reasonable.
[0045] Those skilled in the art know that any method for obtaining the weights of keywords in a semantic network relationship graph based on the PageRank method falls within the protection scope of the present invention.
[0046] In the embodiment of the present invention, the second fusion weight of each keyword in the intermediate keyword set may be the weighted weight of the corresponding first weight and second weight, that is, the second fusion weight W2 of each keyword f =c1×w1+c2×w2, w1 and w2 are respectively the first weight and the second weight of the keyword, c1 and c2 are respectively the weights of the first weight and the second weight of the keyword, in an illustrative embodiment, c1=c2=0.5.
[0047] In the embodiment of the present invention, by comprehensively optimizing the results of multiple keyword extraction methods, the accuracy of keyword importance evaluation can be improved. In addition, by using the semantic network and PageRank algorithm for refinement and fusion, the category identification of text data is more accurate. S250, the first keyword fusion weight set and the second keyword fusion weight set are merged to obtain the third fusion weight of each keyword in the intermediate keyword set, and then the third keyword fusion weight set corresponding to the intermediate keyword set is obtained.
[0048] Further, in S250, the third fusion weight w3 of any keyword i in the intermediate keyword set is f (i) The following conditions are met: w3 f (i)=a×w1(i)+b×w2(i), where w1 f (i) is the first fusion weight of keyword i, a is the weight corresponding to the first fusion weight, w2 f (i) is the second fusion weight of keyword i, b is the weight corresponding to the second fusion weight, i ranges from 1 to q, and q is the number of keywords in the middle keyword set.
[0049] In an embodiment of the present invention, a and b can be set based on actual needs and can be empirical values. In another embodiment, a and b can be obtained through a trained weight determination model. By inputting the first fusion weight and the second fusion weight of all keywords into the trained weight determination model, the corresponding weight can be obtained.
[0050] S260, sorting the keywords in the intermediate keyword set in descending order according to the third fusion weight to obtain the sorted keywords.
[0051] S300, assigning new weights to the sorted keywords as the final weights of the keywords of the text data to be processed;
[0052] In an embodiment of the present invention, the new weight corresponding to each keyword in the sorted keywords is positively correlated with the sorting position of the keyword, that is, the higher the ranking of the keyword, the higher the new weight assigned. The specific assignment method can be determined based on actual conditions.
[0053] S400, based on the final weight corresponding to each keyword, the corresponding context content is obtained from the text data to be processed as the text segment corresponding to the keyword, thereby obtaining the text segment corresponding to the text data to be processed.
[0054] Furthermore, S400 specifically includes:
[0055] S401, dividing the text data to be processed into sentences to obtain a plurality of sentences corresponding to the text data to be processed.
[0056] In an embodiment of the present invention, an NLP tool library such as SpaCy or NLTK can be used to perform sentence segmentation on the data to be processed.
[0057] S402: For each keyword in the sorted keywords, if the final weight W corresponding to the keyword is end <W 01 , the sentence containing the keyword in the text data to be processed is taken as the text segment corresponding to the keyword; if W end >W 02 , the sentence containing the keyword in the text data to be processed, as well as the x1 sentences before the sentence containing the keyword and the x1 sentences after the sentence containing the keyword are taken as the text fragment corresponding to the keyword; if W 01 ≤W end ≤W 02 , the sentence containing the keyword, the x2 sentences before the sentence containing the keyword, and the x2 sentences after the sentence containing the keyword in the text data to be processed are taken as the text segment corresponding to the keyword; and then the text segment corresponding to the text data to be processed is obtained; wherein, W01 Set the weight threshold for the first time, W 02 Set the weight threshold for the second, W 01 <W 02 , 1≤x2<x1.
[0058] In the embodiment of the present invention, W 01 and W 02 Can be set based on actual needs, for example, W 01 Can be set to 0.6≤W 01 ≤0.8,W 02 Can be set to 0.2≤W 02 ≤0.3, preferably, W 01 =0.8,W 02 =0.2.
[0059] In the embodiment of the present invention, x1 and x2 can be set based on actual needs. In an exemplary embodiment, x1=3, x2=1.
[0060] In an embodiment of the present invention, for each identified keyword, the scope of the context extraction is dynamically adjusted according to the weight shown in the document. The higher the weight of the keyword, the more extensive the context content will be extracted and analyzed. For example, for high-weight keywords, not only will the entire sentence containing the keyword be extracted, but multiple sentences will be further extended forward and backward, that is, three sentences before and after, to capture larger text fragments, so as to obtain richer semantic information and related content; while for keywords with medium weights, the scope of the extracted context is relatively small, including only the sentence containing the keyword and one adjacent sentence before and after it; for low-weight keywords, only the sentence where the keyword is located is extracted. In this way, not only the keyword itself but also its surrounding context is captured in the text data, so as to have a more comprehensive understanding of potential abnormal text data.
[0061] S500, predicting the category label of each text segment using n text classification models, n≥3, wherein the category label includes a first label representing that the text data is normal text data and a second label representing that the text data is abnormal text data. In an embodiment of the present invention, the abnormal text data may be, for example, text data of an enterprise suspected of illegal fund-raising.
[0062] Furthermore, in the embodiment of the present invention, S500 specifically includes:
[0063] S510, using n text classification models to predict the label prediction value of each text segment respectively, to obtain n label prediction values corresponding to the text segment.
[0064] In an embodiment of the present invention, the text classification model may include a large language model, a trained BERT model, and a classification model based on keyword determination rules.
[0065] Among them, the large language model can be an existing large language model, such as GPT-4 or YAYI model. The large language model can be based on OpenAI or a dedicated API, by passing text data and generating corresponding prediction results. It is known to those skilled in the art that the large language model can be instructed to predict the label prediction value of the input text segment by constructing a text instruction. The large language model will output the corresponding label prediction value based on the received text segment and text instruction.
[0066] Among them, the trained BERT model can be trained through the following steps:
[0067] S10, obtaining a training sample data set.
[0068] In an embodiment of the present invention, the training sample data set can be obtained by labeling historical text data within a set historical time period using a manual labeling tool such as Label Studio. The historical text data can be a document or a text segment.
[0069] S11, input the training sample data of the current batch into the current BERT model for training, and obtain the corresponding prediction results, that is, the label prediction value of each text data.
[0070] S12, based on the prediction results of the current batch and the corresponding true results, obtain the current loss function value of the current BERT model, and determine whether the current loss function value meets the preset model training end condition. If yes, execute step S14, otherwise, execute step S13.
[0071] In the embodiment of the present invention, the true result is the true label value of the text data.
[0072] S13, update the parameters of the current BERT model based on the current loss function value, and use the next batch of training sample data as the training sample data of the current batch, and execute S11.
[0073] S14, taking the current BERT model as the trained BERT model.
[0074] In an embodiment of the present invention, the loss function value can be calculated based on an existing loss function, such as a cross entropy loss function. The preset model training end condition can be set based on actual needs. For example, the loss is less than or equal to a set loss threshold and remains unchanged within a set time period. In an embodiment of the present invention, the trained BERT model is used to classify new text data, which can extract detailed text features and semantic information.
[0075] The label prediction value predicted by the classification model based on the keyword determination rule satisfies the following conditions: P = ∑ N r=1 W end r ×F r , where P is the label prediction value of any text segment predicted by the classification model based on the keyword determination rule, and W end r is the final weight of the rth keyword, F r is the occurrence ratio of the rth keyword in any text segment, r ranges from 1 to N, and N is the number of keywords in the sorted keywords.
[0076] S520, performing weighted fusion on the n label prediction values of each text segment to obtain a final label prediction value corresponding to the text segment, and determining a category label of the text segment based on the corresponding final label prediction value.
[0077] Among them, in S520, if the final label prediction value P of any text segment is final ≥P0, determine the category label of any text segment as the second label, otherwise, it is the first label, where P final =∑ n h=1 c h ×P h , P h is the label prediction value of any text segment predicted by the h-th text classification model, c h is the weight of the hth text classification model, the value of h is 1 to n, P0 is a set threshold, which can be an empirical value, such as 0.8. In the embodiment of the present invention, the weight of each text classification model can be the same, that is, all are 1 / h.
[0078] In the embodiment of the present invention, the large language model, with its powerful semantic understanding ability, can effectively capture and preliminarily classify a large amount of text data, and determine text fragments that may be abnormal text data. The BERT model has a powerful text feature extraction ability, and can perform more detailed classification and semantic analysis based on the preprocessed text data, thereby improving the recognition accuracy of abnormal text data. At this stage, BERT's fine-grained analysis ability enables the algorithm to better understand the complex context and potential risks in the text. For the rule classification dimension, the final classification judgment can be made according to the predefined keyword judgment rules. Specifically, for each text fragment, the proportion of keywords it contains will be evaluated and the weights of these keywords will be calculated. By accumulating the total weight of these keywords, if the accumulated weight exceeds the preset threshold, the text data will be judged as abnormal text data. Through this method of quantifying the influence of keywords and comprehensively judging the weight score of text data, potential abnormal text data can be accurately identified, ensuring the rigor and reliability of the judgment process. In this way, the classification results of the large language model, BERT and the classification information based on keyword rules are organically integrated to form a multi-dimensional and systematic classification decision. This type of multi-level and multi-angle classification method not only improves the accuracy of abnormal text data identification, but also enhances the generalization ability and robustness of the text classification model.
[0079] In order to ensure the dynamic monitoring capability of text data, real-time data processing technology is introduced to process newly acquired text data in real time through stream processing frameworks such as Apache Kafka and Apache Storm, update the text classification model and perform dynamic warnings.
[0080] In summary, the enterprise text data multi-dimensional processing method based on the large language model provided by the embodiment of the present invention can not only significantly improve the recognition accuracy of abnormal text data, but also enhance the real-time monitoring capability of abnormal text data. The multi-dimensional classification and weighted prediction mechanism ensures the comprehensiveness and accuracy of the comprehensive analysis. In actual application scenarios, it can provide efficient and reliable decision-making support for regulatory authorities, and has extremely high practical value and application prospects.
[0081] An embodiment of the present invention also provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are configured to execute the method described in the embodiment of the present invention.
[0082] The embodiment of the present invention further provides a computer-readable storage medium storing computer-executable instructions, wherein the computer instructions are used to execute the method described in the embodiment of the present invention.
[0083] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution disclosed in the present invention can be achieved, and this document does not limit this.
[0084] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A multi-dimensional processing method for enterprise text data based on a large language model, characterized in that: The method comprises the following steps: S100, obtaining enterprise text data that currently needs to be processed as text data to be processed; S200, obtaining keywords and corresponding weights in the text data to be processed, and reordering the obtained keywords in descending order of weight to obtain ordered keywords; S300, assigning new weights to the sorted keywords as the final weights of the keywords of the text data to be processed; S400, based on the final weight corresponding to each keyword, obtaining corresponding context content from the text data to be processed as the text segment corresponding to the keyword, thereby obtaining the text segment corresponding to the text data to be processed; S500, using n text classification models to predict the category label of each text segment, n≥3, wherein the category label includes a first label representing that the text data is normal text data and a second label representing that the text data is abnormal text data, and the text classification type includes a large language model.
2. The method according to claim 1, characterized in that S200 specifically includes: S210, respectively using m keyword extraction methods to obtain keywords and weights corresponding to the keywords in the text data to be processed, to obtain m keyword sets and m keyword weight sets corresponding to the text data to be processed; m≥2; S220, taking the union of the m keyword sets as the intermediate keyword set of the text data to be processed, and fusing the m keyword weight sets corresponding to the text data to be processed, to obtain a first fusion weight of each keyword in the intermediate keyword set, and then obtaining a first keyword fusion weight set corresponding to the intermediate keyword set; S230, constructing a first semantic network relationship graph based on the co-occurrence matrix corresponding to the intermediate keyword set, and constructing a second semantic network relationship graph based on the adjacency matrix corresponding to the intermediate keyword set; S240, based on the first semantic network relationship graph, obtaining a first weight of each keyword in the intermediate keyword set, and then obtaining a first keyword weight set corresponding to the intermediate keyword set, and based on the second semantic network relationship graph, obtaining a second weight of each keyword in the intermediate keyword set, and then obtaining a second keyword weight set corresponding to the intermediate keyword set; fusing the first keyword weight set and the second keyword weight set to obtain a second fused weight of each keyword in the intermediate keyword set, and then obtaining a second keyword fused weight set corresponding to the intermediate keyword set; S250, fusing the first keyword fusion weight set and the second keyword fusion weight set to obtain a third fusion weight of each keyword in the intermediate keyword set, and then obtaining a third keyword fusion weight set corresponding to the intermediate keyword set; S260, sorting the keywords in the intermediate keyword set in descending order according to the third fusion weight to obtain the sorted keywords.
3. The method according to claim 1, characterized in that: In S220, the first fusion weight W1 of any keyword i in the intermediate keyword set is f (i) The following conditions are met: W1 f (i)=∑ m j=1 k j × j (i), where w j (i) is the weight of keyword i obtained by the j-th keyword extraction method, k j is the weight corresponding to the j-th keyword extraction method, i ranges from 1 to q, q is the number of keywords in the middle keyword set, and j ranges from 1 to m; In S250, the third fusion weight w3 of any keyword i in the intermediate keyword set is f (i) The following conditions are met: w3 f (i)=a×w1(i)+b×w2(i), where w1 f (i) is the first fusion weight of keyword i, a is the weight corresponding to the first fusion weight, w2 f (i) is the second fusion weight of keyword i, b is the weight corresponding to the second fusion weight, i ranges from 1 to q, and q is the number of keywords in the middle keyword set.
4. The method according to claim 2, characterized in that: The keyword extraction methods include the TF-IDF method and the TextRank method.
5. The method according to claim 2, characterized in that: In S240, the first weight and the second weight of each keyword are obtained by using the PageRank method.
6. The method according to claim 1, characterized in that S400 specifically includes: S401, dividing the text data to be processed into sentences to obtain a plurality of sentences corresponding to the text data to be processed; S402: For each keyword in the sorted keywords, if the final weight W corresponding to the keyword is end <W 01 , the sentence containing the keyword in the text data to be processed is taken as the text segment corresponding to the keyword; if W end >W 02 , the sentence containing the keyword, the x1 sentences before the sentence containing the keyword, and the x1 sentences after the sentence containing the keyword in the text data to be processed are taken as the text fragment corresponding to the keyword; if W 01 ≤W end ≤W 02 , the sentence containing the keyword, the x2 sentences before the sentence containing the keyword, and the x2 sentences after the sentence containing the keyword in the text data to be processed are taken as the text segment corresponding to the keyword; and then the text segment corresponding to the text data to be processed is obtained; wherein, W 01 Set the weight threshold for the first time, W 02 Set the weight threshold for the second, W 01 <W 02 , 1≤x2<x1.
7. The method according to claim 1, characterized in that S500 specifically includes: S510, using n text classification models to predict the label prediction value of each text segment respectively, to obtain n label prediction values corresponding to the text segment; S520, performing weighted fusion on the n label prediction values of each text segment to obtain a final label prediction value corresponding to the text segment, and determining a category label of the text segment based on the corresponding final label prediction value; Among them, in S520, if the final label prediction value P of any text segment is final ≥P0, determine the category label of any text segment as the second label, otherwise, it is the first label, where P final =∑ n h=1 c h ×P h , P h is the label prediction value of any text segment predicted by the h-th text classification model, c h is the weight of the h-th text classification model, the value of h ranges from 1 to n, and P0 is the set threshold.
8. The method according to claim 7, characterized in that The text classification model also includes a trained BERT model and a classification model based on keyword determination rules; The label prediction value predicted by the classification model based on the keyword determination rule satisfies the following conditions: P = ∑ N r=1 W end r ×F r , where P is the label prediction value of any text segment predicted by the classification model based on the keyword determination rule, and W end r is the final weight of the rth keyword, F r is the occurrence ratio of the rth keyword in any text segment, r ranges from 1 to N, and N is the number of keywords in the sorted keywords.
9. The method according to claim 1, characterized in that: In S100 , if the amount of text data to be processed in the current database is greater than or equal to a set amount threshold or if the current monitoring time is greater than or equal to a set monitoring time threshold, the text data to be processed in the current database is used as text data to be processed.
10. An electronic device, characterized in that: including a processor and a memory; The processor is used to execute the steps of the method according to any one of claims 1 to 8 by calling the program or instruction stored in the memory.