Artificial Intelligence-Based Text Analysis Method, System, Electronic Device and Medium
By enhancing data and domain identification of the target text, and using multiple model frameworks to train text analysis models, the problems of overfitting and low accuracy of models in the prior art are solved, and more efficient text analysis results are achieved.
Patent Information
- Application Number
- CN202310972661.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-08-03
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2043-08-03
AI Technical Summary
The existing NLP-based sentiment analysis technology has poor robustness due to sample imbalance, and the text analysis accuracy in different application fields is low.
By obtaining the original text related to the target text, performing data augmentation processing, identifying application fields and obtaining multiple model frameworks from the preset model library, using these frameworks to train multiple text analysis models, and selecting the target text analysis model based on multiple evaluation indicators.
It improves the accuracy and robustness of text analysis, avoids model overfitting, and enhances adaptability to different application areas.
Smart Images

Figure CN116956896B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data analysis technology, and in particular to an artificial intelligence-based text analysis method, system, electronic device, and medium. Background Art
[0002] With the progress of society and the development of science and technology, a large amount of comments and opinions will be generated on the Internet. These comments are of vital importance to understanding user needs, social public opinion trends and social expectations. Sentiment analysis technology based on NLP is a technology that uses people's comments on products, services, organizations, individuals, problems, events, topics, etc. to analyze the corresponding opinions, feelings, emotions, evaluations and attitudes.
[0003] Sentiment analysis technology based on NLP usually requires training of artificial intelligence models, but due to sample imbalance, the model is prone to overfitting and has poor robustness, which makes the model less effective in analyzing text. In addition, the focus of text analysis in different application fields is different. If the same model is used to analyze text in different application fields, the analysis accuracy will be low. Summary of the invention
[0004] In view of this, the present application provides a text analysis method, system, electronic device and medium based on artificial intelligence to solve the technical problem of poor accuracy of text analysis.
[0005] The first aspect of the present application provides a text analysis method based on artificial intelligence, the method comprising:
[0006] In response to a user's instruction to analyze a target text, obtaining an original text related to the target text from a text library;
[0007] Performing data enhancement processing on the original text to obtain enhanced text;
[0008] Identify the application field of the target text, and obtain multiple model frameworks corresponding to the application field from a preset artificial intelligence model library;
[0009] Using the multiple model frameworks to perform training based on the original text and the enhanced text to obtain multiple text analysis models;
[0010] Selecting a target text analysis model from the plurality of text analysis models based on a plurality of evaluation indicators;
[0011] The target text analysis model is used to perform data analysis on the target text.
[0012] In a possible implementation manner, performing data enhancement processing on the original text to obtain enhanced text includes:
[0013] Perform word segmentation on the original text to obtain multiple text keywords;
[0014] Calculate the first weight of the text keywords in the text library and calculate the second weight of the text keywords in the original text;
[0015] Perform enhancement processing on the original text according to the first weight and the second weight of the keyword to obtain the enhanced text.
[0016] In a possible implementation manner, the performing enhancement processing on the original text according to the first weight and the second weight of the keyword to obtain the enhanced text includes:
[0017] Compare the first weight with a first preset weight threshold and compare the second weight with a second preset weight threshold;
[0018] When the first weight is less than the first preset weight threshold and the second weight is greater than the second preset weight threshold, obtain a first random probability from a first preset random probability array and perform masking processing on the keyword with the first random probability;
[0019] When the first weight is less than the first preset weight threshold and the second weight is less than the second preset weight threshold, obtain a second random probability from a second preset random probability array and perform deletion processing on the keyword with the second random probability;
[0020] When the first weight is greater than the first preset weight threshold and the second weight is greater than the second preset weight threshold, obtain a third random probability from a third preset random probability array and perform replacement processing on the keyword with the third random probability.
[0021] In a possible implementation manner, the obtaining the original text related to the target text from the text library includes:
[0022] Obtain a preset number of stored texts from the text library;
[0023] Obtain the similarity between the stored text and the target text;
[0024] Obtain the target stored text with a similarity greater than a preset similarity threshold from the stored texts;
[0025] Perform clustering analysis on all the stored texts in the text library to obtain multiple text clusters;
[0026] Determine the text cluster including the target stored text as the target text cluster;
[0027] Determine the stored text in the target text cluster as the original text related to the target text.
[0028] In a possible implementation manner, the training the multiple model frameworks based on the original text and the enhanced text to obtain multiple text analysis models includes:
[0029] Extract the theme information of the original text and the enhanced text;
[0030] Obtain a combined feature vector according to the theme information;
[0031] Use the multiple model frameworks to train based on the combined feature vector to obtain multiple text analysis models.
[0032] In a possible implementation manner, the extracting the theme information of the original text and the enhanced text includes:
[0033] Use the hierarchical Dirichlet process algorithm to perform theme extraction on the original text and the enhanced text to obtain the theme information, where the theme information includes text - theme distribution and theme - word distribution.
[0034] In a possible implementation manner, the selecting a target text analysis model from the multiple text analysis models based on multiple evaluation indicators includes:
[0035] Display each text analysis model and the corresponding multiple evaluation indicator values, and use the text analysis model selected by the user as the target text analysis model;
[0036] Calculate the weighted mean of the multiple evaluation indicator values corresponding to each text analysis model, and use the text analysis model with the largest weighted mean as the target text analysis model, or use the text analysis model with a weighted mean greater than the average weighted mean as the target text analysis model.
[0037] The second aspect of this application provides a text analysis system based on artificial intelligence, and the system includes:
[0038] A text acquisition module, configured to obtain the original text related to the target text from a text library in response to an analysis instruction of the target text by a user;
[0039] An enhancement processing module, configured to perform data enhancement processing on the original text to obtain enhanced text;
[0040] A model acquisition module, configured to identify the application field of the target text and obtain multiple model frameworks corresponding to the application field from a preset model library;
[0041] A model training module, configured to train based on the original text and the augmented text by using the multiple model frameworks to obtain multiple text analysis models;
[0042] An index evaluation module, configured to select a target text analysis model from the multiple text analysis models based on multiple evaluation indexes;
[0043] A text analysis module, configured to perform data analysis on the target text by using the target text analysis model.
[0044] A third aspect of the present application provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the steps of the text analysis method based on artificial intelligence are implemented.
[0045] A fourth aspect of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the text analysis method based on artificial intelligence are implemented.
[0046] For the text analysis method, system, electronic device, and medium based on artificial intelligence provided in the embodiments of the present application, when analyzing a target text, the present application obtains the original text related to the target text and performs data augmentation processing on the original text, ensuring the data quality of the training model and expanding the data volume of the training model, avoiding model overfitting, and improving the analysis effect on the target text; by obtaining multiple model frameworks corresponding to the application field of the target text from a preset model library, the model can be trained targeted, improving the performance of the model. After training multiple text analysis models based on the original text and the augmented text by using the multiple model frameworks, a target text analysis model is selected based on multiple evaluation indexes to perform data analysis on the target text, further improving the analysis effect on the target text. Description of the Drawings
[0047] Figure 1 is a flowchart of the text analysis method based on artificial intelligence shown in the embodiments of the present application;
[0048] Figure 2 is a functional module diagram of the text analysis system based on artificial intelligence shown in the embodiments of the present application;
[0049] Figure 3 is a structural diagram of the electronic device shown in the embodiments of the present application. Detailed Embodiments
[0050] To make the objectives, technical solutions, and advantages of this application clearer, the following will describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application.
[0051] The text analysis method based on artificial intelligence provided in the embodiments of the present invention is executed by an electronic device. Correspondingly, the text analysis system based on artificial intelligence runs in the electronic device.
[0052] Figure 1 It is a flowchart of the text analysis method based on artificial intelligence provided in Embodiment 1 of the present invention. The text analysis method based on artificial intelligence specifically includes the following steps. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted.
[0053] S11. In response to the user's analysis instruction for the target text, obtain the original text related to the target text from the text library.
[0054] Among them, the target text refers to the data that needs to be text-analyzed.
[0055] A text library is pre-set in the electronic device, and a plurality of texts are stored in the text library. For the convenience of description, the texts stored in the text library are referred to as stored texts.
[0056] When the user uploads the target text to the electronic device, it can trigger an analysis instruction for the electronic device to analyze the target text, so as to obtain a plurality of stored texts from the text library. The obtained plurality of stored texts are relevant to the target text.
[0057] In a possible implementation manner, the obtaining the original text related to the target text from the text library includes:
[0058] Obtain a preset number of stored texts from the text library;
[0059] Obtain the similarity between the stored text and the target text;
[0060] Obtain the target stored text with a similarity greater than a preset similarity threshold from the stored texts;
[0061] Perform clustering analysis on all the stored texts in the text library to obtain a plurality of text clusters;
[0062] Determine the text cluster including the target stored text as the target text cluster;
[0063] Determine the stored texts in the target text cluster as the original texts related to the target text.
[0064] Since there is a large amount of stored text in the text library, in order to quickly obtain the original text related to the target text from the large amount of stored text, a small number of stored texts can be randomly obtained from the text library according to a preset quantity, and then the similarity between each obtained stored text and the target text can be calculated through the Euclidean distance, so as to judge the relevance between the obtained stored text and the target text based on the similarity.
[0065] When the similarity of the obtained stored text in the text library is greater than the preset similarity threshold, it indicates that the stored text has a strong relevance to the target text, and then the stored text is saved as the target stored text. When the similarity of the obtained stored text in the text library is less than the preset similarity threshold, it indicates that the stored text has a weak relevance to the target text.
[0066] Next, use the clustering analysis algorithm to analyze all the stored texts in the text library and divide all the stored texts into multiple text clusters. Since the stored texts in the same text cluster have strong relevance and the stored texts in different text clusters have weak relevance, therefore, if a certain text cluster includes the target stored text, it indicates that all the stored texts in this text cluster have strong relevance to the target text, and then this text cluster is determined as the target text cluster, and the stored texts in the target text cluster are determined as the original texts related to the target text.
[0067] Exemplarily, assume that all the stored texts in the text library are divided into 5 text clusters, and 50 stored texts are randomly obtained from the text library. Among them, when the similarity between the stored text D1 and the target text is greater than 0.8, the stored text D1 is saved as the target stored text in the target text set; if the first text cluster includes the stored text D1, then the first text cluster is used as the target text cluster, and the stored texts in the first text cluster are determined as the original texts related to the target text.
[0068] In an alternative embodiment, after determining the text cluster including the target stored text as the target text cluster, the number of target stored texts in the target text cluster can also be calculated, and the stored texts in the target text cluster with the largest number are determined as the original texts related to the target text. Or, the stored texts in the target text clusters with the number exceeding the average number are determined as the original texts related to the target text.
[0069] In the above optional implementation, by obtaining a preset number of stored texts and determining, according to the similarity, the target stored texts with relatively strong relevance to the target text from the obtained stored texts, based on the principle that the stored texts in the same text cluster have relatively strong relevance, all the stored texts in the text library are divided into multiple text clusters, so as to determine the text cluster including the target stored texts as the target text cluster, and further determine the stored texts in the target text cluster as the number of original texts related to the target text. This avoids calculating the similarity between each stored text in the text library and the target text, reduces the amount of data calculation, and thus improves the efficiency of obtaining the original texts related to the target text.
[0070] In a possible implementation, before obtaining the original texts related to the target text from the text library, the target text can also be subjected to cleaning processing. The cleaning processing may include, but is not limited to: deleting irrelevant information in the target text, removing redundant punctuation marks, screening short texts, correcting typos, filling missing values, normalization processing, etc.
[0071] S12. Perform data augmentation processing on the original texts to obtain augmented texts.
[0072] To avoid the problem that the small amount of data of the obtained original texts may cause the text analysis model to overfit, the original texts can be expanded based on data augmentation technology to increase the amount of data of the original texts. Training the text analysis model based on a large amount of texts can improve the generalization ability of the text analysis model and adapt to more application scenarios.
[0073] In a possible implementation, the performing data augmentation processing on the original texts to obtain augmented texts includes:
[0074] Perform word segmentation processing on the original texts to obtain multiple text keywords;
[0075] Calculate the first weight of the text keywords in the text library and calculate the second weight of the text keywords in the original texts;
[0076] Perform augmentation processing on the original texts according to the first weight and the second weight of the keywords to obtain the augmented texts.
[0077] Word segmentation is one of the basic operations of natural language processing, that is, splitting continuous text into individual tokens. Most natural language processing tools and language models process and analyze language texts at the word level, and most raw corpora are presented in the form of string texts. Therefore, it is necessary to perform word segmentation processing on the original texts before analyzing the texts.
[0078] A word segmenter can be used to identify word boundaries based on preset rules, splitting the original text into a series of phrases to obtain multiple text keywords in the original text. These text keywords can provide some summary information or important features of the text content, facilitating subsequent text analysis and processing.
[0079] Each stored text in the text library corresponds to multiple text keywords. Based on the multiple text keywords corresponding to each stored text, a text keyword library can be obtained. The first weight of the text keyword in the text library can be obtained by calculating the TF-IDF value of the text keyword in the text library. The second weight of the text keyword in the original text can be obtained by calculating the TF-IDF value of the multiple text keywords corresponding to the text keyword in the original text.
[0080] Enhanced processing is performed on the keyword in the original text based on the first weight and the second weight to obtain the enhanced text.
[0081] In a possible implementation manner, the enhanced processing of the original text according to the first weight and the second weight of the keyword to obtain the enhanced text includes:
[0082] Compare the first weight with a first preset weight threshold, and compare the second weight with a second preset weight threshold;
[0083] When the first weight is less than the first preset weight threshold and the second weight is greater than the second preset weight threshold, obtain a first random probability from a first preset random probability array, and perform masking processing on the keyword with the first random probability;
[0084] When the first weight is less than the first preset weight threshold and the second weight is less than the second preset weight threshold, obtain a second random probability from a second preset random probability array, and perform deletion processing on the keyword with the second random probability;
[0085] When the first weight is greater than the first preset weight threshold and the second weight is greater than the second preset weight threshold, obtain a third random probability from a third preset random probability array, and perform replacement processing on the keyword with the third random probability.
[0086] When processing the keyword, by comparing the first weight of the keyword with the first preset weight threshold to determine whether the first weight is less than the preset threshold, the importance of the keyword in the text library is judged; and by comparing the second weight of the keyword with the second preset weight threshold to determine whether the second weight is less than the preset threshold, the importance of the keyword in the original text is judged.
[0087] When the first weight is less than the first preset weight threshold and the second weight is greater than the second preset weight threshold, it indicates that the keyword is not important in the text library but is relatively important in the original text. Then, a probability value is randomly selected from the first preset random probability array as the first random probability, and the keyword is masked with the first random probability, that is, one or some words in the keyword are covered to form a new keyword.
[0088] When the first weight is less than the first preset weight threshold and the second weight is less than the second preset weight threshold, it indicates that the keyword is not important in the text library and is not important in the original text. Then, a probability value is randomly selected from the second preset random probability array as the second random probability, and the keyword is deleted with the second random probability, that is, the keyword is removed or erased from the original text.
[0089] When the first weight is greater than the first preset weight threshold and the second weight is greater than the second preset weight threshold, it indicates that the keyword is relatively important in the text library and is relatively important in the original text. Then, a probability value is randomly selected from the third preset random probability array as the third random probability, and the keyword is replaced with the third random probability, that is, the original keyword is replaced with a synonym.
[0090] It should be understood that when the first weight is greater than the first preset weight threshold, there is no situation where the second weight is less than the second preset weight threshold.
[0091] In the above optional implementation, new words can be introduced through random masking and synonym replacement operations, enabling the model to generalize to words not in the training set. This is equivalent to introducing a certain degree of noise into the original text, which helps prevent overfitting of the text analysis model. During the process of enhancing the original text, different enhancement methods are selected according to the importance of keywords in the text library and the original text, making the enhancement of the original text more targeted and improving the effect of data enhancement for the original text, thus contributing to the accuracy of the text analysis model. Although the newly generated enhanced text retains the category label of the original text, it may change the meaning of the sentences in the original text, resulting in sentences with incorrect labels, which may reduce the performance of the text analysis model. Therefore, the original text is enhanced in the form of a random probability to ensure that the performance of the text analysis model does not decline.
[0092] S13. Identify the application field of the target text and obtain multiple model frameworks corresponding to the application field from a preset artificial intelligence model library.
[0093] The preset artificial intelligence model library stores multiple model frameworks. Each model framework corresponds to one or more application fields, and one application field can also correspond to one or more model frameworks.
[0094] Since it is impossible to precisely know which framework of text analysis model is suitable for the target text or which framework of text analysis model has the best effect, in this application, by analyzing the application field where the target text is located, such as the financial field, legal field, sentiment field, medical field, etc., multiple model frameworks related to the application field are selected from the preset artificial intelligence model library for subsequent establishment of the text analysis model.
[0095] S14. Use the multiple model frameworks to train based on the original text and the enhanced text to obtain multiple text analysis models.
[0096] Based on the artificial intelligence model framework, a text analysis model is constructed, and the original text and the enhanced text are used to train the text analysis model, thereby obtaining multiple text analysis models based on different architectures.
[0097] Exemplarily, a text analysis model can be constructed based on a Long Short-Term Memory (LSTM) model, or a text analysis model can also be constructed based on a Recurrent Neural Network (RNN), or a Convolutional Neural Networks (CNN).
[0098] In a possible implementation, training the multiple text analysis models based on the original text and the enhanced text using the multiple model frameworks includes:
[0099] Extracting the topic information of the original text and the enhanced text;
[0100] Obtaining a combined feature vector according to the topic information;
[0101] Training the multiple text analysis models using the multiple model frameworks based on the combined feature vector.
[0102] The method of mining topic models can be used to extract the topic information of the original text and the enhanced text, and each topic information contains a series of related words.
[0103] Generate a feature vector using the extracted topic information, that is, represent each text as a vector, where each dimension represents a topic, and the value represents the number of occurrences or weights of the relevant topics in the text.
[0104] By using a pre-trained word embedding model, convert each word in the original text and the enhanced text into a corresponding word embedding vector, and obtain a combined feature vector by combining the word embedding vector and the topic information to represent the entire sentence. When generating the combined feature vector, the relationships between different words can be considered. For example, multiple word embedding vectors can be combined into a feature vector through simple addition, summation, and averaging operations. In this way, a sentence with multiple words can be represented by a vector with a fixed dimension.
[0105] Exemplarily, when there are 10 topics in a topic model, each text can be represented as a 10-dimensional vector, where each dimension represents a topic, and calculate the number of occurrences or weights of each topic in the text to obtain a combined feature vector composed of topic information.
[0106] Use the combined feature vector as the input of each model framework and train to obtain a text analysis model.
[0107] In a possible implementation, the extracting the topic information of the original text and the enhanced text includes:
[0108] Using the hierarchical Dirichlet process algorithm to perform topic extraction on the original text and the enhanced text to obtain the topic information, where the topic information includes text-topic distribution and topic-word distribution.
[0109] The Hierarchical Dirichlet Process (HDP) algorithm is used to extract the topics of the original text and the augmented text respectively, and the text-topic distribution and topic-word distribution information of the original text and the augmented text are obtained. Each text can contain multiple topics, and each topic is represented by a set of words. The text-topic distribution indicates the topics contained in the text and the weight of each topic, representing the importance and contribution degree of different topics in the text. The topic-word distribution represents the probability distribution of words in each topic. Through the topic-word distribution, the concepts represented by the topic in the corpus or the words dominated by the topic can be understood.
[0110] In the above optional implementation, by extracting the topic information, it helps to understand the topic structure of the original text and the augmented text, which is beneficial to the subsequent tasks.
[0111] S15, Select a target text analysis model from multiple text analysis models based on multiple evaluation metrics.
[0112] After training multiple text analysis models, a test set can be used to evaluate the trained multiple text analysis models, and based on multiple evaluation metrics, such as accuracy, precision, recall, and F1 score, etc., evaluate the performance of each text analysis model, and select the text analysis model with the best performance as the target text analysis model.
[0113] In an optional implementation, each text analysis model and the corresponding multiple evaluation metric values can be displayed, and the user can select one or more text analysis models as the target text analysis model according to actual needs.
[0114] In an optional implementation, the weighted mean of the multiple evaluation metric values corresponding to each text analysis model can be calculated, and the text analysis model with the largest weighted mean is used as the target text analysis model, or the text analysis model with a weighted mean greater than the average weighted mean is used as the target text analysis model.
[0115] S16, Use the target text analysis model to perform data analysis on the target text.
[0116] Input the target text into the target text analysis model, and then the target text analysis model can perform data analysis on the target text to obtain analysis results, such as sentiment category, topic classification, etc.
[0117] When analyzing the target text, this application ensures the data quality of the training model and expands the data volume of the training model by obtaining the original text related to the target text and performing data augmentation on the original text, avoiding model overfitting, and improving the analysis effect of the target text. By obtaining multiple model frameworks corresponding to the application field of the target text from a preset model library, the model can be trained in a targeted manner, improving the performance of the model. After using multiple model frameworks to train multiple text analysis models based on the original text and the enhanced text, the target text analysis model is selected based on multiple evaluation indicators to perform data analysis on the target text, further improving the analysis effect of the target text.
[0118] Figure 2 It is the structural diagram of the text analysis system based on artificial intelligence provided in the second embodiment of the present invention.
[0119] In some embodiments, the text analysis system 20 based on artificial intelligence may include multiple functional modules composed of computer program segments. The computer programs of each program segment in the text analysis system 20 based on artificial intelligence can be stored in the memory of the electronic device and executed by at least one processor to execute (see details in Figure 1 the description) the functions of text analysis based on artificial intelligence.
[0120] In this embodiment, the text analysis system 20 based on artificial intelligence can be divided into multiple functional modules according to the functions it performs. The functional modules may include: a text acquisition module 201, an enhancement processing module 202, a model acquisition module 203, a model training module 204, an index evaluation module 205, and a text analysis module 206. The module referred to in the present invention means a series of computer program segments that can be executed by at least one processor and can complete fixed functions, and are stored in the memory. In this embodiment, the functions of each module will be described in detail in the subsequent embodiments.
[0121] The text acquisition module 201 is used to obtain the original text data related to the target text data from the text database in response to the user's analysis instruction for the target text data.
[0122] In response to the user's analysis instruction for the target text, obtain the original text related to the target text from the text library.
[0123] Wherein, the target text refers to the data that needs to be text-analyzed.
[0124] A text library is pre-set in the electronic device, and multiple texts are stored in the text library. For the convenience of description, the texts stored in the text library are called stored texts.
[0125] When a user uploads target text to an electronic device, an analysis instruction for the electronic device to analyze the target text can be triggered, so as to obtain multiple stored texts from a text library, and the multiple obtained stored texts are relevant to the target text.
[0126] In a possible implementation manner, the obtaining of the original text relevant to the target text from the text library includes:
[0127] Obtain a preset number of stored texts from the text library;
[0128] Obtain the similarity between the stored text and the target text;
[0129] Obtain target stored texts from the stored texts whose similarity is greater than a preset similarity threshold;
[0130] Perform clustering analysis on all the stored texts in the text library to obtain multiple text clusters;
[0131] Determine the text cluster including the target stored text as the target text cluster;
[0132] Determine the stored texts in the target text cluster as the original texts relevant to the target text.
[0133] Since there are a large number of stored texts in the text library, in order to quickly obtain the original texts relevant to the target text from the large number of stored texts, a small number of stored texts can be randomly obtained from the text library according to a preset number first, and then the similarity between each obtained stored text and the target text can be calculated through the Euclidean distance, so as to judge the relevance between the obtained stored text and the target text based on the similarity.
[0134] When the similarity of the stored text obtained from the text library is greater than the preset similarity threshold, it indicates that the stored text has a strong relevance to the target text, and then the stored text is saved as the target stored text. When the similarity of the stored text obtained from the text library is less than the preset similarity threshold, it indicates that the stored text has a weak relevance to the target text.
[0135] Next, use a clustering analysis algorithm to analyze all the stored texts in the text library and divide all the stored texts into multiple text clusters. Since the stored texts in the same text cluster have a strong relevance and the stored texts in different text clusters have a weak relevance, therefore, if a certain text cluster includes the target stored text, it indicates that all the stored texts in this text cluster have a strong relevance to the target text, and then this text cluster is determined as the target text cluster, and the stored texts in the target text cluster are determined as the original texts relevant to the target text.
[0136] Exemplarily, assume that all the stored texts in the text library are divided into 5 text clusters, and 50 stored texts are randomly retrieved from the text library. Among them, when the similarity between the stored text D1 and the target text is greater than 0.8, the stored text D1 is saved as the target stored text in the target text set; if the first text cluster includes the stored text D1, the first text cluster is used as the target text cluster, and the stored texts in the first text cluster are determined as the original texts related to the target text.
[0137] In an alternative embodiment, after determining the text cluster including the target stored text as the target text cluster, the number of target stored texts in the target text cluster can also be calculated, and the stored texts in the target text cluster with the largest number are determined as the original texts related to the target text. Or, the stored texts in the target text clusters with a number exceeding the average number are determined as the original texts related to the target text.
[0138] In the above alternative embodiment, by retrieving a preset number of stored texts and determining the target stored texts with strong relevance to the target text from the retrieved stored texts according to the similarity, based on the principle that the stored texts in the same text cluster have strong relevance, all the stored texts in the text library are divided into multiple text clusters, so as to determine the text cluster including the target stored text as the target text cluster, and further determine the stored texts in the target text cluster as the number of original texts related to the target text. It avoids calculating the similarity between each stored text in the text library and the target text, reduces the data calculation amount, and thus improves the efficiency of obtaining the original texts related to the target text.
[0139] In a possible implementation, before retrieving the original texts related to the target text from the text library, the target text can also be cleaned. The cleaning process may include, but is not limited to: deleting irrelevant information in the target text, removing redundant punctuation marks, screening short texts, correcting spelling mistakes, filling missing values, normalizing, etc.
[0140] The enhancement processing module 202 is used to perform data enhancement processing on the original text data to obtain enhanced text data.
[0141] Perform data enhancement processing on the original text to obtain enhanced text.
[0142] In order to avoid the overfitting phenomenon of the text analysis model caused by the small amount of data of the retrieved original texts, the original texts can be expanded based on data enhancement technology to increase the amount of data of the original texts. Training the text analysis model based on a large amount of texts can improve the generalization ability of the text analysis model and adapt to more application scenarios.
[0143] In a possible implementation, the data enhancement processing of the original text to obtain the enhanced text includes:
[0144] Performing word segmentation on the original text to obtain multiple text keywords;
[0145] Calculating a first weight of the text keyword in the text library and calculating a second weight of the text keyword in the original text;
[0146] Performing enhancement processing on the original text according to the first weight and the second weight of the keyword to obtain the enhanced text.
[0147] Word segmentation is one of the basic operations of natural language processing, that is, splitting continuous text into individual tokens. Most natural language processing tools and language models process and analyze language text at the word level, and most raw corpora are presented in the form of string text. Therefore, it is necessary to perform word segmentation on the original text before analyzing the text.
[0148] A tokenizer can be used based on a preset rule to identify word boundaries, splitting the original text into a series of phrases to obtain multiple text keywords in the original text. These text keywords can provide some summary information or important features of the text content, which helps in subsequent text analysis and processing.
[0149] Each stored text in the text library corresponds to multiple text keywords. According to the multiple text keywords corresponding to each stored text, a text word library can be obtained. The first weight of the text keyword in the text library can be obtained by calculating the TF-IDF value of the text keyword in the text library. The second weight of the text keyword in the original text can be obtained by calculating the TF-IDF value of the multiple text keywords corresponding to the text keyword in the original text.
[0150] Performing enhancement processing on the keyword in the original text based on the first weight and the second weight to obtain the enhanced text.
[0151] In a possible implementation, the performing enhancement processing on the original text according to the first weight and the second weight of the keyword to obtain the enhanced text includes:
[0152] Comparing the first weight with a first preset weight threshold and comparing the second weight with a second preset weight threshold;
[0153] When the first weight is less than the first preset weight threshold and the second weight is greater than the second preset weight threshold, obtain a first random probability from the first preset random probability array, and perform masking processing on the keyword with the first random probability;
[0154] When the first weight is less than the first preset weight threshold and the second weight is less than the second preset weight threshold, obtain a second random probability from the second preset random probability array, and perform deletion processing on the keyword with the second random probability;
[0155] When the first weight is greater than the first preset weight threshold and the second weight is greater than the second preset weight threshold, obtain a third random probability from the third preset random probability array, and perform replacement processing on the keyword with the third random probability.
[0156] When processing the keyword, by comparing the first weight of the keyword with the first preset weight threshold to determine whether the first weight is less than the preset threshold, the importance of the keyword in the text library is judged; and by comparing the second weight of the keyword with the second preset weight threshold to determine whether the second weight is less than the preset threshold, the importance of the keyword in the original text is judged.
[0157] When the first weight is less than the first preset weight threshold and the second weight is greater than the second preset weight threshold, it indicates that the keyword is not important in the text library but is relatively important in the original text. Then, randomly select a probability value from the first preset random probability array as the first random probability, and perform masking processing on the keyword with the first random probability, that is, cover one or some words in the keyword to form a new keyword.
[0158] When the first weight is less than the first preset weight threshold and the second weight is less than the second preset weight threshold, it indicates that the keyword is not important in the text library and is not important in the original text. Then, randomly select a probability value from the second preset random probability array as the second random probability, and perform deletion processing on the keyword with the second random probability, that is, remove or erase the keyword from the original text.
[0159] When the first weight is greater than the first preset weight threshold and the second weight is greater than the second preset weight threshold, it indicates that the keyword is relatively important in the text library and is relatively important in the original text. Then, randomly select a probability value from the third preset random probability array as the third random probability, and perform replacement processing on the keyword with the third random probability, that is, replace the original keyword with a synonym.
[0160] It should be understood that when the first weight is greater than the first preset weight threshold, there is no case where the second weight is less than the second preset weight threshold.
[0161] In the above optional embodiments, new words can be introduced through random masking and synonym replacement operations, enabling the model to generalize to words not in the training set. This is equivalent to introducing a certain degree of noise into the original text, which helps prevent overfitting of the text analysis model. During the process of enhancing the original text, different enhancement methods are selected according to the importance of keywords in the text library and the original text, making the enhancement of the original text more targeted and improving the effect of data enhancement for the original text, thus contributing to the accuracy of the text analysis model. Although the newly generated enhanced text retains the category label of the original text, it may change the meaning of the sentences in the original text, resulting in sentences with incorrect labels, which may reduce the performance of the text analysis model. Therefore, the original text is enhanced in the form of a random probability, ensuring that the performance of the text analysis model does not decline.
[0162] The model acquisition module 203 is configured to identify the application field of the target text data and obtain multiple model frameworks corresponding to the application field from a preset model library.
[0163] Identify the application field of the target text and obtain multiple model frameworks corresponding to the application field from a preset artificial intelligence model library.
[0164] The preset artificial intelligence model library stores multiple model frameworks. Each model framework corresponds to one or more application fields, and one application field can also correspond to one or more model frameworks.
[0165] Since it is impossible to accurately know which framework of text analysis model is suitable for the target text or which framework of text analysis model has the best effect, in this application, by analyzing the application field where the target text is located, such as the financial field, legal field, sentiment field, medical field, etc., multiple model frameworks related to the application field are selected from the preset artificial intelligence model library for subsequent establishment of the text analysis model.
[0166] The model training module 204 is configured to use the multiple model frameworks to train based on the original text data and the enhanced text data to obtain multiple text data analysis models.
[0167] Use the multiple model frameworks to train based on the original text and the enhanced text to obtain multiple text analysis models.
[0168] Based on an artificial intelligence model framework, a text analysis model is constructed, and the original text and the enhanced text are used to train the text analysis model, so as to obtain multiple text analysis models based on different architectures.
[0169] Exemplarily, a text analysis model can be constructed based on a Long Short-Term Memory (LSTM) model, or a text analysis model can be constructed based on a Recurrent Neural Network (RNN), or a Convolutional Neural Networks (CNN).
[0170] In a possible implementation manner, the training using the multiple model frameworks based on the original text and the enhanced text to obtain multiple text analysis models includes:
[0171] Extract the topic information of the original text and the enhanced text;
[0172] Obtain a combined feature vector according to the topic information;
[0173] Use the multiple model frameworks to train based on the combined feature vector to obtain multiple text analysis models.
[0174] The method of mining topic models can be used to extract the topic information of the original text and the enhanced text, and each topic information contains a series of related words.
[0175] Generate a feature vector using the extracted topic information, that is, represent each text as a vector, where each dimension represents a topic, and the value represents the occurrence times or weights of the relevant topics in the text.
[0176] By using a pre-trained word embedding model, each word in the original text and the enhanced text is converted into a corresponding word embedding vector, and by combining the word embedding vector and the topic information, a combined feature vector is obtained to represent the entire sentence. When generating the combined feature vector, the relationships between different words can be considered. For example, multiple word embedding vectors can be combined into a feature vector through simple addition, summation, and averaging operations. In this way, a sentence with multiple words can be represented by a vector with a fixed dimension.
[0177] Exemplarily, when there are 10 topics in a topic model, each text can be represented as a 10-dimensional vector, where each dimension represents a topic, and the occurrence times or weights of each topic in the text are calculated to obtain a combined feature vector composed of topic information.
[0178] Use the combined feature vector as the input for each model framework and train it to obtain a text analysis model.
[0179] In a possible implementation, the extraction of the topic information of the original text and the enhanced text includes:
[0180] Use the Hierarchical Dirichlet Process algorithm to perform topic extraction on the original text and the enhanced text to obtain the topic information, where the topic information includes text-topic distribution and topic-word distribution.
[0181] Use the Hierarchical Dirichlet Process (HDP) algorithm to perform topic extraction on the original text and the enhanced text respectively, and obtain the text-topic distribution and topic-word distribution information of the original text and the enhanced text. Each text can contain multiple topics, and each topic is represented by a set of words. The text-topic distribution indicates the topics contained in the text and the weight of each topic, representing the importance and contribution degree of different topics in the text. The topic-word distribution represents the probability distribution of words in each topic. Through the topic-word distribution, the concepts represented by the topic in the corpus or the words dominated by the topic can be understood.
[0182] The above optional implementation helps to understand the topic structure of the original text and the enhanced text by extracting topic information, which is beneficial to the subsequent tasks.
[0183] The metric evaluation module 205 is used to select a target text data analysis model from multiple text data analysis models based on multiple evaluation metrics.
[0184] Select a target text analysis model from multiple text analysis models based on multiple evaluation metrics.
[0185] After training multiple text analysis models, a test set can be used to evaluate the trained multiple text analysis models, and based on multiple evaluation metrics such as accuracy, precision, recall, and F1 score, etc., evaluate the performance of each text analysis model, and select the text analysis model with the best performance as the target text analysis model.
[0186] In an optional implementation, each text analysis model and the corresponding multiple evaluation metric values can be displayed, and the user can select one or more text analysis models as the target text analysis model according to actual needs.
[0187] In an alternative embodiment, the weighted mean of multiple evaluation metric values corresponding to each text analysis model may be calculated, and the text analysis model with the largest weighted mean is used as the target text analysis model, or the text analysis model with a weighted mean greater than the average weighted mean is used as the target text analysis model.
[0188] The text analysis module 206 is configured to perform data analysis on the target text data by using the target text data analysis model.
[0189] Perform data analysis on the target text by using the target text analysis model.
[0190] Input the target text into the target text analysis model, and then the target text analysis model can perform data analysis on the target text to obtain analysis results, such as sentiment category, topic classification, etc.
[0191] When analyzing the target text in this application, by obtaining the original text related to the target text and performing data augmentation processing on the original text, the data quality of the training model is ensured, and the data volume of the training model is expanded, avoiding overfitting of the model and improving the analysis effect on the target text; by obtaining multiple model frameworks corresponding to the application field of the target text from a preset model library, the model can be trained in a targeted manner to improve the performance of the model. After using multiple model frameworks to train multiple text analysis models based on the original text and the augmented text, the target text analysis model is selected based on multiple evaluation metrics to perform data analysis on the target text, further improving the analysis effect on the target text.
[0192] Refer to Figure 3 As shown, it is a schematic structural diagram of an electronic device provided by an embodiment of this application. In a preferred embodiment of this application, the electronic device 3 includes a memory 31, at least one processor 32, and at least one communication bus 33.
[0193] Those skilled in the art should understand that Figure 3 The structure of the electronic device shown does not constitute a limitation on the embodiments of this application. It can be a bus structure or a star structure. The electronic device 3 may further include more or fewer other hardware or software than shown in the figure, or different component arrangements.
[0194] In some embodiments, the electronic device 3 is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, a microprocessor, an application-specific integrated circuit, a programmable gate array, a digital signal processor, and an embedded device, etc. The electronic device 3 may further include other electronic devices, and the other electronic devices include, but are not limited to, any electronic product that can perform human-computer interaction with a user through means such as a keyboard, a mouse, a remote control, a touchpad, or a voice control device. For example, a personal computer, a tablet computer, a smart phone, a digital camera, etc.
[0195] It should be noted that the electronic device 3 is only an example, and other existing or future electronic products that can be adapted to this application should also be included within the protection scope of this application and are included herein by reference.
[0196] In some embodiments, a computer program is stored in the memory 31, and when the computer program is executed by the at least one processor 32, all or part of the steps in the described text analysis method based on artificial intelligence are implemented. The memory 31 includes a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electrically-erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc memories, magnetic disk memories, tape memories, or any other computer-readable medium capable of carrying or storing data.
[0197] In some embodiments, the at least one processor 32 is the control core of the computer device 3. It connects various components of the entire electronic device 3 through various interfaces and circuits. By running or executing programs or modules stored in the memory 31 and invoking data stored in the memory 31, it performs various functions of the electronic device 3 and processes data. For example, when the at least one processor 32 executes the computer program stored in the memory, it implements all or part of the steps of the text analysis method based on artificial intelligence described in the embodiments of the present application; or implements all or part of the functions of the text analysis method based on artificial intelligence. The at least one processor 32 may be composed of integrated circuits. For example, it may be composed of a single packaged integrated circuit, or may be composed of multiple integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips, etc.
[0198] In some embodiments, the at least one communication bus 33 is configured to enable connection communication between the memory 31 and the at least one processor 32, etc. Although not shown, the electronic device 3 may further include a power supply (such as a battery) for powering each component. Preferably, the power supply can be logically connected to the at least one processor 32 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator. The electronic device 3 may also include a variety of sensors, a Bluetooth module, a Wi-Fi module, etc., which will not be elaborated here.
[0199] The integrated unit implemented in the form of a software function module as described above can be stored in a computer-readable storage medium. The above software function module is stored in a storage medium and includes several instructions for causing an electronic device (which may be a personal computer, an electronic device, or a network device, etc.) or a processor to execute part of the methods described in the embodiments of the present application.
[0200] In several embodiments provided by the present application, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.
[0201] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical unit, and it may be located in one place or distributed across multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0202] The terms used in the embodiments of this application are only for the purpose of describing specific embodiments and are not intended to limit this application. As used in the specification and appended claims of this application, the singular forms "a", "an", "the", "above", "said", "this" are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in this application refers to any or all possible combinations including one or more of the listed items. The terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the embodiments of this application, unless otherwise stated, the meaning of "a plurality" is two or more.
[0203] The above content is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A text analysis method based on artificial intelligence, characterized in that The method includes: In response to a user's analysis instruction for a target text, obtaining the original text related to the target text from a text library; Performing data augmentation processing on the original text to obtain an augmented text; Identifying the application field of the target text and obtaining multiple model frameworks corresponding to the application field from a preset model library; Using the multiple model frameworks to train based on the original text and the augmented text to obtain multiple text analysis models; Selecting a target text analysis model from the multiple text analysis models based on multiple evaluation metrics; Using the target text analysis model to perform data analysis on the target text; Among them, the performing data augmentation processing on the original text to obtain an augmented text includes: Performing word segmentation on the original text to obtain multiple text keywords; Calculating the first weight of the keyword in the text library and calculating the second weight of the keyword in the original text; Performing augmentation processing on the original text according to the first weight and the second weight of the keyword to obtain the augmented text; Among them, the performing augmentation processing on the original text according to the first weight and the second weight of the keyword to obtain the augmented text includes: Comparing the first weight with a first preset weight threshold and comparing the second weight with a second preset weight threshold; When the first weight is less than the first preset weight threshold and the second weight is greater than the second preset weight threshold, obtaining a first random probability from a first preset random probability array and performing a masking process on the keyword with the first random probability; When the first weight is less than the first preset weight threshold and the second weight is less than the second preset weight threshold, obtaining a second random probability from a second preset random probability array and performing a deletion process on the keyword with the second random probability; When the first weight is greater than the first preset weight threshold and the second weight is greater than the second preset weight threshold, obtaining a third random probability from a third preset random probability array and performing a replacement process on the keyword with the third random probability.
2. The text analysis method based on artificial intelligence according to claim 1, wherein The obtaining the original text related to the target text from the text library includes: Obtaining a preset number of stored texts from the text library; Obtaining the similarity between the stored text and the target text; Obtaining a target stored text with a similarity greater than a preset similarity threshold from the stored texts; Performing clustering analysis on all the stored texts in the text library to obtain multiple text clusters; Determining the text cluster including the target stored text as the target text cluster; Determining the stored texts in the target text cluster as the original texts related to the target text.
3. The text analysis method based on artificial intelligence according to claim 2, wherein The using the multiple model frameworks to train based on the original text and the augmented text to obtain multiple text analysis models includes: Extracting the theme information of the original text and the augmented text; Obtaining a combined feature vector according to the theme information; Using the multiple model frameworks to train based on the combined feature vector to obtain multiple text analysis models.
4. The text analysis method based on artificial intelligence according to claim 3, wherein The extraction of the topic information of the original text and the enhanced text includes: Using the hierarchical Dirichlet process algorithm to extract the topics of the original text and the enhanced text to obtain the topic information, where the topic information includes text-topic distribution and topic-word distribution.
5. The text analysis method based on artificial intelligence according to claim 4, wherein The selection of the target text analysis model from multiple text analysis models based on multiple evaluation indicators includes: Displaying each text analysis model and the corresponding multiple evaluation index values, and taking the text analysis model selected by the user as the target text analysis model; or Calculating the weighted mean of the multiple evaluation index values corresponding to each text analysis model, taking the text analysis model with the largest weighted mean as the target text analysis model, or taking the text analysis model with a weighted mean greater than the average weighted mean as the target text analysis model.
6. An artificial intelligence-based text analysis system, characterized in that, For implementing the method according to any one of claims 1 to 5 above, the system includes: A text acquisition module, configured to obtain the original text related to the target text from a text library in response to an analysis instruction of the user for the target text; An enhancement processing module, configured to perform data enhancement processing on the original text to obtain an enhanced text; A model acquisition module, configured to identify the application field of the target text and obtain multiple model frameworks corresponding to the application field from a preset model library; A model training module, configured to use the multiple model frameworks to train based on the original text and the enhanced text to obtain multiple text analysis models; An index evaluation module, configured to select a target text analysis model from multiple text analysis models based on multiple evaluation indicators; A text analysis module, configured to perform data analysis on the target text using the target text analysis model.
7. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the artificial intelligence-based text analysis method according to any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the artificial intelligence-based text analysis method according to any one of claims 1 to 5.
Citation Information
Patent Citations
Abstract extraction method and device based on semantic analysis, equipment and medium
CN113836274A
Data expansion method and device based on keyword recognition, equipment and medium
CN114492390A