Text analysis method and system, electronic equipment and storage medium
By combining semantic similarity and emotional scores, the threshold interval and stability center are dynamically constructed, which solves the problem of ignoring emotional information in the existing technology, and improves the accuracy and robustness of keyword extraction, which is suitable for public opinion analysis, intelligent recommendation and other fields.
Patent Information
- Application Number
- CN202510666220.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-08-26
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing keyword extraction methods mainly rely on semantic similarity calculation or statistical frequency analysis, ignoring the emotional information in the text, resulting in significant shortcomings in expressing the emotional tendencies of the text, especially when dealing with texts with obvious emotional tendencies, they cannot accurately reflect the author's true intention or the reader's emotional reaction.
Using a method combining semantic similarity and emotional score, semantic keywords are extracted through the first model and the second model to obtain emotional keywords, and a threshold interval and stable center are dynamically constructed based on historical data. Multi-dimensional evaluation and correction of keywords are combined with relative stability indicators to screen out comprehensive keywords that have high consistency in semantic expression and emotional tendencies.
It improves the accuracy and robustness of keyword extraction, and enhances the system's understanding of complex text content. It is especially suitable for advanced application scenarios such as public opinion analysis, intelligent recommendation, and emotional recognition that require both semantics and emotions.
Smart Images

Figure CN120542433A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of text analysis technology, and in particular to a text analysis method, system, electronic device, and storage medium. Background Art
[0002] With the rapid development of information technology, text analysis plays a vital role in multiple fields, including information retrieval, public opinion monitoring, and intelligent recommendation. Keyword extraction, as one of the core tasks of text analysis, is of great significance for understanding text content and mining key information. However, most current keyword extraction methods rely primarily on semantic similarity calculations or statistical frequency analysis, often ignoring the emotional information contained in the text. This results in the extracted keywords being significantly inadequate in expressing the text's emotional tendencies. Existing keyword extraction models typically employ a single semantic analysis approach, such as those based on TF-IDF, TextRank, or BERT. While these methods can identify vocabulary related to the text topic to a certain extent, most of these methods fail to effectively integrate the text's emotional characteristics for comprehensive evaluation. Due to a lack of understanding of text sentiment, the extracted keywords may not accurately reflect the author's true intentions or the reader's emotional response. This limitation is particularly prominent when dealing with texts with obvious emotional tendencies (such as positive evaluations or negative criticisms). Summary of the Invention
[0003] The present application aims to solve one of the technical problems in the related art at least to a certain extent.
[0004] To this end, the first purpose of this application is to propose a text analysis method to improve the accuracy of keyword extraction.
[0005] The second objective of this application is to provide a text analysis system.
[0006] The third objective of this application is to provide an electronic device.
[0007] The fourth object of this application is to provide a computer-readable storage medium.
[0008] A fifth object of this application is to provide a computer program product.
[0009] To achieve the above-mentioned purpose, the first embodiment of the present application proposes a text analysis method, comprising: step 1, obtaining a text sample to be analyzed and performing preprocessing;
[0010] Step 2: input the pre-processed text sample into the first model for analysis to obtain the first keyword;
[0011] Step 3: Input the pre-processed text sample into the second model for analysis to obtain a sentiment score, and obtain a second keyword based on the sentiment score;
[0012] Step 4: determining a first similarity between the first keyword and a first related corpus in the text library, and determining a second similarity between the second keyword and a second related corpus in the text library;
[0013] Step 5: Perform comprehensive analysis on the first similarity and the second similarity to obtain comprehensive keywords.
[0014] In some implementations, step 5, performing a comprehensive analysis on the first similarity and the second similarity to obtain a comprehensive keyword, includes the following steps:
[0015] Step 51, extracting the same words in the first keyword and the second keyword, and the first similarity and the second similarity of each of the same words;
[0016] Step 52: determining a difference between the first similarity and the second similarity;
[0017] Step 53: Preset a threshold interval and determine whether the difference value is within the threshold interval. If so, execute step 54; if not, return to step 2.
[0018] Step 54, construct a relative stability index and preset a stability threshold. If the relative stability index is less than the stability threshold, the first keyword or the second keyword is used as a comprehensive keyword; if the relative stability index is greater than or equal to the stability threshold, the first keyword or the second keyword is modified to obtain a comprehensive keyword.
[0019] In some implementations, in addition to obtaining the second keyword, the second model also obtains the sentiment polarity of the preprocessed text sample, where the sentiment polarity includes positive sentiment, negative sentiment, and neutral sentiment, and a +1 label is assigned to the positive sentiment, a -1 label is assigned to the negative sentiment, and a 0 label is assigned to the neutral sentiment.
[0020] In some implementations, presetting a threshold interval includes the following steps:
[0021] Step 531, collecting historical standard text samples and obtaining historical keywords;
[0022] Step 532, determining the text length of the historical standard text sample;
[0023] Step 533: Obtain the historical similarity, historical sentiment score, and historical sentiment polarity of the historical keyword;
[0024] Step 534: obtaining the historical similarity mean and the historical sentiment score standard deviation based on the historical similarity and the historical sentiment score;
[0025] Step 535 : Obtain the upper limit and lower limit of the threshold range according to the text length, the historical similarity mean, the historical sentiment polarity, and the historical sentiment score standard deviation.
[0026] In some implementations, constructing a relative stability index includes the following steps:
[0027] Step 541, defining a relative position factor based on historical sentiment polarity and text length;
[0028] Step 542: Obtain a stable center based on the upper limit, lower limit, text length, historical similarity mean, relative position factor, and historical sentiment score standard deviation;
[0029] Step 543 : Obtain a relative stability index according to the first similarity of the same words, the second similarity of the same words, the stability center, and the sentiment score.
[0030] In some implementations, the first keyword or the second keyword is modified to obtain a comprehensive keyword, including: arranging the relative stability indexes of the identical words from high to low, and taking the first k1 identical words as first relative stability words; taking the k1+1 to k1+n identical words as second relative stability words and marking them as modified words, searching for similar words in a text library based on the modified words, determining the relative stability index of each similar word, searching for the maximum similar word whose relative stability index is greater than the relative stability index of the modified word, replacing the maximum similar word with the modified word, and marking the first relative stability word and the maximum similar word as comprehensive key words.
[0031] To achieve the above-mentioned purpose, a second embodiment of the present application proposes a text analysis system, comprising: a sample processing module, configured to obtain a text sample to be analyzed and perform preprocessing;
[0032] A first analysis module, configured to input the preprocessed text sample into a first model for analysis to obtain a first keyword;
[0033] A second analysis module is configured to input the pre-processed text sample into a second model for analysis to obtain a sentiment score, and obtain a second keyword based on the sentiment score;
[0034] a similarity analysis module, configured to determine a first similarity between the first keyword and a first related corpus in the text library, and to determine a second similarity between the second keyword and a second related corpus in the text library;
[0035] The comprehensive analysis module is used to perform comprehensive analysis on the first similarity and the second similarity to obtain comprehensive keywords.
[0036] To achieve the above-mentioned purpose, the third aspect embodiment of the present application proposes an electronic device, comprising: a processor, and a memory communicatively connected to the processor; the memory stores computer-executable instructions; the processor executes the computer-executable instructions stored in the memory to implement the method described in the first aspect.
[0037] To achieve the above-mentioned purpose, the fourth embodiment of the present application proposes a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the method described in the first aspect.
[0038] To achieve the above-mentioned purpose, the fifth embodiment of the present application proposes a computer program product, including a computer program, which implements the method described in the first aspect when executed by a processor.
[0039] The present application provides a text analysis method, system, electronic device and storage medium, which extracts and optimizes keywords by fusing semantic similarity and sentiment scores, effectively overcoming the keyword bias problem caused by traditional methods ignoring sentiment consistency. The method uses the first model and the second model to extract semantic keywords and sentiment keywords respectively, and dynamically constructs threshold intervals and stability centers based on historical data, and further combines relative stability indicators to perform multi-dimensional evaluation and correction of keywords, thereby screening out comprehensive keywords with high consistency in both semantic expression and sentiment tendency. This mechanism not only improves the accuracy and robustness of keyword extraction, but also enhances the system's ability to understand complex text content. It is especially suitable for advanced application scenarios such as public opinion analysis, intelligent recommendation, and sentiment recognition that require consideration of semantics and emotions. In addition, by introducing auxiliary parameters such as relative position factor, historical sentiment polarity, and text length, the keyword extraction process is made more adaptive and interpretable, significantly improving the technical level and practical value of text analysis.
[0040] Additional aspects and advantages of the present application will be given in part in the description below, and in part will become apparent from the description below, or will be learned through practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0042] Figure 1 A flowchart of a text analysis method provided in an embodiment of the present application;
[0043] Figure 2 A block diagram of a text analysis system provided in an embodiment of the present application;
[0044] Figure 3 A block diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0045] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present application, and should not be construed as limiting the present application.
[0046] The following describes a text analysis method, apparatus, and device according to an embodiment of the present application with reference to the accompanying drawings.
[0047] Figure 1 A flowchart of a text analysis method provided in an embodiment of the present application.
[0048] It should be noted that the execution subject of a text analysis method in an embodiment of the present application is the text analysis device in an embodiment of the present application, and the text analysis device can be configured in an electronic device so that the electronic device can perform a text analysis function.
[0049] Traditional methods fail to fully consider the issue of sentiment consistency during keyword extraction. Even if certain keywords are highly semantically relevant to the text, if they are inconsistent with the overall sentiment of the text, this can lead to misunderstandings or misleading conclusions. More importantly, there are few existing technologies that incorporate sentiment analysis into the keyword extraction process, which makes the system incapable of handling complex and diverse texts. Sentiment analysis can not only help identify the emotional tendencies in the text, but also reveal the deeper meaning hidden behind the text, providing more comprehensive information support for keyword extraction. By combining sentiment scores, it is possible to better screen out keywords that are not only semantically matched but also emotionally consistent, thereby improving the quality and accuracy of keyword extraction.
[0050] like Figure 1 As shown, the text analysis method includes the following steps:
[0051] Step 1: Obtain the text sample to be analyzed and perform preprocessing.
[0052] This embodiment obtains text samples to be analyzed from various sources. These sources may include, but are not limited to, social media platforms, news websites, product review pages, forum posts, etc. To ensure data diversity and representativeness, text samples can be collected from multiple channels and stored in a unified data warehouse.
[0053] The preprocessing steps mainly include the following aspects:
[0054] Remove HTML tags, special symbols (such as emojis, non - ASCII characters), and other irrelevant characters from the text. For example, comments crawled from web pages may contain a large number of HTML tags and JavaScript code, which all need to be cleared. In addition, repeated punctuation marks, extra spaces, and line breaks can also be removed. This step can be achieved through regular expressions or specialized text cleaning tools.
[0055] Converting all letters to lowercase helps standardize text representation and avoid the same word being regarded as different due to case differences. This is particularly important for statistical - based methods, such as TF - IDF calculation.
[0056] Word segmentation refers to splitting a continuous string into meaningful lexical units (tokens). In Chinese, since there are no obvious spaces as delimiters, specialized word segmentation tools, such as the Jieba segmentation library, need to be used. For languages like English, simple splitting can be directly done using spaces, or more complex tokenizers (such as WordPunctTokenizer in NLTK) can be used to handle cases like hyphenated words.
[0057] Stop words refer to those words that frequently appear in the text but contribute little to semantic understanding, such as "的", "是", "我", etc. Removing these words can reduce noise interference and improve analysis efficiency. An appropriate stop word list can be selected according to the specific application scenario, or a custom stop word list can also be defined.
[0058] In step 2, input the preprocessed text sample into the first model for analysis to obtain the first set of keywords.
[0059] The first model used in this embodiment is an algorithm model widely applied in keyword extraction tasks in the prior art, such as TF - IDF, TextRank algorithm, RAKE, or keyword extraction methods based on pre - trained language models. Its goal is to identify the most representative words or phrases from structured or semi - structured text.
[0060] Take, for example, a review of a smartphone user experience: "This phone has a very short battery life, charges slowly, and overheats, but the camera is good." After preprocessing in step 1, this text becomes the following: "The phone has a very short battery life, charges slowly, overheats, and takes good photos." In this example, this text is input into the TextRank model for keyword extraction. TextRank is an unsupervised keyword extraction method based on a graph ranking algorithm. It identifies key terms by constructing a word co-occurrence graph and calculating the importance of each node. Specifically, the model counts the co-occurrence relationships between words and determines which words are more likely to become keywords based on their weight scores in the network. The model outputs the following list of candidate keywords and their weights: mobile phone: 0.32, battery life: 0.29, overheating: 0.27, camera: 0.25, charging: 0.24. The top N highest-weighted words are retained as the final first keyword set.
[0061] Step 3: Input the preprocessed text sample into the second model for analysis to obtain a sentiment score, and obtain a second keyword based on the sentiment score.
[0062] The second model used in this embodiment is a keyword extraction model that integrates sentiment analysis functions, which can not only identify keywords related to the semantics of the text, but also output the sentiment score of each candidate word, and screen out keywords with sentiment tendencies as the second keywords based on this. The text sample is input into a sentiment-aware keyword extraction model based on deep learning, such as a joint model that integrates BERT and a sentiment dictionary. The model first obtains the overall semantic representation of the text through the BERT encoder, and then calculates the semantic relevance and sentiment score of each candidate word based on the local context information of the candidate word and the sentiment value in the sentiment dictionary. The model outputs the candidate word and its corresponding sentiment score. The range of the sentiment score is [-1,1], -1 means strongly negative, 0 means neutral, and +1 means strongly positive. Words with an absolute value of a sentiment score greater than 0.5 are retained as the final second keyword set.
[0063] In addition to obtaining the second keyword, the second model also obtains the sentiment polarity of the preprocessed text sample, where the sentiment polarity includes positive sentiment, negative sentiment, and neutral sentiment. The positive sentiment is assigned a +1 label, the negative sentiment is assigned a -1 label, and the neutral sentiment is assigned a 0 label.
[0064] Step 4: Determine a first similarity between the first keyword and a first related corpus in the text library, and determine a second similarity between the second keyword and a second related corpus in the text library.
[0065] The first relevant corpus is used to match the first keyword, and the second relevant corpus is used to match the second keyword. The first corpus primarily contains general domain terms or high-frequency keyword examples, while the second corpus focuses on sentiment-related vocabulary and their contextual expressions. For example, semantic matching analysis is performed on the first keyword [mobile phone, battery life, heat] and the second keyword [battery life, heat, charging, photo taking]. First, the pre-trained Sentence-BERT model is used to encode each keyword and its corresponding corpus entry into a 768-dimensional vector representation. Then, the cosine similarity formula is used to calculate the degree of semantic matching between the keyword and each corpus entry, and the highest score is selected as the keyword's similarity value. For example, for the keyword "battery life," the most similar corpus in the first corpus is "battery duration," with a similarity score of 0.89; the most similar corpus in the second corpus is "power consumption," with a similarity score of 0.83. Through this step, the system ultimately generates a set of first and second similarity values for each keyword, providing basic data support for subsequent comprehensive stability assessment.
[0066] Step 5: Perform comprehensive analysis on the first similarity and the second similarity to obtain comprehensive keywords.
[0067] Step 51: extract the same words in the first keyword and the second keyword, and the first similarity and the second similarity of each of the same words.
[0068] Step 52: Determine the difference between the first similarity and the second similarity.
[0069] In step 53 , a threshold interval is preset to determine whether the difference value is within the threshold interval. If so, step 54 is executed; if not, the process returns to step 2 .
[0070] When the difference value exceeds the preset threshold, it indicates that the currently extracted first and second keywords are significantly different from the keywords in the historical standard text sample. This may be caused by a variety of factors: for example, a significant change in the text content or sentiment, unreasonable model parameter settings, insufficient data preprocessing, etc. To ensure the consistency and reliability of keyword extraction, the system needs to re-execute step 2, that is, input the preprocessed text sample into the first model for analysis again to obtain a new first keyword.
[0071] Step 531: Collect historical standard text samples and obtain historical keywords.
[0072] Collect a series of historical standard text samples from a database or pre-set corpus and extract historical keywords from them. These historical standard text samples typically include various documents, comments, articles, etc. related to the current analysis task. They have been annotated and verified by experts and have high reference value.
[0073] Step 532: Determine the text length of the historical standard text sample.
[0074] Text length refers to the number of characters or words in a text, reflecting the information content and complexity of the text. For Chinese text, text length is typically measured in characters. For example, if a historical standard text sample contains 100 Chinese characters, its text length is 100. The system can use simple string manipulation functions to count and record the length of each text sample. This text length data will be used in subsequent steps to adjust the threshold calculation formula to ensure fair evaluation of texts of different lengths.
[0075] Step 533: Obtain the historical similarity, historical sentiment score, and historical sentiment polarity of the historical keywords.
[0076] Step 534 , obtaining the historical similarity mean and the historical sentiment score standard deviation based on the historical similarity and the historical sentiment score.
[0077] Step 535 : Obtain the upper limit and lower limit of the threshold range according to the text length, the historical similarity mean, the historical sentiment polarity, and the historical sentiment score standard deviation.
[0078] LB=S avg -r1×L d +r2×1 / T len -r3×E std UB=S avg +r1×L d +r2×1 / T len +r3×E std ;
[0079] In the above formula, LB is the lower limit of the threshold interval, UB is the upper limit of the threshold interval, S avg is the mean historical similarity, L d is the historical sentiment polarity, T len is the text length, E std is the standard deviation of historical sentiment scores, r1 is the adjustment coefficient of historical sentiment polarity, r2 is the adjustment coefficient of text length, and r3 is the adjustment coefficient of the standard deviation of historical sentiment scores, r1=0.1, r2=0.01, r3=0.5.
[0080] Text length reflects the richness of information; longer texts typically contain more detail, thus requiring a wider matching range. The mean historical similarity provides a baseline level of semantic matching between keywords and the overall text, ensuring that keywords are semantically representative. The historical sentiment polarity reflects the emotional orientation of the text, allowing the threshold interval to be dynamically adjusted based on the emotional direction, thereby improving the emotional consistency of keywords. The standard deviation of the historical sentiment score measures the degree of emotional fluctuation, helping the system identify text with significant emotional fluctuations to avoid misjudgments. The threshold interval constructed by integrating these factors not only adapts to texts of different types and lengths, but also effectively improves the accuracy and stability of keyword extraction, making the system more robust and adaptive when faced with diverse texts.
[0081] Step 54, construct a relative stability index and preset a stability threshold. If the relative stability index is less than the stability threshold, the first keyword or the second keyword is used as a comprehensive keyword; if the relative stability index is greater than or equal to the stability threshold, the first keyword or the second keyword is modified to obtain a comprehensive keyword.
[0082] This embodiment determines the stable threshold P by the sentiment score mean, text length, and threshold interval:
[0083] P=(1-|E avg |)×[(UB-LB) / (T len +1)],E avg is the mean sentiment score.
[0084] Step 541: Define a relative position factor based on historical sentiment polarity and text length.
[0085] Y=1+c×L d ×1 / T len ; Y is the relative position factor, c is the position weight.
[0086] Step 542: Obtain a stable center based on the upper limit, lower limit, text length, historical similarity mean, relative position factor, and historical sentiment score standard deviation.
[0087] SC=(a×UB+(a-1)×LB)+(b×S avg +(b-1)×e -Estd )×Y; SC is the stable center, a is the adjustment factor of the threshold interval, b is the adjustment weight, and e is the exponential function.
[0088] The stability center plays a central role in keyword extraction and text analysis. It is a key point determined by comprehensively considering multiple text features. Specifically, the stability center combines multiple dimensions, including the upper and lower threshold limits, text length, mean historical similarity, relative position factor, and standard deviation of historical sentiment scores, to calculate a benchmark point that represents the stability of keywords in the current text. This benchmark point not only reflects the overall semantic and sentimental orientation of the text but also considers the adaptability and stability of keywords in different textual contexts. The stability center serves as an ideal reference point within the threshold range. Keywords closer to this point are more semantically and sentimentally consistent with the overall text characteristics, thus increasing their stability and credibility. Keywords closer to the upper and lower threshold limits deviate from the core expression of the text and may be abnormal or unstable, requiring further evaluation or even correction. Therefore, the stability center can be viewed as a dynamic adjustment mechanism that guides keyword selection and optimization, ensuring that extracted keywords are both semantically representative and sentimentally consistent, thereby improving the accuracy and robustness of keyword extraction.
[0089] Step 543 : Obtain a relative stability index according to the first similarity of the same words, the second similarity of the same words, the stability center, and the sentiment score.
[0090] D i =(1-[|s i -SC| / (UB-LB)])×(1-[|s i -E i | / (UB-LB)]);D i is the relative stability index of the i-th identical word, E i is the second similarity of the i-th identical word, s i is the first similarity of the i-th identical word.
[0091] 1-[|s i -SC| / (UB-LB) measures semantic stability, that is, the degree of proximity between the first similarity of a keyword and its stable center within the threshold interval. The closer to the stable center, the more its semantic expression fits the overall characteristics of the text; 1-[|s i -E i| / (UB-LB) measures sentiment consistency, namely the difference between the first and second similarities, reflecting the semantic and emotional harmony of the keyword. The higher the value of the product of the two, the more semantically accurate and emotionally consistent the keyword is, indicating greater stability. Conversely, there may be a risk of bias or mismatch. Through this mechanism, the system not only identifies high-quality keywords but also dynamically determines which words need to be revised or replaced, thereby improving the accuracy, robustness, and adaptability of keyword extraction. This is particularly effective when processing text with complex semantics and emotions.
[0092] Modifying the first keyword or the second keyword to obtain a comprehensive keyword includes: arranging the relative stability indexes of the identical words from high to low, using the first k1 identical words as first relative stability words; using the k1+1th to k1+nth identical words as second relative stability words and marking them as modified words; searching for similar words in a text library based on the modified words, determining the relative stability index of each similar word, searching for the most similar word whose relative stability index is greater than the relative stability index of the modified word, replacing the most similar word with the modified word, and marking the first relative stability word and the most similar word as comprehensive key words. k is greater than or equal to 2.
[0093] By sorting identical words from high to low based on relative stability, dividing them into first-relative stability words and second-relative stability words, and then replacing similar words with similar words, a comprehensive keyword set is generated, demonstrating clear optimization significance and practical application value. This method ensures that top-ranked, highly stable keywords are retained as core representations of the text's semantics and sentiment. For mid-ranking, less stable words, the text library is used to search for alternatives with similar context and higher stability, thereby improving keyword quality and consistency without disrupting the overall semantic structure. This strategy not only retains the more reliable keywords from the original extraction but also corrects potentially biased or unstable words through a dynamic replacement mechanism, enhancing the overall robustness and expressiveness of the keyword set. The resulting comprehensive keyword set is significantly optimized in terms of semantic accuracy, sentiment consistency, and cross-text adaptability, improving the expressiveness and stability of the keyword extraction system in complex language environments.
[0094] In order to implement the above embodiment, the present application also proposes a text analysis device. Figure 2 This is a structural diagram of a text analysis device provided in an embodiment of the present application. Figure 2 As shown, the text analysis device may include: a sample processing module 401 , a first analysis module 402 , a second analysis module 403 , a similarity analysis module 404 and a comprehensive analysis module 405 .
[0095] The sample processing module 401 is used to obtain the text sample to be analyzed and perform preprocessing;
[0096] A first analysis module 402 is configured to input the pre-processed text sample into a first model for analysis to obtain a first keyword;
[0097] A second analysis module 403 is configured to input the pre-processed text sample into a second model for analysis to obtain a sentiment score, and obtain a second keyword based on the sentiment score;
[0098] A similarity analysis module 404 is configured to determine a first similarity between the first keyword and a first related corpus in the text library, and to determine a second similarity between the second keyword and a second related corpus in the text library;
[0099] The comprehensive analysis module 405 is configured to perform a comprehensive analysis on the first similarity and the second similarity to obtain a comprehensive keyword.
[0100] It should be noted that the above explanation of an embodiment of a text analysis method is also applicable to the text analysis system of this embodiment, and will not be repeated here.
[0101] In order to implement the above embodiment, the present application also proposes an electronic device. Figure 3 , Figure 3 Schematic diagram of the structure of the electronic device provided in the embodiment of the present application. Figure 3 As shown, the electronic device 500 includes: a processor 501, and a memory 502 communicatively connected to the processor 501; the memory 502 stores computer-executable instructions; the processor 501 executes the computer-executable instructions stored in the memory to implement the method provided in the aforementioned embodiment.
[0102] In order to implement the above embodiments, the present application also proposes a computer-readable storage medium, in which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, they are used to implement the methods provided by the above embodiments.
[0103] In order to implement the above embodiments, the present application also proposes a computer program product, including a computer program, which implements the methods provided by the above embodiments when executed by a processor.
[0104] The collection, storage, use, processing, transmission, provision and disclosure of user personal information involved in this application are in compliance with relevant laws and regulations and do not violate public order and good morals.
[0105] It is important to note that personal information collected from users should be used for legitimate and reasonable purposes and should not be shared or sold beyond these legitimate uses. Furthermore, such collection / sharing should be conducted only after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization that includes the relevant user information before using the feature. Furthermore, any necessary steps must be taken to safeguard and secure access to such personal information and ensure that others with access to personal information comply with its privacy policy and procedures.
[0106] This application contemplates providing implementations that allow users to selectively block the use or access of personal information data. Specifically, this disclosure contemplates providing hardware and / or software to prevent or block access to such personal information data. Risks can be minimized by limiting data collection and deleting data once it is no longer needed. Furthermore, where applicable, such personal information can be de-identified to protect user privacy.
[0107] In the descriptions of the foregoing embodiments, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of the different embodiments or examples, unless they are mutually inconsistent.
[0108] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of such features. Throughout the description of this application, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0109] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application belong.
[0110] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.
[0111] It should be understood that various parts of the present application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used to implement: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0112] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0113] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0114] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A text analysis method, characterized in that: The following steps are involved: Step 1: Obtain the text sample to be analyzed and preprocess it; Step 2: input the pre-processed text sample into the first model for analysis to obtain the first keyword; Step 3: Input the pre-processed text sample into the second model for analysis to obtain a sentiment score, and obtain a second keyword based on the sentiment score; Step 4: determining a first similarity between the first keyword and a first related corpus in the text library, and determining a second similarity between the second keyword and a second related corpus in the text library; Step 5: Perform comprehensive analysis on the first similarity and the second similarity to obtain comprehensive keywords.
2. The method according to claim 1, characterized in that Step 5, performing a comprehensive analysis on the first similarity and the second similarity to obtain a comprehensive keyword, includes the following steps: Step 51, extracting the same words in the first keyword and the second keyword, and the first similarity and the second similarity of each of the same words; Step 52: determining a difference between the first similarity and the second similarity; Step 53: Preset a threshold interval and determine whether the difference value is within the threshold interval. If so, execute step 54; if not, return to step 2. Step 54, construct a relative stability index and preset a stability threshold. If the relative stability index is less than the stability threshold, the first keyword or the second keyword is used as a comprehensive keyword; if the relative stability index is greater than or equal to the stability threshold, the first keyword or the second keyword is modified to obtain a comprehensive keyword.
3. The method according to claim 2, characterized in that In addition to obtaining the second keyword, the second model also obtains the sentiment polarity of the preprocessed text sample, where the sentiment polarity includes positive sentiment, negative sentiment, and neutral sentiment. The positive sentiment is assigned a +1 label, the negative sentiment is assigned a -1 label, and the neutral sentiment is assigned a 0 label.
4. The method according to claim 3, characterized in that Presetting a threshold range includes the following steps: Step 531, collecting historical standard text samples and obtaining historical keywords; Step 532, determining the text length of the historical standard text sample; Step 533: Obtain the historical similarity, historical sentiment score, and historical sentiment polarity of the historical keyword; Step 534: obtaining the historical similarity mean and the historical sentiment score standard deviation based on the historical similarity and the historical sentiment score; Step 535 : Obtain the upper limit and lower limit of the threshold range according to the text length, the historical similarity mean, the historical sentiment polarity, and the historical sentiment score standard deviation.
5. The method according to claim 4, characterized in that Constructing a relative stability indicator includes the following steps: Step 541, defining a relative position factor based on historical sentiment polarity and text length; Step 542: Obtain a stable center based on the upper limit, lower limit, text length, historical similarity mean, relative position factor, and historical sentiment score standard deviation; Step 543 : Obtain a relative stability index according to the first similarity of the same words, the second similarity of the same words, the stability center, and the sentiment score.
6. The method according to claim 5, characterized in that The first keyword or the second keyword is modified to obtain a comprehensive keyword, including: arranging the relative stability indexes of the identical words from high to low, and taking the first k1 identical words as first relative stability words; taking the k1+1th to k1+nth identical words as second relative stability words and marking them as modified words, searching for similar words in a text library based on the modified words, determining the relative stability index of each similar word, searching for the maximum similar word whose relative stability index is greater than the relative stability index of the modified word, replacing the maximum similar word with the modified word, and marking the first relative stability word and the maximum similar word as comprehensive key words.
7. A text analysis system, characterized in that: include: The sample processing module is used to obtain the text samples to be analyzed and perform preprocessing; A first analysis module, configured to input the preprocessed text sample into a first model for analysis to obtain a first keyword; A second analysis module is configured to input the pre-processed text sample into a second model for analysis to obtain a sentiment score, and obtain a second keyword based on the sentiment score; a similarity analysis module, configured to determine a first similarity between the first keyword and a first related corpus in the text library, and to determine a second similarity between the second keyword and a second related corpus in the text library; The comprehensive analysis module is used to perform comprehensive analysis on the first similarity and the second similarity to obtain comprehensive keywords.
8. An electronic device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 7.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 7 when executed by a processor.