A search engine system based on text sentiment analysis

By introducing the sentiment polarization index and expectation violation index, the problem of misjudgment of sarcastic language in existing technologies is solved, the accuracy of sentiment analysis and the timeliness of corporate response are achieved, and the healthy development of brand image is ensured.

CN119577110BActive Publication Date: 2025-09-23INFORMATION & COMM BRANCH OF STATE GRID JIANGSU ELECTRIC POWER
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411617129.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-13
Publication Date
2025-09-23
Estimated Expiration
2044-11-13

AI Technical Summary

Technical Problem

Existing search engines based on text sentiment analysis are prone to misjudgment when dealing with sarcastic language, causing companies to mistakenly believe that market feedback is good and ignore potential problems, which may trigger a public relations crisis.

Method used

The emotional polarization index and expectation violation index are introduced to accurately capture complex emotions and negative emotions in sarcastic language through real-time parameter capture and storage, anomaly analysis and model comparison, risk assessment and response measures modules to avoid misjudgment.

Benefits of technology

It improves the accuracy of sentiment analysis, reduces misjudgments, ensures that companies can respond to market feedback in a timely manner, avoids damage to brand image, and achieves rational allocation of resources and rapid response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119577110B_ABST
    Figure CN119577110B_ABST
Patent Text Reader

Abstract

The present invention discloses a search engine system based on text sentiment analysis, which relates to the field of search engine technology and includes a real-time parameter capture and storage module, an anomaly analysis and model comparison module, a risk assessment module, and a countermeasure module: In the real-time parameter capture and storage module, during the sentiment analysis process, each text will generate a series of parameters, and the parameters generated when the sentiment analysis model performs text analysis are captured and stored in real time to ensure low latency and integrity of data flow. The present invention enables the system to accurately capture complex emotions and avoid misjudgment by introducing the sentiment polarization index and the expectation violation index. Real-time parameter capture ensures low latency, and multi-level analysis of machine learning improves robustness. Through the classification of low, medium, and high risk levels, the system implements on-demand intervention and resource optimization to avoid business losses and damage to brand image, and ensure that enterprises can efficiently respond to market feedback and uncertainty.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of search engine technology, and in particular to a search engine system based on text sentiment analysis. Background Art

[0002] A search engine based on text sentiment analysis is an intelligent system that leverages natural language processing (NLP) and sentiment analysis techniques to help users obtain search results relevant to sentiment. It not only recognizes and understands user queries but also analyzes the sentiment of online documents, reviews, and articles, such as positive, negative, or neutral sentiment. Unlike traditional search engines, which primarily rely on keyword matching and page relevance ranking, sentiment analysis search engines focus on subjective information in text, such as user emotions, opinions, and contextual evaluations. This enables them to provide more targeted results when users query for sentiment-related information, such as product reviews or event sentiment. Their advantage lies in providing a more precise and personalized experience. For example, when a user wants to learn about a product's reputation, the system not only displays relevant information but also filters positive and negative reviews to present more valuable feedback. This type of engine is particularly useful in e-commerce and social media public opinion monitoring, helping users quickly obtain the sentiment information they need. It can also be applied in the news industry to filter reports with specific sentiments or, in crisis management, to help companies identify and respond to negative public sentiment. Therefore, search engines based on sentiment analysis not only expand the depth of information retrieval but also significantly improve the user search experience and efficiency.

[0003] Search engines based on text sentiment analysis are implemented using natural language processing (NLP) and machine learning models. Their working principle involves several key steps. First, when a user enters a query, the system tokenizes the query, removes meaningless words, and uses intent recognition models to analyze the user's actual needs. For example, if a user enters "well-reviewed mobile phones," the system can recognize that the user is looking for products with positive reviews. Next, the search engine crawls massive amounts of data from the internet, such as reviews, news, and social media updates, and cleans and formats the data, converting the text into structured data for subsequent analysis. The core component is the application of sentiment analysis models. Using models such as BERT or Transformer, the system identifies the sentiment of the text, such as positive, negative, or neutral sentiment, and further analyzes the intensity and type of sentiment in advanced systems. After analysis, the system generates a sentiment tag for each text entry and stores it, along with keywords, in a database to create an index. After the user query matches the database, the system ranks the results based on keyword matching and sentiment tags, and optimizes the ranking based on information such as time, click-through rate, and user feedback. Finally, the system displays the sentiment results that meet the user's needs and annotates the sentiment on the interface to facilitate quick judgment. Furthermore, the search engine continuously optimizes its models and algorithms based on user feedback to enhance the user experience. This sentiment analysis-driven search engine is not only applicable to e-commerce and social media monitoring, but also plays an important role in crisis public relations and market intelligence.

[0004] The existing technology has the following deficiencies:

[0005] When search engines perform sentiment analysis on crawled text data, sarcastic expressions, while using positive vocabulary, actually convey negative sentiment. For example, "This phone is so great, it broke after just two days." In this case, sentiment analysis models may misidentify such sentences as positive, ignoring their potential negative connotations. This misjudgment can cause companies monitoring product or brand reputation to mistakenly classify sarcastic negative reviews as positive feedback, leading them to mistakenly believe that market satisfaction with their products is high, thereby overlooking hidden product issues or customer dissatisfaction. If companies fail to identify these issues and implement corrective measures promptly, they may miss opportunities to optimize their products and services, and the accumulation of these issues may lead to public relations crises.

[0006] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention

[0007] The purpose of the present invention is to provide a search engine system based on text sentiment analysis. By introducing the sentiment polarization index and the expectation violation index, the system accurately captures complex emotions and negative emotions in sarcastic language to avoid misjudging them as positive feedback. Real-time parameter capture ensures low latency and integrity, and multi-level analysis of machine learning models enhances system robustness. Under the three risk levels of low, medium and high, the system implements on-demand intervention and resource optimization. Log monitoring is performed when the risk is low, alarms are triggered and backup model adjustments are made when the risk is medium, and models are quickly switched and audited and repaired when the risk is high, avoiding business losses and damage to brand image, ensuring that enterprises can efficiently respond to market feedback and uncertainty, and solving the problems in the above-mentioned background technology.

[0008] To achieve the above objectives, the present invention provides the following technical solutions: a search engine system based on text sentiment analysis, comprising a real-time parameter capture and storage module, an anomaly analysis and model comparison module, a risk assessment module, and a response measure module:

[0009] Real-time parameter capture and storage module: During the sentiment analysis process, each text will generate a series of parameters. The parameters generated when the sentiment analysis model performs text analysis are captured and stored in real time to ensure low latency and integrity of data flow;

[0010] The anomaly analysis and model comparison module analyzes the acquired parameter information for anomalies and inputs the analyzed parameter information into a pre-trained machine learning model to determine whether there are potential errors in the current sentiment analysis through feature comparison.

[0011] The risk assessment module, based on the analysis results of the machine learning model, further comprehensively analyzes the analysis process with potential errors and identifies the risk level of potential errors;

[0012] The response module takes different response measures according to different risk levels.

[0013] Preferably, the parameters generated by the sentiment analysis model when performing text analysis include sentiment complexity information and expectation consistency information. Sentiment complexity information refers to the situation where multiple emotions appear simultaneously in the text and conflict or intertwine with each other. Expectation consistency information is used to measure whether the text expression meets the user's expectations or conventional cognition, reflecting the degree of consistency between the user experience and their original expectations.

[0014] Preferably, after obtaining the sentiment complexity information and expected consistency information generated by the sentiment analysis model when performing text analysis, the sentiment complexity information is subjected to an abnormal analysis to generate a sentiment polarization index, and the expected consistency information is subjected to an abnormal analysis to generate an expected violation index. The generated sentiment polarization index and expected violation index are input into a pre-learned machine learning model, and an analysis abnormality index is generated through the machine learning model. The analysis abnormality index is used to determine whether there are potential errors in the current sentiment analysis.

[0015] Preferably, the analysis anomaly index generated when the sentiment analysis model performs text analysis is compared with a preset analysis anomaly index reference threshold to determine whether there are potential errors in the current sentiment analysis. The analysis results are as follows:

[0016] If the analysis anomaly index is greater than or equal to a preset analysis anomaly index reference threshold, the text analysis state of the sentiment analysis model is classified as a potential error analysis state;

[0017] If the analysis anomaly index is less than a preset analysis anomaly index reference threshold, the text analysis state of the sentiment analysis model is classified as a correct analysis state.

[0018] Preferably, when the sentiment analysis model is divided into a potential error analysis state during the text analysis process, a plurality of analysis anomaly indices generated by the subsequent process of the sentiment analysis model for text analysis are continuously obtained to establish a data set for comprehensive analysis, and the analysis anomaly indices in the analysis set are compared with the first-level reference threshold, the second-level reference threshold, and the analysis anomaly index reference threshold, wherein the second-level reference threshold is greater than the first-level reference threshold, and the first-level reference threshold is greater than the analysis anomaly index reference threshold, and the analysis anomaly index is compared and analyzed with the second-level reference threshold, the first-level reference threshold, and the analysis anomaly index reference threshold, and the number of analysis anomaly indices that are less than the first-level reference threshold and greater than or equal to the analysis anomaly index reference threshold is calibrated as La, the number of analysis anomaly indices that are less than the second-level reference threshold and greater than or equal to the first-level reference threshold is calibrated as Lb, and the number of analysis anomaly indices that are greater than or equal to the second-level reference threshold is calibrated as Lc;

[0019] Comprehensively analyze La, Lb, and Lc to generate the risk level coefficient RLC based on the following formula:

[0020]

[0021] , where k1, k2, and k3 are the preset proportional coefficients of La, Lb, and Lc, respectively, and k1, k2, and k3 are all greater than 0.

[0022] Preferably, the generated risk level coefficient is compared and analyzed with a preset first risk level coefficient reference threshold and a preset second risk level coefficient reference threshold to determine the risk level of the potential error analysis state. The comparison and analysis results are as follows:

[0023] If the risk level coefficient is less than a preset risk level coefficient reference threshold, the potential error analysis state is classified as a low-risk potential error analysis;

[0024] If the risk level coefficient is greater than or equal to the first risk level coefficient reference threshold and less than the second risk level coefficient reference threshold, the potential error analysis state is classified as a medium risk potential error analysis;

[0025] If the risk level coefficient is greater than the second risk level coefficient reference threshold, the potential error analysis state is classified as a high-risk potential error analysis.

[0026] Preferably, within the detection window, the logic for performing abnormal analysis on the emotional complexity information to generate the emotional polarization index is as follows:

[0027] Under the detection window, the text to be analyzed is broken down into multiple emotion categories, and each emotion category is assigned an emotion intensity score through the emotion classification model. The emotion distribution vector generated for each text is calibrated as V s , where s∈{pos,neg,neu}, pos represents positive emotion, neg represents negative emotion, neu represents neutral emotion, and the value of each element of the emotion distribution vector is the emotion intensity score, which is expressed as follows: V s ={P pos , P neg , P neu}, where P pos is the intensity score of positive emotion, P neg is the intensity score of negative sentiment, P neu is the intensity score of neutral emotion;

[0028] The opposition function is defined to measure the difference between two emotion categories. The calculation expression is as follows:

[0029]

[0030] , where D(s1, s2) is the degree of opposition between two emotion categories, represents the intensity score of sentiment category s1 in the text, It represents the intensity score of sentiment category s2 in the text. s1 and s2 correspond to any sentiment category in the sentiment distribution vector respectively. ∈ is a small constant that prevents the denominator from being zero to ensure calculation stability.

[0031] The D emotion category opposition degree D(s1, s2) is accumulated to generate the emotion polarization index. The calculation expression is as follows:

[0032]

[0033] , where EPI is the sentiment polarization index, (s1, s2)∈C is the combination of sentiment categories, and C is the set of all sentiment category combinations.

[0034] Preferably, within the detection window, the logic for performing anomaly analysis on the expected consistency information to generate an expected violation index is as follows:

[0035] In the detection window, the similarity analysis is performed between the input text and the preset expected model. The calculation expression is as follows:

[0036]

[0037] , where S match is the similarity score between the text and the expected model, T i is the i-th keyword in the input text, M i is the i-th keyword in the expected model, n is the total number of keywords, w i It is the weight of the keyword;

[0038] For each keyword T in the text i Calculate the degree of deviation from the expected model in the semantic space. The calculation expression is as follows:

[0039]

[0040] , where D shift Indicates the total outlier offset distance, E(T i ) is the keyword T in the input text i The vector embedding, E(M i ) is the keyword M in the expected model i The vector embedding of ||E(T i )-E(M i )|| is the Euclidean distance;

[0041] In order to capture the impact of emotional or contextual transitions in sentences on expectancy violation, the contextual transition influence coefficient is introduced. The calculation expression is as follows:

[0042]

[0043] , where C shift is the context transition influence coefficient, α is the adjustment coefficient, which is used to control the influence weight of transition words on the results, N contrast is the number of occurrences of transition words in the sentence, N totalis the total number of words in the sentence;

[0044] Combine the similarity score S between the text and the expected model match , total outlier offset distance D shift and contextual transition influence coefficient C shift , generate the expected violation index, the calculation expression is as follows:

[0045]

[0046] , where EVI is the expectation violation index.

[0047] In the above technical solution, the technical effects and advantages provided by the present invention are:

[0048] By introducing the sentiment polarization index and the expectation violation index, the present invention enables the system to more accurately capture the negative emotions hidden in complex emotions and sarcastic language, avoiding misjudgment as positive feedback. The introduction of real-time parameter capture and storage modules ensures low-latency transmission and integrity of data, enabling the model to quickly capture potential errors and perform anomaly analysis. At the same time, the use of machine learning models for multi-level comparison and analysis further enhances the robustness of the model. Even when faced with complex emotional expressions or deviations from expectations, the system can respond promptly when errors first appear, avoiding the accumulation of misjudgments. Comprehensive risk assessment and dynamic response mechanisms ensure that errors in the model's daily operations are no longer ignored, providing companies with more reliable sentiment analysis results and helping them gain a deeper understanding of market feedback.

[0049] The present invention divides risks into three categories: low, medium, and high, to ensure that enterprise resources can be reasonably allocated, to achieve on-demand intervention, and to avoid unnecessary manual intervention and waste of resources. In low-risk situations, log records and data accumulation are used for continuous monitoring to provide data support for subsequent model optimization. In medium-risk scenarios, the system will trigger an alarm and conduct manual review, while using the backup model to make temporary adjustments to avoid excessive impact on the business. In high-risk situations, the system quickly switches to the backup model and generates a detailed audit report to ensure that model problems can be fixed in a timely manner. This multi-level response mechanism reduces the possibility of business losses, improves the flexibility and efficiency of the system in dealing with uncertainty, and at the same time avoids damage to the brand image due to analytical errors, ensuring the healthy development of the relationship between the company and its customers. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction to the drawings required for use in the embodiments will be given below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.

[0051] Figure 1 This is a system flow chart of a search engine system based on text sentiment analysis of the present invention. DETAILED DESCRIPTION

[0052] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these example embodiments are provided so that the description of this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art.

[0053] The present invention provides Figure 1 The search engine system based on text sentiment analysis is characterized by including a real-time parameter capture and storage module, an anomaly analysis and model comparison module, a risk assessment module, and a response module:

[0054] Real-time parameter capture and storage module: During the sentiment analysis process, each text will generate a series of parameters. The parameters generated when the sentiment analysis model performs text analysis are captured and stored in real time to ensure low latency and integrity of data flow;

[0055] The parameters generated by the sentiment analysis model during text analysis include sentiment complexity information and expectation consistency information. Sentiment complexity information refers to the situation where multiple emotions appear simultaneously in the text and conflict or intertwine with each other. Expectation consistency information is used to measure whether the text expression meets the user's expectations or conventional cognition, reflecting the degree of consistency between the user experience and their original expectations.

[0056] The anomaly analysis and model comparison module analyzes the acquired parameter information for anomalies and inputs the analyzed parameter information into a pre-trained machine learning model to determine whether there are potential errors in the current sentiment analysis through feature comparison.

[0057] After obtaining the sentiment complexity information and expectation consistency information generated by the sentiment analysis model during text analysis, the sentiment complexity information is analyzed for anomalies to generate a sentiment polarization index. The expectation consistency information is analyzed for anomalies to generate an expectation violation index. The generated sentiment polarization index and expectation violation index are input into a pre-learned machine learning model. The machine learning model is used to generate an analysis anomaly index. The analysis of the anomaly index is used to determine whether there are potential errors in the current sentiment analysis.

[0058] During text analysis by sentiment analysis models, changes in emotional complexity often indicate that a text contains multiple emotions or conflicting expressions, a common characteristic of sarcastic language. Sarcasm may superficially use positive words (such as "awesome"), but the context or subsequent parts of the sentence convey implicit negativity (such as "it broke after two days"). Because sentiment analysis models typically rely on keyword weights, sentiment scores, and simple syntactic structures to determine sentiment, they are prone to misjudgment when faced with such mixed emotions. Models may mistakenly assign high positive weights to seemingly positive words, while ignoring turning points or implicit negativity. Especially without incorporating contextual variations, syntactic reversals, or sarcasm detection mechanisms, models may classify such sentences as positive reviews, ignoring their negative connotations. This misjudgment is particularly dangerous in product reviews and user feedback monitoring, as companies may mistakenly believe that market satisfaction is high while ignoring potential problems and customer dissatisfaction.

[0059] The logic for generating the sentiment polarization index by performing anomaly analysis on sentiment complexity information within the detection window is as follows:

[0060] Under the detection window, the text to be analyzed is broken down into multiple emotion categories (such as positive, negative, and neutral), and each emotion category is assigned an emotion intensity score through the emotion classification model. The emotion distribution vector generated for each text is calibrated as V s , where s∈{pos,neg,neu}, pos represents positive emotion, neg represents negative emotion, neu represents neutral emotion, and the value of each element of the emotion distribution vector is the emotion intensity score, which is expressed as follows: V s ={P pos , P neg , P neu}, where P pos is the intensity score of positive emotions (such as the sum of the weights of praise or commendation words), P neg is the intensity score of negative emotions (such as the sum of the word weights of criticism and dissatisfaction), P neu is the intensity score of neutral emotions (e.g., descriptive words without emotional tendencies);

[0061] The opposition function is defined to measure the difference between two emotion categories (such as the conflict between positive and negative emotions). The calculation expression is as follows:

[0062]

[0063] , where D(s1,s2) is the degree of opposition between two emotion categories (such as positive and negative emotions), represents the intensity score of sentiment category s1 in the text, It represents the intensity score of sentiment category s2 in the text. s1 and s2 correspond to any sentiment category in the sentiment distribution vector, respectively. They are determined according to the context in the actual text. ∈ is a small constant that prevents the denominator from being zero to ensure calculation stability.

[0064] In sentiment analysis, the difference between two emotion categories refers to conflicts, contradictions, or inconsistencies in the intensity, tendency, and semantic expression of different emotion types. This difference is not only reflected in the numerical difference in emotion intensity, but also in the relationship between different emotion categories in the linguistic context. This difference is particularly critical for detecting mixed emotions, sarcastic expressions, and emotional transitions.

[0065] The core dimension of the difference between two emotion categories

[0066] 1. Differences in emotional intensity

[0067] Definition: The degree to which two sentiment categories (e.g., positive and negative sentiment) differ in their intensity scores.

[0068] Function: When the scores of two emotion categories are close (such as positive emotion P pos =0.6 and negative emotions P neg =0.5), the system will consider the emotional complexity to be high, which may be sarcastic language or contradictory evaluation.

[0069] Example: "These headphones sound great, but the battery life is too short."

[0070] The positive emotion intensity in this sentence is high (good sound quality), but the negative emotion is also significantly present (bad battery), showing emotional conflict.

[0071] 2. Semantic opposition of emotion categories

[0072] Definition: The semantic properties of different sentiment categories are inherently opposing. For example, positive sentiment and negative sentiment are inherently in conflict.

[0073] Function: When positive and negative emotions appear in the same text at the same time, they will show emotional contradiction or opposition. This semantic opposition is an important basis for judging emotional complexity in sentiment analysis.

[0074] 3. Contextual Relevance of Emotional Categories

[0075] Definition: Whether the association between two emotion categories is consistent with the contextual logic. If an emotion category appears in a context that is inconsistent with it, the difference increases.

[0076] Purpose: For example, if a compliment is mixed with a derogatory description, or a positive modifier is added to a negative review, the system needs to determine whether there is irony or contradiction based on the context.

[0077] Example: "The service was 'terrific'; we waited two hours for our food to arrive."

[0078] The emotional difference between the surface positive emotion ("great") and the actual experience ("wait two hours") is huge, which means that the sentiment analysis system needs to capture this semantic contradiction.

[0079] 4. Temporal or syntactic transitions

[0080] Definition: The sentiment within a text changes and shifts across time periods or syntactic structures. For example, the first half of a text is positive, while the second half is negative.

[0081] Effect: This transition will increase the differences between emotional categories. The system needs to recognize the existence of this transition to avoid misjudgment.

[0082] Example: "This phone looked great, but it broke after a few days of use."

[0083] The first half of the sentence has positive sentiment, while the second half turns negative. The difference between the sentiment categories indicates that the user is actually dissatisfied.

[0084] The D emotion category opposition degree D(s1,s2) is accumulated to generate the emotion polarization index. The calculation expression is as follows:

[0085]

[0086] , where EPI is the emotional polarization index, (s1,s2)∈C is the combination of emotional categories, and C is the set of all emotional category combinations;

[0087] It can be seen from the calculation expression of the sentiment polarization index that the larger the sentiment polarization index expression value generated after the abnormal analysis of the sentiment complexity information, the higher the sentiment complexity in the text, that is, there is an obvious emotional conflict or reversal, such as sarcastic language that uses positive words on the surface, but the emotions it actually expresses are negative. In this case, because the model is easily disturbed by superficial positive words and ignores turning points or implicit negative emotions, the probability of misjudging it as a positive evaluation will increase significantly. A high polarization value reflects the fierce opposition between positive and negative emotions in a sentence. If the model does not have sensitivity to such complex emotions, the risk of misidentification will also increase. On the contrary, when the sentiment polarization index expression value is smaller, the expression of emotions is more single and consistent, and the possibility of sarcastic language or irony is lower, so the risk of the model mistakenly judging negative emotions as positive evaluations is also relatively small.

[0088] During text analysis by sentiment analysis models, changes in expectation consistency indicate that users' actual experiences or emotional expressions deviate from conventional expectations, a key characteristic of sarcastic language. While sarcastic language superficially uses positive words (such as "awesome" and "excellent"), these words convey negative sentiment in the context (e.g., "This phone is awesome, but it broke after just two days"). This expression violates the typical logic of user emotional expression, making it difficult for sentiment analysis models to accurately determine emotional tendencies through keywords or simple sentiment classification. If the model relies solely on the surface positive or negative sentiment of words while ignoring the semantics and emotional transitions of the context, it will misclassify sarcastic expressions as positive. Consequently, the system overlooks the underlying negative connotations and fails to capture user dissatisfaction and disappointment. This misjudgment can cause companies to mistakenly believe that market feedback is positive when monitoring brand public opinion or analyzing product reviews, ignoring underlying negative sentiment. This leads to missed opportunities to identify and address issues, potentially leading to greater customer churn or public relations crises.

[0089] The logic for generating the expected violation index by performing anomaly analysis on expected consistency information within the detection window is as follows:

[0090] In the detection window, the similarity analysis is performed between the input text and the preset expected model. The calculation expression is as follows:

[0091]

[0092] , where S match is the similarity score between the text and the expected model, T i is the i-th keyword in the input text, M i is the i-th keyword in the expected model, n is the total number of keywords, w i It is the weight of the keyword;

[0093] An expectation model is a reference template built based on historical data and general cognition, describing users' reasonable expectations for a specific scenario (such as a product or service).

[0094] For each keyword T in the text i Calculate the degree of deviation from the expected model in the semantic space. The calculation expression is as follows:

[0095]

[0096] , where D shift Indicates the total outlier offset distance, E(T i ) is the keyword T in the input text i The vector embedding, E(M i ) is the keyword M in the expected model iThe vector embedding of ||E(T i )-E(M i )|| is the Euclidean distance, which is used to calculate the keyword T in the text i The keyword M in the period model i The Euclidean distance in the semantic vector space quantifies the degree of semantic deviation between the two words;

[0097] Vector embedding is the process of converting words, phrases, or sentences in a text into high-dimensional continuous vectors through mathematical models, allowing them to represent semantic features and relationships in vector space. Embedding vectors capture the semantic similarities and differences between words in context, mapping previously discrete language data into a structure that can be subjected to mathematical operations and machine learning analysis. Common vector embedding methods include Word2Vec, GloVe, and BERT. These methods generate embedding representations for each word through large-scale corpus training, resulting in semantically similar words (such as "king" and "queen") being close in distance in vector space, while semantically unrelated words (such as "king" and "car") are farther apart. This enables the model to perform deeper semantic analysis, such as sentiment classification, similarity matching, and anomaly detection. The dimensions of the embedding vectors are typically tens to hundreds, enabling the system to efficiently process complex language information.

[0098] In order to capture the impact of emotional or contextual transitions in sentences on expectancy violation, the contextual transition influence coefficient is introduced. The calculation expression is as follows:

[0099]

[0100] , where C shift is the context transition influence coefficient, α is the adjustment coefficient, which is used to control the influence weight of transition words on the results, N contrast is the number of occurrences of transition words in the sentence, N total is the total number of words in the sentence;

[0101] In sentiment analysis, shifts in sentiment or context can significantly influence the judgment of expectation violations. These shifts are often reflected in specific sentence structures (such as transition words or causal sentence patterns), suggesting that user feedback may not meet expectations. The following details the key factors that influence sentiment shifts and how to capture them:

[0102] 1. The impact of transition words

[0103] describe:

[0104] Transition words in sentences (such as "but," "however," and "although") often indicate a reversal of emotion. These words often connect the semantics and emotions of the preceding and following sentences, making the overall expression of true emotion not as positive as it appears. For example:

[0105] Example: "This phone looks great, but the battery life is terrible."

[0106] On the surface, there is a positive evaluation ("the appearance is beautiful"), but the transition word "but" reveals the user's dissatisfaction and the main emotional tendency is negative.

[0107] Influence:

[0108] If these transition words are ignored, the sentiment analysis model may mistakenly classify the sentence as positive. Therefore, the system needs to introduce transition word detection and increase their weight in the overall sentiment judgment to avoid misclassification.

[0109] 2. Contextual Coherence and Contextual Change

[0110] describe:

[0111] In long or multi-sentence expressions, the coherence and changes in the preceding and following contexts can affect the interpretation of sentiment. For example, a user might express satisfaction first and then raise a serious complaint. To account for such contextual changes, the system needs to weight the sentiment of the preceding and following sentences to capture the user's true feelings:

[0112] Example: "I was very pleased with the hotel's ambiance, but the service was far below my expectations."

[0113] Influence:

[0114] If the system only analyzes the sentiment of the first half of a sentence, it might mistakenly conclude that the user is satisfied. However, the emotional expression of the second half of a sentence is the real core. Therefore, the system needs to analyze the contextual transitions between the first and second sentences and improve its sensitivity to contextual changes.

[0115] 3. The impact of causal sentences

[0116] describe:

[0117] Some sentences use cause-and-effect structures to show that an event or experience brings about a specific result, and the hidden emotions are often revealed through the result part. For example:

[0118] Example: "I have a very bad impression of this company because of their slow customer service response."

[0119] Although a neutral statement appears in the sentence ("customer service response"), the result expresses a strong negative emotion.

[0120] Influence:

[0121] The system must identify the core sentiment in causal relationships to avoid being misled by preceding descriptions (such as customer service responses). Negative sentiment in causal sentences is often concentrated in the result part and should be given more weight.

[0122] 4. Adjust the weight of emotional intensity

[0123] describe:

[0124] In sentences with emotional transitions, the emotional intensity of different parts is often inconsistent. The system needs to weight the overall emotion according to the emotional intensity after the transition. For example:

[0125] Example: "While the room was nice, I would never stay at this hotel again."

[0126] Although the first half of the sentence gives positive feedback, the second half has a stronger negative sentiment and implies the user's final attitude.

[0127] Influence:

[0128] The system needs to dynamically adjust its overall judgment based on the intensity of the emotion. For example, it can calculate the emotional intensity of each part of a sentence through semantic embedding and assign higher weights to important parts to ensure that the analysis results accurately reflect the user's true emotions.

[0129] To capture emotional or contextual shifts, the system analyzes transition words, contextual coherence, causal sentence patterns, and changes in sentiment intensity. Appropriate weighting and parameter settings are then used to achieve accurate sentiment assessment. Capturing these transitions is crucial for generating the Expectation Violation Index, helping the system identify hidden emotions within complex expressions and avoid misjudgments. Furthermore, leveraging vector embedding and parametric models, the system can efficiently process multiple layers of sentiment and contextual transitions, improving both the accuracy and sensitivity of the analysis.

[0130] Combine the similarity score S between the text and the expected model match , total outlier offset distance D shift and contextual transition influence coefficient C shift , generate the expected violation index, the calculation expression is as follows:

[0131]

[0132] , where EVI is the expectation violation index;

[0133] The expression for calculating the expectancy violation index (EVI) indicates that a larger EVI value, generated after anomaly analysis of expectancy consistency information, indicates a significant expectancy violation in the text. This means that positive words appear to be used, but the context conveys a negative sentiment. In such cases, the sentiment analysis model may misjudge the evaluation as positive due to its reliance on the sentiment of the surface words, increasing the probability of overlooking the underlying negative connotations. For example, in a sentence like "This phone is awesome! It broke after just two days," "Awesome" is a positive word, but the overall sentiment is negative when combined with the subsequent context. A lower EVI value indicates that the text's emotional expression is more direct and aligns with user expectations (e.g., "The phone is excellent" or "The phone is terrible"), reducing the likelihood of the model misjudging the text. Therefore, the EVI value can serve as an important signal for identifying sarcastic language. A higher EVI value indicates a greater probability of misjudgment, requiring the model to further analyze the context and sentiment transitions. Conversely, a lower EVI value indicates a more accurate model analysis, making it less likely to overlook potential sentiment reversals.

[0134] The machine learning model is not limited here. Any machine learning model that can comprehensively analyze the emotional polarization index EPI and the expectation violation index EVI to generate the analysis anomaly index AAI can be used. To implement the technical solution of the present invention, the present invention provides a specific implementation method:

[0135] The calculation formula for the analysis anomaly index AAI is: AAI = ln(EPI·f1+EVI·f2), where f1 and f2 are the preset proportional coefficients of the emotional polarization index EPI and the expectation violation index EVI, respectively, and both f1 and f2 are greater than 0.

[0136] It can be seen from the calculation expression of the analysis anomaly index that, within the detection window, the larger the sentiment polarization index performance value generated after the abnormal analysis of the sentiment complexity information, the larger the expected violation index performance value generated after the abnormal analysis of the expected consistency information. That is, the larger the generated analysis anomaly index performance value, the greater the probability of misjudgment when the sentiment analysis model performs text analysis. Conversely, the smaller the probability of prediction when the sentiment analysis model performs text analysis.

[0137] The analysis anomaly index generated by the sentiment analysis model during text analysis is compared with the pre-set analysis anomaly index reference threshold to determine whether there are potential errors in the current sentiment analysis. The analysis results are as follows:

[0138] If the analysis anomaly index is greater than or equal to a preset analysis anomaly index reference threshold, the text analysis state of the sentiment analysis model is classified as a potential error analysis state;

[0139] If the analysis anomaly index is less than the preset analysis anomaly index reference threshold, the sentiment analysis model is classified as a correct analysis state for the text analysis;

[0140] The risk assessment module, based on the analysis results of the machine learning model, further comprehensively analyzes the analysis process with potential errors and identifies the risk level of potential errors;

[0141] After the sentiment analysis model is divided into a potential error analysis state during the text analysis process, several analysis anomaly indices generated by the subsequent process of the sentiment analysis model are continuously obtained to establish a data set for comprehensive analysis, and the analysis anomaly indices in the analysis set are compared with the first-level reference threshold, the second-level reference threshold, and the analysis anomaly index reference threshold, wherein the second-level reference threshold is greater than the first-level reference threshold, and the first-level reference threshold is greater than the analysis anomaly index reference threshold. The analysis anomaly index is compared and analyzed with the second-level reference threshold, the first-level reference threshold, and the analysis anomaly index reference threshold, and the number of analysis anomaly indices that are less than the first-level reference threshold and greater than or equal to the analysis anomaly index reference threshold is calibrated as La, the number of analysis anomaly indices that are less than the second-level reference threshold and greater than or equal to the first-level reference threshold is calibrated as Lb, and the number of analysis anomaly indices greater than or equal to the second-level reference threshold is calibrated as Lc;

[0142] Comprehensively analyze La, Lb, and Lc to generate the risk level coefficient RLC based on the following formula:

[0143]

[0144] , where k1, k2, and k3 are the preset proportional coefficients of La, Lb, and Lc, respectively, and k1, k2, and k3 are all greater than 0.

[0145] From the risk level coefficient calculation expression, it can be seen that the larger the risk level coefficient value, the more serious the potential error analysis state when the sentiment analysis model performs text analysis, and vice versa, the less serious the potential error analysis state when the sentiment analysis model performs text analysis;

[0146] The generated risk level coefficient is compared and analyzed with the pre-set first risk level coefficient reference threshold and the second risk level coefficient reference threshold to determine the risk level of the potential error analysis state. The comparison and analysis results are as follows:

[0147] If the risk level coefficient is less than a preset risk level coefficient reference threshold, the potential error analysis state is classified as a low-risk potential error analysis;

[0148] If the risk level coefficient is greater than or equal to the first risk level coefficient reference threshold and less than the second risk level coefficient reference threshold, the potential error analysis state is classified as a medium risk potential error analysis;

[0149] If the risk level coefficient is greater than the second risk level coefficient reference threshold, the potential error analysis state is classified as a high-risk potential error analysis;

[0150] Response measures module, taking different response measures according to different risk levels;

[0151] Countermeasures for low-risk potential error analysis

[0152] For low-risk potential error analysis, due to the low risk level coefficient, it means that the model's errors have a minor impact on the overall system and may be within the normal error range. Therefore, logging and continuous monitoring can be adopted. Specifically, the system will record the error analysis status and its corresponding analysis anomaly index in the log and track it through periodic data analysis tools (such as monitoring panels). At this time, there is no need to take immediate intervention measures. The system will analyze the model performance based on the accumulated error data and provide data support for future model optimization. This approach not only avoids unnecessary waste of resources, but also ensures the robustness of the system in long-term operation.

[0153] Countermeasures for Medium-Risk Potential Error Analysis

[0154] When the error status is classified as medium risk, it indicates that the error may have a certain business impact on the system and requires semi-automatic intervention and manual review. The system triggers a preset alarm mechanism and sends detailed information about the analysis error (such as the anomaly index and related text content) to a designated review queue for review by a human analyst. The system also uses a backup model or preset error handling rules to temporarily adjust the subsequent sentiment analysis process to ensure that the accuracy of the analysis results does not deteriorate further. This type of measure is suitable for scenarios where serious business damage has not yet occurred, but risk control is necessary.

[0155] Countermeasures for high-risk potential error analysis

[0156] When the error status is determined to be high risk, it indicates that the model's incorrect analysis could have a serious impact on the business, requiring urgent intervention and a comprehensive response. In this case, the system immediately switches to a backup sentiment analysis model or rolls back to the previous stable version, suspending the use of the current model. The system generates a detailed analysis report and submits it to the technical team for rapid remediation. Expert review may also be implemented to re-label specific datasets and fine-tune the model. In addition, a comprehensive audit of the model's historical data and analysis process is required to prevent similar issues from recurring. This emergency response ensures that the system can be restored to stability as quickly as possible, avoiding significant damage to the business and brand image.

[0157] By introducing the sentiment polarization index and the expectation violation index, the present invention enables the system to more accurately capture the negative emotions hidden in complex emotions and sarcastic language, avoiding misjudgment as positive feedback. The introduction of real-time parameter capture and storage modules ensures low-latency transmission and integrity of data, enabling the model to quickly capture potential errors and perform anomaly analysis. At the same time, the use of machine learning models for multi-level comparison and analysis further enhances the robustness of the model. Even when faced with complex emotional expressions or deviations from expectations, the system can respond promptly when errors first appear, avoiding the accumulation of misjudgments. Comprehensive risk assessment and dynamic response mechanisms ensure that errors in the model's daily operations are no longer ignored, providing companies with more reliable sentiment analysis results and helping them gain a deeper understanding of market feedback.

[0158] The present invention divides risks into three categories: low, medium, and high, to ensure that enterprise resources can be reasonably allocated, to achieve on-demand intervention, and to avoid unnecessary manual intervention and waste of resources. In low-risk situations, log records and data accumulation are used for continuous monitoring to provide data support for subsequent model optimization. In medium-risk scenarios, the system will trigger an alarm and conduct manual review, while using the backup model to make temporary adjustments to avoid excessive impact on the business. In high-risk situations, the system quickly switches to the backup model and generates a detailed audit report to ensure that model problems can be fixed in a timely manner. This multi-level response mechanism reduces the possibility of business losses, improves the flexibility and efficiency of the system in dealing with uncertainty, and at the same time avoids damage to the brand image due to analytical errors, ensuring the healthy development of the relationship between the company and its customers.

[0159] The above description is merely illustrative of certain exemplary embodiments of the present invention. It goes without saying that those skilled in the art will be able to modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of protection of the claims.

Claims

1. A search engine system based on text sentiment analysis, characterized in that: It includes real-time parameter capture and storage module, anomaly analysis and model comparison module, risk assessment module and response module: Real-time parameter capture and storage module: During the sentiment analysis process, each text will generate a series of parameters. The parameters generated when the sentiment analysis model performs text analysis are captured and stored in real time to ensure low latency and integrity of data flow; The anomaly analysis and model comparison module analyzes the acquired parameter information for anomalies and inputs the analyzed parameter information into a pre-trained machine learning model to determine whether there are potential errors in the current sentiment analysis through feature comparison. The risk assessment module, based on the analysis results of the machine learning model, further comprehensively analyzes the analysis process with potential errors and identifies the risk level of potential errors; Response measures module, taking different response measures according to different risk levels; The parameters generated by the sentiment analysis model during text analysis include sentiment complexity information and expectation consistency information. Sentiment complexity information refers to the situation where multiple emotions appear simultaneously in the text and conflict or intertwine with each other. Expectation consistency information is used to measure whether the text expression meets the user's expectations or conventional cognition, reflecting the degree of consistency between the user experience and their original expectations. After obtaining the sentiment complexity information and expected consistency information generated by the sentiment analysis model during text analysis, the sentiment complexity information is analyzed for anomalies to generate a sentiment polarization index. After performing an anomaly analysis on the expected consistency information, an expected violation index is generated. The generated sentiment polarization index and expected violation index are input into a pre-learned machine learning model. The analysis anomaly index is generated through the machine learning model. The analysis anomaly index is used to determine whether there are potential errors in the current sentiment analysis.

2. A search engine system based on text sentiment analysis according to claim 1, characterized in that: The analysis anomaly index generated by the sentiment analysis model during text analysis is compared with the pre-set analysis anomaly index reference threshold to determine whether there are potential errors in the current sentiment analysis. The analysis results are as follows: If the analysis anomaly index is greater than or equal to a preset analysis anomaly index reference threshold, the text analysis state of the sentiment analysis model is classified as a potential error analysis state; If the analysis anomaly index is less than a preset analysis anomaly index reference threshold, the text analysis state of the sentiment analysis model is classified as a correct analysis state.

3. The search engine system based on text sentiment analysis according to claim 2, characterized in that: After the sentiment analysis model is divided into a potential error analysis state during the text analysis process, several analysis anomaly indices generated by the subsequent process of the sentiment analysis model are continuously obtained to establish a data set for comprehensive analysis, and the analysis anomaly indices in the analysis set are compared with the first-level reference threshold, the second-level reference threshold, and the analysis anomaly index reference threshold, wherein the second-level reference threshold is greater than the first-level reference threshold, and the first-level reference threshold is greater than the analysis anomaly index reference threshold. The analysis anomaly index is compared and analyzed with the second-level reference threshold, the first-level reference threshold, and the analysis anomaly index reference threshold, and the number of analysis anomaly indices that are less than the first-level reference threshold and greater than or equal to the analysis anomaly index reference threshold is calibrated as La, the number of analysis anomaly indices that are less than the second-level reference threshold and greater than or equal to the first-level reference threshold is calibrated as Lb, and the number of analysis anomaly indices greater than or equal to the second-level reference threshold is calibrated as Lc; Comprehensively analyze La, Lb, and Lc to generate the risk level coefficient RLC based on the following formula: , Wherein, k1, k2, and k3 are preset proportional coefficients of La, Lb, and Lc, respectively, and k1, k2, and k3 are all greater than 0.

4. The search engine system based on text sentiment analysis according to claim 3, characterized in that: The generated risk level coefficient is compared and analyzed with the pre-set first risk level coefficient reference threshold and the second risk level coefficient reference threshold to determine the risk level of the potential error analysis state. The comparison and analysis results are as follows: If the risk level coefficient is less than a preset risk level coefficient reference threshold, the potential error analysis state is classified as a low-risk potential error analysis; If the risk level coefficient is greater than or equal to the first risk level coefficient reference threshold and less than the second risk level coefficient reference threshold, the potential error analysis state is classified as a medium risk potential error analysis; If the risk level coefficient is greater than the second risk level coefficient reference threshold, the potential error analysis state is classified as a high-risk potential error analysis.

5. A search engine system based on text sentiment analysis according to claim 1, characterized in that: The logic for generating the sentiment polarization index by performing anomaly analysis on sentiment complexity information within the detection window is as follows: Under the detection window, the text to be analyzed is broken down into multiple emotion categories, and each emotion category is assigned an emotion intensity score through the emotion classification model. The emotion distribution vector generated for each text is calibrated as V s , where s∈{pos,neg,neu}, pos represents positive emotion, neg represents negative emotion, neu represents neutral emotion, and the value of each element of the emotion distribution vector is the emotion intensity score, which is expressed as follows: V s ={P pos ,P neg ,P neu }, where P pos is the intensity score of positive emotion, P neg is the intensity score of negative sentiment, P neu is the intensity score of neutral emotion; The opposition function is defined to measure the difference between two emotion categories. The calculation expression is as follows: , Where D(s1,s2) is the degree of opposition between two emotion categories, represents the intensity score of sentiment category s1 in the text, It represents the intensity score of sentiment category s2 in the text. s1 and s2 correspond to any sentiment category in the sentiment distribution vector respectively. ∈ is a small constant that prevents the denominator from being zero to ensure calculation stability. The D emotion category opposition degree D(s1,s2) is accumulated to generate the emotion polarization index. The calculation expression is as follows: , Where EPI is the sentiment polarization index, (s1,s2)∈C is the combination of sentiment categories, and C is the set of all sentiment category combinations.

6. A search engine system based on text sentiment analysis according to claim 1, characterized in that: The logic for generating the expected violation index by performing anomaly analysis on expected consistency information within the detection window is as follows: In the detection window, the similarity analysis is performed between the input text and the preset expected model. The calculation expression is as follows: , Where S match is the similarity score between the text and the expected model, T i is the i-th keyword in the input text, M i is the i-th keyword in the expected model, n is the total number of keywords, w i It is the weight of the keyword; For each keyword T in the text i Calculate the degree of deviation from the expected model in the semantic space. The calculation expression is as follows: , Where D shift Indicates the total outlier offset distance, E(T i ) is the keyword T in the input text i The vector embedding, E(M i ) is the keyword M in the expected model i The vector embedding of ||E(T i )-E(M i )|| is the Euclidean distance; In order to capture the impact of emotional or contextual transitions in sentences on expectancy violation, the contextual transition influence coefficient is introduced. The calculation expression is as follows: , Where C shift is the context transition influence coefficient, α is the adjustment coefficient, which is used to control the influence weight of transition words on the results, N contrast is the number of occurrences of transition words in the sentence, N total is the total number of words in the sentence; Combine the similarity score S between the text and the expected model match , total outlier offset distance D shift and contextual transition influence coefficient C shift , generate the expected violation index, the calculation expression is as follows: , Where EVI is the expectation violation index.

Citation Information

Patent Citations

  • Marketing platform risk early warning management system for Internet promotion

    CN118674278A

  • Data mining method, data mining apparatus, electronic device and storage medium

    US20230004613A1