Big Data Information Background Real-time Precision Analysis and Processing System

Through natural language processing technology based on deep learning, preprocessing and semantic encoding of evaluation data, the problem that traditional sentiment analysis methods cannot adapt to user expression habits and changes in the language environment is solved, and more accurate product sentiment analysis is achieved.

CN119477442BActive Publication Date: 2025-07-22BEIJING MENGSHEN INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411521564.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-29
Publication Date
2025-07-22
Estimated Expiration
2044-10-29

AI Technical Summary

Technical Problem

Traditional sentiment analysis methods are difficult to dynamically adjust, cannot adapt to changes in user expression habits and language environment, and cannot effectively capture semantic information and context dependencies in the text, resulting in inaccurate analysis results.

Method used

The evaluation data is preprocessed by deep learning-based natural language processing technology, and the reconstruction and semantic encoding of product public opinion question template questions are introduced. Product public opinion is evaluated through product emotional tendency reasoning and significant aggregation representation, and automatically adjusts to adapt to changes in user expression habits and language environment.

Benefits of technology

It significantly improves the accuracy and generalization ability of product sentiment analysis, can better understand the semantic information and context dependencies in the text, and adapt to changes in user expression habits and locale environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119477442B_ABST
    Figure CN119477442B_ABST
Patent Text Reader

Abstract

The present application provides a real-time precise analysis and processing system for big data information background, which relates to the field of data analysis. It uses natural language processing technology based on deep learning to preprocess each evaluation data in the set of evaluation data of the target product, introduces the reconstruction processing and semantic coding of the product public opinion question template questions, and conducts product sentiment tendency reasoning on each reconstructed evaluation data. Based on this, it automatically evaluates whether the product public opinion is positive or negative according to the semantic significant aggregation representation between the semantic inference coding features of each product sentiment tendency obtained by the reasoning, better adapts to the changes in users' expression habits and language environments, and at the same time can better understand the semantic information and context dependence relationship in the text, thus significantly improving the accuracy and generalization ability of product sentiment analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data analysis, and more specifically, to a real-time precise analysis and processing system for big data information in the background. Background Art

[0002] With the rapid development of the Internet and the popularity of social media, the amount of data generated by users has increased exponentially. This data is not only huge in quantity but also covers a wide range of fields, including posts and comments on social media, as well as user reviews on e-commerce platforms. Each piece of data carries the true feelings and opinions of users. Therefore, using this data for sentiment analysis can help enterprises better understand and meet the needs of users, thereby enhancing the market competitiveness of products and the user experience.

[0003] However, traditional methods for sentiment analysis of products often establish a model once and it is difficult to dynamically adjust according to new data. Over time, changes in users' expression habits and language environments may cause the model to become invalid. For example, on social media, new popular words and expressions emerge continuously, while traditional models are usually static and unable to update and adapt to these changes in a timely manner. This lack of dynamic adjustment ability causes the model to become inaccurate when faced with new data. In addition, traditional methods often rely on keyword matching or simple statistical methods, which cannot effectively capture semantic information and context dependencies in the text. Keyword matching can only identify isolated words and cannot understand the meaning of words in specific contexts, resulting in inaccurate sentiment analysis results and being unable to fully reflect the true evaluation of products.

[0004] Therefore, a real-time precise analysis and processing system for big data information in the background is needed to solve the above problems. Summary of the Invention

[0005] To solve the above technical problems, this application is proposed. An embodiment of this application provides a real-time precise analysis and processing system for big data information in the background.

[0006] According to one aspect of this application, a real-time precise analysis and processing system for big data information in the background is provided, which includes:

[0007] An evaluation data acquisition and processing module, configured to collect a set of evaluation data of a target product and perform preprocessing to obtain a set of preprocessed evaluation data;

[0008] An evaluation data reconstruction module, configured to add product public opinion question template questions to the tails of each preprocessed evaluation data in the set of preprocessed evaluation data to obtain a set of reconstructed evaluation data;

[0009] A product sentiment tendency inference module, which is used to perform semantic encoding on the set of the reconstructed evaluation data to obtain a set of reconstructed evaluation data semantic encoding vectors, and then perform product sentiment tendency inference to obtain a set of product sentiment tendency semantic inference encoding vectors;

[0010] A product sentiment tendency aggregation module, which is used to perform product sentiment tendency semantic feature aggregation on the set of the product sentiment tendency semantic inference encoding vectors to obtain a significant aggregation representation of the product sentiment tendency;

[0011] Among them, the product sentiment tendency aggregation module includes: a product static energy factor calculation unit, which is used to calculate the static energy factors of each product sentiment tendency semantic inference encoding vector in the set of the product sentiment tendency semantic inference encoding vectors to obtain a set of product sentiment tendency semantic static energy factors; an energy distribution significant aggregation unit, which is used to perform significant aggregation based on energy distribution on the set of the product sentiment tendency semantic inference encoding vectors based on the set of the product sentiment tendency semantic static energy factors to obtain the significant aggregation representation of the product sentiment tendency;

[0012] An analysis result generation module, which is used to obtain an analysis result based on the significant aggregation representation of the product sentiment tendency.

[0013] The present application has at least the following technical effects: Compared with the prior art, the big data information background real-time precise analysis and processing system provided by the present application uses natural language processing technology based on deep learning to preprocess each evaluation data in the set of evaluation data of the target product, introduces the reconstruction processing and semantic encoding of the product public opinion question template questions, and performs product sentiment tendency inference on each reconstructed evaluation data, so as to automatically evaluate whether the product public opinion is positive or negative according to the semantic significant aggregation representation between the product sentiment tendency semantic inference encoding features obtained by the inference, and can be automatically adjusted step by step as new data flows in to better adapt to the changes in the user's expression habits and language environment, and at the same time can better understand the semantic information and context dependence relationship in the text, thereby significantly improving the accuracy and generalization ability of product sentiment analysis. Description of the Drawings

[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained according to these drawings without creative efforts. In the drawings:

[0015] Figure 1It is a system block diagram of a big data information background real-time precise analysis and processing system according to an embodiment of the present application.

[0016] Figure 2 It is a schematic diagram of data flow of a big data information background real-time precise analysis and processing system according to an embodiment of the present application.

[0017] Figure 3 It is a block diagram of a product sentiment tendency inference module in a big data information background real-time precise analysis and processing system according to an embodiment of the present application.

[0018] Figure 4 It is a block diagram of a product sentiment tendency aggregation module in a big data information background real-time precise analysis and processing system according to an embodiment of the present application.

[0019] Figure 5 It is a block diagram of a product static energy factor calculation unit in a big data information background real-time precise analysis and processing system according to an embodiment of the present application.

[0020] Figure 6 It is a block diagram of an energy distribution significant aggregation unit in a big data information background real-time precise analysis and processing system according to an embodiment of the present application. Detailed implementation manners

[0021] Next, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. It should be understood that the present application is not limited by the exemplary embodiments described herein.

[0022] With the rapid development of the Internet and the wide penetration of social media, the amount of user-generated data is growing at an alarming rate. These data are not only huge in quantity, but also cover numerous fields, such as posts and comments on social media, and user feedback on e-commerce platforms, etc. These data contain the true emotions and opinions of users, providing a valuable opportunity for enterprises to gain insights into user needs through sentiment analysis, which can enhance the market competitiveness of products and optimize the user experience.

[0023] However, traditional sentiment analysis methods often have limitations. They are usually built based on one-time models and lack the adaptability to new data and the ability of dynamic adjustment. Over time, users' language habits and expression patterns are constantly evolving. For example, there are an endless stream of new words and expressions on social media, while traditional models often fail to capture these changes in a timely manner, resulting in a decrease in the accuracy and effectiveness of the models over time. In addition, traditional methods often rely on keyword matching or simple statistical analysis, which cannot deeply understand the semantic content and context dependence of the text. Keyword matching can only identify isolated words and cannot grasp the specific meanings of these words in a specific context, which limits the accuracy of sentiment analysis and makes it unable to fully reflect users' true evaluations of products.

[0024] To address the above technical problems, the technical concept of this application is to collect a set of evaluation data of the target product, and use natural language processing technology based on deep learning to preprocess each evaluation data, introduce the reconstruction processing and semantic encoding of product public opinion question template questions, and perform product sentiment tendency inference on each reconstructed evaluation data. Based on the semantic significant aggregation representation between the semantic inference coding features of each product sentiment tendency obtained from the inference, it automatically evaluates whether the product public opinion is positive or negative. Compared with traditional keyword matching or simple statistical methods, it can automatically adjust gradually as new data flows in to better adapt to the changes in users' expression habits and language environments. At the same time, it can better understand the semantic information and context dependence in the text, thus significantly improving the accuracy and generalization ability of product sentiment analysis.

[0025] Figure 1 It is a system block diagram of the big data information background real-time precise analysis and processing system according to an embodiment of this application. Figure 2 It is a schematic diagram of data flow of the big data information background real-time precise analysis and processing system according to an embodiment of this application. As Figure 1 and Figure 2As shown, in the real-time precise analysis and processing system 100 for big data information background, it includes: an evaluation data acquisition and processing module 110, which is used to collect a set of evaluation data of the target product and perform preprocessing to obtain a set of preprocessed evaluation data; an evaluation data reconstruction module 120, which is used to add product public opinion question template questions to the tails of each preprocessed evaluation data in the set of preprocessed evaluation data to obtain a set of reconstructed evaluation data; a product sentiment tendency inference module 130, which is used to perform semantic encoding on the set of reconstructed evaluation data to obtain a set of reconstructed evaluation data semantic encoding vectors, and then perform product sentiment tendency inference to obtain a set of product sentiment tendency semantic inference encoding vectors; a product sentiment tendency aggregation module 140, which is used to perform product sentiment tendency semantic feature aggregation on the set of product sentiment tendency semantic inference encoding vectors to obtain a significant aggregation representation of product sentiment tendency; an analysis result generation module 150, which is used to obtain an analysis result based on the significant aggregation representation of product sentiment tendency.

[0026] In the embodiment of the present application, the evaluation data acquisition and processing module 110 is used to collect a set of evaluation data of the target product and perform preprocessing to obtain a set of preprocessed evaluation data. Specifically, in the embodiment of the present application, the evaluation data acquisition and processing module is used to: collect the set of evaluation data of the target product; perform preprocessing on each evaluation data in the set of evaluation data to obtain the set of preprocessed evaluation data.

[0027] Specifically, each piece of evaluation data in the set of evaluation data collected for the target product usually contains the specific feelings, experiences, and opinion information of users about the product. These information are the core data for judging the direction of product public opinion. By collecting the evaluation data of different users on the target product and performing sentiment semantic analysis on it, it can help the enterprise to more accurately grasp the user sentiment, and then provide support for enterprise decision-making. Here, the evaluation data of the target product is obtained from the database. As the user evaluation data changes continuously, the evaluation data in the database also changes correspondingly. Accordingly, considering that each piece of evaluation data in the set of evaluation data may contain content unrelated to the product, such as advertisements, links, or meaningless characters, these information will interfere with the sentiment analysis. At the same time, there may also be spelling mistakes or typing mistakes when users input evaluations in the evaluation data, which can easily cause the model to fail to correctly understand the text. Based on this, in order to improve the quality of the data and enable the subsequent sentiment analysis model to more accurately identify the true intentions and sentiment tendencies of users, in the technical solution of this application, each piece of evaluation data in the set of evaluation data is preprocessed to obtain a set of preprocessed evaluation data. Specifically, in a specific implementable manner of the embodiment of this application, the steps of preprocessing each piece of evaluation data in the set of evaluation data may include: a. Using regular expressions to delete the links, advertisement content, special characters, and punctuation marks in the evaluation data; b. Using a spelling correction tool to check and correct the rating data; c. Unifying the format of the evaluation data and converting all of them into text form.

[0028] In the embodiment of this application, the evaluation data reconstruction module 120 is used to add product public opinion question template questions at the end of each preprocessed evaluation data in the set of preprocessed evaluation data to obtain a set of reconstructed evaluation data. Specifically, the set of preprocessed evaluation data expresses the evaluations and feelings of different users about the product. In order to guide and structure the user evaluation data and help the model better understand the context in the comments to more accurately capture the user sentiment tendency, in the technical solution of this application, product public opinion question template questions are added at the end of each preprocessed evaluation data in the set of preprocessed evaluation data to obtain a set of reconstructed evaluation data. Specifically, in a specific implementable manner of the embodiment of this application, the following product public opinion question template questions can be added at the end of the preprocessed evaluation data: "Are you satisfied with the overall quality of the product?", "In which aspects do you think the product needs to be improved?", etc. By adding the question template questions, additional context information is provided, which helps the model better understand the background of the user evaluation and then can more accurately capture the user sentiment tendency towards the product.

[0029] In the embodiment of the present application, the product sentiment tendency inference module 130 is configured to perform semantic encoding on the set of the reconstructed evaluation data to obtain a set of reconstructed evaluation data semantic encoding vectors, and then perform product sentiment tendency inference to obtain a set of product sentiment tendency semantic inference encoding vectors. Specifically, Figure 3 It is a block diagram of the product sentiment tendency inference module in the big data information background real-time precision analysis and processing system according to the embodiment of the present application. As Figure 3 shown, the product sentiment tendency inference module 130 includes: a reconstructed evaluation data semantic encoding unit 131, configured to perform semantic encoding on each of the reconstructed evaluation data in the set of the reconstructed evaluation data to obtain the set of the reconstructed evaluation data semantic encoding vectors; a product sentiment tendency semantic feature encoding unit 132, configured to input each of the reconstructed evaluation data semantic encoding vectors in the set of the reconstructed evaluation data semantic encoding vectors into a product sentiment tendency inference device based on an RNN model to obtain the set of the product sentiment tendency semantic inference encoding vectors.

[0030] Specifically, in order to further understand and process the semantic meaning between the contexts of each of the reconstructed evaluation data in the set of the reconstructed evaluation data, so as to more accurately judge whether the user's attitude towards the product is positive or negative, in the technical solution of the present application, semantic encoding is performed on each of the reconstructed evaluation data in the set of the reconstructed evaluation data to capture and refine the key implicit evaluation semantic information, and a set of reconstructed evaluation data semantic encoding vectors is obtained. Subsequently, considering that each of the reconstructed evaluation data semantic encoding vectors in the set of the reconstructed evaluation data semantic encoding vectors contains the user evaluation semantic information and the hidden sentiment tendency in different evaluation data. The RNN model captures the sequential dependence relationship between words in the text, and uses the internal memory mechanism (such as the hidden state) to remember the previous lexical information and utilize it when processing the current lexical, so as to more accurately understand the semantics of the text. Therefore, in order to be able to more accurately capture and quantify the sentiment tendency in the user evaluation, in the technical solution of the present application, each of the reconstructed evaluation data semantic encoding vectors in the set of the reconstructed evaluation data semantic encoding vectors is input into a product sentiment tendency inference device based on an RNN model to obtain a set of product sentiment tendency semantic inference encoding vectors. That is, through the RNN model, the sentiment tendency implicit in each evaluation data can be inferred through the semantic information of the evaluation data, so as to better understand the complex sentiment expression of the user in each evaluation.

[0031] In the embodiment of the present application, the product sentiment tendency aggregation module 140 is configured to perform product sentiment tendency semantic feature aggregation on the set of the product sentiment tendency semantic inference encoding vectors to obtain a significant aggregation representation of the product sentiment tendency. Specifically, Figure 4Block diagram of the product sentiment tendency aggregation module in the big data information background real-time precise analysis and processing system according to an embodiment of the present application. As Figure 4 shown, the product sentiment tendency aggregation module 140 includes: a product static energy factor calculation unit 141, configured to calculate the static energy factors of each product sentiment tendency semantic inference coding vector in the set of product sentiment tendency semantic inference coding vectors to obtain a set of product sentiment tendency semantic static energy factors; an energy distribution significant aggregation unit 142, configured to perform significant aggregation based on energy distribution on the set of product sentiment tendency semantic inference coding vectors based on the set of product sentiment tendency semantic static energy factors to obtain the product sentiment tendency significant aggregation representation.

[0032] Correspondingly, considering that the set of product sentiment tendency semantic inference coding vectors expresses the semantic feature representation of each user evaluation at the sentiment analysis level, the semantic information carried in its vector also implies the characteristics of the positive or negative sentiment tendency of the user towards the product. In order to further refine and aggregate these sentiment characteristics, so as to obtain a more significant and concentrated sentiment representation, thereby improving the accuracy and interpretability of product public opinion sentiment analysis, in the technical solution of the present application, product sentiment tendency semantic feature aggregation is performed on the set of product sentiment tendency semantic inference coding vectors to obtain a product sentiment tendency significant aggregation representation vector as the product sentiment tendency significant aggregation representation.

[0033] Specifically, first, calculate the static energy factors of each product sentiment tendency semantic inference coding vector in the set of product sentiment tendency semantic inference coding vectors to obtain a set of product sentiment tendency semantic static energy factors. Specifically, by calculating the static energy factors, the importance of each sentiment inference coding vector in sentiment analysis can be quantified, which helps to identify which feature vectors have a greater impact on the overall sentiment tendency and provides decision support for subsequent selection of the initial clustering center.

[0034] Then, based on the set of product sentiment semantic static energy factors, perform significant aggregation based on energy distribution on the set of product sentiment semantic inference encoding vectors to obtain the product sentiment significant aggregation representation, which can identify those features that occupy important positions in the sentiment distribution, thus better reflecting the overall sentiment tendency. Specifically, select the product sentiment semantic inference encoding vector corresponding to the largest product sentiment semantic static energy factor from the set of product sentiment semantic static energy factors as the initial center vector for product sentiment semantic inference clustering. In this way, a most representative initial center point can be determined, making the final clustering result more in line with the actual sentiment distribution and enhancing the ability to capture and understand real features. Next, calculate the dynamic aggregation energy factors of each inference encoding vector in the set of product sentiment semantic inference encoding vectors to obtain the set of product sentiment semantic inference dynamic aggregation energy factors. In particular, the dynamic aggregation energy factor is calculated based on the static energy factor of the initial center vector and each product sentiment semantic inference encoding vector, as well as the spatial span between the initial center vector and each product sentiment semantic inference encoding vector. This can make the obtained dynamic aggregation energy factor not only consider the proximity in position but also reflect the importance difference of the internal attributes of each point, thereby being able to more accurately reflect the interaction between different samples. Then, in order to better focus on the key features that can best represent the sentiment tendency, input each dynamic aggregation energy factor into a gated mask unit for adaptive adjustment and attention of weights, so as to strengthen the attention to significant features and reduce the influence of irrelevant information at the same time, obtaining the set of product sentiment semantic inference dynamic aggregation weight factors. Finally, perform weighted aggregation on the set of product sentiment semantic inference encoding vectors based on the mixed weights in the set of weight factors to generate the final product sentiment significant aggregation representation vector. This can synthesize the importance of each sentiment inference encoding vector and generate a more accurate and representative comprehensive sentiment representation to assist in subsequent product sentiment analysis.

[0035] More specifically, Figure 5 is a block diagram of the product static energy factor calculation unit in the big data information background real-time precise analysis and processing system according to an embodiment of the present application. As Figure 5As shown, the product static energy factor calculation unit 141 includes: a product sentiment semantic inference mean variance calculation sub-unit 1411, configured to calculate the mean and variance of the product sentiment semantic inference coding vector to obtain the product sentiment semantic inference mean and the product sentiment semantic inference variance; a product sentiment semantic inference difference modulation vector generation sub-unit 1412, configured to subtract the product sentiment semantic inference coding vector from the product sentiment semantic inference mean in a position-by-position manner, and calculate the fourth power of each position of the subtracted feature vector to obtain a product sentiment semantic inference difference modulation vector; a product sentiment semantic inference expected value calculation sub-unit 1413, configured to calculate the expected value of the product sentiment semantic inference difference modulation vector to obtain a product sentiment semantic inference expected value; and a product sentiment semantic static energy factor generation sub-unit 1414, configured to divide the product sentiment semantic inference expected value by the square of the product sentiment semantic inference variance, and input the obtained value into a sigmoid function to obtain the product sentiment semantic static energy factor.

[0036] More specifically, Figure 6 is a block diagram of an energy distribution significant aggregation unit in a big data information background real-time precise analysis and processing system according to an embodiment of the present application. As Figure 6As shown, the energy distribution significant aggregation unit 142 includes: a product sentiment tendency semantic reasoning clustering initial center vector determination subunit 1421, configured to select the product sentiment tendency semantic reasoning coding vector corresponding to the largest product sentiment tendency semantic static energy factor in the set of product sentiment tendency semantic static energy factors as the product sentiment tendency semantic reasoning clustering initial center vector; a product sentiment tendency dynamic aggregation energy factor calculation subunit 1422, configured to calculate the dynamic aggregation energy factor of each product sentiment tendency semantic reasoning coding vector in the set of product sentiment tendency semantic reasoning coding vectors based on the spatial span between each product sentiment tendency semantic reasoning coding vector in the set of product sentiment tendency semantic reasoning coding vectors and the product sentiment tendency semantic reasoning clustering initial center vector, as well as the static energy factor of each product sentiment tendency semantic reasoning coding vector and the static energy factor of the product sentiment tendency semantic reasoning clustering initial center vector, to obtain a set of product sentiment tendency semantic reasoning dynamic aggregation energy factors; a dynamic aggregation energy factor gating mask subunit 1423, configured to input the set of product sentiment tendency semantic reasoning dynamic aggregation energy factors into a gating mask unit to obtain a set of product sentiment tendency semantic reasoning dynamic aggregation weight factors; and a product sentiment tendency significant aggregation representation generation subunit 1424, configured to calculate the weighted sum of the set of product sentiment tendency semantic reasoning coding vectors with the set of product sentiment tendency semantic reasoning dynamic aggregation weight factors to obtain a product sentiment tendency semantic reasoning coding vector as the product sentiment tendency significant aggregation representation.

[0037] More specifically, in the embodiment of the present application, the product sentiment tendency dynamic aggregation energy factor calculation subunit is configured to: multiply the static energy factor of the product sentiment tendency semantic reasoning coding vector and the static energy factor of the product sentiment tendency semantic reasoning clustering initial center vector by a first weighting parameter to obtain a first product sentiment tendency semantic reasoning dynamic aggregation energy factor; multiply the square of the spatial span between the product sentiment tendency semantic reasoning coding vector and the product sentiment tendency semantic reasoning clustering initial center vector by a second weighting parameter to obtain a second product sentiment tendency semantic reasoning dynamic aggregation energy factor; and divide the first product sentiment tendency semantic reasoning dynamic aggregation energy factor by the second product sentiment tendency semantic reasoning dynamic aggregation energy factor to obtain the product sentiment tendency semantic reasoning dynamic aggregation energy factor.

[0038] More specifically, in the embodiments of the present application, the dynamic aggregation energy factor gating mask subunit includes: a product sentiment dynamic aggregation energy factor normalization secondary subunit, configured to perform normalization processing on the set of product sentiment tendency semantic inference dynamic aggregation energy factors to obtain a set of normalized product sentiment tendency semantic inference dynamic aggregation energy factors; and a normalized product sentiment dynamic aggregation energy factor mask secondary subunit, configured to perform masking processing on the set of normalized product sentiment tendency semantic inference dynamic aggregation energy factors to obtain a set of product sentiment tendency semantic inference dynamic aggregation weight factors.

[0039] More specifically, in the embodiments of the present application, the product sentiment dynamic aggregation energy factor normalization secondary subunit is configured to: use the sigmoid function to perform normalization processing on the set of product sentiment tendency semantic inference dynamic aggregation energy factors to obtain the set of normalized product sentiment tendency semantic inference dynamic aggregation energy factors.

[0040] More specifically, in the embodiments of the present application, the normalized product sentiment dynamic aggregation energy factor mask secondary subunit is configured to: in response to each normalized product sentiment tendency semantic inference dynamic aggregation energy factor in the set of normalized product sentiment tendency semantic inference dynamic aggregation energy factors being greater than a predetermined threshold, set the normalized product sentiment tendency semantic inference dynamic aggregation energy factor to its original value and set the rest to zero to obtain the set of product sentiment tendency semantic inference dynamic aggregation weight factors.

[0041] In the embodiments of the present application, specifically, the product sentiment tendency aggregation module is configured to: perform product sentiment tendency semantic feature aggregation processing on the set of product sentiment tendency semantic inference encoding vectors according to the following formula; where the formula is:

[0042] X = {x1, x2,..., x i ,..., x n}

[0043]

[0044] x c = x m

[0045]

[0046] w si = mask(w i )

[0047]

[0048] Among them, X is the set of product sentiment tendency semantic inference coding vectors, and x1, x2,..., x i ,..., x n are respectively the 1st, 2nd,..., ith,..., nth product sentiment tendency semantic inference coding vectors in the set of product sentiment tendency semantic inference coding vectors. is the eigenvalue at each position in the ith product sentiment tendency semantic inference coding vector, E(·) is for calculating the expected value, μ i and σ i 2 are respectively the mean and variance of X i , sigmoid(·) is the sigmoid function. is the static energy factor of product sentiment tendency semantics corresponding to x i , argmax j (·) is for selecting the j value corresponding to the maximum value, m is the maximum matching value, x c is the initial center vector of product sentiment tendency semantic inference clustering. is the static energy factor of product sentiment tendency semantics corresponding to x m , Count(x i →x m ) represents the spatial span between x i and x m , γ and δ are weighting parameters. is the dynamic aggregation energy factor of product sentiment tendency semantic inference corresponding to X i , W i is the normalized dynamic aggregation energy factor of product sentiment tendency semantic inference corresponding to X i , mask(·) is for masking processing, θ is a predetermined threshold, w si is the dynamic aggregation weight factor of product sentiment tendency semantic inference corresponding to x i , n is the number of vectors in the set of product sentiment tendency semantic inference coding vectors, and x < is the product sentiment tendency semantic inference coding vector.

[0049] In the embodiment of the present application, the analysis result generation module 150 is configured to obtain an analysis result based on the significant aggregation representation of the product sentiment tendency. Specifically, in the embodiment of the present application, the analysis result generation module is configured to: input the significant aggregation representation vector of the product sentiment tendency into the real-time product public opinion analysis module based on a classifier to obtain the analysis result, and the analysis result is used to indicate that the product public opinion is positive or the product public opinion is negative. Specifically, classification processing is performed on the significant aggregation representation of the product sentiment tendency obtained by significantly aggregating the set of product sentiment tendency semantic inference coding vectors, so as to automatically evaluate whether the product public opinion is positive or the product public opinion is negative. Compared with the traditional keyword matching or simple statistical method, it can be automatically adjusted step by step as new data flows in to better adapt to the changes in user expression habits and language environments, and at the same time can better understand the semantic information and context dependence relationships in the text, thereby significantly improving the accuracy and generalization ability of product sentiment analysis.

[0050] Particularly, the present application considers that each product sentiment tendency semantic inference coding vector in the set of product sentiment tendency semantic inference coding vectors respectively represents the product sentiment tendency semantic inference coding feature representation obtained by text semantic feature encoding and text semantic feature decoding of each preprocessed evaluation data in the set of preprocessed evaluation data. When performing product sentiment tendency semantic feature aggregation on the set of product sentiment tendency semantic inference coding vectors, the text semantic encoding features and text semantic decoding differences of each preprocessed evaluation data in the set of preprocessed evaluation data will cause the sentiment semantic fusion representation of the significant aggregation representation vector of the product sentiment tendency to have a complex aggregation space structure. Therefore, it is expected to improve its classification regression convergence and generalization effect under the complex aggregation space structure.

[0051] Preferably, in an example of the present application, passing the product sentiment tendency significantly aggregated representation vector through the real-time product public opinion analysis module based on a classifier to obtain an analysis result includes: calculating the sum of the absolute values of the respective eigenvalues of the product sentiment tendency significantly aggregated representation vector to obtain a first product sentiment tendency significantly aggregated representation space structure value, and calculating the square root of the sum of the squares of the respective eigenvalues of the product sentiment tendency significantly aggregated representation vector to obtain a second product sentiment tendency significantly aggregated representation space structure value; multiplying each eigenvalue of the product sentiment tendency significantly aggregated representation vector by the first product sentiment tendency significantly aggregated representation space structure value and the second product sentiment tendency significantly aggregated representation space structure value respectively to obtain a first product sentiment tendency significantly aggregated representation structure reference value and a second product sentiment tendency significantly aggregated representation structure reference value corresponding to each eigenvalue; multiplying each eigenvalue of the product sentiment tendency significantly aggregated representation vector by the length of the product sentiment tendency significantly aggregated representation vector and the square root of the length respectively to obtain a first product sentiment tendency significantly aggregated representation scale transformation value and a second product sentiment tendency significantly aggregated representation scale transformation value corresponding to each eigenvalue; dividing the first product sentiment tendency significantly aggregated representation structure reference value by the difference between the first product sentiment tendency significantly aggregated representation space structure value and the first product sentiment tendency significantly aggregated representation scale transformation value to obtain a first product sentiment tendency significantly aggregated representation transformation adjustment value; dividing the second product sentiment tendency significantly aggregated representation structure reference value by the difference between the second product sentiment tendency significantly aggregated representation space structure value and the second product sentiment tendency significantly aggregated representation scale transformation value to obtain a second product sentiment tendency significantly aggregated representation transformation adjustment value; calculating the weighted sum of the first product sentiment tendency significantly aggregated representation transformation adjustment value and the second product sentiment tendency significantly aggregated representation transformation adjustment value to obtain each eigenvalue of the optimized product sentiment tendency significantly aggregated representation vector; passing the optimized product sentiment tendency significantly aggregated representation vector through the real-time product public opinion analysis module based on a classifier to obtain the analysis result.

[0052] Among them, the optimized representation of the product sentiment tendency significantly aggregated representation vector is:

[0053]

[0054] v 1i =(α×v i ) / (α - L×v i )

[0055]

[0056] v i ∈V∈R 1×L

[0057] v 1i ∈ V1 ∈ R 1×L

[0058] v 2i ∈ V2 ∈ R 1×L

[0059] where V is the product sentiment tendency significantly aggregated representation vector, L is the number of eigenvalues in the product sentiment tendency significantly aggregated representation vector, R is the real number field, v i is the eigenvalue at the i-th position in the product sentiment tendency significantly aggregated representation vector, α is the first product sentiment tendency significantly aggregated representation space structure value, β is the second product sentiment tendency significantly aggregated representation space structure value, v 1i represents the first product sentiment tendency significantly aggregated representation transformation adjustment value at the i-th position in the first product sentiment tendency significantly aggregated representation transformation adjustment vector, v 2i is the second product sentiment tendency significantly aggregated representation transformation adjustment value at the i-th position in the second product sentiment tendency significantly aggregated representation transformation adjustment vector, V1 is the first product sentiment tendency significantly aggregated representation transformation adjustment vector, represents pointwise addition by position, ω represents the weighted hyperparameter, ⊙ represents pointwise multiplication by position, V2 represents the second product sentiment tendency significantly aggregated representation transformation adjustment vector, and V′ represents the optimized product sentiment tendency significantly aggregated representation vector.

[0060] Specifically, in the preferred example, for the spatial structure information of the feature set of the product sentiment tendency significantly aggregated representation vector in the high-dimensional space, by using the class norm space structured representation of the product sentiment tendency significantly aggregated representation vector as a reference window, a scale-based box transformation is performed on each eigenvalue of the product sentiment tendency significantly aggregated representation vector, and a box attention weight adjustment based on the spatial structure of each eigenvalue of the product sentiment tendency significantly aggregated representation vector is realized to ensure the spatial transformation (translation, scaling, and rotation) invariance of the product sentiment tendency significantly aggregated representation vector under the interaction of the feature space, thereby improving the convergence and generalization effects of the classification and regression of the feature set of the product sentiment tendency significantly aggregated representation vector under the complex spatial structure representation, improving the accuracy of the analysis result obtained by passing it through the product public opinion real-time analysis module based on the classifier. Compared with the traditional keyword matching or simple statistical methods, it can be automatically adjusted step by step as new data flows in to better adapt to the changes in the user's expression habits and language environment, and at the same time can better understand the semantic information and context dependence relationship in the text, thereby significantly improving the accuracy and generalization ability of product sentiment analysis.

[0061] In summary, the real-time precise analysis and processing system 100 of big data information based on the embodiments of the present application is elucidated. It uses natural language processing technology based on deep learning to preprocess each evaluation data in the set of evaluation data of the target product, introduces the reconstruction processing and semantic encoding of the product public opinion question template questions, and performs product sentiment tendency reasoning on each reconstructed evaluation data. Based on this, it automatically evaluates whether the product public opinion is positive or negative according to the semantic significant aggregation representation between the semantic inference coding features of each product sentiment tendency obtained by the reasoning, and can be automatically adjusted step by step as new data flows in to better adapt to the changes in user expression habits and language environments. At the same time, it can better understand the semantic information and context dependence relationship in the text, thus significantly improving the accuracy and generalization ability of product sentiment analysis.

[0062] As described above, the real-time precise analysis and processing system 100 of big data information based on the embodiments of the present application can be implemented in various terminal devices, such as servers for real-time precise analysis and processing of big data information. In one example, the real-time precise analysis and processing system 100 of big data information based on the embodiments of the present application can be integrated into the terminal device as a software module and / or a hardware module. For example, the real-time precise analysis and processing system 100 of big data information can be a software module in the operating system of the terminal device, or can be an application program developed for the terminal device; of course, the real-time precise analysis and processing system 100 of big data information can also be one of the many hardware modules of the terminal device.

[0063] Alternatively, in another example, the real-time precise analysis and processing system 100 of big data information and the terminal device can also be separate devices, and the real-time precise analysis and processing system 100 of big data information can be connected to the terminal device through a wired and / or wireless network and transmit interaction information in accordance with a predefined data format.

Claims

1. A real-time precise analysis and processing system for big data information in the background, characterized in that, Including: An evaluation data acquisition and processing module, configured to collect a set of evaluation data of a target product and perform preprocessing to obtain a set of preprocessed evaluation data; An evaluation data reconstruction module, configured to add product public opinion question template questions to the tails of the preprocessed evaluation data in the set of preprocessed evaluation data to obtain a set of reconstructed evaluation data; A product sentiment tendency reasoning module, configured to perform semantic encoding on the set of reconstructed evaluation data to obtain a set of reconstructed evaluation data semantic encoding vectors, and then perform product sentiment tendency reasoning to obtain a set of product sentiment tendency semantic reasoning encoding vectors; A product sentiment tendency aggregation module, configured to perform product sentiment tendency semantic feature aggregation on the set of product sentiment tendency semantic reasoning encoding vectors to obtain a significant aggregation representation of product sentiment tendency; Wherein, the product sentiment tendency aggregation module includes: a product static energy factor calculation unit, configured to calculate the static energy factors of the product sentiment tendency semantic reasoning encoding vectors in the set of product sentiment tendency semantic reasoning encoding vectors to obtain a set of product sentiment tendency semantic static energy factors; an energy distribution significant aggregation unit, configured to perform significant aggregation based on energy distribution on the set of product sentiment tendency semantic reasoning encoding vectors based on the set of product sentiment tendency semantic static energy factors to obtain the significant aggregation representation of product sentiment tendency; An analysis result generation module, configured to obtain an analysis result based on the significant aggregation representation of product sentiment tendency; Wherein, the product static energy factor calculation unit includes: A product sentiment tendency semantic reasoning mean and variance calculation sub-unit, configured to calculate the mean and variance of the product sentiment tendency semantic reasoning encoding vectors to obtain a product sentiment tendency semantic reasoning mean and a product sentiment tendency semantic reasoning variance; A product sentiment tendency semantic reasoning difference modulation vector generation sub-unit, configured to subtract the product sentiment tendency semantic reasoning encoding vectors from the product sentiment tendency semantic reasoning mean in a position-by-position manner, and calculate the fourth power of each position of the subtracted feature vector to obtain a product sentiment tendency semantic reasoning difference modulation vector; A product sentiment tendency semantic reasoning expected value calculation sub-unit, configured to calculate the expected value of the product sentiment tendency semantic reasoning difference modulation vector to obtain a product sentiment tendency semantic reasoning expected value; A product sentiment tendency semantic static energy factor generation sub-unit, configured to divide the product sentiment tendency semantic reasoning expected value by the square of the product sentiment tendency semantic reasoning variance, and input the obtained value into a sigmoid function to obtain the product sentiment tendency semantic static energy factor.

2. The real-time precise analysis and processing system for big data information background according to claim 1, wherein The evaluation data acquisition and processing module is configured to: Collect a set of evaluation data of the target product; Perform preprocessing on each evaluation data in the set of evaluation data to obtain the set of preprocessed evaluation data.

3. The real-time precise analysis and processing system for big data information background according to claim 2, characterized in that, The product sentiment tendency reasoning module includes: A reconstructed evaluation data semantic encoding unit, configured to perform semantic encoding on each reconstructed evaluation data in the set of reconstructed evaluation data to obtain the set of reconstructed evaluation data semantic encoding vectors; The product sentiment semantic feature encoding unit is used to input each reconstructed evaluation data semantic encoding vector in the set of reconstructed evaluation data semantic encoding vectors into the product sentiment inference engine based on the RNN model to obtain the set of product sentiment semantic inference encoding vectors.

4. The real-time precise analysis and processing system for big data information background according to claim 3, wherein The energy distribution significant aggregation unit includes: The product sentiment semantic inference clustering initial center vector determination subunit is used to select the product sentiment semantic inference encoding vector corresponding to the largest product sentiment semantic static energy factor in the set of product sentiment semantic static energy factors as the product sentiment semantic inference clustering initial center vector; The product sentiment dynamic aggregation energy factor calculation subunit is used to calculate the dynamic aggregation energy factor of each product sentiment semantic inference encoding vector in the set of product sentiment semantic inference encoding vectors based on the spatial span between each product sentiment semantic inference encoding vector in the set of product sentiment semantic inference encoding vectors and the product sentiment semantic inference clustering initial center vector, as well as the static energy factor of each product sentiment semantic inference encoding vector and the static energy factor of the product sentiment semantic inference clustering initial center vector, to obtain the set of product sentiment semantic inference dynamic aggregation energy factors; The dynamic aggregation energy factor gating mask subunit is used to input the set of product sentiment semantic inference dynamic aggregation energy factors into the gating mask unit to obtain the set of product sentiment semantic inference dynamic aggregation weight factors; The product sentiment significant aggregation representation generation subunit is used to calculate the weighted sum of the set of product sentiment semantic inference encoding vectors with the set of product sentiment semantic inference dynamic aggregation weight factors to obtain the product sentiment semantic inference encoding vector as the product sentiment significant aggregation representation.

5. The real-time precise analysis and processing system for big data information background according to claim 4, characterized in that, The product sentiment dynamic aggregation energy factor calculation subunit is used to: Multiply the static energy factor of the product sentiment semantic inference encoding vector and the static energy factor of the product sentiment semantic inference clustering initial center vector by the first weighting parameter to obtain the first product sentiment semantic inference dynamic aggregation energy factor; Multiply the square of the spatial span between the product sentiment semantic inference encoding vector and the product sentiment semantic inference clustering initial center vector by the second weighting parameter to obtain the second product sentiment semantic inference dynamic aggregation energy factor; Divide the first product sentiment semantic inference dynamic aggregation energy factor by the second product sentiment semantic inference dynamic aggregation energy factor to obtain the product sentiment semantic inference dynamic aggregation energy factor.

6. The real-time precise analysis and processing system for big data information background according to claim 5, characterized in that, The dynamic aggregation energy factor gating mask subunit includes: The product sentiment dynamic aggregation energy factor normalization secondary subunit is used to perform normalization processing on the set of product sentiment semantic inference dynamic aggregation energy factors to obtain the set of normalized product sentiment semantic inference dynamic aggregation energy factors; The normalized product sentiment dynamic aggregation energy factor masking secondary subunit is used to perform masking processing on the set of normalized product sentiment tendency semantic inference dynamic aggregation energy factors to obtain the set of product sentiment tendency semantic inference dynamic aggregation weight factors.

7. The real-time precise analysis and processing system for big data information background according to claim 6, characterized in that, The product sentiment dynamic aggregation energy factor normalization secondary subunit is used to: perform normalization processing on the set of product sentiment tendency semantic inference dynamic aggregation energy factors by using the sigmoid function to obtain the set of normalized product sentiment tendency semantic inference dynamic aggregation energy factors.

8. The real-time precise analysis and processing system for big data information background according to claim 7, characterized in that The normalized product sentiment dynamic aggregation energy factor masking secondary subunit is used to: in response to each normalized product sentiment tendency semantic inference dynamic aggregation energy factor in the set of normalized product sentiment tendency semantic inference dynamic aggregation energy factors being greater than a predetermined threshold, set the normalized product sentiment tendency semantic inference dynamic aggregation energy factor to its original value and set the rest to zero to obtain the set of product sentiment tendency semantic inference dynamic aggregation weight factors.

9. The real-time precise analysis and processing system for big data information background according to claim 8, characterized in that, The analysis result generation module is used to: input the product sentiment tendency significant aggregation representation vector into the product public opinion real-time analysis module based on a classifier to obtain the analysis result, and the analysis result is used to indicate that the product public opinion is positive or the product public opinion is negative.

Citation Information

Patent Citations

  • Fine tuning method and device for a language model, computing equipment and storage medium

    CN113468877A

  • Computer data processing method and system based on big data

    CN118134529A