Efficient text sentiment analysis method and system based on Transform architecture

By introducing data preprocessing, adversarial attack detection and defense mechanism units into the Transformer architecture, the anti-interference index, adversarial attack index and defense mechanism index are calculated, and the problem of the Transformer architecture being vulnerable to adversarial attacks in text sentiment analysis is solved, achieving more accurate and reliable sentiment analysis results.

CN120124618AInactive Publication Date: 2025-06-10ZHONGKE XINCHEN INTELLIGENT TECHNOLOGY (HUNAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510202735.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-10
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing Transformer architecture is vulnerable to confrontational attacks in text sentiment analysis, resulting in inaccurate sentiment analysis output.

Method used

An efficient text sentiment analysis method based on Transformer architecture is designed, including emotion data collection, feature extraction, emotion calculation, evaluation and execution modules, and through data preprocessing, adversarial attack detection and defense mechanism units, anti-interference index, adversarial attack index and defense mechanism index are calculated to improve the reliability of the analysis.

Benefits of technology

Effectively detect and defend against adversarial attacks, improve the accuracy and reliability of text sentiment analysis and ensure the stability of sentiment analysis results in practical applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124618A_ABST
    Figure CN120124618A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing, and discloses an efficient text sentiment analysis method and system based on a Transform architecture, by establishing a sentiment data collection module, a Transform architecture recognition data module, a Transform architecture sentiment calculation module, a sentiment data evaluation module and a sentiment data execution module, data are collected in the sentiment data collection module, and the sentiment data are extracted from the sentiment data evaluation module; the method comprises the following steps that: a Transform architecture data identification module performs feature extraction and classification according to collected data, a Transform architecture model is used for judging the emotion tendency of a text, whether the emotion is positive, negative or neutral emotion is judged, a Transform architecture emotion calculation module performs data calculation, and an emotion data evaluation module evaluates the emotion tendency of the text according to a calculation result in the Transform architecture emotion calculation module. And the emotion data execution module evaluates the Transform architecture and provides a modification scheme, the emotion data execution module performs final review on the received emotion demand according to the modification scheme provided by the emotion data evaluation module, and in the review process, the accuracy and integrity of the data are rechecked.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to an efficient text sentiment analysis method and system based on a Transformer architecture. Background Art

[0002] In the field of natural language processing, the Transformer architecture mainly consists of two parts: the encoder and the decoder. In natural language processing tasks, the encoder is responsible for feature extraction and encoding of the input sequence, and the decoder generates the target sequence based on the encoder's output. Compared with the traditional recurrent neural network (RNN) and its variants, the long short-term memory network (LSTM) and the gated recurrent unit (GRU), the Transformer can process sequence data in parallel, greatly improving the speed of training and reasoning, and is particularly suitable for large-scale data processing. Although the application of the Transformer architecture has achieved results, it also has some defects. One of them is that it is vulnerable to adversarial attacks, that is, the model is sensitive to slight modifications to the input text. Attackers can incorporate subtle noise or disturbances into the original input. Although these changes are difficult to be detected by the human eye, they can cause errors in the model's sentiment analysis output, thereby affecting the accuracy and reliability of sentiment analysis. This has become one of the important issues that need to be solved in this technical field. Summary of the invention

[0003] 1. Technical issues to be resolved

[0004] In view of the shortcomings of the prior art, the present invention provides an efficient text sentiment analysis method and system based on the Transformer architecture, which has the advantages of adversarial attack detection, defense mechanism and data preprocessing, and solves the problem that the Transformer architecture in the existing system is vulnerable to adversarial attacks.

[0005] (II) Technical solution

[0006] To achieve the above object, the present invention provides the following technical solution: an efficient text sentiment analysis method based on Transformer architecture, comprising the following steps:

[0007] Step 1: Establish an emotion data collection module, a Transformer architecture recognition data module, a Transformer architecture emotion calculation module, an emotion data evaluation module, and an emotion data execution module;

[0008] Step 2: Collect the original emotional demand data in the emotional data collection module, and pre-process and number the data to provide a basis for subsequent analysis and processing;

[0009] Step 3: The Transformer architecture recognition data module identifies the original emotional demand data collected by the emotional data collection module, and promptly provides emotional analysis. It extracts features and classifies the collected data, and uses the model of the Transformer architecture to judge the emotional tendency of the text, determining whether it is a positive, negative, or neutral emotion;

[0010] Step 4: The Transformer architecture emotion calculation module receives the original emotional demand data from the emotional data collection module and the preliminary judgment results of the text features and emotional tendencies extracted by the Transformer architecture recognition data module, and performs data calculations;

[0011] Step 5: The emotional data evaluation module makes an evaluation of the Transformer architecture based on the calculation results in the Transformer architecture emotion calculation module, and proposes a modification plan;

[0012] Step 6: The emotional data execution module performs a final review of the received emotional demands according to the modification plan provided by the emotional data evaluation module. During the review process, the accuracy and integrity of the data are checked again.

[0013] Preferably, an efficient text emotion analysis system based on the Transformer architecture, the system includes an emotional data collection module, a Transformer architecture recognition data module, a Transformer architecture emotion calculation module, an emotional data evaluation module, and an emotional data execution module.

[0014] Preferably, the emotional data collection module includes a text data collection unit, an image data collection unit, and a domain-specific data collection unit. The text data collection unit obtains the original text data set in real time through sources such as social media and comment systems. The image data collection unit is used to obtain an image data set related to the text. The domain-specific data collection unit obtains a domain-specific emotional data set through information in the medical and financial fields. After numbering the internal data of the text data collection unit, the image data collection unit, and the domain-specific data collection unit, they are connected to the Transformer architecture recognition data module and the Transformer architecture emotion calculation module through a network interface.

[0015] Preferably, the original text data set is numbered The image data collection unit numbers the noise level during the image acquisition process according to the characteristics of the image data set related to the text. The noise level number during the image acquisition process is Y 1 、Y 2 、Y 3 、…Y v, the domain-specific data collection unit numbers the data according to the characteristics of the specific sentiment dataset, and the specific sentiment dataset number is

[0016] Preferably, the Transformer architecture recognition data module records the sentiment tendency score according to the preliminary judgment result of the sentiment tendency and numbers the data. The sentiment tendency score number is E 1 、E 2 、E 3 、…E v 。

[0017] Preferably, the Transformer architecture sentiment calculation module includes a data preprocessing unit, an adversarial attack detection unit, and a defense mechanism unit.

[0018] Preferably, the data preprocessing unit calculates the anti-interference index Kz of the preprocessing process according to the original text data, and its calculation formula is:

[0019]

[0020] In the formula, Kz represents the anti-interference index of the preprocessing process, represents the original text data set, p represents the anti-interference index of the original text data set in the system, n 1 represents the number of noise samples detected in the original text data set in the system, n represents the total number of samples in the original text data set in the system, m 1 represents the number of noise feature dimensions detected in the original text data set in the system, m represents the total number of feature dimensions in the original text data set in the system, a and b respectively represent the weights of the number of noise samples and the number of noise feature dimensions in the anti-interference index, represents the proportion of non-noise samples in the original text data set, k 1 represents the coefficient of environmental interference of non-noise samples estimated by the system in the original text data set, represents the proportion of non-noise feature dimensions in the original text data set, k 2 represents the coefficient of spatial dimension interference of non-noise feature dimensions estimated by the system in the original text data set.

[0021] Preferably, the adversarial attack detection unit calculates the Transformer architecture adversarial attack index Gz according to the image dataset related to the text and the preliminary judgment data of the sentiment tendency, and its calculation formula is:

[0022]

[0023] In the formula, Gz represents the Transformer architecture adversarial attack index, E 1 、E2 , E 3 , …E v represents the sentiment tendency score, E i represents the sentiment tendency score of the i-th word, obtained by preliminarily judging the data according to the sentiment tendency, Y 1 , Y 2 , Y 3 , …Y v represents the noise level during the image acquisition process, Y i represents the noise level during the i-th image acquisition process, which can be estimated by statistically calculating the coefficient of variation of the image in texts of different sentiment categories, F i represents the attention weight of the i-th image, indicating the importance of the word in the text. v represents the number of acquired images, and μ is a constant minimum value.

[0024] Preferably, the defense mechanism unit calculates the defense mechanism index Wz according to a specific sentiment dataset, and its calculation formula is:

[0025] Wz = T(i) + ΔT

[0026] In the formula, Wz represents the defense mechanism index, represents the specific sentiment dataset, T(i) represents the i-th specific sentiment dataset processed by the Transformer architecture for the specific sentiment dataset, and ΔT represents the correction value calculated according to the type and degree of the adversarial attack of the original Transformer architecture in the system. The new data after defense is obtained through this formula.

[0027] Preferably, the sentiment data evaluation module makes evaluations on the Transformer architecture respectively according to the anti-interference index Kz, the adversarial attack index Gz of the Transformer architecture, and the defense mechanism index Wz during the preprocessing process, and proposes corresponding modification schemes.

[0028] Compared with the prior art, the present invention provides an efficient text sentiment analysis method and system based on the Transformer architecture, having the following beneficial effects:

[0029] 1. The present invention evaluates through the emotional data evaluation module based on the numerical magnitudes of the anti-interference index Kz, the adversarial attack index Gz of the Transformer architecture, and the defense mechanism index Wz during the preprocessing process, and real-time monitors the magnitudes of the calculated values. When the module identifies abnormal values, by pre-correcting the parameters of the preprocessing process and adjusting the parameters of the Transformer architecture, the system can continuously improve the efficient text sentiment analysis performance based on the Transformer architecture while ensuring the reliability of the sentiment analysis recognized by the system in practical applications, solving the problem that the Transformer architecture in the existing system is vulnerable to adversarial attacks. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0031] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0032] Please refer to Figure 1 , an efficient text sentiment analysis method based on the Transformer architecture, including the following steps:

[0033] Step 1: Establish an emotional data collection module, a Transformer architecture recognition data module, a Transformer architecture sentiment calculation module, an emotional data evaluation module, and an emotional data execution module;

[0034] Step 2: Collect the original emotional requirement data in the emotional data collection module, preprocess and number the data, providing a basis for subsequent analysis and processing;

[0035] Step 3: The Transformer architecture recognition data module identifies the original emotional requirement data collected in the emotional data collection module and provides timely sentiment analysis. Feature extraction and classification are performed according to the collected data, and the model of the Transformer architecture is used to judge the sentiment tendency of the text, judging whether it is a positive, negative, or neutral sentiment;

[0036] Step 4: The Transformer architecture sentiment calculation module receives the original emotional requirement data from the emotional data collection module and the preliminary judgment results of the text features and sentiment tendencies extracted from the Transformer architecture recognition data module, and performs data calculation;

[0037] Step 5: The Emotion Data Evaluation Module evaluates the Transformer architecture based on the results calculated by the Transformer Architecture Emotion Calculation Module and proposes a modification plan.

[0038] Step 6: The Emotion Data Execution Module performs a final review of the received emotion requirements according to the modification plan provided by the Emotion Data Evaluation Module. During the review process, the accuracy and integrity of the data are checked again to prevent omissions and ensure the reliability of the final emotion analysis result.

[0039] An efficient text emotion analysis system based on the Transformer architecture, which includes an Emotion Data Collection Module, a Transformer Architecture Recognition Data Module, a Transformer Architecture Emotion Calculation Module, an Emotion Data Evaluation Module, and an Emotion Data Execution Module.

[0040] The Emotion Data Collection Module includes a Text Data Collection Unit, an Image Data Collection Unit, and a Domain-Specific Data Collection Unit. The Text Data Collection Unit obtains the original text data set in real time through sources such as social media and comment systems. The Image Data Collection Unit is used to obtain the image data set related to the text (such as emojis or pictures, which can enhance the accuracy of emotion analysis). The Domain-Specific Data Collection Unit obtains the domain-specific emotion data set through medical and financial domain information to ensure that the model can understand the specific emotion expressions within the domain. After numbering the internal data of the Text Data Collection Unit, the Image Data Collection Unit, and the Domain-Specific Data Collection Unit, they are connected to the Transformer Architecture Recognition Data Module and the Transformer Architecture Emotion Calculation Module through the network interface.

[0041] The original text data set is numbered The Text Data Collection Unit numbers the number of noise samples detected in the original text data set in the system, the total number of samples in the original text data set in the system, the number of noise feature dimensions detected in the original text data set in the system, and the total number of feature dimensions in the original text data set in the system according to the characteristics of the original text data set. The number of noise samples detected in the original text data set in the system, the total number of samples in the original text data set in the system, the number of noise feature dimensions detected in the original text data set in the system, and the total number of feature dimensions in the original text data set in the system are numbered as n 1 , n, m 1 , m. The Image Data Collection Unit numbers the noise level during the image acquisition process according to the characteristics of the image data set related to the text. The noise level during the image acquisition process is numbered as Y 1 , Y 2 , Y 3 , …Yv , the domain-specific data collection unit numbers the data according to the characteristics of the specific sentiment dataset, and the specific sentiment dataset number is

[0042] The Transformer architecture recognition data module records the sentiment tendency score based on the preliminary judgment result of the sentiment tendency and numbers the data. The sentiment tendency score number is E 1 , E 2 , E 3 , …E v .

[0043] The Transformer architecture sentiment calculation module includes a data preprocessing unit, an adversarial attack detection unit, and a defense mechanism unit.

[0044] The data preprocessing unit calculates the anti-interference index Kz of the preprocessing process according to the original text data. The calculation formula is:

[0045]

[0046] In the formula, Kz represents the anti-interference index of the preprocessing process, represents the original text dataset, p represents the anti-interference index of the original text dataset in the system, n 1 represents the number of noise samples detected in the original text dataset in the system, n represents the total number of samples in the original text dataset in the system, m 1 represents the number of noise feature dimensions detected in the original text dataset in the system, m represents the total number of feature dimensions in the original text dataset in the system, a and b respectively represent the weights of the number of noise samples and the number of noise feature dimensions in the anti-interference index, represents the proportion of non-noise samples in the original text dataset, k 1 represents the coefficient of environmental interference of non-noise samples estimated by the system in the original text dataset, represents the proportion of non-noise feature dimensions in the original text dataset, k 2 represents the coefficient of spatial dimension interference of non-noise feature dimensions estimated by the system in the original text dataset.

[0047] The advantages are: by calculating the anti-interference index Kz of the preprocessing process, the sentiment data evaluation module evaluates according to the value of the anti-interference index Kz of the preprocessing process. When the anti-interference index Kz of the preprocessing process is high, it indicates that the data preprocessing effect is good, the data purity and robustness are high, and the processing efficiency is improved by reducing the complexity of the preprocessing steps. When the anti-interference index Kz of the preprocessing process is low, it indicates that there is more noise in the data, and noise detection and filtering algorithms need to be added to improve the data quality.

[0048] The adversarial attack detection unit preliminarily judges the data based on the image dataset related to the text and the sentiment tendency, and calculates the adversarial attack index Gz of the Transformer architecture. Its calculation formula is:

[0049]

[0050] In the formula, Gz represents the adversarial attack index of the Transformer architecture, and E 1 、E 2 、E 3 、…E v represents the sentiment tendency score. E i represents the sentiment tendency score of the i-th word, which is obtained based on the data of the preliminary judgment of the sentiment tendency. Y 1 、Y 2 、Y 3 、…Y v represents the noise level during the image acquisition process. Y i represents the noise level during the i-th image acquisition process, which can be estimated by statistically calculating the coefficient of variation of the image in texts of different sentiment categories, etc. F i represents the attention weight of the i-th image, indicating the importance of the word in the text. v represents the number of acquired images, and μ is a constant minimum value to prevent the denominator from being zero.

[0051] The advantages are as follows: By calculating the adversarial attack index Gz of the Transformer architecture, the sentiment data evaluation module evaluates according to the value of the adversarial attack index Gz of the Transformer architecture. When the adversarial attack index Gz of the Transformer architecture is high, it indicates that the model is more sensitive to adversarial attacks, and more adversarial training samples need to be introduced to enhance the defense mechanism, thereby optimizing the robustness of the model. When the adversarial attack index Gz of the Transformer architecture is low, it indicates that the model has a strong defense ability against adversarial attacks. Keep the current model structure, but the adversarial attack detection algorithm needs to be checked and updated regularly.

[0052] The defense mechanism unit calculates the defense mechanism index Wz according to the specific sentiment dataset. Its calculation formula is:

[0053] Wz = T(i) + ΔT

[0054] In the formula, Wz represents the defense mechanism index, represents the specific sentiment dataset, T(i) represents the i-th specific sentiment dataset processed by the Transformer architecture of the specific sentiment dataset, and ΔT represents the correction value calculated from the type and degree of the original adversarial attack of the Transformer architecture in the system. The new data after defense is obtained through this formula.

[0055] The advantages are as follows: By calculating the defense mechanism index Wz, the emotional data evaluation module evaluates according to the value of the defense mechanism index Wz. When the defense mechanism index Wz is high, it indicates that the defense mechanism is effective and can better correct the impact of adversarial attacks. When the calculated defense mechanism index Wz is low, it indicates that the defense mechanism has poor effect, and it is necessary to adjust the model structure and redesign the defense strategy to improve the prevention ability.

[0056] The emotional data evaluation module evaluates the Transformer architecture according to the anti-interference index Kz in the preprocessing process, the adversarial attack index Gz of the Transformer architecture, and the defense mechanism index Wz, and proposes corresponding modification schemes.

[0057] The advantages are as follows: The emotional data evaluation module evaluates according to the values of the anti-interference index Kz in the preprocessing process, the adversarial attack index Gz of the Transformer architecture, and the defense mechanism index Wz, and monitors the calculated values in real time. When the module identifies abnormal values, by pre-correcting the parameters of the preprocessing process and adjusting the parameters of the Transformer architecture, the system can continuously improve the efficient text sentiment analysis performance based on the Transformer architecture while ensuring the reliability of the sentiment analysis identified by the system in practical applications.

[0058] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An efficient text sentiment analysis method based on Transformer architecture, characterized in that: The following steps are involved: Step 1: Establish an emotion data collection module, a Transformer architecture recognition data module, a Transformer architecture emotion calculation module, an emotion data evaluation module, and an emotion data execution module; Step 2: Collect the original emotional demand data in the emotional data collection module, and pre-process and number the data to provide a basis for subsequent analysis and processing; Step 3: The Transformer architecture recognition data module identifies the original emotional demand data collected in the emotional data collection module and provides emotional analysis in a timely manner. It extracts and classifies features based on the collected data and uses the Transformer architecture model to judge the emotional tendency of the text, whether it is positive, negative or neutral. Step 4: The Transformer architecture emotion calculation module receives the original emotion demand data from the emotion data collection module and the text features and preliminary judgment results of emotion tendency extracted from the Transformer architecture recognition data module, and performs data calculation; Step 5: The sentiment data evaluation module evaluates the Transformer architecture based on the results calculated in the Transformer architecture sentiment calculation module and proposes a modification plan; Step 6: The emotional data execution module conducts a final review of the received emotional requirements based on the modification plan provided by the emotional data evaluation module. During the review process, the accuracy and completeness of the data are checked again.

2. An efficient text sentiment analysis system based on Transformer architecture, characterized by: The system includes an emotion data collection module, a Transformer architecture recognition data module, a Transformer architecture emotion calculation module, an emotion data evaluation module and an emotion data execution module.

3. The efficient text sentiment analysis system based on Transformer architecture according to claim 2, characterized in that: The emotion data collection module includes a text data collection unit, an image data collection unit and a domain-specific data collection unit. The text data collection unit acquires the original text data set in real time through the sources of social media and comment systems. The image data collection unit is used to acquire image data sets related to the text. The domain-specific data collection unit acquires domain-specific emotion data sets through medical and financial field information. The text data collection unit, the image data collection unit and the domain-specific data collection unit number the internal data of the units and then connect to the Transformer architecture recognition data module and the Transformer architecture emotion calculation module through a network interface.

4. The efficient text sentiment analysis system based on Transformer architecture according to claim 2, characterized in that: The original text dataset is numbered as The image data collection unit performs data numbering on the noise level in the image acquisition process according to the text-related image data set features, and the noise level numbering in the image acquisition process is Y1, Y2, Y3, ...Y v The domain-specific data collection unit performs data numbering according to the characteristics of the specific emotion data set, and the specific emotion data set numbering is 5. The efficient text sentiment analysis system based on Transformer architecture according to claim 2, characterized in that: The Transformer architecture recognition data module records the sentiment tendency score according to the preliminary judgment result of sentiment tendency and performs data numbering. The sentiment tendency score numbering is E1, E2, E3, ... E v .

6. The efficient text sentiment analysis system based on Transformer architecture according to claim 2, characterized in that: The Transformer architecture emotion computing module includes a data preprocessing unit, an adversarial attack detection unit and a defense mechanism unit.

7. The efficient text sentiment analysis system based on Transformer architecture according to claim 6, characterized in that: The data preprocessing unit calculates the anti-interference index Kz of the preprocessing process according to the original text data, and the calculation formula is: In the formula, Kz represents the anti-interference index of the pretreatment process, represents the original text dataset, p represents the anti-interference index of the original text dataset in the system, n1 represents the number of noise samples detected in the original text dataset in the system, n represents the total number of samples in the original text dataset in the system, m1 represents the number of noise feature dimensions detected in the original text dataset in the system, m represents the total number of feature dimensions in the original text dataset in the system, a and b represent the weights of the number of noise samples and the noise feature dimension in the anti-interference index respectively. represents the proportion of non-noise samples in the original text dataset, k1 represents the coefficient of environmental interference of non-noise samples in the original text dataset estimated by the system, It represents the proportion of non-noise feature dimensions in the original text dataset, and k2 represents the coefficient of the non-noise feature dimensions in the original text dataset estimated by the system to be disturbed by the spatial dimension.

8. The efficient text sentiment analysis system based on Transformer architecture according to claim 6, characterized in that: The adversarial attack detection unit calculates the Transformer architecture adversarial attack index Gz based on the text-related image data set and the preliminary judgment data of emotional tendency, and its calculation formula is: In the formula, Gz represents the Transformer architecture anti-attack index, E1, E2, E3, …E v represents the sentiment tendency score, E i represents the sentiment tendency score of the i-th word, which is obtained based on the preliminary judgment data of sentiment tendency, Y1, Y2, Y3, ...Y v represents the noise level during image acquisition, Y i represents the noise level during the acquisition of the i-th image, which can be estimated by counting the coefficient of variation of the image in different sentiment categories, etc. i represents the attention weight of the i-th image, indicating the importance of the word in the text, v represents the number of images obtained, and μ is a constant minimum value.

9. The efficient text sentiment analysis system based on Transformer architecture according to claim 6, characterized in that: The defense mechanism unit calculates the defense mechanism index Wz according to the specific emotion data set, and the calculation formula is: Wz=T(i)+ΔT In the formula, Wz represents the defense mechanism index, Represents a specific emotion dataset, T(i) represents the i-th specific emotion dataset processed by the Transformer architecture, ΔT represents the correction value calculated by the type and degree of the original Transformer architecture in the system to counter the attack, and the new data after defense is obtained through this formula.

10. The efficient text sentiment analysis system based on Transformer architecture according to claim 9, characterized in that: The sentiment data evaluation module evaluates the Transformer architecture according to the anti-interference index Kz of the preprocessing process, the Transformer architecture anti-attack index Gz and the defense mechanism index Wz, and proposes corresponding modification plans.