Deep Learning-Based Text Semantic Analysis System

By dividing text information into two categories: simple and complex, and adopting corresponding fusion strategies to dynamically adjust the modal weight, the semantic conflict problem in multi-modal situations is solved, and the accuracy and robustness of emotion analysis and emotion recognition are improved.

CN119721046BActive Publication Date: 2025-07-18NANJING CAIGUO TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411771041.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-04
Publication Date
2025-07-18
Estimated Expiration
2044-12-04

AI Technical Summary

Technical Problem

The prior art cannot effectively handle semantic conflicts between text information in multimodal situations, resulting in a decrease in the accuracy of sentiment analysis and emotion recognition.

Method used

By dividing text information into two categories: simple and complex, combining different fusion strategies, a fixed fusion strategy is used to process simple text information, a dynamic weighted fusion strategy is used to process complex text information, and a dynamic adjustment of modal weights is used to resolve conflicts.

Benefits of technology

It improves the accuracy of emotion analysis and emotion recognition, reduces the risk of misjudgment, and enhances the robustness and adaptability of the model in complex scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119721046B_ABST
    Figure CN119721046B_ABST
Patent Text Reader

Abstract

The present invention discloses a text semantic analysis system based on deep learning, which relates to the technical field of text semantic analysis, and includes a semantic extraction module, a feature extraction and evaluation module, a text complexity division module, a fixed fusion strategy module, and a dynamic weighted fusion module: The semantic extraction module first extracts semantic information from the generated text data, mines the semantic relationships and structural information in the text, and captures the key themes, emotional tendencies, and intentions in the text. By dividing the text into simple and complex categories and combining different fusion strategies, the system adopts fixed fusion when the information is consistent to efficiently integrate multi-modal information; in conflict scenarios, dynamic weighted fusion is introduced to accurately focus on the most relevant modality and avoid semantic conflicts from affecting the accuracy of sentiment analysis. The dynamic weighted fusion module enables the model to adaptively adjust the modality weights, enhancing the robustness and adaptability of the system in complex multi-modal scenarios, and is applicable to applications such as emotion recognition and sentiment analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text semantic analysis, and particularly to a text semantic analysis system based on deep learning. Background Art

[0002] Text semantic analysis based on deep learning is a method of using deep neural networks to understand and process text data. By analyzing semantic information at levels such as words, sentences, and paragraphs, it achieves a deeper level of language understanding. Compared with traditional analysis that relies on keywords or word frequencies, deep learning automatically learns semantic patterns in text through training with large corpora, enhancing the ability to understand implicit meanings. It can be used in tasks such as sentiment analysis, automatic abstract generation, information retrieval, and machine translation. In addition, in the processing of multi-modal data sources (such as images, audio, and video), deep learning can also fuse the features of different types of data with text semantics through multi-modal learning to achieve cross-modal understanding. For example, in video analysis, it associates audio and image information, or in image processing, it generates text descriptions, thereby enhancing the machine's overall perception ability of multi-source data.

[0003] Text semantic analysis based on deep learning involves multi-modal semantic fusion, especially in scenarios where multiple information sources (such as text, images, audio, video, etc.) need to be integrated to achieve a more comprehensive semantic understanding. Multi-modal semantic fusion refers to integrating data from different modalities and jointly analyzing them through deep learning models to capture richer semantic relationships. For example, in sentiment analysis, not only the sentiment of the text is analyzed, but also the expressions in the image or the intonation in the audio are combined to generate a more accurate sentiment judgment. Multi-modal semantic fusion enhances the depth and breadth of semantic analysis, enabling the model to identify multi-dimensional information in complex environments and improving the understanding accuracy.

[0004] The existing technologies have the following deficiencies:

[0005] The prior art usually integrates and collaboratively analyzes the text information (such as text, images, audio, video, etc.) extracted from different types of data sources through a single fusion strategy, so as to comprehensively understand the semantics from multiple perspectives. This approach is applicable in most scenarios, but when the text information of different modalities conflicts with each other, a single fusion strategy may not be able to effectively handle the conflict. For example, in a multimodal context, the news headline is "The local park welcomes a peak of tourists and the environment is becoming more beautiful", conveying a positive emotion; the audio commentary also describes the scene as an "active" community life in a peaceful tone, further strengthening the positive atmosphere. However, the accompanying picture shows a large amount of garbage and a chaotic scene, suggesting a negative emotion about the environmental problem. This semantic conflict may cause the model to misunderstand in emotion judgment, because the positive information in the text and audio masks the negative hint in the picture. If the fixed fusion strategy fails to recognize and adjust this modal conflict, the model may make a judgment based on a wrong assumption, resulting in a semantic understanding deviation. This misjudgment will make the model draw inaccurate conclusions in sentiment analysis, such as misjudging as neutral or mixed emotions, ignoring the emotional turns or background information in the text and images, directly affecting the accuracy of the task, especially in applications such as sentiment analysis, emotion recognition, and situation inference.

[0006] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure, and thus it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0007] The object of the present invention is to provide a text semantic analysis system based on deep learning. By dividing the text into simple and complex categories and combining different fusion strategies, the system adopts fixed fusion when the information is consistent to efficiently integrate multimodal information; in conflict scenarios, dynamic weighted fusion is introduced to accurately focus on the most relevant modality and avoid the influence of semantic conflicts on the accuracy of sentiment analysis. The dynamic weighted fusion module enables the model to adaptively adjust the modal weights, improving the robustness and adaptability of the system in complex multimodal scenarios, and is applicable to applications such as emotion recognition and sentiment analysis to solve the problems in the above background art.

[0008] To achieve the above object, the present invention provides the following technical solutions: A text semantic analysis system based on deep learning, including a semantic extraction module, a feature extraction and evaluation module, a text complexity division module, a fixed fusion strategy module, and a dynamic weighted fusion module:

[0009] The semantic extraction module, first, extracts semantic information from the generated text data, mines the semantic relationships and structural information in the text, and captures the key themes, emotional tendencies, and intentions in the text;

[0010] A feature extraction and evaluation module extracts features from the obtained text semantic information. Based on the extracted features, the obtained text semantic information is intelligently evaluated through a pre-trained machine learning model.

[0011] A text complexity division module divides the obtained text semantic information into simple text information and complex text information based on the evaluation results of the machine learning model.

[0012] A fixed fusion strategy module selects a fixed fusion strategy that matches its characteristics to process simple text information. By fusing information from various modalities, the semantics of the text can be more comprehensively understood from multiple perspectives.

[0013] A dynamic weighted fusion module uses an attention mechanism to process complex text information. According to specific task requirements and the current context, it identifies the contribution of modal information to the current decision and accordingly increases its influence, dynamically adjusting the weights between different modalities.

[0014] Preferably, the specific steps for extracting semantic information from the generated text data, mining semantic relationships and structural information in the text, and capturing key themes, sentiment tendencies, and intentions in the text are as follows:

[0015] First, preprocess the generated text data and convert it into a format that the model can process.

[0016] After completing the preprocessing, the text needs to be converted into a numerical representation.

[0017] For the text represented by vectors, next, analyze the semantic relationships and sentiment information.

[0018] Finally, extract the key themes of the text through a topic model and further refine the intentions of the text.

[0019] Preferably, when extracting features from the obtained text semantic information, the extracted features include the syntactic complexity of the text and the degree of dispersion of different viewpoints in the text. Under the detection window, after analyzing the syntactic complexity of the text and the degree of dispersion of different viewpoints in the text, a syntactic depth index and a viewpoint dispersion index are generated respectively. The syntactic depth index and the viewpoint dispersion index are input into a pre-trained machine learning model, and a text semantic complexity coefficient is generated through the machine learning model. The obtained text semantic information is intelligently evaluated through the text semantic complexity coefficient.

[0020] Preferably, under the detection window, the specific steps for analyzing the syntactic complexity of the text and generating a syntactic depth index are as follows:

[0021] Under the detection window, perform syntactic parsing on the input text, and for each sentence Convert it into its corresponding syntactic tree , where the sentence set , is the th sentence, and the syntactic tree set , is the syntactic tree of the th sentence. Each sentence corresponds to a syntactic tree, is the total number of sentences;

[0022] In the generated syntactic tree, calculate the maximum hierarchical nesting depth of each syntactic tree, which reflects the hierarchical complexity of the sentence. The calculation expression is as follows:

[0023] , where is the maximum hierarchical nesting depth of the syntactic tree , that is, the maximum nesting depth of the syntactic tree corresponding to each sentence , is the set of all nodes of the syntactic tree , is the th node of the syntactic tree , , represents the path length from the root node to the node ;

[0024] Calculate the dependency path complexity of each sentence. The dependency path complexity comprehensively considers the strength and path depth of the dependency relationship between words within the sentence, and reflects the semantic density of the path through the cumulative multiplication of the weights of the dependency relationships on the path. The calculation expression is as follows:

[0025] , where is the dependency path complexity, indicating the dependency path complexity of the sentence , is the th sentence 's path number, that is, the number of all paths from the root node to the leaf node in the syntactic tree of the sentence , is the index number of the path, is the cumulative multiplication of the weights of the dependency relationships in the path, is the dependency relationship weight, that is, the weight of the th dependency relationship in the path ;

[0026] Combine the maximum hierarchical nesting depth and the dependency path complexity of each sentence to generate the syntactic depth index of the entire text. The calculation expression is as follows:

[0027] , where is the syntactic depth index, is the dependency path complexity adjustment parameter, is the overall adjustment parameter.

[0028] Preferably, under the detection window, the specific steps for analyzing the dispersion degree of different viewpoints in the text and generating the viewpoint dispersion index are as follows:

[0029] Under the detection window, the text is divided into multiple paragraphs, each paragraph is regarded as an independent viewpoint unit, and it is labeled as , the viewpoint unit set , is the viewpoint unit of the paragraph,

[0030] For each viewpoint unit , use a deep learning model to convert it into a high-dimensional vector representation, and the calculation expression is as follows: , where is the viewpoint vector representation, which is the vector representation of the viewpoint unit , is a pre-trained language model based on the Transformer architecture, which captures the semantic information of the text through bidirectional context understanding;

[0031] To measure the dispersion degree between each viewpoint unit, calculate the cosine similarity between every two viewpoint vectors, and then obtain the semantic distance. The calculation expression is as follows:

[0032] , where is the semantic distance between viewpoints, is the viewpoint vector representation of the viewpoint unit , is the viewpoint vector representation and the viewpoint vector representation cosine similarity;

[0033] Based on the calculated semantic distance between viewpoints , construct a dispersion matrix, and the expression is as follows:

[0034] , where is the viewpoint dispersion matrix, and the semantic distance between viewpoints is one of the elements in the viewpoint dispersion matrix;

[0035] Based on the viewpoint dispersion matrix The eigenvalue decomposition result is used to generate the opinion dispersion index, and the calculation expression is as follows:

[0036] , where is the opinion dispersion index, is the opinion dispersion matrix is the largest eigenvalue of is the opinion dispersion matrix is the th eigenvalue of is the total number of eigenvalues.

[0037] Preferably, the text semantic complexity coefficient generated after analyzing the text semantic information under the detection window is compared with the preset reference threshold of the text semantic complexity coefficient, and the obtained text semantic information is divided. The division steps are as follows:

[0038] If the text semantic complexity coefficient is greater than or equal to the preset reference threshold of the text semantic complexity coefficient, the text semantic information is divided into complex text information;

[0039] If the text semantic complexity coefficient is less than the preset reference threshold of the text semantic complexity coefficient, the text semantic information is divided into simple text information.

[0040] Preferably, for complex text information, the attention mechanism is used for processing. According to the specific task requirements and the current situation, the contribution of modal information to the current decision is identified, and its influence is increased accordingly. The specific steps for dynamically adjusting the weights between different modalities are as follows:

[0041] First, based on the attention mechanism, calculate the contribution degree of the text information extracted by each modality to the decision in the current situation. The calculation expression is as follows: , where is the initial attention score of the text information extracted by the modality is the input feature of the text information extracted by the modality is the adaptation function, which combines the features of the modality and the text semantic complexity coefficient to dynamically generate the attention score;

[0042] According to the text semantic complexity coefficient and the preset reference threshold of the text semantic complexity coefficient , adjust the attention score of each modality. If , it indicates complex text information, and the weight is increased to highlight the modal information with large contribution. The calculation expression is as follows:

[0043] ​​, where is the text information extracted by modal extraction is the adjusted attention score, is the adjustment coefficient, which is used to control the influence of the complexity coefficient on the attention score;

[0044] Using the syntactic depth index and the opinion dispersion index , further adjust the text information extracted by each modality The weights are calculated as follows:

[0045] , where is the text information extracted by the adjusted modality is the weight of is the syntactic depth index is the adjustment coefficient of, which is used to control the influence degree of the syntactic depth index on the weight of the text information extracted by the modality, is the opinion dispersion index is the adjustment coefficient of, which controls the influence degree of the opinion dispersion index on the weight of the text information extracted by the modality, is the text information extracted by the modality is the adjusted attention score, is the total number of text information extracted by the modality;

[0046] After obtaining the final weight of each modality , the features of each modality are weighted and fused according to the weights to obtain the output of multi-modal semantic fusion. The calculation expression is as follows:

[0047] , where is the final multi-modal fusion result, which represents the sum of the weighted feature information of each modality according to its weight.

[0048] In the above technical solution, the technical effects and advantages provided by the present invention:

[0049] By classifying text information into two categories, simple text and complex text, and then combining different fusion strategies, the present invention can more accurately process emotions and semantics when there are conflicts in multi-modal information. This classification enables the system to use an efficient fixed fusion strategy when the information is consistent, while introducing dynamic weighted fusion in complex text scenarios with information conflicts. Especially in sentiment analysis and context inference, fixed fusion is adopted for simple text information to quickly integrate multi-modal information, ensuring accuracy and efficiency. For complex text, dynamic weighted fusion focuses more attention on the most relevant modality by adaptively adjusting the weights of each modality, avoiding the masking of key information by emotional and semantic conflicts, and thus more accurately understanding the emotional turning points and background information of the text. This greatly reduces the risk of misjudgment as neutral or mixed emotions, making sentiment analysis more reliable and precise in semantic conflict situations.

[0050] Through the dynamic weighted fusion module, the present invention enables the model to adaptively adjust the information weights of each modality according to different task requirements and contexts, solving the limitation that fixed fusion strategies are difficult to adapt to diverse contexts. The system can identify the most influential modality in the current decision according to the context and task requirements, and thus dynamically adjust its influence to ensure optimal semantic understanding under multi-modal conflicts. For example, when the image conflicts with the text and audio information, the negative information in the image is identified through an adaptive attention mechanism, and the weights are reasonably allocated to avoid the masking of positive text and audio information. This flexible adjustment strategy not only improves the performance of the model in complex scenarios but also enhances its wide applicability in different applications (such as emotion recognition, sentiment analysis, context understanding, etc.), making the system more robust and adaptable to cope with the complex and variable real application scenarios of multi-modal information. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required in the embodiments. Obviously, the drawings described below are only some embodiments recorded in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0052] Figure 1 It is a schematic diagram of the modules of the text semantic analysis system based on deep learning of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0053] Now, example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these example embodiments are provided so that this disclosure will be more comprehensive and complete, and will fully convey the concept of the example embodiments to those skilled in the art.

[0054] The present invention provides a text semantic analysis system based on deep learning as shown in Figure 1 Figure 2, including a semantic extraction module, a feature extraction and evaluation module, a text complexity division module, a fixed fusion strategy module, and a dynamic weighted fusion module:

[0055] The semantic extraction module, first, extracts semantic information from the generated text data, mines the semantic relationships and structural information in the text, and captures the key themes, sentiment tendencies, and intentions in the text;

[0056] The specific steps for extracting semantic information from the generated text data, mining the semantic relationships and structural information in the text, and capturing the key themes, sentiment tendencies, and intentions in the text are as follows:

[0057] First, preprocess the generated text data and convert it into a format that the model can process;

[0058] The preprocessing includes steps such as word segmentation, stop word removal, and part-of-speech tagging. Word segmentation splits the text into independent word or sub-word units, which is the basis for subsequent analysis; removing stop words (such as "de", "shi", etc.) helps to retain the key information in the text; part-of-speech tagging can help identify the part of speech of nouns, verbs, etc., providing assistance for analyzing the text structure. Through text preprocessing, the model can focus more on the core content of the text and ensure that effective semantic information is extracted in subsequent steps.

[0059] After completing the preprocessing, the text needs to be converted into a numerical representation;

[0060] Word embedding (such as Word2Vec, GloVe) is a commonly used method, which represents words as vectors, enabling the model to process text semantics in a mathematical space. Word embedding technology captures the semantic similarity between words. For example, "cat" and "dog" are placed in a similar vector space. For scenarios where context is important, more advanced context word embedding technologies (such as BERT, ELMo) can be used. These models dynamically adjust the representation of words according to the context. This step enables the model to understand the similarity and relevance between words and lays the foundation for subsequent semantic relationship extraction.

[0061] For the text represented by vectors, next, analyze the semantic relationships and sentiment information;

[0062] Deep learning models such as recurrent neural network LSTM or self-attention mechanism can be used to identify the context relationships, sentiment tendencies, and intentions in the text. These models can identify the subject-predicate-object structure, modification relationships, etc. in the text to help understand the core semantics of the sentence. Sentiment analysis determines the sentiment tendency of the text, such as positive, negative, or neutral, through specific sentiment words and context features. This step provides strong support for the model to identify the theme, sentiment, and intention of the text.

[0063] Finally, key themes of the text are extracted through topic models such as LDA, BERT-based topic models to further refine the intention of the text;

[0064] Topic extraction can help the model identify the main discussion content of the text, while intention recognition infers the purpose of the text by analyzing sentence patterns, keywords, etc. For example, in customer feedback, the theme may be product features, and the intention may be to express dissatisfaction or make suggestions. Topic extraction and intention recognition provide the overall framework of the text for the model, enabling it to understand the text semantics more comprehensively and accurately and providing high-quality semantic information for subsequent tasks such as sentiment analysis and classification.

[0065] The feature extraction and evaluation module extracts features from the obtained text semantic information. Based on the extracted features, the obtained text semantic information is intelligently evaluated through a pre-trained machine learning model;

[0066] Features are extracted from the obtained text semantic information. The extracted features include the syntactic complexity of the text and the dispersion degree of different viewpoints in the text. Under the detection window, after analyzing the syntactic complexity of the text and the dispersion degree of different viewpoints in the text, a syntactic depth index and a viewpoint dispersion index are generated respectively. The syntactic depth index and the viewpoint dispersion index are input into a pre-trained machine learning model, and a text semantic complexity coefficient is generated through the machine learning model. The obtained text semantic information is intelligently evaluated through the text semantic complexity coefficient;

[0067] A sharp increase in the syntactic complexity of a text usually indicates a more complex semantic information structure. Syntactic complexity includes the frequent use of nested clauses, multiple modifiers, long sentences, and compound sentences, which make the logical relationships between sentences more difficult to parse. For example, nested clauses and multiple modifiers add layers of understanding, such that each clause or modifier may have an independent meaning or provide additional background information. Complex syntactic structures also mean that the dependency relationships between words are more intensive, making it difficult to locate the core components such as the subject, verb, and object, increasing the challenges for the model in deconstructing and understanding the text. In addition, complex syntax is usually accompanied by richer semantic relationships, such as causal, contrastive, concessive, and other logical associations, which, if not properly handled, can result in semantic deviations or misunderstandings. Therefore, the increase in syntactic complexity directly affects the structural complexity of the text, indicating that its semantic information contains richer and more hierarchical content, requiring more meticulous analysis and higher comprehension ability to accurately parse.

[0068] Under the detection window, the specific steps for analyzing the syntactic complexity of the text and generating the syntactic depth index are as follows:

[0069] Under the detection window, the input text is syntactically parsed, and each sentence is converted into its corresponding syntactic tree , where the sentence set , is the th sentence, , is the syntactic tree of the th sentence, and each sentence corresponds to a syntactic tree, is the total number of sentences;

[0070] A syntactic tree (ParseTree) is a tree diagram representing the internal grammatical structure of a sentence. By showing the hierarchical relationships and dependency relationships between words and phrases in a sentence, it helps to understand the overall grammatical construction of the sentence. In a syntactic tree, nodes represent the various components of a sentence, such as words, phrases, or clauses, and the lines connecting these nodes represent grammatical dependency relationships, such as subject-verb relationships, modification relationships, etc. Syntactic trees can be divided into two main types: dependency parse trees, which focus on the dependency relationships between words; and constituency parse trees, which focus on the phrase-level structure of sentences. By constructing syntactic trees, the model can conduct in-depth grammatical analysis of sentences, identify semantic levels, and provide a basis for semantic analysis and sentiment analysis in natural language processing.

[0071] In the generated syntactic trees, calculate the maximum hierarchical nesting depth of each syntactic tree to reflect the hierarchical complexity of the sentence. The calculation expression is as follows:

[0072] , where is the maximum hierarchical nesting depth of the syntactic tree, that is, for each sentence the corresponding syntactic tree the maximum nesting depth, is the set of all nodes of the syntactic tree is the syntactic tree of all nodes, is the syntactic tree the th node, represents the path length from the root node to the node ;

[0073] The hierarchical nesting depth of a syntactic tree refers to the number of levels from the root node to the deepest leaf node in the tree structure, reflecting the syntactic complexity of the sentence. The greater the hierarchical nesting depth, the more nested structures (such as clauses and modifiers) the sentence contains, and the more complex the syntactic relationships. For example, a simple sentence may only contain a subject and a predicate, with a smaller nesting depth; while a sentence containing multiple clauses and modifiers will have a deeper nesting and a higher hierarchical nesting depth. This depth metric is used to measure the structural complexity of a sentence. The more nested, the greater the difficulty of parsing and understanding the sentence usually is.

[0074] The path length from the root node to a node refers to the number of edges from the root node of the syntactic tree to a specific node, reflecting the hierarchical position of the node in the sentence structure. The longer the path length, the farther the node is from the core components (such as the subject-predicate structure) in deeper nesting. Usually, the path length from the root node to the leaf node can be used to calculate the relative importance and dependency of each component. For example, in a sentence, the subject and predicate are usually on a shorter path, while modifiers or clauses may be on a longer path. The path length helps to parse the hierarchy and semantic dependencies of the sentence, providing a basis for semantic understanding in natural language processing.

[0075] Calculate the dependency path complexity for each sentence. The dependency path complexity represents the complexity and strength of each dependency relationship in the path from the root node to the leaf node. The dependency path complexity comprehensively considers the strength and path depth of the dependency relationships between the words within the sentence, and reflects the semantic density of the path through the weighted multiplication of the dependency relationships on the path. The calculation is expressed as follows:

[0076] , where is the dependency path complexity, representing the dependency path complexity of the sentence , is the th sentence the number of paths, that is, the sentence The number of all paths from the root node to the leaf node in the syntactic tree, is the index number of the path, used to traverse each path from the root node to the leaf node, is the cumulative product of the weights of the dependency relations in the path, which is for the path The cumulative product of the weights of each dependency relation on it represents the complexity of the dependency relations of the path. is the dependency relation weight, that is, the path in the weight of the [ordinal] dependency relation, representing the strength or complexity of the dependency relation between nodes;

[0077] The dependency path complexity is an index to measure the complexity and strength of each dependency relation on the path from the root node to the leaf node in the syntactic tree. The dependency path refers to the path in the syntactic tree that extends from the root node to the leaf node step by step along the dependency relations, usually including the dependency relations between multiple words. The path complexity is represented by the weights of these dependency relations. The weight of each dependency relation reflects the strength of the relation (such as the subject-predicate relation or the modification relation), as well as the importance of the dependency relation to the semantics. The higher the path complexity, the more complex the levels and associations of the dependency relations on the path, and the corresponding increase in the structural and comprehension difficulty of the sentence.

[0078] The root node and the leaf node are two key node types in the syntactic tree. The root node is located at the top of the syntactic tree, usually representing the core or the subject-predicate structure of the sentence, controlling the entire grammatical hierarchy of the sentence, and is the starting point of all dependency paths. The leaf node is located at the end of the syntactic tree and has no child nodes, usually representing the specific words in the sentence (such as nouns, verbs, etc.). The path from the root node to the leaf node represents the hierarchical dependency relations of the sentence from the core structure to the specific words, and through these paths, the hierarchical relations of the sentence's grammatical structure and semantic dependencies can be fully demonstrated.

[0079] Combining the maximum hierarchical nesting depth and the dependency path complexity of each sentence, a syntactic depth index for the entire text is generated, and the calculation expression is as follows:

[0080] , where, is the syntactic depth index, is the dependency path complexity is a tuning parameter used to control the influence degree of the dependency path complexity on the syntactic depth index , is the overall tuning parameter used to control the non-linear adjustment of the overall result;

[0081] Under the detection window, the larger the performance value of the syntactic depth index generated after analyzing the syntactic complexity of the text, the more complex the semantic information structure of the obtained text usually indicates. The syntactic depth index reflects the syntactic complexity by measuring factors such as the nesting level of syntactic structures, the number of clauses, and complex modifiers in the text. When the syntactic depth index is high, it means that there are more nested structures, compound sentences, and multi-level semantic relationships in the text, which makes the dependency relationships between sentences more complex and thus increases the difficulty of understanding. Therefore, the larger the syntactic depth index generated under the monitoring window, the more abundant semantic levels and complex structural logics the text usually contains; on the contrary, a lower index indicates that the text structure is relatively simple and the semantic information is more direct.

[0082] When the dispersion degree of different viewpoints in the text increases, it indicates that the semantic information structure of the obtained text is relatively complex. This is because a text with a high dispersion degree usually contains multiple independent or contradictory viewpoints, covering different topics or positions, making the text no longer revolve around a single core but presenting a multi-level and multi-directional discussion. The differences between viewpoints, potential conflict relationships, and cross-topic relevance increase the complexity of the text semantic structure. The model needs to understand and integrate different viewpoints during analysis, distinguish the subtle differences among them, and identify the emotional tendencies and logical relationships of each viewpoint. In addition, a text with a high dispersion degree often contains implicit hierarchical expressions, such as some viewpoints not being explicitly stated but presented in ways such as metaphors and contrasts, which makes the difference between the surface semantics and the deep semantics of the text, further increasing the difficulty of semantic analysis. Therefore, the increase in the dispersion degree of viewpoints indicates that the text contains a complex semantic structure, and the model requires stronger semantic understanding and reasoning abilities to accurately analyze the multiple information therein.

[0083] Under the detection window, the specific steps for generating the viewpoint dispersion index by analyzing the dispersion degree of different viewpoints in the text are as follows:

[0084] Under the detection window, divide the text into multiple paragraphs, regard each paragraph as an independent viewpoint unit, and label it as , , is the viewpoint unit of the paragraph, is the total number of paragraphs. The text length and complexity determine the segmentation granularity. Long texts can be further divided to ensure that each paragraph can independently express a viewpoint;

[0085] The overall length of the text and the complexity of its content will affect how we segment it so as to more accurately extract independent viewpoints in the analysis. If the text is long or its content contains multiple levels of semantics, we can divide it into smaller parts (such as sentences or short paragraphs), and each part can express a complete viewpoint independently. This way of refined segmentation helps ensure that the model can accurately capture the independence and characteristics of each viewpoint, avoiding confusion or loss of semantic information, especially when dealing with complex or diverse content.

[0086] For each viewpoint unit , it is transformed into a high-dimensional vector representation using a deep learning model (such as BERT), and the calculation expression is as follows: , where is the viewpoint vector representation, is the vector representation of the viewpoint unit , is a pre-trained language model based on the Transformer architecture, which captures the semantic information of the text through bidirectional context understanding. The BERT model encodes the given text and transforms it into a high-dimensional vector representation suitable for semantic analysis;

[0087] The viewpoint vector representation means that each viewpoint unit in the text is transformed into a high-dimensional vector through a deep learning model to capture its semantic information. This vector representation quantifies the semantic features of the text, enabling the model to analyze the relationships between different viewpoints in the vector space. Here, the role of the viewpoint vector representation is to facilitate the model to compare and calculate the semantics of different paragraphs or viewpoints in the text. For example, by calculating the semantic distance between each viewpoint vector, the degree of dispersion or semantic conflict of the viewpoints in the text can be identified, thus more accurately analyzing the overall structural complexity of the text. This process is crucial for semantic parsing in tasks such as multi-modal semantic fusion.

[0088] To measure the degree of dispersion between each viewpoint unit, calculate the cosine similarity between every two viewpoint vectors, and then obtain the semantic distance. The calculation expression is as follows:

[0089] , where is the semantic distance between viewpoints, is the viewpoint vector representation of the viewpoint unit , is the viewpoint vector representation and the viewpoint vector representation of the cosine similarity;

[0090] Cosine similarity is a way to measure the similarity between two vectors. It represents the degree of similarity by calculating the cosine value of the angle between the two vectors, with a value range from -1 to 1. The closer the value is to 1, the more similar the two vectors are. Here, the role of cosine similarity is to measure the semantic similarity between different opinion units in the text, and to identify semantic consistency or difference by judging the angle between opinion vectors. This method helps the model analyze the semantic distance between different opinions, provides a basis for the calculation of the opinion dispersion index, and thus better understands the structural complexity and internal semantic relationship of the text.

[0091] Based on the calculated semantic distance between opinions , a dispersion matrix is constructed, and the expression is as follows:

[0092] , where, is the opinion dispersion matrix, and the semantic distance between opinions is one of the elements in the opinion dispersion matrix;

[0093] The role of the opinion dispersion matrix is to quantify and represent the semantic distance or dispersion degree between each opinion unit in the text. By recording the semantic differences between each opinion, it helps the model identify and analyze the internal structure and complexity of the text. When the distance values between the opinion units in the matrix are large, it indicates that the opinions in the text are more dispersed or conflicting, with a higher semantic complexity; while when the distance values are small, it indicates that the opinions are concentrated and the semantic consistency is strong. The opinion dispersion matrix provides a basis for multi-modal semantic analysis, enabling the model to more accurately understand the multi-level semantic relationship of the text and dynamically adjust the analysis strategy to cope with complex text content.

[0094] Based on the eigenvalue decomposition result of the opinion dispersion matrix , an opinion dispersion index is generated. The opinion dispersion index is defined as the ratio of the spectral radius (i.e., the largest eigenvalue) of the matrix to the harmonic mean of all eigenvalues, in order to highlight the most significant dispersion characteristics and at the same time reduce the influence of outliers. The calculation expression is as follows:

[0095] , where, is the opinion dispersion index, is the largest eigenvalue of the opinion dispersion matrix , representing the dispersion degree in the most significant direction of the opinion dispersion, is the opinion dispersion matrix 's -th eigenvalue, and the eigenvalue is the result obtained after eigenvalue decomposition of the matrix , and each Corresponds to an eigenvector, which reflects the degree of variation of the opinion dispersion matrix in a certain direction. Is the total number of eigenvalues, representing the opinion dispersion matrix The number of all eigenvalues in it;

[0096] The largest eigenvalue Represents the most significant opinion dispersion direction in the text. The larger the value, the more significant the opinion difference in a certain semantic direction. This value is in the numerator, making the dispersion index more sensitive to the semantically significant dispersion direction.

[0097] Under the detection window, the larger the performance value of the opinion dispersion index generated after analyzing the dispersion degree of different opinions in the text, the higher the dispersion degree of different opinions in the text, and thus the more complex the semantic information structure of the text is reflected. Specifically, when the opinion dispersion index is relatively high, the text usually contains multiple positions, cross-topic content, and even contradictory opinions, which increases the difficulty of semantic integration and differentiation for the model during the analysis process. Therefore, a higher opinion dispersion index indicates that the text requires stronger semantic parsing and reasoning capabilities, suggesting its complex structure; while when the opinion dispersion index is relatively low, the semantics of the text is more concentrated, the opinions are consistent, and the information structure is relatively simple, making it easier for the model to understand and process.

[0098] The machine learning model is not limited here. Any machine learning model that can achieve the comprehensive analysis of the syntactic depth index and the opinion dispersion index to generate the text semantic complexity coefficient is acceptable. To implement the technical solution of the present invention, the present invention provides a specific implementation method:

[0099] The text semantic complexity coefficient The calculation expression is as follows:

[0100] , where, , Are the preset proportionality coefficients of the syntactic depth index and the opinion dispersion index , and , Are both greater than 0.

[0101] It can be seen from the calculation expression of the text semantic complexity coefficient that under the detection window, the larger the performance value of the syntactic depth index generated after analyzing the syntactic complexity of the text, and the larger the performance value of the opinion dispersion index generated after analyzing the dispersion degree of different opinions in the text, the larger the performance value of the text semantic complexity coefficient generated after analyzing the semantic information of the text under the detection window, indicating that the semantic information structure of the text is relatively complex; on the contrary, it indicates that the semantic information structure of the text is relatively simple.

[0102] A text complexity classification module divides the obtained text semantic information into simple text information and complex text information based on the evaluation results of a machine learning model;

[0103] Compare and analyze the text semantic complexity coefficient generated after analyzing the text semantic information under the detection window with a pre-set reference threshold of the text semantic complexity coefficient, and divide the obtained text semantic information. The division steps are as follows:

[0104] If the text semantic complexity coefficient is greater than or equal to the pre-set reference threshold of the text semantic complexity coefficient, the text semantic information is divided into complex text information;

[0105] If the text semantic complexity coefficient is less than the pre-set reference threshold of the text semantic complexity coefficient, the text semantic information is divided into simple text information;

[0106] Simple text information refers to text with clear semantics, clear structure, and consistent sentiment. Complex text information refers to text containing multiple viewpoints, cross-topic content, or implicit semantics. Its structure is usually more rich and variable, and the semantic expression may involve multiple emotional tendencies, and even there are contradictory viewpoints.

[0107] A fixed fusion strategy module selects a fixed fusion strategy that matches its characteristics to process simple text information. By fusing information from various modalities, it can understand the semantics of the text more comprehensively from multiple perspectives;

[0108] Since the semantics of simple text information is clear and the information between modalities is usually consistent, traditional fusion methods such as early fusion and late fusion can be used to directly integrate information from modalities such as text, images, and audio. These strategies can efficiently combine multi-modal data and improve the model's understanding ability.

[0109] Adopt the selected fusion strategy to integrate and co-analyze simple text and data from other modalities. By fusing information from various modalities, the model can understand the semantics of the text more comprehensively from multiple perspectives. For example, in news reports, combining text content and relevant pictures can convey the whole picture of the event more accurately. Co-analysis helps the model capture details and improve the accuracy and reliability of the analysis.

[0110] A dynamic weighted fusion module uses an attention mechanism to process complex text information. According to specific task requirements and the current context, it identifies the contribution of modal information to the current decision and accordingly increases its influence, dynamically adjusting the weights between different modalities;

[0111] For complex text information, the attention mechanism is used for processing. According to specific task requirements and the current context, identify the contribution of modal information to the current decision, and accordingly increase its influence, and dynamically adjust the weights between different modalities. The specific steps are as follows:

[0112] First, based on the attention mechanism, calculate the contribution degree of the text information extracted by each modality to the decision in the current context. The calculation expression is as follows: , where is the text information extracted by the modality 's initial attention score, is the input feature of the text information extracted by the modality ; is the adaptation function, which combines the features of the modality and the text semantic complexity coefficient to dynamically generate the attention score;

[0113] The adaptation function (AdaptiveFunction) is a function that dynamically adjusts model parameters or weights, and is used to automatically adapt the output of the model according to input features or context information. Here, the adaptation function 's role is to calculate the attention score according to the input features of the text information extracted by each modality and the text semantic complexity coefficient

[0114] to measure the contribution degree of this modality to the current decision. This adaptive adjustment enables the model to flexibly handle complex situations. By dynamically assigning higher weights to more important modalities, the semantic analysis is made more accurate. Especially in the case of multi-modal information conflicts or complex texts, the adaptation function helps the model focus on the most informative modality. According to the text semantic complexity coefficient and the preset reference threshold of the text semantic complexity coefficient , adjust the attention score of each modality. If

[0115] , it indicates complex text information, and increase the weight to highlight the modality information with greater contribution. The calculation expression is as follows: where is the adjusted attention score of the text information extracted by the modality is the adjustment coefficient, which is used to control the influence of the complexity coefficient on the attention score;

[0116] The role of the adjustment coefficient is to control the influence strength of the text semantic complexity coefficient on the attention score. When the text complexity is high, the adjustment coefficient can amplify the influence of this complexity on the model's attention allocation, ensuring that the model pays more attention to key modal information; when the complexity is low, the adjustment coefficient can weaken the influence of complexity and avoid unnecessary weight adjustments. This adjustment mechanism enables the model to flexibly cope with complexity changes in different contexts, ensuring the accuracy and stability of the attention mechanism, thereby improving the accuracy of semantic analysis.

[0117] Using the syntactic depth index and the opinion dispersion index , further adjust the text information extracted from each modality weights, and the calculation formula is as follows:

[0118] , where is the weight of the text information extracted from the modality after adjustment , is the adjustment coefficient of the syntactic depth index , used to control the influence degree of the syntactic depth index on the weight of the text information extracted from the modality, is the adjustment coefficient of the opinion dispersion index , controlling the influence degree of the opinion dispersion index on the weight of the text information extracted from the modality, is the text information extracted from the modality after adjusted attention score, is the total number of text information extracted from the modality;

[0119] After obtaining the final weight of each modality, the features of each modality are weighted and fused according to the weight to obtain the output of multi-modal semantic fusion, and the calculation formula is as follows:

[0120] , where is the final multi-modal fusion result, representing the sum of the weighted feature information of each modality according to its weight.

[0121] Through this weighted fusion, the model can focus on the modality information with higher contribution, while taking into account the context features of complex texts, making the final output more interpretable and accurate.

[0122] By classifying text information into two categories: simple text and complex text, and then combining different fusion strategies, the present invention can more accurately process emotions and semantics when there are conflicts in multi-modal information. This classification enables the system to use an efficient fixed fusion strategy when the information is consistent, while introducing a dynamic weighted fusion in complex text scenarios where information conflicts. Especially in sentiment analysis and context inference, fixed fusion is adopted for simple text information to quickly integrate multi-modal information, ensuring accuracy and efficiency. For complex text, dynamic weighted fusion focuses more attention on the most relevant modality by adaptively adjusting the weights of each modality, avoiding the masking of key information by emotional and semantic conflicts, and thus more accurately understanding the emotional turning points and background information of the text. This greatly reduces the risk of misjudgment as neutral or mixed emotions, making sentiment analysis more reliable and precise in semantic conflict scenarios.

[0123] Through the dynamic weighted fusion module, the present invention enables the model to adaptively adjust the information weights of each modality according to different task requirements and contexts, solving the limitation that fixed fusion strategies are difficult to adapt to diverse contexts. The system can identify the most influential modality in the current decision according to the context and task requirements, and thus dynamically adjust its influence to ensure optimal semantic understanding under multi-modal conflicts. For example, when the image conflicts with the text and audio information, the negative information in the image is identified through an adaptive attention mechanism, and the weights are reasonably allocated to avoid the masking of positive text and audio information. This flexible adjustment strategy not only improves the performance of the model in complex scenarios but also enhances its wide applicability in different applications (such as emotion recognition, sentiment analysis, context understanding, etc.), making the system more robust and adaptable to cope with the complex and changing real application scenarios of multi-modal information.

[0124] Only some exemplary embodiments of the present invention have been described by way of illustration. Undoubtedly, for those of ordinary skill in the art, the described embodiments can be modified in various different ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

Claims

1. A text semantic analysis system based on deep learning, characterized in that It includes a semantic extraction module, a feature extraction and evaluation module, a text complexity division module, a fixed fusion strategy module, and a dynamic weighted fusion module: The semantic extraction module first extracts semantic information from the generated text data, mines the semantic relationships and structural information in the text, and captures the key themes, sentiment tendencies, and intentions in the text; The feature extraction and evaluation module extracts features from the obtained text semantic information, and based on the extracted features, intelligently evaluates the obtained text semantic information through a pre-trained machine learning model; The text complexity division module divides the obtained text semantic information into simple text information and complex text information based on the evaluation results of the machine learning model; The fixed fusion strategy module selects a fixed fusion strategy that matches its characteristics to process simple text information. By fusing information from multiple modalities, it comprehensively understands the semantics of the text from multiple perspectives; The dynamic weighted fusion module uses the attention mechanism to process complex text information. According to specific task requirements and the current situation, it identifies the contribution of modal information to the current decision and accordingly increases its influence, dynamically adjusting the weights between different modalities; Extract features from the obtained text semantic information. The extracted features include the syntactic complexity of the text and the degree of dispersion of different viewpoints in the text. Under the detection window, after analyzing the syntactic complexity of the text and the degree of dispersion of different viewpoints in the text, a syntactic depth index and a viewpoint dispersion index are generated respectively. The syntactic depth index and the viewpoint dispersion index are input into a pre-trained machine learning model, and a text semantic complexity coefficient is generated through the machine learning model. The obtained text semantic information is intelligently evaluated through the text semantic complexity coefficient; Compare and analyze the text semantic complexity coefficient generated after analyzing the text semantic information under the detection window with a pre-set reference threshold of the text semantic complexity coefficient, and divide the obtained text semantic information. The division steps are as follows: If the text semantic complexity coefficient is greater than or equal to the pre-set reference threshold of the text semantic complexity coefficient, the text semantic information is divided into complex text information; If the text semantic complexity coefficient is less than the pre-set reference threshold of the text semantic complexity coefficient, the text semantic information is divided into simple text information.

2. The text semantic analysis system based on deep learning according to claim 1, characterized in that, The specific steps for extracting semantic information from the generated text data, mining the semantic relationships and structural information in the text, and capturing the key themes, sentiment tendencies, and intentions in the text are as follows: First, preprocess the generated text data and convert it into a format that the model can process; After completing the preprocessing, convert the text into a numerical representation; Represent the text as a vector, and then analyze the semantic relationships and sentiment information; Finally, extract the key themes of the text through a topic model and further refine the intentions of the text.

3. The text semantic analysis system based on deep learning according to claim 1, characterized in that Under the detection window, the specific steps for analyzing the syntactic complexity of the text and generating a syntactic depth index are as follows: Under the detection window, the input text is syntactically parsed, and each sentence is converted into its corresponding syntactic tree , where the sentence set , is the th sentence, and the syntactic tree set , is the syntactic tree of the th sentence. Each sentence corresponds to a syntactic tree, is the total number of sentences; In the generated syntactic tree, calculate the maximum hierarchical nesting depth of each syntactic tree, which reflects the hierarchical complexity of the sentence. The calculation expression is as follows: , where is the maximum hierarchical nesting depth of the syntactic tree , that is, the maximum nesting depth of the syntactic tree corresponding to each sentence ; is the set of all nodes of the syntactic tree ; is the -th node of the syntactic tree ; ; represents the path length from the root node to node ; Calculate the dependency path complexity of each sentence. The dependency path complexity comprehensively considers the strength of the dependency relationship between words within a sentence and the path depth. The semantic density of the path is reflected by the cumulative multiplication of the weights of the dependency relationships on the path. The calculation is expressed as follows: , where is the dependency path complexity, representing the dependency path complexity of the sentence . is the number of paths of the th sentence , that is, the number of all paths from the root node to the leaf node in the syntactic tree of the sentence . is the index number of the path, is the cumulative product of the weights of the dependency relationships in the path, is the dependency relationship weight, that is, the weight of the th dependency relationship in the path; The maximum hierarchical nesting depth of each sentence and the dependency path complexity are combined to generate the syntactic depth index of the entire text. The calculation formula is as follows: , where is the syntactic depth index, is the dependency path complexity adjustment parameter, is the overall adjustment parameter.

4. The text semantic analysis system based on deep learning according to claim 1, characterized in that Under the detection window, analyze the dispersion degree of different viewpoints in the text. The specific steps for generating the viewpoint dispersion index are as follows: Under the detection window, the text is divided into multiple paragraphs. Each paragraph is regarded as an independent view unit and labeled as , the set of view units , is the view unit of the paragraph, is the total number of paragraphs; For each opinion unit , it is converted into a high-dimensional vector representation using a deep learning model, and the calculation expression is as follows: , where is the opinion vector representation, which is the vector representation of the opinion unit , is a pre-trained language model based on the Transformer architecture, which captures the semantic information of the text through bidirectional context understanding; To measure the dispersion degree between each viewpoint unit, calculate the cosine similarity between every two viewpoint vectors, and then obtain the semantic distance. The calculation expression is as follows: , where is the semantic distance between viewpoints, is the viewpoint unit 's viewpoint vector representation, is the viewpoint vector representation and the viewpoint vector representation 's cosine similarity; Semantic distance between computational viewpoints , construct a scatter matrix, and the expression is as follows: , where is the opinion dispersion matrix, and the semantic distance between opinions is one of the elements in the opinion dispersion matrix; Based on the eigenvalue decomposition result of the opinion dispersion matrix generate the opinion dispersion index, and the calculation expression is as follows: , where is the opinion dispersion index, is the opinion dispersion matrix of the largest eigenvalue, is the opinion dispersion matrix of the th eigenvalue, is the total number of eigenvalues.

5. The text semantic analysis system based on deep learning according to claim 1, characterized in that: For complex text information, use the attention mechanism for processing. According to specific task requirements and the current context, identify the contribution of modal information to the current decision, and accordingly increase its influence, and dynamically adjust the weights between different modalities. The specific steps are as follows: First, based on the attention mechanism, calculate the contribution degree of the text information extracted from each modality to the decision-making in the current context. The calculation formula is as follows: , where is the initial attention score of the text information extracted by the modality , is the input feature of the text information extracted by the modality , is the adaptation function, which dynamically generates the attention score by combining the features of the modality and the text semantic complexity coefficient . According to the text semantic complexity coefficient and the pre-set reference threshold of the text semantic complexity coefficient , adjust the attention score for each modality. If , it indicates complex text information. Increase the weight to highlight the modality information with great contribution. The calculation expression is as follows: , where is the text information extracted by modal extraction is the adjusted attention score is the adjustment coefficient used to control the influence of the complexity coefficient on the attention score; Using the syntactic depth index and the opinion dispersion index , further adjust the weights of the text information extracted from each modality , and the calculation expression is as follows: , where is the text information extracted from the adjusted modality weight of is the syntactic depth index is the adjustment coefficient of is the opinion dispersion index is the adjustment coefficient of is the text information extracted from the modality is the adjusted attention score of is the total number of text information extracted from the modality; After obtaining the final weight of each modality After that, the features of each modality are weighted and fused according to the weights to obtain the output of multi-modal semantic fusion. The calculation formula is as follows: , where is the final multi-modal fusion result, representing the sum of the weighted feature information of each modality according to its weight.

Citation Information

Patent Citations

  • Multi-modal sentiment analysis method based on double-flow attention and gating fusion

    CN117010407A

  • Cross-modal positive and negative semantic classification method based on text emotion and image content perception

    CN118690259A