Intelligent analysis system for English long and difficult sentence structure in combination with context characteristics
By constructing a three-layer contextual feature extraction model and dynamic weight adjustment, the problem of multi-dimensional contextual association in the parsing of long and difficult English sentences is solved. It realizes the synchronous linkage between syntactic structure and contextual semantics, improves the parsing accuracy, and is suitable for scenarios such as machine translation and academic literature reading.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 吕冰
- Filing Date
- 2026-01-16
- Publication Date
- 2026-05-01
AI Technical Summary
Existing English long and complex sentence structure analysis technologies ignore the multi-dimensional contextual relationships between syntactic component features, contextual discourse logic features, and sentence expression function features. This leads to errors in splitting nested clauses, confusion of main and subordinate clause logical relationships, and deviations in locating core syntactic components. As a result, they cannot achieve synchronous linkage between syntactic structure analysis and contextual semantics, and the analysis accuracy is insufficient.
A three-layer contextual feature extraction model is constructed, consisting of a syntactic structure layer, a discourse semantic layer, and a sentence function layer. Through a dynamic coupling and association module and an adaptive weight adjustment module, the multi-dimensional contextual features of long and complex English sentences are dynamically coupled and associated, and the parsing weight is optimized, thereby improving the parsing accuracy.
It improves the accuracy of analyzing long and complex English sentences, solves problems such as errors in splitting nested clauses, confusion between main and subordinate clauses, and deviations in locating core components, and provides technical support for machine translation, academic literature reading, language teaching and other scenarios.
Smart Images

Figure CN121960441A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of English parsing technology, specifically to an intelligent parsing system for complex English sentence structures that incorporates contextual features. Background Technology
[0002] As the most widely used lingua franca globally, English plays an irreplaceable bridging role in academic exchange, business communication, cultural dissemination, and international cooperation. Accurately understanding the semantic meaning of English texts is a core prerequisite for efficient cross-linguistic information transmission and is of great significance for promoting global knowledge sharing, economic cooperation, and cultural integration. Long and complex English sentences are a special and crucial type of linguistic unit in English texts. They typically contain complex syntactic structures such as nested clauses, multiple modifiers, and intricate logical connections, accurately carrying rich semantic information. Structural analysis of long and complex English sentences refers to the process of breaking down and organizing their syntactic components, main-subordinate clause relationships, and modifier relationships. This process is a fundamental step in deeply understanding the semantics of texts and directly determines the accuracy and efficiency of information retrieval. It plays a vital role in practical applications such as machine translation, academic literature reading, language teaching, and cross-linguistic information retrieval.
[0003] However, existing technologies for parsing the structure of long and complex English sentences still have certain shortcomings. These technologies generally rely on static grammatical rules, mathematical matrix models, or single-clause splitting, adhering to the traditional approach of "structure first, context-detached." They focus solely on the syntactic structure of the long and complex sentence itself, neglecting the multi-dimensional contextual relationships between syntactic components, the logical features of the surrounding text, and the expressive functions of the sentence. This approach leads to frequent problems when processing complex sentences, such as errors in splitting nested clauses, confusion of main and subordinate clause logical relationships, and misalignment of core syntactic components. It fails to achieve synchronous linkage between syntactic structure parsing and contextual semantics, and the parsing accuracy is insufficient for practical applications. This severely restricts the application effectiveness of technologies that rely on long and complex sentence parsing. Therefore, developing an intelligent English long and complex sentence structure parsing system that incorporates contextual features is of great significance. Summary of the Invention
[0004] The purpose of this invention is to overcome the shortcomings of existing technologies and provide an intelligent parsing system for complex English sentences that incorporates contextual features. This system constructs a three-layered contextual feature extraction model—comprising a syntactic structure layer, a discourse semantic layer, and a sentence function layer—to dynamically couple and correlate the multi-dimensional contextual features of complex English sentences. Furthermore, an adaptive weight adjustment module optimizes the feature parsing weights of each layer in real time based on the distribution of contextual features, achieving synchronous linkage between syntactic structure parsing and contextual semantics. This improves the accuracy of complex English sentence parsing and provides technical support for applications such as machine translation, academic literature reading, and language teaching that rely on complex sentence parsing, meeting the practical needs for efficient and accurate parsing of complex sentences.
[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: an intelligent parsing system for English long and difficult sentence structures that combines contextual features, the system comprising: a contextual feature hierarchical extraction module, a dynamic coupling and association module, an adaptive weight adjustment module, and a syntactic structure parsing module; The context feature layer extraction module extracts syntactic structure layer features, discourse semantic layer features, and sentence function layer features of long and difficult English sentences, and transmits the extracted three layers of features to the dynamic coupling association module and the adaptive weight adjustment module, respectively. After receiving the three-layer features, the dynamic coupling association module establishes a dynamic mapping relationship between the three-layer features and performs deep association fusion, then transmits the fused features to the syntactic structure parsing module. The adaptive weight adjustment module adjusts the parsing weight ratio of the three layers of features based on the received three-layer features and the distribution of contextual features of long and difficult sentences, and transmits the adjusted weight data to the syntactic structure parsing module. The syntactic structure parsing module combines the received fusion features and weighted data to perform main and subordinate clause splitting of long and difficult sentences, core syntactic component localization and logical relationship identification, and outputs the parsing results.
[0006] Furthermore, the context feature hierarchical extraction module performs the following operations when extracting the three-layer core context features of long and complex English sentences: The input long and complex English sentences are preprocessed to remove redundant characters and perform word segmentation, part-of-speech tagging, and preliminary syntactic analysis. Based on the preliminary syntactic analysis results, the main and subordinate clause relationships, the distribution of modifiers and the collocation rules of grammatical components in long and complex sentences are identified and integrated to form a set of syntactic structure layer features. By combining the contextual text information of long and complex sentences, semantic connection nodes, semantic pointing relationships and topic consistency features are mined, and a set of semantic layer features of the text is formed after systematic sorting. Identify logical connectors, sentence structure, and semantic tendencies in long and complex sentences; clarify the sentence function types of statement, contrast, cause and effect, concession, condition, and parallelism; and construct a feature set of sentence function layers.
[0007] Furthermore, the dynamic coupling and association module performs the following operations when establishing a dynamic mapping relationship between the three layers of features and performing deep association fusion: The syntactic structure layer, discourse semantic layer, and sentence function layer features output by the context feature layer extraction module are standardized to unify the feature data format and dimensions. Construct a three-layer feature association matrix and calculate the association strength value between features at each layer based on a semantic relevance algorithm; Dynamic mapping rules are established based on the association strength value to associate the features of the syntactic structure layer with the semantic association features in the discourse semantic layer and the logical function features in the sentence function layer. A feature fusion algorithm is used to deeply integrate the three layers of associated features to form a unified fusion feature set, so that syntactic components and discourse semantics and sentence function are organically linked. The correlation strength value is calculated using the following formula: ,in Indicates the first The first feature term and the first The correlation strength value of each feature term, , , These represent feature terms in the syntactic structure layer, discourse semantic layer, and sentence function layer, respectively. This represents the function for calculating cosine similarity. , , The feature correlation coefficient is derived from the three-layer feature cross-correlation coefficient obtained by training on a large-scale English long and difficult sentence corpus. It is determined by optimizing the co-occurrence frequency and semantic matching degree of features in different text styles.
[0008] Furthermore, the adaptive weight adjustment module performs the following operations when adjusting the parsing weight ratio of the three-layer features: A quantitative analysis of the distribution of contextual features of long and complex sentences was conducted to determine the syntactic complexity index, semantic association strength index, and sentence function type index, and a weight adjustment evaluation system was constructed. Based on the preset feature importance benchmark and in combination with the specific circumstances of the above evaluation indicators, the initial weights of features in the syntactic structure layer, discourse semantic layer, and sentence function layer are initially assigned. Real-time monitoring of intermediate results during the parsing of long and complex sentences; assessment of the support provided by each layer of features for the parsing operation; and adjustment of the weight ratio of that layer of features when ambiguity or deviation occurs in the parsing process corresponding to a certain feature. Based on the adjustment results of the feature weights, a weight configuration file is generated and fed back to the syntactic structure parsing module; The initial weights are allocated using the following formula: ,in Indicates the first Initial weights for layer features, These correspond to the syntactic structure layer, the discourse semantic layer, and the sentence function layer, respectively. This is a syntactic complexity metric. This is a semantic association strength index value. This is a value representing the sentence structure function type index. , , The contribution coefficient of the indicator is derived from the statistical analysis of the contribution of each indicator to the accuracy of the analysis in historical analysis data. The statistical objects cover the analysis cases of long and difficult sentences in various text styles such as academic, legal, and business.
[0009] Furthermore, the syntactic structure layer features include the component boundary features of subject, verb, object, attributive, adverbial, and complement, the syntactic function features of non-finite verbs, and the association range features of conjunctions. Among them, the component boundary features clarify the start and end positions of each grammatical component, the syntactic function features of non-finite verbs distinguish the different syntactic functions of infinitives, gerunds, and present participles, and the association range features of conjunctions define the boundaries of the syntactic components connected by the conjunctions.
[0010] Furthermore, the extraction of semantic features of the text is based on the thematic framework of the text containing long and difficult sentences. It combines the semantic connection relationship within paragraphs and the logical correspondence between sentences before and after. The semantic similarity is used to determine the semantic association strength of the context. At the same time, the frequency and distribution of core topic words are identified to enhance the pertinence of topic consistency features and reflect the semantic positioning and association logic of long and difficult sentences in the whole text.
[0011] Furthermore, the identification of the sentence structure function layer features adopts a combination of rule matching and semantic reasoning. First, explicit logical connectors in long and difficult sentences are matched through a preset logical connector rule library to preliminarily determine the sentence structure function type. Then, semantic reasoning is performed based on the semantic layer features of the text to verify the preliminary judgment result. At the same time, implicit logical relationships without explicit connectors are identified to supplement implicit sentence structure function features and cover the function recognition scenarios of complex sentences.
[0012] Furthermore, the dynamic coupling and association module introduces a feature priority ranking mechanism during feature fusion. Priorities are set based on the influence of each layer of features on the parsing results. Features at the syntactic structure layer, which directly affect the splitting of main and subordinate clauses, have a higher priority than features at the discourse semantic layer. Features at the sentence function layer, which directly affect the identification of logical relationships, have a priority that matches the discourse semantic layer features. This priority adjustment focuses on the features required for the core parsing operation. The feature priority is calculated using the following formula: ,in Indicates the first Priority scores for class features This indicates the degree of influence of this type of feature on the analysis results. This represents the average correlation strength between this type of feature and features from other layers. , The weighting coefficients are derived from feature priority adjustment coefficients trained on a domain-specific corpus, and are determined by comparing the consistency of parsing results under different priority settings.
[0013] Furthermore, when performing core syntactic component localization, the syntactic structure parsing module adopts a combination of component boundary detection algorithm and semantic association verification. First, it preliminarily locates the range of core syntactic components based on component boundary features in the syntactic structure layer features, and then verifies the component localization results in combination with discourse semantic layer features to eliminate component confusion or localization deviation.
[0014] Compared with existing technologies, this intelligent English sentence structure parsing system that combines contextual features has the following advantages: This invention constructs a three-layer contextual feature extraction model consisting of a syntactic structure layer, a discourse semantic layer, and a sentence function layer. It dynamically couples and correlates the multi-dimensional contextual features of long and complex English sentences, and uses an adaptive weight adjustment module to optimize the feature parsing weights of each layer in real time based on the distribution of contextual features. This achieves synchronous linkage between syntactic structure parsing and contextual semantics, solving problems such as errors in nested clause splitting, confusion of main and subordinate clause logic, and deviations in locating core components caused by neglecting multi-dimensional contextual correlations in existing technologies. This improves the accuracy of long and complex English sentence structure parsing, providing technical support for applications such as machine translation, academic literature reading, and language teaching that rely on long and complex sentence parsing. It meets the practical needs for efficient and accurate parsing of long and complex sentences, expanding the application boundaries and practical value of long and complex English sentence parsing technology.
[0015] Other advantages, objectives and features of the invention will be set forth in part in the description which follows, and in part will be apparent to those skilled in the art from the following examination or study, or may be learned from the practice of the invention. Attached Figure Description
[0016] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.
[0017] Figure 1A schematic diagram of the structure of an intelligent English long and complex sentence structure analysis system that incorporates contextual features; Figure 2 A flowchart of the workflow for an intelligent English sentence structure analysis system that incorporates contextual features; Figure 3 The flowchart shows the process of the layered extraction module for contextual features when extracting the three-layer core contextual features of long and complex English sentences. Detailed Implementation
[0018] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0019] The intelligent English long and complex sentence structure parsing system provided by this invention, which combines contextual features, aims to solve the problems of nested clause splitting errors, main and subordinate clause logical confusion, and core component positioning deviation caused by neglecting multi-dimensional contextual relationships in existing technologies. It improves the accuracy of long and complex sentence parsing and provides technical support for machine translation, academic literature reading, language teaching and other scenarios.
[0020] See Figure 1 and Figure 2 The system consists of four core modules, the details of which are as follows: The contextual feature layer extraction module first preprocesses the input long and complex English sentences, removing redundant characters and performing word segmentation, part-of-speech tagging, and preliminary syntactic analysis. Based on the preliminary analysis results, it identifies the main and subordinate clause relationships, the distribution of modifiers, etc., to form syntactic structure layer features (including subject-verb-object-adverb-complement boundaries, non-finite verb syntactic functions, etc.). Combining contextual text information, it mines semantic connection nodes, topic consistency, and other features to form discourse semantic layer features. It identifies logical connectors, sentence expression forms, etc., clarifies sentence function types, constructs sentence function layer features, and synchronously transmits the three layers of features to subsequent modules.
[0021] The dynamic coupling and association module standardizes the three-layer features, unifying the data format and dimensions; it constructs an association matrix, calculates the association strength between features at each layer, establishes dynamic mapping rules, and realizes the corresponding association between the syntactic structure layer and the discourse semantic layer and sentence function layer features; it introduces a feature priority ranking mechanism, focuses on the features required for core parsing, and uses a fusion algorithm to deeply integrate the associated features to form a unified fusion feature set, realizing the organic linkage between syntactic components and discourse semantics and sentence function.
[0022] The adaptive weight adjustment module performs quantitative analysis on the distribution of contextual features of long and complex sentences, determines three types of indicators: syntactic complexity, semantic relevance strength, and sentence function type, and constructs a weight adjustment evaluation system. Based on the preset feature importance benchmark and evaluation indicators, it initially allocates the initial weights of the three layers of features. It monitors the intermediate parsing results in real time, and when the parsing process corresponding to a certain layer of features becomes ambiguous or deviates, it dynamically increases the weight ratio of that layer, generates a weight configuration file, and provides feedback.
[0023] The syntactic structure parsing module combines the received fusion features and weighted data to perform main and subordinate clause splitting, core syntactic component localization, and logical relationship identification. When localizing core syntactic components, it uses a combination of component boundary detection algorithm and semantic association verification to eliminate component confusion or localization deviation, and finally outputs accurate parsing results, realizing synchronous linkage between syntactic structure parsing and contextual semantics.
[0024] Example 1 This embodiment applies an intelligent English sentence structure parsing system that incorporates contextual features to academic literature review scenarios. It addresses the need for parsing complex English sentences in international engineering and technology journals, which often contain multiple nested clauses, complex modifiers, and implicit logical relationships. This solves problems such as low efficiency in literature comprehension and biased extraction of core information caused by the ambiguity of complex sentence structures. Through the system's hierarchical contextual feature extraction, dynamic association fusion, and adaptive weight adjustment mechanisms, it achieves accurate segmentation of main and subordinate clauses, precise location of core components, and clear identification of logical relationships in complex sentences. This provides technical support for researchers to quickly grasp the core viewpoints of literature and conduct efficient interdisciplinary academic exchanges.
[0025] See Figure 1 , Figure 2 and Figure 3 The specific implementation process of this embodiment is as follows: First, start the system and input the target long and complex English sentence selected from an international engineering and technology journal. The system will automatically trigger the context feature layering extraction module to perform feature extraction. The first step is to preprocess the input long and complex sentence, automatically filter redundant special characters in the text, use NLP word segmentation algorithm to complete word segmentation, label the part-of-speech tagging model for each word, and use basic syntactic analysis tools to complete the preliminary syntactic framework construction.
[0026] Based on the preliminary syntactic analysis results, the system further identifies the main and subordinate clause relationships in long and complex sentences, clarifies the nesting levels of the main clause, relative clauses, and adverbial clauses, locates the distribution of modifiers such as adjective phrases and prepositional phrases, sorts out the collocation rules of grammatical components such as subject-verb collocation and verb-object collocation, and integrates them to form a set of syntactic structure features, which includes key features such as the component boundaries of subject-verb-object, relative clauses, adverbial clauses, and complements, the syntactic functions of non-finite verbs, and the scope of association of conjunctions.
[0027] Subsequently, combining the contextual text information of the journal article containing the complex sentence, and focusing on the thematic framework of "Mechanical Performance Testing of New Materials," the semantic connections within the paragraph and the logical correspondence between preceding and following sentences were analyzed. Semantic similarity calculations were used to determine the strength of the semantic association between the context, and statistical analysis was conducted on the sentence structure. " By analyzing the frequency and distribution of core topic words such as "", the specificity of topic consistency features is enhanced, and a set of semantic features of the text is formed after systematic analysis.
[0028] Finally, the sentence function layer features are identified by combining rule matching and semantic reasoning. First, the explicit logical connectors in long and difficult sentences are matched using a pre-set logical connector rule base to preliminarily determine the sentence function type. Then, the preliminary judgment results are verified by semantic reasoning based on the semantic layer features of the text. At the same time, implicit logical relationships without explicit connectors are identified to supplement implicit sentence function features, clarify the function types of long and difficult sentences such as statement, transition, and causation, construct a complete set of sentence function layer features, and transmit the three layers of features to the dynamic coupling association module and the adaptive weight adjustment module, respectively.
[0029] After receiving the three layers of features, the dynamic coupling and association module first standardizes the features at the syntactic structure layer, discourse semantic layer, and sentence function layer, unifying the feature data format and dimensions to ensure the features are correlated. Next, it constructs a correlation matrix for the three layers of features and calculates the correlation strength value between each layer based on a semantic relevance algorithm. In the specific implementation of this embodiment, the correlation strength value is calculated using the following formula: ,in Indicates the first The first feature term and the first The correlation strength value of each feature term, , , These respectively represent the feature terms in the syntactic structure layer, discourse semantic layer, and sentence function layer. This represents the function for calculating cosine similarity. , , The feature correlation coefficient is derived from the three-layer feature cross-correlation coefficient obtained by training on a large-scale English long and difficult sentence corpus. It is determined by optimizing the co-occurrence frequency and semantic matching degree of features in different text styles.
[0030] Based on the calculated association strength values, dynamic mapping rules are established to associate syntactic structure layer features with semantic association features in the discourse semantic layer and logical function features in the sentence function layer. Subsequently, a feature priority ranking mechanism is introduced. In this specific implementation, feature priority is calculated using the following formula: ,in Indicates the first Priority scores for class features This indicates the degree of influence of this type of feature on the analysis results. This represents the average correlation strength between this type of feature and features from other layers. , The weighting coefficients are derived from feature priority adjustment coefficients trained on a domain-specific corpus, and are determined by comparing the consistency of parsing results under different priority settings.
[0031] According to the priority ranking rules, the syntactic structure layer features that directly affect the splitting of main and subordinate clauses have a higher priority than the discourse semantic layer features. The sentence function layer features that directly affect the identification of logical relations have a priority that matches the discourse semantic layer features. By adjusting the priority, the features required for the core parsing operation are focused. Finally, the feature fusion algorithm is used to deeply integrate the three layers of features after association to form a unified fusion feature set, so that the syntactic components and discourse semantic sentence function form an organic linkage. The fusion features are then transmitted to the syntactic structure parsing module.
[0032] After receiving the three layers of features simultaneously, the adaptive weight adjustment module performs quantitative analysis on the distribution of contextual features in long and complex sentences, determines the syntactic complexity index, semantic association strength index, and sentence function type index, and constructs a complete weight adjustment evaluation system. Based on the preset feature importance benchmark and the specific circumstances of the above evaluation indicators, the initial weights of the three layers of features are initially allocated.
[0033] In the specific implementation of this embodiment, the initial weights are allocated using the following formula: ,in Indicates the first Initial weights for layer features, These correspond to the syntactic structure layer, the discourse semantic layer, and the sentence function layer, respectively. Syntactic complexity index value This is a semantic association strength index value. This is a value representing the sentence structure function type index. , , The contribution coefficient of the indicator is derived from the statistical analysis of the contribution of each indicator to the accuracy of the analysis in historical analysis data. The statistical objects cover the analysis of long and difficult sentences in various genres such as academic, legal and business.
[0034] The system monitors the intermediate results during the parsing of long and complex sentences in real time, judges the support of each layer of features for the parsing operation, and automatically adjusts and increases the weight ratio of the feature corresponding to a certain layer when it finds that there is ambiguity or deviation in the parsing link. Based on the adjustment result of the feature weight, a weight configuration file is generated and fed back to the syntactic structure parsing module.
[0035] After receiving the fused feature and weight data, the syntactic structure parsing module initiates the parsing process. First, combining the syntactic structure layer features and weight configuration, it performs clause splitting, clarifying the boundaries and hierarchical relationships of nested clauses based on the dynamically coupled association features to avoid splitting errors. In the core syntactic component localization stage, a combination of component boundary detection algorithms and semantic association verification is used. First, based on the component boundary features in the syntactic structure layer features, the scope of core syntactic components such as subject, verb, object, attributive, adverbial, and complement is preliminarily located. Then, the component localization results are verified by combining the discourse semantic layer features, eliminating cases of component confusion or localization deviation.
[0036] Finally, based on the functional features of the sentence structure and the fused semantic association information, the logical relationship between the main and subordinate clauses is identified, the logical types such as adversation, cause and effect, concession, condition and parallel are clearly stated, and after completing the complete syntactic structure parsing, standardized parsing results are output.
[0037] In summary, this embodiment, by applying the system to the scenario of parsing long and complex sentences in academic literature, achieves simultaneous linkage between syntactic structure analysis and contextual semantics. It effectively solves the problems of incorrect nested clause splitting, logical confusion between main and subordinate clauses, and deviations in core component location caused by neglecting multi-dimensional contextual relationships in existing technologies, significantly improving the accuracy of parsing the structure of long and complex English sentences. Researchers can use this system to quickly clarify the syntactic framework, core components, and logical relationships of long and complex sentences in academic literature, efficiently extract core information from the literature, and greatly improve the efficiency of academic reading.
[0038] Example 2 This embodiment applies an intelligent English sentence structure parsing system that incorporates contextual features to product description translation scenarios on cross-border e-commerce platforms. Specifically, it addresses the challenges of inaccurate parsing of complex English sentences in imported home appliances and machinery descriptions, which often contain multiple conditional restrictions, nested functional descriptions, and dense technical terms. This resolves issues such as ambiguity in functional expressions, confusion of usage conditions, and incorrect terminology caused by inaccurate sentence parsing in machine translation. Building upon the aforementioned embodiment, the system adapts to the stylistic features of product description texts. Through hierarchical feature extraction, dynamic association, and weight adjustment, it achieves accurate parsing of the core components, conditional logic, and terminological relationships of complex sentences. This provides structured syntactic data for machine translation, ensuring the accuracy and readability of product description translations and helping cross-border e-commerce users quickly understand product information.
[0039] See Figure 1 , Figure 2 and Figure 3 The specific implementation process of this implementation is as follows: After starting the system, input the target complex English sentence from the product description document of home appliances on the cross-border e-commerce platform, and the system will trigger a contextual feature layered extraction process. Based on the preprocessing operations of the aforementioned embodiments, and considering the high concentration of technical terms in product descriptions, a technical terminology dictionary matching step is added to extract the relevant information. " " We perform precise word segmentation and part-of-speech tagging on specialized terms in the home appliance industry to ensure the accuracy of terminology component identification in the preliminary syntactic analysis.
[0040] Based on the preliminary syntactic analysis results, continuing the logic of syntactic structure layer feature extraction, the focus is on identifying the subject-verb-object-adjective-adverbial-complement components in long and complex sentences, distinguishing the syntactic role of non-finite verbs in product function descriptions, clarifying the scope of association of conditional conjunctions such as "if" and "when," and integrating them to form a set of syntactic structure layer features adapted to the product description style. The semantic layer extraction is centered on the overall framework of the product description: "function introduction - usage method - safety tips." It analyzes the semantic connection between product function descriptions and operation steps within paragraphs, strengthens the association strength of core topic words such as "function," "operation," and "maintenance" through semantic similarity calculation, and sorts out the semantic orientation relationships of product components and operation actions in preceding and following sentences, forming a set of semantic layer features focused on the product description logic.
[0041] In the sentence function layer recognition, the system prioritizes matching explicit conditional and declarative connectives based on a pre-defined logical connective rule library. It also combines the semantic tendencies of modal verbs such as "must" and "should" in the product description text to verify the preliminary judgment results of sentence function through semantic reasoning. The system supplements implicit conditional logical relationships without explicit connectives, constructs a feature set of sentence function layers that mainly describes product functions and clarifies usage conditions, and transmits the three layers of features synchronously to subsequent modules.
[0042] After receiving the three-layer features, the dynamic coupling association module first standardizes the feature data format and dimensions according to the standardized processing method described in the previous embodiment to ensure the associativity of product description terminology features and logical features. Then, it constructs an association matrix of the three-layer features and calculates the association strength value between each layer of features based on a semantic relevance algorithm. In the specific implementation of this embodiment, the association strength value is calculated using the following formula: Based on the correlation strength value, a dynamic mapping rule is established, which focuses on accurately mapping the correlation range features of conditional connectors in the syntactic structure layer to the semantic correlation features of product operations in the discourse semantic layer and the conditional sentence function features in the sentence function layer.
[0043] A feature priority ranking mechanism is introduced. In the specific implementation of this embodiment, the feature priority is calculated using the following formula: According to the sorting rules, the sentence function layer features that directly affect the conditional logic recognition are highly adapted to the discourse semantic layer features. The term-related component boundary features in the syntactic structure layer are given high priority. By adjusting the priority, the core needs of product description parsing are focused. The feature fusion algorithm is used to deeply integrate the three layers of features after association, forming a unified fusion feature set and transmitting it to the syntactic structure parsing module.
[0044] After receiving the three layers of features, the adaptive weight adjustment module performs quantitative analysis on the contextual feature distribution of long and complex sentences in product descriptions. Based on the evaluation system of the aforementioned embodiment, it focuses on optimizing the quantitative standard of condition nesting level in the syntactic complexity index, the judgment rule of product terminology consistency in the semantic association strength index, and the weight ratio of conditional sentence types in the sentence function type index, and constructs a weight adjustment evaluation system adapted to the product description style.
[0045] Based on the preset feature importance benchmark and the optimized evaluation index, initial weights are initially assigned to the three layers of features. In the specific implementation of this embodiment, the initial weights are assigned using the following formula: The system monitors and analyzes intermediate results in real time, focusing on the splitting of conditional sentences and the parsing of product terms. When it finds that a certain feature causes ambiguity in conditional logic or deviation in the positioning of term components, it automatically increases the weight of that feature, generates a weight configuration file adapted to the parsing of product descriptions, and feeds it back to the syntactic structure parsing module.
[0046] After receiving the fused features and weighted data, the syntactic structure parsing module initiates a targeted parsing process. Building upon the main-subordinate clause splitting logic of the aforementioned embodiments, it focuses on splitting multiple nested conditional clauses in product descriptions, clarifying the hierarchical relationship between the main clause and conditional adverbial clauses, and avoiding confusion regarding conditional boundaries. In the core syntactic component localization stage, it continues the approach of combining component boundary detection algorithms with semantic association verification, prioritizing the localization of core product operation verbs such as "operate," "install," and "maintain," as well as professional terminology such as "component," "parameter," and "standard." The localization results are verified through the consistency features of product terminology in the discourse semantic layer, ensuring that core components are unbiased. In the logical relationship identification stage, it focuses on conditional and causal logic, clarifying the correspondence between "usage conditions - product functions" and "operation steps - expected results," forming structured syntactic parsing results. This adapts to the structured data requirements of machine translation for product description texts, ultimately outputting standardized parsing results.
[0047] In summary, this embodiment, by applying the system to the product description translation scenario in cross-border e-commerce, optimizes the text style based on the technical solutions of the aforementioned embodiments, achieving precise linkage between the parsing of long and complex sentences in product descriptions and contextual semantics. It effectively solves the problems of terminology positioning deviation and conditional logic confusion existing in the parsing of long and complex sentences in professional texts, significantly improving the accuracy of the syntactic data required for machine translation. With the help of this system, cross-border e-commerce platforms can achieve efficient and accurate parsing of long and complex sentences in product descriptions, providing reliable support for subsequent translation stages and avoiding user misunderstandings or product usage risks caused by translation ambiguities.
[0048] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. An intelligent system for analyzing the structure of complex English sentences by incorporating contextual features, characterized in that: The system includes: a context feature hierarchical extraction module, a dynamic coupling and association module, an adaptive weight adjustment module, and a syntactic structure parsing module; The context feature layer extraction module extracts syntactic structure layer features, discourse semantic layer features, and sentence function layer features of long and difficult English sentences, and transmits the extracted three layers of features to the dynamic coupling association module and the adaptive weight adjustment module, respectively. After receiving the three-layer features, the dynamic coupling association module establishes a dynamic mapping relationship between the three-layer features and performs deep association fusion, then transmits the fused features to the syntactic structure parsing module. The adaptive weight adjustment module adjusts the parsing weight ratio of the three layers of features based on the received three-layer features and the distribution of contextual features of long and difficult sentences, and transmits the adjusted weight data to the syntactic structure parsing module. The syntactic structure parsing module combines the received fusion features and weighted data to perform main and subordinate clause splitting of long and difficult sentences, core syntactic component localization and logical relationship identification, and outputs the parsing results.
2. The intelligent English long and complex sentence structure parsing system combining contextual features as described in claim 1, characterized in that, The context feature hierarchical extraction module performs the following operations when extracting the three core context features of long and complex English sentences: The input long and complex English sentences are preprocessed to remove redundant characters and perform word segmentation, part-of-speech tagging, and preliminary syntactic analysis. Based on the preliminary syntactic analysis results, the main and subordinate clause relationships, the distribution of modifiers and the collocation rules of grammatical components in long and complex sentences are identified and integrated to form a set of syntactic structure layer features. By combining the contextual text information of long and complex sentences, semantic connection nodes, semantic pointing relationships and topic consistency features are mined, and a set of semantic layer features of the text is formed after systematic sorting. Identify logical connectors, sentence structure, and semantic tendencies in long and complex sentences; clarify the sentence function types of statement, contrast, cause and effect, concession, condition, and parallelism; and construct a feature set of sentence function layers.
3. The intelligent English long and complex sentence structure parsing system combining contextual features as described in claim 1, characterized in that, The dynamic coupling and association module performs the following operations when establishing a dynamic mapping relationship between the three layers of features and performing deep association fusion: The syntactic structure layer, discourse semantic layer, and sentence function layer features output by the context feature layer extraction module are standardized to unify the feature data format and dimensions. Construct a three-layer feature association matrix and calculate the association strength value between features at each layer based on a semantic relevance algorithm; Dynamic mapping rules are established based on the association strength value to associate the features of the syntactic structure layer with the semantic association features in the discourse semantic layer and the logical function features in the sentence function layer. The feature fusion algorithm is used to deeply integrate the three layers of features after association to form a unified fusion feature set, so that the syntactic components and discourse semantics and sentence function are organically linked.
4. The intelligent English long and complex sentence structure parsing system combining contextual features as described in claim 1, characterized in that, The adaptive weight adjustment module performs the following operations when adjusting the parsing weight ratio of the three-layer features: A quantitative analysis of the distribution of contextual features of long and complex sentences was conducted to determine the syntactic complexity index, semantic association strength index, and sentence function type index, and a weight adjustment evaluation system was constructed. Based on the preset feature importance benchmark and in combination with the specific circumstances of the above evaluation indicators, the initial weights of features in the syntactic structure layer, discourse semantic layer, and sentence function layer are initially assigned. Real-time monitoring of intermediate results during the parsing of long and complex sentences; assessment of the support provided by each layer of features for the parsing operation; and adjustment of the weight ratio of that layer of features when ambiguity or deviation occurs in the parsing process corresponding to a certain feature. Based on the adjustment results of the feature weights, a weight configuration file is generated and fed back to the syntactic structure parsing module.
5. The intelligent English long and complex sentence structure parsing system combining contextual features as described in claim 1, characterized in that, The syntactic structure features include the component boundary features of subject, verb, object, attributive, adverbial, and complement; the syntactic function features of non-finite verbs; and the association range features of conjunctions. Among them, the component boundary features clarify the start and end positions of each grammatical component; the syntactic function features of non-finite verbs distinguish the different syntactic functions of infinitives, gerunds, and present participles; and the association range features of conjunctions define the boundaries of the syntactic components connected by the conjunctions.
6. The intelligent English long and complex sentence structure parsing system combining contextual features as described in claim 1, characterized in that, The extraction of semantic features of the text is based on the thematic framework of the text containing long and difficult sentences. It combines the semantic connection relationship within paragraphs and the logical correspondence between sentences. The semantic similarity is used to determine the semantic association strength of the context. At the same time, the frequency and distribution of core topic words are identified to enhance the pertinence of topic consistency features and reflect the semantic positioning and association logic of long and difficult sentences in the whole text.
7. The intelligent English long and complex sentence structure parsing system combining contextual features as described in claim 1, characterized in that, The recognition of the sentence structure function layer features adopts a combination of rule matching and semantic reasoning. First, explicit logical connectors in long and difficult sentences are matched by a preset logical connector rule library to preliminarily determine the sentence structure function type. Then, semantic reasoning is performed based on the semantic layer features of the text to verify the preliminary judgment results. At the same time, implicit logical relationships without explicit connectors are identified to supplement implicit sentence structure function features and cover the function recognition scenarios of complex sentences.
8. The intelligent English long and complex sentence structure parsing system combining contextual features as described in claim 1, characterized in that, When performing feature fusion, the dynamic coupling and association module introduces a feature priority ranking mechanism, which sets the priority according to the degree of influence of each layer of features on the parsing results. Among them, the syntactic structure layer features that directly affect the splitting of main and subordinate clauses have a higher priority than the text semantic layer features, and the sentence function layer features that directly affect the identification of logical relations have a priority that is adapted to the text semantic layer features. By adjusting the priority, the module focuses on the features required for the core parsing operation.
9. The intelligent English long and complex sentence structure parsing system combining contextual features as described in claim 1, characterized in that, When performing core syntactic component localization, the syntactic structure parsing module adopts a combination of component boundary detection algorithm and semantic association verification. First, it preliminarily locates the range of core syntactic components based on component boundary features in the syntactic structure layer features. Then, it verifies the component localization results by combining the semantic layer features of the discourse, thus eliminating cases of component confusion or localization deviation.