A patent document intelligent classification method and system based on semantic understanding

By constructing a multi-level semantic representation system and standardized classification decision rules, the challenges of semantic understanding and cross-language retrieval in patent classification have been solved, enabling more accurate identification of technical fields and cross-language retrieval.

CN121833954BActive Publication Date: 2026-05-22HUNAN KUANGCHU TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUNAN KUANGCHU TECH CO LTD
Filing Date
2026-03-12
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing patent classification methods struggle to deeply understand the technical semantics of patent documents, accurately identify ambiguous semantic boundaries and technical relationships, and suffer from insufficient recall and precision in cross-language retrieval.

Method used

By constructing a multi-layered semantic representation system, identifying technical field boundaries and semantic association patterns, establishing standardized classification decision rules, and supporting cross-language patent retrieval.

Benefits of technology

It achieves deep semantic understanding of patent documents, improves the accuracy of technical field identification, and expands the application scenarios of cross-language retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121833954B_ABST
    Figure CN121833954B_ABST
Patent Text Reader

Abstract

The application discloses a patent literature intelligent classification method and system based on semantic understanding, acquires patent literature data and classification system configuration parameters, recognizes high-density semantic areas through text segmentation and information density analysis, and constructs a classification index library; adopts a pre-training model to perform vectorization coding to form a semantic vector space, identifies ambiguous feature points through bidirectional semantic detection to form a theme clustering space; implements label matching analysis on the theme clustering space, and performs priority sorting and disambiguation processing to establish a semantic classification rule library; decomposes a classification matching strategy into core feature and auxiliary feature matching sequences, extracts field attributes of a candidate classification set, and implements weight proportioning to determine classification attribution parameters; extracts semantic mapping rules in combination with source language identification, and realizes cross-language retrieval through semantic alignment, thereby realizing accurate understanding of patent technology semantics and supporting multilingual retrieval.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of patent information processing technology, and in particular to a method and system for intelligent classification of patent documents based on semantic understanding. Background Technology

[0002] With the continuous growth of global patent applications, patent documents have become a core carrier of technological innovation and intellectual property protection. However, the vast amount of patent documents is characterized by scattered technical content, diverse expression methods, and blurred technical boundaries, posing serious challenges to patent retrieval, technical analysis, and knowledge mining. Traditional patent classification methods mainly rely on manual annotation or shallow keyword matching, which makes it difficult to deeply understand the technical connotation of patent documents, resulting in insufficient accuracy and consistency of classification results.

[0003] While existing automated classification technologies incorporate statistical learning and text mining methods, they still have significant shortcomings when handling complex scenarios such as overlapping technical fields, ambiguous terminology, and cross-language retrieval. On the one hand, these methods lack a deep understanding of the semantics of patent technologies, failing to accurately identify ambiguous semantic boundaries and technical relationships. On the other hand, cross-language patent retrieval relies on simple word translation mapping, ignoring the semantic differences of technical concepts in different language systems, resulting in retrieval recall and precision rates that fall short of practical requirements. Therefore, there is an urgent need for a novel patent document classification method that can deeply understand the semantics of patent technologies, accurately identify technical field affiliations, and support cross-language retrieval scenarios. Summary of the Invention

[0004] This invention discloses a method and system for intelligent classification of patent documents based on semantic understanding. It aims to construct a multi-level semantic representation system by performing deep semantic analysis on patent documents, identifying technical field boundaries and semantic association patterns, establishing standardized classification decision rules, and supporting cross-language patent retrieval, thereby providing accurate and reliable classification results for patent retrieval, technical analysis, and knowledge management.

[0005] The first aspect of this invention proposes an intelligent classification method for patent documents based on semantic understanding, comprising the following steps:

[0006] Obtain patent literature data and classification system configuration parameters, and construct a semantic classification index library based on the patent literature data and the classification system configuration parameters;

[0007] Based on the semantic classification index library, core semantic extraction is performed to form a semantic vector space. Semantic ambiguity feature points are identified using the semantic vector space, and a topic clustering space is constructed based on the semantic ambiguity feature points.

[0008] A label matching analysis is performed on the topic clustering space to determine the set of technical topic labels. The set of technical topic labels is then transformed into a classification knowledge system to form a semantic classification rule base. A classification matching strategy is generated based on the semantic classification rule base.

[0009] The classification matching strategy is decomposed into a core feature matching sequence and an auxiliary feature matching sequence. Semantic difference analysis is performed on the auxiliary feature matching sequence to extract candidate classification sets. Domain attribute extraction is performed on the candidate classification sets to form classification parameters.

[0010] The source language identifier is obtained by triggering cross-language retrieval through the core feature matching sequence and the classification attribution parameter. Based on the source language identifier, the corresponding semantic mapping rule is extracted from the semantic classification rule base to form a cross-language mapping relationship. The cross-language mapping relationship is semantically aligned with the core feature matching sequence to generate retrieval classification results.

[0011] A second aspect of this invention proposes a patent document intelligent classification system based on semantic understanding, comprising:

[0012] The data acquisition module is used to acquire patent document data and classification system configuration parameters, and to perform semantic parsing based on the patent document data and the classification system configuration parameters to construct a semantic classification index library;

[0013] The feature extraction module is used to extract core semantics based on the semantic classification index library to form a semantic vector space, identify semantically ambiguous feature points with the help of the semantic vector space, and construct a topic clustering space based on the semantically ambiguous feature points.

[0014] The rule generation module is used to perform label matching analysis on the topic clustering space to determine the technical topic label set, transform the technical topic label set into a classification knowledge system to form a semantic classification rule base, and generate a classification matching strategy based on the semantic classification rule base;

[0015] The classification processing module is used to decompose the classification matching strategy into a core feature matching sequence and an auxiliary feature matching sequence, perform semantic difference analysis on the auxiliary feature matching sequence to extract a candidate classification set, and perform domain attribute extraction on the candidate classification set to form classification attribution parameters.

[0016] The result output module is used to trigger cross-language retrieval to obtain the source language identifier by the core feature matching sequence and the classification attribution parameter, extract the corresponding semantic mapping rule from the semantic classification rule base according to the source language identifier to form a cross-language mapping relationship, and perform semantic alignment between the cross-language mapping relationship and the core feature matching sequence to generate retrieval classification results.

[0017] The beneficial effects of this invention are reflected in the following points: 1. By performing text segmentation and information density analysis on patent document data, high-density semantic regions are identified and a classification index library is constructed. A pre-trained model is then used to vectorize text fragments to form a semantic vector space. Bidirectional semantic similarity detection identifies semantically ambiguous feature points, forming a topic clustering space. This achieves a deep semantic understanding of patent documents and improves the accuracy of technical field identification. 2. By implementing label matching and overlap analysis on the topic clustering space, prioritizing and disambiguating labels, a set of technical topic labels is formed. This is then transformed into a classification knowledge system to constitute a semantic classification rule library, establishing complete classification decision rules and achieving standardization and systematization of the classification process. 3. By decomposing the classification matching strategy into core feature matching sequences and auxiliary feature matching sequences, extracting domain attributes and weighting candidate classification sets, and combining language feature vector extraction and semantic mapping rules, cross-language patent retrieval is achieved, expanding the linguistic scope and application scenarios of the retrieval. Attached Figure Description

[0018] The accompanying drawings illustrate specific examples of the technical solutions described in this invention and, together with the detailed embodiments, form part of the specification, serving to explain the technical solutions, principles, and effects of this invention.

[0019] Unless otherwise specified or defined, the same reference numerals in different figures represent the same or similar technical features, and different reference numerals may be used to represent the same or similar technical features.

[0020] Figure 1 This is a flowchart illustrating an intelligent patent document classification method based on semantic understanding according to the present invention.

[0021] Figure 2 This is a structural block diagram of a patent document intelligent classification system based on semantic understanding, according to the present invention. Detailed Implementation

[0022] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.

[0023] It should be understood that, when used in this application specification, the term "comprising" indicates the presence of the described feature, integral, step, operation, element, and / or component, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof.

[0024] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of this application include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0025] The technical solutions of the embodiments of this application will be described below.

[0026] like Figure 1 As shown, this embodiment of the invention provides a method for intelligent classification of patent documents based on semantic understanding, including the following steps S110-S150:

[0027] Step S110: Obtain patent document data and classification system configuration parameters, and construct a semantic classification index library based on semantic parsing of the patent document data and classification system configuration parameters.

[0028] Specifically, patent document data and classification system configuration parameters are acquired. The patent document data centers on the claims text, which serves as the core expression of the technical solution and carries the description of the invention's technical features. Patent document data is acquired in batches from patent databases, with a single acquisition quantity set at 10,000 documents to ensure sufficient samples for semantic analysis. The claims text in the patent document data is organized in the form of independent claims and dependent claims. Independent claims typically contain 200 to 800 Chinese characters, while dependent claims contain 50 to 300 Chinese characters. The classification system configuration parameters define the mapping relationship between IPC classification numbers and technical fields, forming a hierarchical classification node structure. The IPC classification number field in the patent document data corresponds to the classification nodes in the classification system configuration parameters. For example, the IPC classification number "G06F" corresponds to the "Electrical Digital Data Processing" node in the classification system configuration parameters, and the keyword set associated with this node includes terms such as "computation," "processor," "algorithm," and "data structure." The text encoding of the patent document data uniformly adopts UTF-8 format, and the classification system configuration parameters are stored in JSON format.

[0029] In some embodiments, the step of constructing a semantic classification index library by semantic parsing based on the patent document data and the classification system configuration parameters includes: generating a text segmentation unit set based on the patent document data; performing information density analysis on the text segmentation unit set to form a density distribution map; identifying high-density semantic regions within the density distribution map; and constructing a semantic classification index library by mapping the high-density semantic regions to the classification system configuration parameters.

[0030] A text segmentation unit set is generated based on patent document data. The claim text is organized in the patent document data as independent claims and dependent claims. Independent claims typically contain 200 to 800 Chinese characters, while dependent claims contain 50 to 300 Chinese characters. The text segmentation unit set breaks down each claim into multiple semantically complete technical feature description segments, based on the punctuation marks "semicolon" and "comma" in the patent document data, as well as the boundaries of technical terms. In the patent text "A data processing method, including: acquiring raw data; preprocessing the raw data; inputting the preprocessed data into a model," a semicolon separates three independent technical steps, each describing a complete technical action. After splitting, independent semantic analysis can be performed. Without splitting, the entire sentence contains multiple technical features, and the accuracy of TF-IDF calculations will be reduced due to feature mixing. In the patent document data, the sentence "A data processing method, including: acquiring raw data; preprocessing the raw data; inputting the preprocessed data into a model" is split into three segments: "acquiring raw data," "preprocessing the raw data," and "inputting the preprocessed data into a model." Each segment in the text segmentation unit set records its start and end character positions in the patent document data, forming a segment position index. This index records the precise location of the segment within the original text. A patent claim typically generates 15 to 60 segmented segments, and the text segmentation unit set aggregates approximately 350,000 segmented segments from 10,000 patents.

[0031] Information density analysis was performed on the text segmentation unit set to generate a density distribution map. The information density values ​​of the segmented text fragments in the text segmentation unit set were calculated using the TF-IDF algorithm. The TF-IDF value reflects the importance of technical terms in the fragment, and the calculation formula is TF-IDF = TF × log(N / DF), where TF is the term frequency (frequency of the term in a single patent document), N is the total number of patent documents, and DF is the number of patent documents containing the term. The term "preprocessing" appearing in the segmented text fragment "preprocessing the original data" has a frequency of 0.8% in the entire text segmentation unit set of 350,000 fragments and a frequency of 15% in a single patent document, resulting in a TF-IDF value of 2.73. The density distribution map displays the spatial distribution of information density in a two-dimensional coordinate system. The horizontal axis corresponds to the fragment position index in the text segmentation unit set, and the vertical axis corresponds to the information density value of the fragment. The fragment position indices in the text segmentation unit set are arranged sequentially from 0 to 350,000. The density distribution map maps the TF-IDF values ​​of these 350,000 fragments onto the coordinate system to form a density value matrix. The density value matrix has a size of 3500×100. The 350,000 segments are grouped into 3500 intervals based on their position indices. Each interval contains 100 consecutive segments. The average TF-IDF value of the segments within an interval is used as the density representative value for that interval. The row indices of the density value matrix correspond to the interval numbers, and the column indices correspond to the position offsets of the segments within the intervals. The interval with position indices from 8000 to 8100 in the text segmentation unit cluster corresponds to the data in the density value matrix for that interval, and the average TF-IDF value of this interval reaches 3.15.

[0032] High-density semantic regions are identified within the density distribution map. Intervals in the density matrix whose values ​​exceed the global average density plus one standard deviation are identified as high-density candidate intervals. The global average density is 1.82, and the density standard deviation is 0.96. The threshold for identification is calculated as 1.82 + 1 × 0.96 = 2.78. Of the 3500 intervals in the density distribution map, 420 intervals have a density value exceeding the 2.78 threshold. The corresponding segment index ranges in the location coordinate mapping of these intervals are extracted to form a list of candidate region locations. High-density semantic regions require that candidate intervals be spatially continuous and span at least five intervals. Continuity requires that the identified high-density semantic regions correspond to paragraphs in the patent text that centrally describe a specific technical topic. If only high density is required without continuity, scattered technical terminology fragments in the text will be incorrectly grouped into the same topic. For example, a patent may mention "data processing" multiple times but discuss different processing methods; without continuity, these fragments will be incorrectly aggregated. The region with location indices 7800 to 8300 in the density distribution map contains 52 consecutive high-density intervals, satisfying the requirements of continuity and span. After continuity screening, the 420 intervals in the candidate region location list were merged into one spatially continuous region. This merging operation was based on the spatial distance between adjacent intervals, ultimately resulting in 65 high-density semantic regions. Each region recorded its boundary coordinates, which identified the start and end positions of the high-density semantic region in the density numerical matrix. Segmented segments within the high-density semantic regions were extracted from the density distribution map for semantic aggregation. The aggregation method employed was LDA topic modeling to identify the dominant technological topics of the region. The high-density semantic region with location indices 7800 to 8300 was analyzed using LDA to identify the dominant technological topics, which were recorded as regional semantic features in the form of topic vectors. These topic vectors contain the distribution of technological topics and topic weights within the region.

[0033] A semantic classification index is constructed based on a mapping between high-density semantic regions and classification system configuration parameters. The classification system configuration parameters contain a hierarchical structure of 8 subcategories, 120 major categories, and 640 minor categories. Each classification node in the configuration parameters is associated with a set of technical field keywords and a semantic feature vector for that node. Cosine similarity is calculated between the regional semantic features of the high-density semantic regions and the semantic feature vectors of each classification node in the classification system configuration parameters. The cosine similarity calculation formula is as follows: In this system, A represents the semantic feature vector of the region, and B represents the semantic feature vector of the classification node. Nodes with a cosine similarity greater than 0.7 are selected as mapping candidate nodes. The feature vector of the node "G06F 17 / 30 (Data Retrieval)" in the region semantic features and classification system configuration parameters is calculated using cosine similarity, resulting in a cosine similarity of 0.82. This node becomes the primary mapping node for this region. Mapping candidate nodes are sorted according to their cosine similarity from high to low to determine the primary and secondary mapping relationships. The fragments corresponding to the boundary coordinates of the high-density semantic regions are extracted and uniformly labeled with their corresponding classification categories in the semantic classification index library. A mapping relationship is formed between the region boundary coordinates and the fragment position index. In the classification system configuration parameters, 640 subclass nodes are matched one by one with 65 high-density semantic regions. 62 regions are successfully mapped, and the 3 unmatched regions are classified as "unclassified" due to unclear semantic features. The semantic classification index library is organized using classification nodes as keys and all fragment indices mapped by that node as values. In the classification system configuration parameters, the "G06F 17 / 30" node is associated with a total of 82,000 fragment indexes from 18 high-density semantic regions in the semantic classification index library.

[0034] Step S120: Extract core semantics based on the semantic classification index library to form a semantic vector space, identify semantically ambiguous feature points using the semantic vector space, and construct a topic clustering space based on the semantically ambiguous feature points.

[0035] Specifically, a semantic vector space is constructed by performing core semantic extraction based on a semantic classification index. The text fragments corresponding to the 82,000 fragment indices associated with the "G06F 17 / 30 (Data Retrieval)" node in the semantic classification index are extracted as input samples for core semantic extraction. These fragments are labeled with a unified classification in the semantic classification index. Core semantic extraction encodes the text fragments, converting each fragment into a 768-dimensional vector representation. Each dimension of the vector corresponds to the semantic feature activation of the model's hidden layer. The 62 high-density semantic regions in the semantic classification index, which have already been mapped, contain a total of 325,000 fragment indices, and these fragments are vectorized one by one. The semantic vector space organizes the data in matrix form, with a matrix size of 325,000 rows × 768 columns. Each row corresponds to the vector representation of a fragment in the semantic classification index, and each column corresponds to a dimensional feature of the vector space. For the patent fragments "Query Processing Method Based on Inverted Index" and "Data Retrieval Technology Using B-Tree Structure," the vector distance between them in the semantic vector space is relatively close, reflecting the semantic similarity within the "data retrieval" category. The dimensionality features of the semantic vector space are reduced using principal component analysis, retaining the first 128 principal component dimensions. In the semantic classification index, each classification node corresponds to a different vector cluster in the semantic vector space, and the spatial distance between clusters reflects the degree of semantic difference between classification nodes.

[0036] In some embodiments, identifying semantically ambiguous feature points using the semantic vector space includes: dividing the semantic vector space into a main feature subspace and an auxiliary feature subspace; performing bidirectional semantic similarity detection on the main feature subspace and the auxiliary feature subspace to form a forward detection path and a reverse detection path; comparing the semantic differences between the forward detection path and the reverse detection path to form a bidirectional conflict point list; and labeling points in the bidirectional conflict point list that have semantic deviations in both directions as semantically ambiguous feature points.

[0037] The semantic vector space is divided into a principal feature subspace and an auxiliary feature subspace. The 128 principal component dimensions of the semantic vector space are sorted according to their variance contribution rate. The top 64 dimensions with a cumulative contribution rate of 80% are assigned to the principal feature subspace, and the remaining 64 dimensions are assigned to the auxiliary feature subspace. The principal feature subspace carries the most significant semantic distinguishing ability within the semantic vector space. In patent text classification, the first few principal components of the principal feature subspace typically correspond to core differences in technical fields. For example, the first principal component might distinguish between "hardware technology" and "software technology," and the second principal component might distinguish between "data processing" and "signal processing." Using only the auxiliary feature subspace would lose this macro-level technical field distinction information, leading to cross-field patents being incorrectly classified as similar. The auxiliary feature subspace captures subtle semantic differences and ambiguous boundary regions in the semantic vector space, playing a crucial role in identifying detailed features of semantic boundaries. The fragment "data retrieval optimization algorithm" is explicitly classified into the "data retrieval" cluster in the principal feature subspace, but its projected coordinates in the auxiliary feature subspace show a trend of shifting towards the "algorithm optimization" cluster. This shift reveals potential semantic cross-features. Vectors located near the classification boundary in the semantic vector space are projected onto the cluster edge in the main feature subspace, while their projection onto the auxiliary feature subspace exhibits a large coordinate offset. This offset feature is an important clue for identifying semantic ambiguity.

[0038] Bidirectional semantic similarity detection is performed on the main feature subspace and auxiliary feature subspace to form forward and reverse detection paths. Unidirectional detection may miss semantic ambiguities due to different projection directions of the factor space. For example, the fragment "neural network optimization algorithm" might be detected as "neural network" in the main feature subspace, but as "optimization algorithm" in the reverse detection from the auxiliary feature subspace. Bidirectional detection is necessary to detect such semantic discrepancies. A center vector is selected in the main feature subspace as the detection starting point. The center vector is the average coordinate of all vectors in the subspace, representing the semantic centroid of the main feature subspace. The forward detection path starts from the center vector of the main feature subspace and gradually probes the similarity to the corresponding positions in the auxiliary feature subspace, using a fixed-step spatial sampling strategy. Vectors in the main feature subspace whose distance from the center vector is within a set threshold are selected as the candidate point set for forward detection. These vectors form the core semantic region around the center in the main feature subspace. For each vector in the candidate point set, the forward detection path calculates its projected coordinates in the auxiliary feature subspace and compares them with other vectors near that location using cosine similarity. The cosine similarity score is recorded as the detection score of that vector on the forward path. The formula for calculating cosine similarity is: The fragment "fast retrieval based on hash table" has a cosine similarity of 0.78 in the forward path, indicating that the semantic expression of this fragment is relatively consistent in the two subspaces. The reverse detection path starts from the center vector of the auxiliary feature subspace, which is also the mean of the coordinates of all vectors in the subspace. The reverse path performs reverse similarity detection towards the main feature subspace. Vectors in the auxiliary feature subspace whose distance from the center vector is within a small threshold constitute the candidate point set for reverse detection. These vectors correspond to the core distribution area of ​​fine-grained semantic features in the auxiliary feature subspace. The reverse detection path calculates the projected coordinates of each vector in the candidate point set in the main feature subspace. The cosine similarity value at this position is denoted as sim_backward, reflecting the semantic consistency evaluation result of the vector on the reverse path.

[0039] For example, the step of comparing the semantic differences between the forward detection path and the reverse detection path to form a bidirectional conflict point list includes: extracting semantic similarity change sequences based on the forward detection path and the reverse detection path; performing gradient analysis on the semantic similarity change sequences to determine gradient abrupt change locations; extracting semantic feature vectors for each abrupt change location at the gradient abrupt change location; and arranging the semantic feature vectors according to the difference intensity to generate a bidirectional conflict point list.

[0040] Semantic similarity change sequences are extracted based on the forward and reverse detection paths. During the forward detection path traversal, similarity values ​​are recorded at fixed spatial intervals, forming discrete similarity sampling points, totaling approximately 240 points, covering the complete trajectory of the forward path from start to finish. The semantic similarity change sequence arranges these 240 sampling points in the detection order, forming a one-dimensional array where each element corresponds to the similarity value of a sampling point. The starting point of the forward detection path shows high similarity, indicating a high degree of consistency in the semantic expression of the vector at that position within the main and auxiliary feature subspaces. Similarity gradually decreases as the path extends outwards, reaching a low level in the middle section, reflecting the uncertainty of the semantic boundary region. The similarity at the end of the path shows a rebound trend. The reverse detection path uses a smaller spatial interval, generating approximately 160 sampling points. The semantic similarity change sequence of the reverse path is also stored as an array with a length of 160. The sampling points of the forward and reverse detection paths are spatially aligned and interpolated to ensure that both paths have similarity values ​​at the same locations for comparison. The semantic similarity change sequence is supplemented with the missing similarity estimates using an interpolation method. The length of the interpolated sequence is uniformly expanded to 300 points, and the uniform sequence length provides a consistent data foundation for subsequent gradient analysis.

[0041] Gradient analysis is performed on semantic similarity change sequences to determine gradient abrupt change locations. The similarity gradient reflects the rate of change of semantic similarity along the detection path. In the 300 sampling points of the semantic similarity change sequence, the gradient at each location is calculated using the similarity difference between that location and its adjacent locations. A positive gradient value indicates increasing similarity, while a negative gradient value indicates decreasing similarity. The semantic similarity change sequence exhibits a sudden reversal of the gradient sign at certain locations, with the gradient changing from negative to positive or vice versa, indicating a reversal of the similarity change trend. Gradient sign reversal corresponds to abrupt changes at semantic boundaries. For example, when the detection path passes through the boundary between "data retrieval" and "data mining," the similarity first decreases and then increases, and the gradient changes from negative to positive; this location is the semantic boundary. The criterion for determining gradient abrupt change locations is that the difference between adjacent gradient values ​​exceeds a set threshold and their signs are opposite; the threshold is set to 1.5 times the standard deviation of the gradient. There are 18 locations in the semantic similarity change sequence that meet this criterion in the forward path and 15 in the reverse path; these locations correspond to regions of drastic change in semantic boundaries. The 18 gradient mutation positions in the forward path and the 15 gradient mutation positions in the reverse path are spatially paired and matched. The pairing process determines whether two mutation positions correspond to the same semantic boundary based on the spatial distance between the mutation positions. The gradient mutation positions in the reverse path are spatially offset from those in the forward path. Some gradient mutation positions are close to overlapping on the two paths, while others are far apart.

[0042] Semantic feature vectors are extracted at each gradient abrupt change location. The spatial coordinates corresponding to the gradient abrupt change location are obtained through reverse mapping of the sampling point locations. This mapping process utilizes the correspondence between sampling points and vector space coordinates to obtain the precise coordinates of the abrupt change location in the main feature subspace and auxiliary feature subspace. The full-dimensional representation of the semantic feature vector at this location is read from the original semantic vector space, which has a vector dimension of 128, containing all feature information of the vector at that location in the dimensionality-reduced space. The 18 gradient abrupt change locations on the forward path correspond to 18 semantic feature vectors, and the 15 gradient abrupt change locations on the reverse path correspond to 15 semantic feature vectors. The two sets of vectors are paired and matched in spatial location, based on the proximity of spatial distance. The pairing condition for gradient abrupt change locations is that the spatial distance between the two locations is less than a set threshold. There are 12 pairs of abrupt change locations that meet the pairing condition. These paired locations exhibit gradient abrupt change features in both the forward and reverse detection directions. The semantic feature vectors are used to calculate the vector difference at the paired locations. The difference vector is defined as the difference between the forward and reverse vectors, and the magnitude of the difference vector reflects the degree of semantic deviation between the two directions at that location. Unpaired gradient mutation positions were classified into the unidirectional mutation category. There were a total of 9 unidirectional mutation positions, including 6 unpaired positions in the forward path and 3 unpaired positions in the reverse path. These positions only showed gradient mutations in one detection direction, while the similarity changed gradually in the other detection direction.

[0043] A bidirectional conflict point list is generated by arranging the positions according to the intensity of their semantic feature vector differences. The intensity of the difference is measured by the magnitude of the difference vector, which is obtained by subtracting the semantic vectors of the forward and reverse positions. A larger magnitude indicates higher conflict intensity, and a smaller magnitude indicates lower conflict intensity. The 12 pairs of paired positions are sorted from highest to lowest intensity of difference, with the positions with the highest conflict intensity at the beginning of the list and those with lower conflict intensity at the end. This sorting method ensures that the positions with the most significant semantic conflicts are addressed first. The bidirectional conflict point list stores conflict information in a table format. Table columns include the conflict point number, forward position coordinates, reverse position coordinates, difference intensity, and semantic feature vector. The table structure facilitates the identification of ambiguous points and the construction of cluster spaces. Unpaired gradient abrupt change locations in the bidirectional conflict point list are categorized as unidirectional conflicts. These locations exhibit semantic abrupt changes only in one detection direction. Unidirectional conflict locations are individually marked in the list. For example, in the patent excerpt "Deep Learning-Based Image Segmentation Technology," a gradient abrupt change is detected in the forward path, but no corresponding abrupt change is detected in the reverse path at the corresponding location. This location is marked as a unidirectional conflict, indicating that the semantic boundary is more sensitive in a certain detection direction. The 12 successfully paired conflict points are marked as bidirectional conflicts in the bidirectional conflict point list, while the 9 unpaired locations are marked as unidirectional conflicts.

[0044] Points exhibiting semantic deviation in both directions within the bidirectional conflict point list are designated as semantically ambiguous feature points. The bidirectional conflict point list contains 12 conflict points, which show abrupt changes in similarity gradients in both the forward and reverse detection directions. The semantic deviation criterion is that the difference intensity exceeds a set threshold and the difference in similarity values ​​between the two directions is greater than another set threshold. Seven points in the bidirectional conflict point list meet this condition, accounting for 58% of the total number of bidirectional conflict points. The conflict point with the highest difference intensity corresponds to the segment "Multimodal Data Fusion Index," which is classified as "Data Retrieval" in the main feature subspace but tends towards "Data Fusion" in the auxiliary feature subspace, indicating a significant discrepancy in semantic classification. The semantically ambiguous feature points record the spatial coordinates and deviation values ​​of these seven locations. The deviation value comprehensively considers both the difference intensity and the similarity difference, calculated through a weighted summation; a higher value indicates a more severe degree of semantic ambiguity at that location. Five bidirectional conflict points that did not meet the deviation judgment criteria in the list of bidirectional conflict points were classified as weak ambiguity points. These points have relatively low semantic ambiguity and have little impact on semantic classification. For example, although the patent fragment "query optimization and index structure" showed gradient mutations in both bidirectional detection, its difference intensity was lower than the threshold. The semantic attribution tendencies in both directions were basically consistent, both pointing to the "data retrieval" category, and it was judged as a weak ambiguity point.

[0045] A topic clustering space is constructed based on semantically ambiguous feature points. The seven coordinates of the semantically ambiguous feature points define the boundary range of highly ambiguous regions in the main feature subspace. The boundary range forms a closed region around these ambiguous points in the form of a convex hull, which is determined by calculating the smallest convex polygon containing all ambiguous points. The topic clustering space spatially divides the semantic vector space into ambiguous and unambiguous regions. Vectors in ambiguous regions require additional semantic disambiguation processing, while vectors in unambiguous regions can be directly classified into topics. The deviation value of the semantically ambiguous feature points serves as the boundary weight parameter of the topic clustering space. Positions with higher deviation values ​​correspond to stricter ambiguity judgment conditions, reducing the risk of misclassification of vectors near that position. The topic clustering space uses the DBSCAN clustering algorithm to perform topic clustering on vectors in unambiguous regions. The neighborhood radius and minimum sample number in the clustering parameters are adaptively adjusted according to the vector density to ensure that reasonable clustering structures can be formed in regions with different densities. Vectors near the fragment "Multimodal Data Fusion Index" form an independent ambiguous cluster in the topic clustering space. This cluster simultaneously contains semantic features of both the topics "data retrieval" and "data fusion". The topic clustering space identifies 25 topic clusters, of which 3 are ambiguous topic clusters, accounting for 12%. These ambiguous clusters require specialized disambiguation processing in the label matching analysis. The spatial coordinates of the semantically ambiguous feature points serve as anchor points for inter-cluster boundaries in the topic clustering space.

[0046] Step S130: Perform label matching analysis on the topic clustering space to determine the technical topic label set, transform the technical topic label set into a classification knowledge system to form a semantic classification rule base, and generate a classification matching strategy based on the semantic classification rule base.

[0047] In some embodiments, the step of performing label matching analysis on the topic clustering space to determine the technical topic label set includes: performing overlap analysis on the labels of each cluster topic in the topic clustering space to generate label cross-regions; performing technical boundary evaluation based on the label cross-regions to form a boundary clarity index; ranking the labels by priority using the boundary clarity index to generate a candidate label set; and performing disambiguation processing based on the candidate label set to generate the technical topic label set.

[0048] Overlap analysis of the labels of each cluster theme in the topic clustering space is performed to generate label cross regions. The candidate label term "data retrieval" of cluster 1 has a complete match with the candidate label terms of clusters 2, 5, and 8. These four clusters form label cross region A in the topic clustering space. The criterion for determining a label cross region is that at least two clusters share the same candidate label term. Multiple label cross regions are identified in the topic clustering space, and each cross region contains a different number of clusters. The spatial extent of label cross region A is determined by the convex hull of the boundaries of its containing clusters. The spatial region enclosed by the convex hull defines the geometric boundary of the label cross region. Label cross region B in the topic clustering space covers clusters 3 and 4, sharing the label "feature extraction". Cluster 3 contains segments related to "image feature extraction" and "video feature extraction", while cluster 4 contains segments related to "text feature extraction" and "speech feature extraction". Although the two clusters share the label "feature extraction", cluster 3 focuses on visual data processing, while cluster 4 focuses on symbolic data processing. The technical means and algorithm principles are fundamentally different, so they need to be kept as independent clusters rather than merged. The semantic representation of shared labels within label crossover regions may exhibit subtle differences across different clusters, necessitating boundary clarity analysis to determine whether merging is appropriate. Vector density within label crossover regions is measured as the ratio of the number of samples within a cluster to its spatial volume; crossover regions with high vector density indicate concentrated and consistent semantic representation within that region. Label crossover regions record the contribution weight of each participating cluster to the shared label. This contribution weight is defined as the proportion of vectors containing the shared label within that cluster to the total number of vectors in the crossover region. The weight distribution reveals the differences in the semantic contributions of different clusters to the shared label.

[0049] A boundary clarity index is formed by assessing the technical boundaries of the tag intersection area. Multiple clusters in tag intersection area A are spatially adjacent but not completely merged, with measurable boundary intervals between clusters. The boundary intervals between different cluster pairs show significant differences, and the size of the boundary interval directly determines the clarity of cluster separation. The boundary clarity index quantifies the spatial separation degree of clusters within the tag intersection area. The core formula is C=d_min / (d_max+ε), where d_min is the minimum boundary distance between clusters, d_max is the maximum boundary distance between clusters, and ε is a small constant to prevent division by zero. The closer the boundary clarity index value of the tag intersection area is to 1, the clearer the cluster separation; the closer the value is to 0, the more blurred the cluster boundaries. When the boundary clarity index of the tag intersection area is 0.85, it indicates that the boundaries of different technical themes within the area are clear, and it is possible to accurately determine whether the patent fragment belongs to "data retrieval" or "data mining"; when the index is 0.3, it indicates that the boundaries are blurred, and the fragment may involve two themes simultaneously. The boundary clarity index, combined with the vector density of the label intersection area, forms a comprehensive boundary quality score. The scoring formula is Q = C × log(ρ + 1), where C is the clarity index, ρ is the normalized dimensionless relative vector density value, and log is the natural logarithm. The comprehensive boundary quality score takes into account both boundary clarity and semantic density. The comprehensive boundary quality score of the label intersection area is used for weight allocation in label priority ranking. Labels corresponding to intersection areas with higher comprehensive boundary quality scores have higher reliability in technical classification. The cluster contribution weight is used to adjust the influence of different clusters on the semantics of the labels.

[0050] A candidate label set is generated by prioritizing labels using a boundary clarity index. The boundary clarity index of the label intersection area serves as the priority weight for shared labels within that intersection area; higher clarity label intersection areas correspond to higher priority labels. The shared label "feature extraction" is associated with multiple label intersection areas, each with a different clarity index. The average priority weight of this label is obtained by averaging the comprehensive boundary quality scores of all associated intersection areas. The average comprehensive boundary quality score reflects the label's overall discriminative ability across multiple technical scenarios. Technical topic labels in the topic clustering space are sorted from high to low priority weight, with labels possessing clear semantic boundaries and strong discriminative ability ranked first. The candidate label set selects labels with high priority weights, considering both semantic representativeness and boundary clarity requirements. The label "neural network" is prioritized for inclusion in the candidate label set due to its clear boundaries in associated label intersection areas and concentrated cluster contribution weights. The cluster contribution weight of the label intersection area is used to adjust the subcategorization of labels in the candidate label set; clusters with high contribution weights have a stronger influence on the semantic definition of the label. The candidate label set records the label intersection area number and priority weight associated with each label. The labels generated by ambiguous topic clusters in the topic clustering space are marked with ambiguity markers in the candidate label set. The ambiguity markers are used to trigger subsequent dedicated disambiguation processes.

[0051] A set of technical topic tags is generated based on the candidate tag set through disambiguation processing. Ambiguous tags in the candidate tag set are disambiguated through contextual semantic expansion, which extracts the tags of neighboring clusters as disambiguation context. The ambiguous tag "data processing" is associated with multiple clusters in the candidate tag set. Neighboring clusters of different clusters contain different tag sets, and the differences between neighboring cluster tags indicate that the ambiguous tag refers to different technical fields in different clusters. For example, the ambiguous tag "data processing" is found to contain the tags "batch processing" and "streaming computing" in the neighboring clusters of cluster 7, and the tags "signal processing" and "filtering algorithm" in the neighboring clusters of cluster 9. The differences between neighboring cluster tags provide disambiguation context; for example, the neighboring clusters of "batch data processing" contain "batch processing" and "streaming computing," while the neighboring clusters of "signal data processing" contain "signal processing" and "filtering algorithm." The significant differences in contextual semantics support tag splitting, and the disambiguation process splits it into two sub-tags: "batch data processing" and "signal data processing." After disambiguation processing, the total number of tags in the candidate tag set increases, and the increase in the number of tags equals the number of sub-tags generated by the split ambiguous tags. The technical topic tag set summarizes all tags after disambiguation, and each tag records its disambiguation context. The tags in the technical topic tag set are mapped to the clusters in the topic clustering space.

[0052] The technical topic tag set is transformed into a classification knowledge system to form a semantic classification rule base. The tags in the technical topic tag set are organized into a three-level classification tree structure according to the technical hierarchy. The first-level nodes correspond to the major technical categories, the second-level nodes correspond to the intermediate technical categories, and the third-level nodes correspond to the subcategories of technical categories. The first-level node "Data Processing Technology" includes the second-level nodes "Data Retrieval", "Data Mining", and "Data Visualization". The tag "Batch Data Processing" in the technical topic tag set belongs to the "Data Processing Technology - Data Mining" path. The classification knowledge system defines semantic association rules between tags. The association rules are expressed in the form of "IF tag A THEN may be associated with tag B (confidence X)". When the tag "feature extraction" appears in the patent fragment, the association rule "IF feature extraction THEN classifier training (0.82)" of the classification knowledge system is triggered, indicating that the fragment may also involve classifier training technology. The classification knowledge system obtains tag association rules by statistically analyzing the co-occurrence patterns of intra-cluster vectors in the topic clustering space. Tag pairs with high co-occurrence frequencies form high-confidence association rules. The semantic classification rule base stores the classification knowledge system in a rule engine format. Each rule contains four fields: antecedent label, consequent label, confidence, and support. Support reflects the frequency of the rule's occurrence in the training data; rules with high support correspond to common association patterns between technical labels. The disambiguation contextual information in the technical topic label set is transformed into contextual constraints for the classification knowledge system. These constraints limit the activation of labels only when they appear in specific contexts.

[0053] A classification matching strategy is generated based on a semantic classification rule base. The association rules in the semantic classification rule base are organized into rule chains, which describe the reasoning path from input text to classification results. The classification matching strategy defines the triggering order and matching threshold of the rule chains. The triggering order is arranged from high to low rule confidence, with high-confidence rules participating in classification reasoning first. The rule with the highest confidence in the semantic classification rule base, "IF Feature Extraction THEN Classifier Training," has the highest priority in the classification matching strategy. Triggering this rule requires the input text to contain the label "Feature Extraction" and the semantic feature vector similarity to exceed the matching threshold of 0.6. The classification matching strategy categorizes the rules in the semantic classification rule base into strong rules and weak rules. Strong rules have a confidence level exceeding 0.7, while weak rules have a confidence level between 0.5 and 0.7. The strategy prioritizes applying strong rules for classification reasoning, using them for the first round of classification reasoning, and using weak rules for supplementary classification of segments not covered by strong rules. When the patent document "An Image Processing Method Based on Deep Learning" is input, the strategy identifies the "feature extraction" label and triggers a strong rule chain. Through three steps of reasoning—"feature extraction → deep learning → neural network"—the classification node "G06N 3 / 08" is obtained. The classification matching strategy limits the activation conditions of rules based on contextual constraints, ensuring that the corresponding association rule is activated only when the label appears in a specific context. The classification matching strategy defines a rule conflict resolution mechanism, selecting the rule with the highest confidence result when multiple rules derive different classification results. The classification matching strategy implements the sequential execution of the rule chain in the form of a state machine.

[0054] Step S140: The classification matching strategy is decomposed into core feature matching sequence and auxiliary feature matching sequence. Semantic difference analysis is performed on the auxiliary feature matching sequence to extract candidate classification sets. Domain attribute extraction is performed on the candidate classification sets to form classification parameters.

[0055] Specifically, the classification matching strategy is decomposed into a core feature matching sequence and an auxiliary feature matching sequence. During the execution of the state machine of the classification matching strategy, the trigger sequences of rule matching states and threshold check states are recorded. These trigger sequences are arranged in an ordered linked list according to rule confidence from high to low. The core feature matching sequence extracts strong rule matching records with a confidence level exceeding 0.7 from the trigger sequences. Each record contains three fields: the antecedent label, the consequent label, and the matching confidence level of the trigger rule. In the processing of the patent document "A Speech Recognition System Based on a Neural Network," the core feature matching sequence captures a three-step strong rule chain: "Neural Network → Deep Learning → Speech Processing." The matching confidence level of each step exceeds 0.75. Specifically, the confidence level of the first step rule "Neural Network → Deep Learning" is 0.85, and the confidence level of the second step rule "Deep Learning → Speech Processing" is 0.78. This three-step rule chain forms a complete technological evolution from basic algorithm theory to practical application scenarios. The auxiliary feature matching sequence extracts weak rule matching records with confidence levels between 0.5 and 0.7 from the trigger sequence. The number of weak rule matching records is typically greater than that of strong rule matching records, reflecting the fuzzy region characteristics of the classification boundary and providing crucial clues for identifying cross-domain technology fusion. The rule conflict resolution mechanism of the classification matching strategy has already completed conflict screening when generating the core feature matching sequence. The auxiliary feature matching sequence retains some conflict records, which correspond to boundary detection information for candidate classifications.

[0056] Semantic dissimilarity analysis is performed on the auxiliary feature matching sequence to extract candidate classification sets. The weak rule chain of the auxiliary feature matching sequence contains multiple possible classification paths, with different paths pointing to different candidate classification nodes. Semantic dissimilarity analysis calculates the semantic distance between the endpoint classification nodes of each classification path in the auxiliary feature matching sequence. The semantic distance is measured by the hierarchical span of the classification tree structure. Candidate nodes "G06N 3 / 08 (Neural Network)" and "G10L 15 / 00 (Speech Recognition)" are both at the same hierarchical depth (3 levels): the former is located at the 6th position in the first-level category G06 (Physics) and the 14th position in the second-level category N (Computer Science), while the latter is located at the 10th position in the first-level category G10 (Musical Instruments) and the 12th position in the second-level category L (Electroacoustics). Although they are at the same hierarchical depth (3 levels), their position indices (P) are significantly different. Semantic distance calculation shows that they are far apart in the classification tree. The auxiliary feature matching sequence contains two candidate classification nodes, "G06N 3 / 08 (Neural Network)" and "G10L 15 / 00 (Speech Recognition)". These two nodes belong to different first-level categories in the IPC classification tree and have a large semantic distance, indicating a correlation between auxiliary features across different technical fields. The candidate classification set summarizes all endpoint classification nodes in the auxiliary feature matching sequence, with the number of nodes ranging from 3 to 8. When there are two technical topic groups within the candidate classification set, such as patent documents involving both "neural network" and "speech recognition", the candidate nodes form two peaks in the G06N and G10L groups, with the distance between the peaks reflecting the breadth of technical fields. Semantic dissimilarity analysis clusters the nodes in the candidate classification set, grouping nodes with a semantic distance less than a set threshold into the same cluster. The weak rule chain of the auxiliary feature matching sequence forms a confidence distribution of the classification candidates in the candidate classification set, with candidate nodes having higher confidence receiving higher weight in domain attribute extraction. The candidate classification set records the matching confidence of each candidate node.

[0057] In some embodiments, the step of extracting domain attributes from the candidate classification set to form classification parameters includes: constructing an attribute feature space based on the candidate classification set; mapping domain relevance through the attribute feature space to form a domain relevance identifier; dividing the attribute feature space into a strong relevance interval and a weak relevance interval using the domain relevance identifier as a boundary; and comparing the attribute distribution characteristics of the strong relevance interval and the weak relevance interval to form classification parameters.

[0058] An attribute feature space is constructed based on the candidate classification set. Each of the 3 to 8 candidate nodes in the candidate classification set is associated with a set of domain attribute features. These domain attribute features are extracted from the candidate node's parent and sibling nodes in the classification tree, covering nodes two levels upstream and the associated sibling nodes. The domain attribute features of the candidate node "G06N 3 / 08 (Neural Network)" include the technical keywords "computational model," "artificial intelligence," and "learning algorithm" from the parent node "G06N (Computer System Based on a Specific Computational Model)" and the keywords "network topology" and "neuronal connection" from the sibling node "G06N 3 / 04 (Neural Network Structure)". Each candidate node is associated with 12 to 20 attribute keywords. The attribute feature space is constructed as a matrix with candidate nodes as rows and domain attribute keywords as columns. The matrix elements represent the association strength between the candidate node and the attribute keywords, with values ​​ranging from 0 to 1, reflecting the degree of correlation between the candidate node's technical features and the attribute keywords. The association strength is derived from the co-occurrence relationship between the candidate node and the keywords in the classification tree. When the candidate classification set contains 5 candidate nodes, the total number of extracted domain attribute keywords is approximately 60 to 100. Keywords extracted from different candidate nodes are deduplicated and merged to form the column dimensions of the matrix. The deduplication process retains shared keywords among candidate nodes, which form the basis of the domain relevance mapping. The candidate node "G10L 15 / 00 (speech recognition)" shows a high correlation with attribute keywords such as "acoustic model," "speech signal," and "feature extraction" in the attribute feature space, and a moderate correlation with the "neural network" attribute, reflecting the characteristics of cross-domain technology applications.

[0059] Domain association identifiers are formed by mapping domain association through attribute feature space. The domain association between candidate nodes is obtained by calculating row vector similarity in the matrix of attribute feature space. The row vectors in the attribute feature space contain information on the association strength of candidate nodes in each attribute dimension, providing a complete data foundation for domain association mapping. The similarity calculation uses the cosine similarity method, and the core formula is: Here, A_i and A_j are the row vectors of candidate nodes i and j, respectively. The cosine similarity of the row vectors of candidate nodes "G06N 3 / 08" and "G10L 15 / 00" is 0.52. The similarity value is between 0 and 1, with a higher value indicating that the two candidate nodes are more similar in domain attributes. Domain association mapping transforms the similarity values ​​of candidate node pairs into domain association identifiers. Node pairs with a similarity greater than 0.6 are identified as strongly associated, while those with a similarity less than 0.6 are identified as weakly associated. This hierarchical method transforms the degree of association of candidate nodes from numerical values ​​into interpretable category identifiers. Multiple pairs of node combinations are formed among candidate nodes in the candidate classification set, and domain association mapping generates an association identifier for each pair. When patent literature involves both neural network and speech recognition technologies, the weak association identifiers of candidate nodes "G06N 3 / 08" and "G10L 15 / 00" reflect the cross-functional characteristics of the two technical fields. The weak association indicates that the two nodes do not belong to the same core technology cluster, nor are they completely unrelated independent domains. Domain association identifiers are stored in the form of an association matrix. A diagonal element of 1 indicates that a node is fully associated with itself, and the off-diagonal elements are the association degree values ​​of the node pairs. The cosine similarity calculation results are filled in the corresponding positions of the association matrix.

[0060] The attribute feature space is divided into strongly correlated and weakly correlated intervals based on domain association identifiers. The association matrix of the domain association identifiers is used to determine the boundaries between these intervals through threshold segmentation, with a threshold set to 0.6. Domain association identifiers transform the numerical similarity of candidate node pairs into operable classification identifiers, facilitating subsequent interval segmentation. This threshold is determined based on statistical analysis of extensive patent classification practice data and effectively distinguishes between close and loose associations within technical fields. Candidate node pairs with an association degree exceeding 0.6 in the attribute feature space are classified into the strongly correlated interval, while those with an association degree below 0.6 are classified into the weakly correlated interval. Among the five candidate nodes in the candidate classification set, nodes "G06N 3 / 08" and "G06N 3 / 04" have an association degree of 0.78 and are classified into the strongly correlated interval, forming a neural network technology cluster. Nodes within this cluster exhibit highly overlapping attribute distributions across keyword dimensions such as "artificial intelligence," "learning algorithm," and "neuron." Candidate nodes in strongly correlated intervals are typically located in the same secondary or tertiary category in the classification tree, exhibiting highly similar technical features and a relatively clear classification affiliation. These nodes represent the core technical fields of the patent documents. Candidate nodes in weakly correlated intervals span different primary or secondary categories in the classification tree, exhibiting significant differences in technical fields. For example, the correlation between nodes "G10L 15 / 00" and "G06N 3 / 08" is 0.52, placing them in a weakly correlated interval. The weakly correlated node "G10L 15 / 00" shows significant differences in attribute distribution across keywords such as "acoustic model" and "speech signal" compared to strongly correlated nodes, but exhibits some overlap across keywords such as "feature extraction" and "pattern recognition." The segmentation of strongly and weakly correlated intervals divides candidate nodes into a core technology layer and a peripheral technology layer based on technical similarity. The core technology layer corresponds to the main innovative points of the patent, while the peripheral technology layer corresponds to the patent's technical application scenarios or auxiliary technologies.

[0061] For example, the step of comparing the attribute distribution characteristics of the strongly correlated intervals and the weakly correlated intervals to form classification parameters includes: mapping the attribute distribution density of the strongly correlated intervals and the weakly correlated intervals to form a dual-interval density map; determining the attribute distribution intersection region based on the dual-interval density map; performing boundary ambiguity assessment on the attribute distribution intersection region to construct a boundary strength identifier; and applying differentiated weighting to the strongly correlated intervals and the weakly correlated intervals according to the boundary strength identifier to form classification parameters.

[0062] A dual-interval density map is formed by mapping the attribute distribution density of strongly correlated and weakly correlated intervals. The row vectors of candidate nodes in strongly correlated intervals are projected onto the attribute keyword dimension in the attribute feature space. The strong correlation density is formed by summing the correlation strength of each attribute keyword column across all nodes in the strongly correlated interval. This summation process superimposes the correlation strength of multiple nodes on the same keyword dimension, reflecting the overall importance of the keyword in the strongly correlated interval. The density of weakly correlated intervals is calculated using the same method. The dual-interval density map plots the density curves of strongly correlated and weakly correlated intervals with attribute keywords on the horizontal axis and cumulative correlation strength on the vertical axis. The two curves present a comparative relationship on the same coordinate system. The density curve of strongly correlated intervals shows a peak at a specific attribute keyword position. The peak corresponds to the common technical features of the strongly correlated candidate nodes. For example, keywords such as "neural network" and "deep learning" show obvious density peaks in strongly correlated intervals, with peak values ​​typically 2 to 5 times higher than those at the same keyword position in weakly correlated intervals. This density difference visually demonstrates the significant difference in the distribution of technical features between strongly and weakly correlated intervals. The density curves of weakly correlated intervals exhibit a relatively flat distribution along the attribute keyword dimension, with lower peak heights and smaller fluctuation ranges, indicating that the technical characteristics of weakly correlated intervals are more dispersed. The horizontal axis of the dual-interval density map contains 60 to 100 attribute keywords, which cover the core technical areas of candidate classifications. The values ​​of the density curves at these keyword positions constitute the density data.

[0063] The attribute distribution intersection region is determined based on a dual-interval density map. The two density curves of the dual-interval density map intersect at some attribute keyword positions. These intersection positions correspond to shared technical attributes in strongly correlated and weakly correlated intervals, reflecting the fusion points of different technical fields. The attribute distribution intersection region is defined as the set of attribute keywords where the absolute value of the difference between the density of the strongly correlated interval and the density of the weakly correlated interval is less than a set threshold. This threshold is dynamically determined based on the statistical distribution of the overall density data, typically set to 0.5 times the density standard deviation. Among the 60 to 100 attribute keywords in the dual-interval density map, 8 to 15 keywords satisfy the intersection condition; these keywords constitute the attribute distribution intersection region. The patent document "Speech Feature Extraction Method Based on Neural Networks" includes keywords such as "feature extraction," "signal processing," and "pattern recognition" in the attribute distribution intersection region. These keywords connect the two fields of neural network technology and speech processing technology, reflecting the characteristics of cross-domain applications. Specifically, the density of the keyword "feature extraction" is 2.3 in the strongly correlated interval and 2.1 in the weakly correlated interval, with a difference of only 0.2, satisfying the intersection condition. The proportion of keywords in the attribute distribution intersection area to the total number of keywords reflects the degree of technical overlap between strong and weak association intervals. The higher the proportion, the more blurred the domain boundaries of the candidate classification, and the more cautious the classification decision needs to be.

[0064] A boundary ambiguity assessment is performed on the intersection areas of attribute distributions to construct boundary strength identifiers. Boundary ambiguity is defined as the ratio of the number of keywords in the intersection area to the total number of keywords, ranging from 0 to 1. The size of the intersection area directly reflects the degree of technical overlap between strong and weak association intervals. A higher boundary ambiguity value indicates a more blurred boundary between strong and weak association intervals, and a greater degree of technical overlap; a lower value indicates a clearer boundary and more distinct technical features. The boundary strength identifier is obtained through a reverse mapping of boundary ambiguity, with the core formula S=1-M, where M is the boundary ambiguity and S is the boundary strength identifier. The reverse mapping ensures that the identifier value is high when the boundary is clear and low when the boundary is ambiguous. A boundary strength identifier value close to 1 indicates a clear boundary between strong and weak association intervals and a relatively certain classification, while a value close to 0 indicates an ambiguous boundary, requiring a comprehensive assessment of multiple factors. When the boundary strength indicator is higher than 0.8, the overlap area between the attributes of strong and weak correlation intervals accounts for only 20%, and the technical boundaries are clear. At this time, the technical features of strong correlation candidate nodes (such as G06N 3 / 08) are dominant, while the features of weak correlation candidate nodes (such as G10L 15 / 00) are auxiliary features. The classification tends to be based on the candidate classification of the strong correlation interval, and the strong correlation candidate node can be directly used as the primary classification. When the boundary strength indicator is lower than 0.5, the technical overlap area exceeds 50%. The classification needs to comprehensively consider the candidate classifications of strong and weak correlation intervals and achieve accurate classification through weighting. At this time, both strong and weak correlation nodes play an important role in the final classification result.

[0065] The classification parameters are constructed by applying differentiated weights to strongly correlated and weakly correlated intervals based on boundary strength indicators. The boundary strength indicator, serving as the base weight for strongly correlated intervals, directly reflects the classification confidence of strongly correlated candidate nodes. The base weight for weakly correlated intervals is 1 minus the boundary strength indicator, ensuring that the weights of the two intervals are complementary. This complementary weight mechanism guarantees that all candidate nodes are considered in the classification decision. The differentiated weight allocation adds the matching confidence adjustment of candidate nodes to the base weights. The average confidence of candidate nodes in strongly correlated intervals is higher than that in weakly correlated intervals, typically by 0.1 to 0.3. This confidence adjustment gives high-confidence nodes an extra boost in weight allocation. The weights of strongly correlated and weakly correlated intervals are normalized after confidence adjustment to ensure that the sum of the weights of the two intervals is 1, maintaining the integrity and consistency of the weight system. The classification parameters are defined as structured data of "strongly correlated candidate node set - strong correlation weight, weakly correlated candidate node set - weak correlation weight," clearly expressing the classification tendency and confidence level. The classification parameters of the patent documents are "{G06N 3 / 08,G06N 3 / 04}:0.876,{G10L 15 / 00}:0.124". The parameters clearly indicate the classification tendency and weight allocation. The strongly correlated candidate nodes "G06N 3 / 08" and "G06N 3 / 04" together have a weight of 0.876, which is dominant and reflects the core attribute of neural network technology. The weakly correlated candidate node "G10L 15 / 00" has a weight of 0.124, which indicates that speech recognition technology belongs to the auxiliary feature at the application scenario level in the patent.

[0066] Step S150: Cross-language retrieval is triggered by the core feature matching sequence and classification parameters to obtain the source language identifier. Based on the source language identifier, the corresponding semantic mapping rules are extracted from the semantic classification rule base to form a cross-language mapping relationship. The cross-language mapping relationship is semantically aligned with the core feature matching sequence to generate retrieval classification results.

[0067] In some embodiments, the step of triggering cross-language retrieval to obtain the source language identifier by the core feature matching sequence and the classification attribution parameter includes: extracting language feature vectors from the core feature matching sequence to form a feature vector set; identifying language type markers in the feature vector set; performing language matching between the language type markers and the classification attribution parameter to form a candidate language set; and confirming the language of the candidate language set to generate the source language identifier.

[0068] Extract language feature vectors from the core feature matching sequence to form a feature vector set. The label terms included in the 3-to-5-step strong rule chain of the core feature matching sequence carry language feature information, and the language features are reflected in the character encoding, character combination pattern, and lexical structure of the terms. The label term "neural network" uses UTF-8 encoding, the character combination pattern is continuous arrangement of Chinese characters, and the lexical structure conforms to the word formation rules of Chinese technical terms. The extraction of language feature vectors performs character-level n-gram analysis on each label term in the core feature matching sequence, and the n-gram window size is set to 2 to 4 characters. The 2-gram features of the label term "neural network" include "neural", "network", and "neural network", and the 3-gram features include "neural network" and "neural network". The feature vector set summarizes the n-gram features of all label terms in the core feature matching sequence, and the total number of features ranges from 50 to 200. The language feature vector calculates the occurrence frequency of each n-gram feature in the known language corpus, and the occurrence frequency constitutes the dimension value of the feature vector. When the core feature matching sequence contains Chinese terms "deep learning" and "feature extraction" and English term "neural network", the feature vector set contains both Chinese n-gram features and English n-gram features. The 2-gram features "deep", "learning", and "learning" of the Chinese term "deep learning" have a high matching rate in the Chinese corpus, and the character combination patterns "ne", "eu", and "ur" of the English term "neural network" have a high matching rate in the English corpus. The language attribution of the features is determined by corpus matching.

[0069] Identify language type tags in the feature vector set. The n-gram features of the feature vector set determine the language type attribution through the matching rate with the multilingual corpus. Features with a matching rate exceeding 80% are marked as the corresponding language type. The Chinese n-gram feature "neural" has a matching rate of 95% in the Chinese corpus and 0% in the English corpus, and is marked as a Chinese feature. The language type tags adopt the ISO 639-1 language code standard, with Chinese marked as "zh", English marked as "en", and Japanese marked as "ja". The 50 to 200 n-gram features of the feature vector set are marked with language types one by one, and the marking results form a language type frequency distribution. When the core feature matching sequence mainly contains Chinese terms, the proportion of features marked as "zh" in the feature vector set exceeds 70%. The frequency distribution of the language type tags presents a hierarchical structure of the dominant language and the secondary language, and the dominant language corresponds to the main expression language of the core feature matching sequence. When the patent document is written in Chinese but quotes English technical terms, the frequency distribution of the language type tags is 75% for "zh" and 25% for "en", and the dominant language is Chinese.

[0070] Implement language matching between language type tags and classification attribution parameters to form a candidate language set. The frequency distribution of language type tags provides a basis for ranking primary and secondary languages, and provides accurate language priority information for subsequent language matching. The candidate classification nodes in the classification attribution parameters are associated with the corresponding node identifiers of the node in different language classification systems, and the corresponding relationships are stored in a cross-language classification mapping table. The candidate classification node "G06N 3 / 08" corresponds to "Neural Network Computing" in the Chinese classification system, "Neural Network Computing" in the English classification system, and the corresponding Japanese terms in the Japanese classification system. The dominant language "zh" of the language type tag is matched with the language mapping of the candidate nodes in the classification attribution parameters. The matching process queries the cross-language classification mapping table. The cross-language classification mapping table maintains a complete mapping relationship of multi-language classification systems for each candidate classification node. The query operation uses the candidate node code "G06N 3 / 08" as the index key and returns the corresponding terms and classification paths of the node in each language classification system. The language matching result determines whether there is a corresponding node for the candidate classification node in the dominant language classification system, and the languages with corresponding nodes are included in the candidate language set. The candidate classification node "G06N 3 / 08" has corresponding nodes in the Chinese, English, and Japanese classification systems, and the candidate language set includes three languages: "zh", "en", and "ja". The secondary language features of the language type tag also participate in the matching, and the secondary languages are marked as auxiliary languages in the candidate language set. When the patent retrieval requirement is cross-language retrieval, all languages in the candidate language set are used as the target language range for retrieval.

[0071] Perform language confirmation on the candidate language set to generate a source language identifier. The multiple languages in the candidate language set determine the primary and secondary order through the frequency distribution of language type tags, and the language with the highest frequency is confirmed as the source language. When the frequency distribution of language type tags is 75% for "zh" and 25% for "en", the source language is confirmed as "zh". When the frequencies of two languages are close (the difference is less than 10%), the language confirmation introduces a weight adjustment mechanism for the classification attribution parameters. The language preference of strongly associated candidate nodes in the classification attribution parameters is used as the confirmation basis. For example, when the matching degree of a strongly associated node in the Chinese classification system is 0.85 and in the English classification system is 0.72, the source language is preferentially confirmed as Chinese. The source language identifier adopts a combination form of ISO 639-1 code plus region code. The Chinese source language identifier "zh-CN" represents Simplified Chinese, the English source language identifier "en-US" represents American English, and "en-GB" represents British English. The language confirmation process combines the language preference of strongly associated candidate nodes in the classification attribution parameters. When the matching degree of a strongly associated node in the Chinese classification system is higher than that in the English classification system, the source language tends to be confirmed as Chinese. The source language identifier records the confirmed language code and region code. The source language identifier "zh-CN" indicates that the source language is Simplified Chinese.

[0072] In some embodiments, extracting corresponding semantic mapping rules from the semantic classification rule library according to the source language identifier to form a cross-language mapping relationship includes: extracting initial semantic mapping rules from the semantic classification rule library according to the source language identifier; performing sparsity analysis on the initial semantic mapping rules to form sparse region identifiers; performing rule compensation screening based on the sparse region identifiers to form candidate mapping rules; and concatenating the candidate mapping rules step by step according to semantic integrity to form a cross-language mapping relationship.

[0073] Extract initial semantic mapping rules from the semantic classification rule library according to the source language identifier. The source language identifier determines the language range for rule extraction and guides the selection of the corresponding language rule set from the partitioned rule library for matching. The association rules in the semantic classification rule library are stored in partitions according to language types, and the Chinese rule set, English rule set, and Japanese rule set correspond to different language technical term systems respectively. The source language identifier "zh-CN" triggers the query operation of the Chinese rule set, and the query condition is the label match between the rule antecedent label and the label of the core feature matching sequence. When the core feature matching sequence contains the labels "neural network" and "deep learning", the rules "IF neural network THEN artificial intelligence (0.78)" and "IF deep learning THEN machine learning (0.82)" are matched in the Chinese rule set. The initial semantic mapping rules extract all the matched rules, and the number of rules ranges from 5 to 20. The rule set covers the main technical paths of the core feature matching sequence. The rules in the semantic classification rule library include two types: intra-language mapping and cross-language mapping. The antecedent and consequent labels of the intra-language mapping rules belong to the same language, and the antecedent and consequent labels of the cross-language mapping rules belong to different languages. The initial semantic mapping rules preferentially extract cross-language mapping rules, such as the mapping rule "neural network → Neural Network (0.95)" where the Chinese label "neural network" corresponds to the English label "Neural Network".

[0074] Sparsity analysis was performed on the initial semantic mapping rules to form sparse region identifiers. The coverage of the initial 5 to 20 semantic mapping rules in the technical topic space was evaluated by the distribution density of the rule consequent labels; regions with low distribution density were identified as sparse regions. The technical topic space adopted 25 topic clusters from the topic clustering space as a reference framework, and the consequent labels of the initial semantic mapping rules were mapped to the distribution statistics of the topic clusters. Topic cluster 1 "Data Retrieval" contains two consequent labels: "Information Retrieval" and "Query Processing"; topic cluster 2 "Neural Networks" contains two consequent labels: "Deep Learning" and "Convolutional Network"; and topic cluster 3 "Speech Processing" contains no consequent labels. A sparse region identifier is defined as a set of topic clusters with zero or only one consequent label; topic cluster 3 "Speech Processing" was identified as a sparse region. Among the 25 topic clusters of the initial semantic mapping rules, the sparse region identifiers cover 8 to 12 topic clusters. The existence of sparse regions indicates that the initial semantic mapping rules do not adequately cover some technical topics. These sparse topics may not receive effective mapping support in cross-language retrieval, leading to the omission of relevant classification nodes in the search results. Rule compensation is needed to improve the mapping relationship. The sparse region identifier records the number of each sparse topic cluster and the number of current consequent tags.

[0075] Candidate mapping rules are formed through rule compensation screening based on sparse region identifiers. The sparse region identifiers indicate the scope of technical topics requiring compensation, and rule compensation screening searches for mapping rules in relevant technical fields based on these identifiers. Rule compensation filters mapping rules related to sparse region topics from the full rule set of the semantic classification rule base, with the filtering condition being that the rule's consequent label belongs to the technical field of the sparse topic cluster. For sparse topic cluster 3, "speech processing," the compensation rule screening searches the rule base for rules whose consequent labels contain the keywords "speech," "audio," and "acoustics." The screening process simultaneously checks whether the rule's antecedent label has a semantic association with the label of the core feature matching sequence. For example, if the antecedent label "feature extraction" is adjacent to the "neural network" label in the core feature matching sequence on the technical path, it satisfies the semantic association condition. The compensation screening yields two rules: "IF Feature Extraction THEN Speech Features (0.65)" and "IF Signal Processing THEN Audio Signal (0.68)." These rules fill the gap in the initial semantic mapping rules for the speech processing field. The candidate mapping rules are a compilation of the initial semantic mapping rules and compensation filtering rules, expanding the total number of rules from the initial 5-20 to 12-35. The confidence threshold for rule compensation filtering is set lower than the initial rule extraction threshold, allowing rules with medium confidence (0.5-0.7) to be added to the candidate set. The candidate mapping rules in the patent document "Neural Network Speech Recognition Method" simultaneously include high-confidence rules from the neural network field and medium-confidence compensation rules from the speech processing field, significantly expanding the technical coverage of the rule set.

[0076] Candidate mapping rules are sequentially concatenated according to semantic completeness to form cross-linguistic mapping relationships. The 12th to 35th candidate mapping rules are sorted by the semantic distance between the rule's consequent label and the core feature matching sequence, with rules closer in semantic distance appearing first. Semantic completeness is defined as the degree to which the rule chain covers all labels of the core feature matching sequence; a rule chain with high completeness contains fewer rules and has strong semantic coherence. For sequential concatenation, the first rule from the candidate mapping rules is selected as the starting node, with the rule "Neural Network → Neural Network (0.95)" serving as the first-level mapping. The consequent label "Neural Network" of the first-level rule serves as a candidate antecedent label for the second-level rule, and rules with the antecedent label "Neural Network" are matched from the candidate mapping rules. The second-level match finds the rule "Neural Network → Deep Learning (0.88)," extending the rule chain to "Neural Network → Neural Network → Deep Learning." The sequential concatenation process continues until the rule chain covers all labels of the core feature matching sequence or no subsequent matching rule can be found. The concatenation terminates under two conditions: first, the current rule chain has covered all labels of the core feature matching sequence, indicating a complete mapping path; second, there are no rules among the candidate mapping rules that match the current consequent label, indicating an interrupted mapping path. Cross-language mapping relationships form multiple parallel rule chains, each corresponding to a possible cross-language mapping path. The cross-language mapping relationship corresponding to the core feature matching sequence "Neural Network → Deep Learning → Speech Recognition" in the patent document includes the rule chain "Neural Network → Deep Learning → Speech Recognition". The number of rule chains in the cross-language mapping relationship is 2 to 5, and the average confidence of the rule chain is calculated as the geometric mean of the confidence of all rules in the chain.

[0077] Semantic alignment is performed between cross-language mapping relationships and core feature matching sequences to generate retrieval classification results. The 2nd to 5th rule chains of the cross-language mapping relationships are aligned with the 3rd to 5th strong rule chains of the core feature matching sequences. The alignment process matches the semantic equivalence between the nodes of the matching rule chains and the labels of the strong rule chains. An alignment mapping is established between the label "Neural Network" of the core feature matching sequence and the node "Neural Network" of the cross-language mapping relationship rule chain, recorded as "Neural Network ⟷ Neural Network". The semantic alignment matching success rate is defined as the proportion of successfully aligned node pairs to the total number of labels in the core feature matching sequence. Rule chains with a matching success rate exceeding 80% are retained. When the core feature matching sequence contains 4 labels and the cross-language mapping relationship rule chain successfully aligns 3 labels, the matching success rate is 75%. This rule chain does not meet the retention threshold and is excluded from the retrieval classification results. Alignment failure usually occurs because the semantic path of the rule chain deviates from the technical topic of the core feature matching sequence; for example, the rule chain contains labels from technical fields not covered by the core feature matching sequence. The retrieval classification results extract the final cross-language classification nodes from the preserved rule chains. Each classification node corresponds to the classification code of the endpoint label of the rule chain in the target language classification system. For example, the endpoint label "SpeechRecognition" of the cross-language mapping rule chain "Neural Network → Deep Learning → Speech Recognition" corresponds to "G10L 15 / 00" in the English IPC classification system. The retrieval classification results summarize all endpoint classification nodes of the preserved rule chains, with 1 to 3 nodes. The confidence level of each node is inherited from the average confidence level of the rule chain. The retrieval classification results for the patent document "A Speech Recognition Method Based on Neural Networks" are "G06N 3 / 08 (Neural Network, 0.88)" and "G10L 15 / 00 (Speech Recognition, 0.76)", which includes classifications for both the Chinese source language and the English target language. The retrieval classification results combine the weighted allocation of classification parameters, assigning higher ranking priority to strongly correlated candidate classification nodes in the retrieval results.

[0078] To implement the above-described method embodiments, a semantic understanding-based intelligent patent document classification method is proposed to achieve the corresponding functions and technical effects. See also... Figure 2 , Figure 2 This diagram illustrates a structural block diagram of a semantic understanding-based intelligent patent document classification system 200 provided in an embodiment of this application. For ease of explanation, only the parts relevant to this embodiment are shown. The semantic understanding-based intelligent patent document classification system 200 provided in this embodiment includes:

[0079] Data acquisition module 201 is used to acquire patent document data and classification system configuration parameters, and to perform semantic parsing based on the patent document data and the classification system configuration parameters to construct a semantic classification index library;

[0080] Feature extraction module 202 is used to extract core semantics based on the semantic classification index library to form a semantic vector space, identify semantically ambiguous feature points with the help of the semantic vector space, and construct a topic clustering space based on the semantically ambiguous feature points;

[0081] Rule generation module 203 is used to perform label matching analysis on the topic clustering space to determine the technical topic label set, transform the technical topic label set into a classification knowledge system to form a semantic classification rule base, and generate a classification matching strategy based on the semantic classification rule base;

[0082] The classification processing module 204 is used to decompose the classification matching strategy into a core feature matching sequence and an auxiliary feature matching sequence, perform semantic difference analysis on the auxiliary feature matching sequence to extract a candidate classification set, and perform domain attribute extraction on the candidate classification set to form a classification attribution parameter.

[0083] The result output module 205 is used to trigger cross-language retrieval to obtain the source language identifier through the core feature matching sequence and the classification attribution parameter, extract the corresponding semantic mapping rule from the semantic classification rule base according to the source language identifier to form a cross-language mapping relationship, and perform semantic alignment between the cross-language mapping relationship and the core feature matching sequence to generate retrieval classification results.

[0084] The aforementioned intelligent patent document classification system 200 based on semantic understanding can implement one of the intelligent patent document classification methods based on semantic understanding described in the above method embodiments. The options in the above method embodiments are also applicable to this embodiment and will not be detailed here. The remaining content of this application embodiment can be referred to the content of the above method embodiments, and will not be repeated in this embodiment.

[0085] The purpose of the above embodiments is to reproduce and derive the technical solution of the present invention by way of example, and to fully describe the technical solution, purpose and effect of the present invention. The purpose is to enable the public to have a more thorough and comprehensive understanding of the disclosure of the present invention, and not to limit the scope of protection of the present invention.

[0086] The above embodiments are not an exhaustive list based on the present invention, and there may be many other embodiments not listed. Any substitutions and improvements made without departing from the concept of the present invention are within the protection scope of the present invention.

Claims

1. A method for intelligent classification of patent documents based on semantic understanding, characterized in that, include: Obtain patent literature data and classification system configuration parameters, and construct a semantic classification index library based on the patent literature data and the classification system configuration parameters; Based on the semantic classification index, core semantic extraction is performed to construct a semantic vector space. The semantic vector space is then used to identify semantically ambiguous feature points, including: dividing the semantic vector space into a main feature subspace and an auxiliary feature subspace; performing bidirectional semantic similarity detection on the main feature subspace and the auxiliary feature subspace to form a forward detection path and a reverse detection path; comparing the semantic differences between the forward detection path and the reverse detection path to form a bidirectional conflict point list; identifying points in the bidirectional conflict point list that exhibit semantic deviation in both directions as semantically ambiguous feature points; and constructing a topic clustering space based on the semantically ambiguous feature points. The step of comparing the semantic differences between the forward detection path and the reverse detection path to form the bidirectional conflict point list includes: extracting semantic similarity change sequences based on the forward detection path and the reverse detection path; performing gradient analysis on the semantic similarity change sequences to determine gradient abrupt change locations; extracting semantic feature vectors at each abrupt change location; and arranging the semantic feature vectors according to the difference intensity to generate the bidirectional conflict point list. The process of performing label matching analysis on the topic clustering space to determine the technical topic label set includes: performing overlap analysis on the labels of each cluster topic in the topic clustering space to generate label intersection areas; evaluating the technical boundaries based on the label intersection areas to form a boundary clarity index; prioritizing the labels based on the boundary clarity index to generate a candidate label set; performing disambiguation processing on the candidate label set to generate a technical topic label set; transforming the technical topic label set into a classification knowledge system to form a semantic classification rule base; and generating a classification matching strategy based on the semantic classification rule base. The classification matching strategy is decomposed into a core feature matching sequence and an auxiliary feature matching sequence. Semantic difference analysis is performed on the auxiliary feature matching sequence to extract candidate classification sets. Domain attribute extraction is performed on the candidate classification sets to form classification parameters. The source language identifier is obtained by triggering cross-language retrieval through the core feature matching sequence and the classification attribution parameter. Based on the source language identifier, the corresponding semantic mapping rule is extracted from the semantic classification rule base to form a cross-language mapping relationship. The cross-language mapping relationship is semantically aligned with the core feature matching sequence to generate retrieval classification results.

2. The method according to claim 1, characterized in that, The step of constructing a semantic classification index library based on the patent document data and the classification system configuration parameters includes: Generate a set of text segmentation units based on the patent document data; Information density analysis is performed on the text segmentation unit set to generate a density distribution map; Identify high-density semantic regions within the density distribution map; A semantic classification index library is constructed by mapping the high-density semantic regions to the classification system configuration parameters.

3. The method according to claim 1, characterized in that, The step of extracting domain attributes from the candidate classification set to form classification parameters includes: Construct an attribute feature space based on the candidate classification set; Domain association identifiers are formed by mapping domain association degree through the attribute feature space; The attribute feature space is divided into a strong association interval and a weak association interval, with the domain association identifier as the boundary; The attribute distribution characteristics of the strongly correlated intervals and the weakly correlated intervals are compared to form the classification and attribution parameters.

4. The method according to claim 1, characterized in that, The step of obtaining the source language identifier by triggering cross-language retrieval through the core feature matching sequence and the classification attribution parameter includes: The core feature matching sequence is subjected to language feature vector extraction to form a feature vector set; Language type markers are identified in the feature vector set; The language type label is matched with the classification parameters to form a candidate language set; The candidate language set is used to confirm the language and generate source language identifiers.

5. The method according to claim 1, characterized in that, The step of extracting corresponding semantic mapping rules from the semantic classification rule base based on the source language identifier to form a cross-language mapping relationship includes: Based on the source language identifier, initial semantic mapping rules are extracted from the semantic classification rule base; Sparsity analysis is performed on the initial semantic mapping rules to form sparse region identifiers; Candidate mapping rules are formed by rule compensation and filtering based on the sparse region identifiers; The candidate mapping rules are concatenated step by step according to semantic completeness to form a cross-language mapping relationship.

6. The method according to claim 3, characterized in that, The attribute distribution features of the strongly correlated intervals and the weakly correlated intervals constitute the classification and attribution parameters, including: A dual-interval density map is formed by mapping the attribute distribution density of the strongly correlated intervals and the weakly correlated intervals. The cross-region of attribute distribution is determined based on the dual-interval density map. Boundary ambiguity assessment is performed on the intersection regions of the attribute distributions to construct boundary strength identifiers; Based on the boundary strength identifier, differentiated weighting ratios are applied to the strongly correlated intervals and the weakly correlated intervals to form classification parameters.

7. A patent document intelligent classification system based on semantic understanding, characterized in that, include: The data acquisition module is used to acquire patent document data and classification system configuration parameters, and to perform semantic parsing based on the patent document data and the classification system configuration parameters to construct a semantic classification index library; The feature extraction module is used to perform core semantic extraction based on the semantic classification index to construct a semantic vector space, and to identify semantically ambiguous feature points using the semantic vector space. This includes: dividing the semantic vector space into a main feature subspace and an auxiliary feature subspace; performing bidirectional semantic similarity detection on the main feature subspace and the auxiliary feature subspace to form a forward detection path and a reverse detection path; comparing the semantic differences between the forward detection path and the reverse detection path to form a bidirectional conflict point list; labeling points in the bidirectional conflict point list that exhibit semantic deviation in both directions as semantically ambiguous feature points; and constructing a topic clustering space based on the semantically ambiguous feature points. The step of comparing the semantic differences between the forward detection path and the reverse detection path to form a bidirectional conflict point list includes: extracting semantic similarity change sequences based on the forward detection path and the reverse detection path; performing gradient analysis on the semantic similarity change sequences to determine gradient abrupt change locations; extracting semantic feature vectors at each abrupt change location; and arranging the semantic feature vectors according to the difference intensity of the semantic feature vectors to generate a bidirectional conflict point list. The rule generation module is used to perform label matching analysis on the topic clustering space to determine the technical topic label set, including: performing overlap analysis on the labels of each cluster topic in the topic clustering space to generate label intersection areas; evaluating the technical boundaries based on the label intersection areas to form a boundary clarity index; prioritizing the labels based on the boundary clarity index to generate a candidate label set; performing disambiguation processing on the candidate label set to generate a technical topic label set; converting the technical topic label set into a classification knowledge system to form a semantic classification rule base; and generating a classification matching strategy based on the semantic classification rule base. The classification processing module is used to decompose the classification matching strategy into a core feature matching sequence and an auxiliary feature matching sequence, perform semantic difference analysis on the auxiliary feature matching sequence to extract a candidate classification set, and perform domain attribute extraction on the candidate classification set to form classification attribution parameters. The result output module is used to trigger cross-language retrieval to obtain the source language identifier by the core feature matching sequence and the classification attribution parameter, extract the corresponding semantic mapping rule from the semantic classification rule base according to the source language identifier to form a cross-language mapping relationship, and perform semantic alignment between the cross-language mapping relationship and the core feature matching sequence to generate retrieval classification results.

Citation Information

Patent Citations

  • RAG intelligent retrieval question-answering system and method based on enhanced metadata

    CN121256007A

  • Text classification method and system based on semantic analysis

    CN121327136A