Knowledge graph construction method and device for intellectual property retrieval and storage medium

By mining implicit associations in intellectual property text datasets, generating weakly supervised signals, and automatically labeling entities and semantic relationships, the high cost and bias issues caused by external annotation are resolved, thereby improving the accuracy and efficiency of intellectual property retrieval.

CN121745264APending Publication Date: 2026-03-27HENAN UNIV OF ANIMAL HUSBANDRY & ECONOMY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-13
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies in the field of intellectual property rely on externally labeled supervisory data, resulting in high costs and uncontrollable biases. They also struggle to fully cover the multi-dimensional semantic relationships between texts, impacting retrieval efficiency and accuracy.

Method used

By mining implicit associations in target intellectual property text datasets, weakly supervised signals are generated, cross-validation and deduplication are performed, entities and semantic relationships are automatically labeled, and a knowledge graph is constructed by combining semantic consistency verification and domain adaptability screening.

Benefits of technology

It eliminates the need for external annotation resources, saves costs, improves the professionalism and accuracy of knowledge representation, and enhances the precision and response efficiency of intellectual property retrieval.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121745264A_ABST
    Figure CN121745264A_ABST
Patent Text Reader

Abstract

The invention provides an intellectual property retrieval-oriented knowledge graph construction method and device and a storage medium, and the method comprises the steps: reading a target intellectual property text data set and a parallel corpus and a reference associated text in the target intellectual property text data set, mining synonymous mapping clues, reference traceability clues and technical theme associated clues implied in the texts, and constructing the intellectual property retrieval-oriented knowledge graph by the synonymous mapping clues, the reference traceability clues and the technical theme associated clues; sorting to obtain a weak supervision signal set, extracting synonymous expression pairs in parallel corpora and semantic association pairs in a reference association text, carrying out cross validation and duplicate removal to obtain a text alignment reference set, and carrying out global association matching on the text alignment reference set and the target intellectual property text data set to obtain an initial structured knowledge unit set; the method comprises the steps of obtaining a standardized knowledge unit set, performing credibility regularization to obtain a standardized knowledge unit set, performing entity classification and relation association organization according to hierarchical requirements of intellectual property retrieval, and constructing to obtain the intellectual property retrieval-oriented knowledge graph. The intellectual property retrieval accuracy and response efficiency can be effectively improved through the intellectual property retrieval method and device.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of knowledge graph technology, and in particular to a method, apparatus and storage medium for constructing a knowledge graph for intellectual property retrieval. Background Technology

[0002] In the field of intellectual property, organizing textual information in a structured form can provide precise correlation support for intellectual property information retrieval. Existing technologies typically rely on externally annotated supervisory data to drive key construction stages. Some solutions only mine clues for specific categories of correlation information within intellectual property texts. In the entity and relationship labeling stage, they often directly extract information from the target text. Knowledge unit organization often only performs verification processes for specific dimensions. Finally, in the graph organization stage, they are often classified and integrated according to the inherent technical attributes of entities. This construction method not only incurs high manpower and time costs due to reliance on external annotation but may also introduce uncontrollable biases due to the professional limitations of the annotators. Mining only specific categories of correlation information is insufficient to comprehensively cover the multi-dimensional semantic relationships between texts; directly extracting entities and relationships is easily affected by differences in textual expression, leading to insufficient correlation accuracy; performing only specific-dimensional verification cannot guarantee that the knowledge content conforms to the professional standards of the intellectual property field; and organizing the graph according to inherent technical attributes is difficult to meet the actual needs of the retrieval scenario, ultimately affecting the efficiency and experience of the retrieval process. Summary of the Invention

[0003] In view of this, the present invention provides a method, apparatus, and storage medium for constructing a knowledge graph for intellectual property retrieval. The technical solution of the embodiments of the present invention is implemented as follows: On one hand, embodiments of the present invention provide a knowledge graph construction method for intellectual property retrieval. The method includes: reading a target intellectual property text dataset and its inherent parallel corpus and reference-related texts; mining implicit synonym mapping clues, reference tracing clues, and technical topic association clues between texts; and organizing these into a weakly supervised signal set, which contains various implicit marker information used to indicate semantic associations between texts; based on the weakly supervised signal set, extracting synonym pairs from the parallel corpus and semantic association pairs from the reference-related texts; performing cross-validation and deduplication on the extracted synonym pairs and semantic association pairs to obtain a text alignment benchmark set; and performing global association matching between the text alignment benchmark set and the target intellectual property text dataset. The process involves automatically associating and marking intellectual property entities and their semantic relationships within a text by mapping and associating them using a text alignment benchmark set. This yields an initial set of structured knowledge units, which includes entity and relationship representations that have not undergone credibility verification. The initial set of structured knowledge units is then normalized for credibility, and entities and relationship representations that meet preset credibility requirements are retained through semantic consistency verification and domain adaptability filtering, resulting in a standardized set of knowledge units. Finally, the standardized set of knowledge units is organized into entity categories and relationship associations according to the hierarchical requirements of intellectual property retrieval, constructing a knowledge graph for intellectual property retrieval. This knowledge graph includes a technical topic hierarchy adapted to retrieval requirements, entity association links, and semantic mapping paths.

[0004] On the other hand, embodiments of the present invention provide a knowledge graph construction apparatus, including a memory and a processor. The memory stores a computer program that can run on the processor, and the processor executes the program to implement the steps in the above-described method.

[0005] Thirdly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the methods described above.

[0006] The knowledge graph construction method for intellectual property retrieval provided by this invention does not rely on external annotation resources. It only mines the associated text inherent in the target dataset to generate weak supervision signals, which saves annotation costs and deeply reuses the data's own association information. The text alignment benchmark set generated by cross-validation and deduplication integration can effectively avoid semantic conflicts and provide accurate references for subsequent labeling. By using the mapping association of alignment benchmarks to deduce automatically labeled entities and semantic relationships, the semantic bias caused by direct extraction can be reduced. The standardized knowledge units, which are verified by semantic consistency and selected by domain adaptability, can improve the professionalism and accuracy of knowledge representation. Finally, the knowledge graph constructed according to the hierarchical requirements of retrieval can fit the retrieval logic and effectively improve the accuracy and response efficiency of intellectual property retrieval. Attached Figure Description

[0007] Figure 1 This is a schematic diagram illustrating the implementation process of a knowledge graph construction method for intellectual property retrieval provided in an embodiment of the present invention.

[0008] Figure 2 This is a schematic diagram of the hardware entity of a knowledge graph construction device provided in an embodiment of the present invention. Detailed Implementation

[0009] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on the present invention. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0010] This invention provides a method for constructing a knowledge graph for intellectual property retrieval, which can be executed by the processor of a knowledge graph construction device. The knowledge graph construction device can refer to a device with data processing capabilities, such as a server, laptop, tablet, or desktop computer.

[0011] Figure 1 This is a schematic diagram illustrating the implementation process of a knowledge graph construction method for intellectual property retrieval provided in an embodiment of the present invention, as shown below. Figure 1 As shown, the method includes: Step S100: Read the target intellectual property text dataset and its built-in parallel corpus and reference-related texts, and mine the implicit synonym mapping clues, reference tracing clues and technical topic association clues between texts. Organize them to obtain a weak supervision signal set, which contains various implicit marker information used to indicate the semantic association of texts.

[0012] Target intellectual property text datasets are collections of texts containing various intellectual property-related information, such as patent document databases. Parallel corpora are text content expressing similar or identical semantics in different texts. For example, different patent documents may use different expressions for the same technical feature, but their core semantics are consistent; these different expressions constitute parallel corpora. Citation-related texts are portions of text where citation relationships exist. For instance, if one patent document cites a technical solution from another patent document, then there is a citation relationship between the two documents, and the related text content is the citation-related text.

[0013] Synonym mapping clues are clues that reveal synonymous relationships between texts. By exploring these clues, we can find content that is expressed differently but has the same meaning in different texts. For example, in different patent documents, "automatic control device" and "automated control equipment" may express the same technical concept; this synonymous relationship is discovered through synonym mapping clues. Citation tracing clues are clues used to trace the source of text citations. By analyzing these clues, we can clarify the citation transmission chain between texts. For example, if a patent document cites multiple other documents, citation tracing clues can find the source and transmission path of these citations. Technical topic association clues are clues that reveal the technical topic connections between different texts. For example, if different patent documents simultaneously mention "artificial intelligence image recognition technology," then there is a technical topic association between these texts, and relevant clues can help identify this association.

[0014] For example, step S100 may specifically include the following steps S110 to S160: Step S110: Read the target intellectual property text dataset and its built-in parallel corpus and referenced text. According to the original paragraph division marks of the text, classify the parallel corpus and referenced text into independent storage partitions. The text in each partition retains the original page number, paragraph number and reference mark information, and obtains the partitioned storage of parallel corpus partition text set and referenced text partition text set.

[0015] In this embodiment of the invention, reading the target intellectual property text dataset and its parallel and cited texts is for preliminary data separation and organization. The original paragraph division markers are structural information inherent in the text itself; for example, in patent documents, texts are typically divided by chapters and paragraphs. These markers help to accurately locate and process text content. Separating parallel texts and cited texts into independent storage partitions facilitates subsequent targeted processing of different types of text. These storage partitions can be different tables in a database or different folders in a file system.

[0016] The text within each partition retains its original page numbers, paragraph numbers, and citation markers. This information is crucial for subsequent text analysis and correlation. Page numbers and paragraph numbers help accurately locate the text within the original document, while citation markers help trace the text's sources. For example, in a patent document database, text containing parallel corpora is stored in a table called "Parallel Corpus Partition," and text containing cited related text is stored in a table called "Citation Related Text Partition." Each table records the text's page numbers, paragraph numbers, and citation markers.

[0017] Step S120: Traverse each text segment in the parallel corpus partitioned text set, compare it sentence by sentence with all other texts in the set, identify sentence combinations that have the same core meaning but different sentence structures, add cross-text association tags to each sentence combination, and obtain a set of labeled parallel corpus sentence combinations.

[0018] In this embodiment of the invention, traversing each segment of text in the parallel corpus partitioned text set is for the purpose of comprehensive processing of all texts in the set. Sentence-by-sentence content comparison involves splitting each segment of text into sentences, then comparing each sentence with all other sentences in the set to identify sentence combinations that express the same core meaning but differ in sentence structure. Consistent core meaning means that the core content expressed by the sentences, such as the technical meaning, applicable scenario, and scope, is the same, while differences in sentence structure refer to differences in structural elements such as the subject position, predicate type, and modifier order.

[0019] For example, step S120 may specifically include the following steps S121 to S126: Step S121: Read the labeled set of parallel corpus sentence combinations, extract the content of the two sentences in each sentence combination and the source information of the parallel corpus text to which they belong, split the sentence content into a continuous sequence of individual words, retain the original word order and part-of-speech information of the words, and obtain a set of sentence word sequences.

[0020] In this embodiment of the invention, reading the labeled set of parallel corpus sentence combinations is to obtain the previously processed sentence combinations containing labeled information. Extracting the content of the two sentences in each sentence combination and the source information of the parallel corpus text to which they belong is to clarify the specific content and source of each sentence. Breaking the sentence content into a continuous sequence of individual words is to perform a more detailed analysis of the sentences. By preserving the original word order and part-of-speech information of the words, the semantics and structure of the sentences can be better understood.

[0021] Step S122: Perform word-by-word matching on the two word sequences in each sentence word sequence set, mark the completely overlapping words and their positions in the sequence, count the proportion of overlapping words to the total number of words in the two sentences, record the continuous distribution of matching words, and obtain the sentence word matching detail set.

[0022] In this embodiment of the invention, word-by-word matching is performed on the word sequences of two clauses in each clause word sequence set to find identical words in the two clauses. Marking completely overlapping words and their positions in the sequences clarifies the specific location of the identical words in the clauses. Statistical analysis of the proportion of overlapping words to the total number of words in the two clauses helps determine the similarity between the two clauses; a higher proportion indicates a higher similarity. Recording the continuous distribution of matched words analyzes the distribution pattern of identical words in the clauses, such as whether there are consecutive segments of identical words. In practice, this can be achieved by iterating through the word sequences of the two clauses and comparing whether the words at each position are identical. The positions of overlapping words can be recorded using lists or dictionaries. The proportion of overlapping words can be calculated by counting the number of overlapping words and the total number of words in the two clauses. Recording the continuous distribution of matched words can be implemented using pointers or state machines; when consecutive matching words are encountered, their start and end positions are recorded.

[0023] Step S123: For sentence combinations in the sentence word matching detail set where the continuous distribution ratio of matching words meets the requirements, restore the complete semantic expression of the two sentences, compare the technical meaning, applicable scenarios, and limited scope expressed by the two, confirm that the core expression is completely consistent, and obtain a semantically consistent sentence combination subset.

[0024] For example, step S123 may specifically include the following steps S1231 to S1236: Step S1231: Read the sentence combinations in the sentence word matching detail set where the continuous distribution ratio of matching words meets the requirements, extract the complete text content of the two sentences in each combination, including the contextual auxiliary expressions before and after the sentences, and obtain the sentence content set with context.

[0025] In this embodiment of the invention, reading sentence combinations whose continuous distribution ratio of matching words in the sentence word matching detail set meets the requirements is to filter out sentence combinations that meet the conditions for further processing. Extracting the complete text content of the two sentences in each combination, including the contextual auxiliary expressions before and after the sentences, is because the semantic understanding of sentences often requires the combination of contextual information, and contextual auxiliary expressions can help to more accurately understand the complete semantics of the sentences.

[0026] Step S1232: Semantically restore the content of each clause with context, sort out the core information such as the technical viewpoint, the object of description, and the implementation method expressed by the clause, and organize the core information into structured semantic description items according to a fixed logical order to obtain a set of structured semantic descriptions of the clauses.

[0027] In practical implementation, semantic reconstruction can utilize natural language processing techniques, such as semantic role labeling and event extraction. Semantic role labeling identifies the various semantic roles in a sentence, such as agent, patient, and instrument, leading to a better understanding of the sentence's semantics. Event extraction extracts information about the events described in the sentence, such as the subject, object, and action of the event. For organizing core information, information extraction algorithms can be used, such as rule-based or machine learning-based algorithms. Finally, data structures, such as lists and dictionaries, can be used to organize the core information into structured semantic descriptions.

[0028] Step S1233: Compare the structured semantic description entries of the two clauses in each clause combination item by item, check whether the contents of each core field such as technical viewpoint, description object, and implementation method are completely identical, record the comparison results of each field, and obtain the semantic field comparison result set.

[0029] For example, consider two clause combinations. One clause's structured semantic description entry is "intelligent sensors - capable of improving production efficiency - employing novel algorithms," and the other clause's structured semantic description entry is "intelligent sensors - capable of improving production efficiency - utilizing advanced algorithms." During comparison, the description object and technical viewpoint fields completely overlap, but the implementation method fields differ. In practical implementation, a loop can be used to iterate through the two structured semantic description entries in each clause combination, comparing each core field one by one. During comparison, string comparison methods can be used to determine if the field content is the same. The comparison results for each field can be recorded using a list or dictionary; for example, a dictionary can be used to store the comparison results.

[0030] Step S1234: For sentence combinations where all core fields in the semantic field comparison result set are completely overlapping, further check whether there are semantic conflicts in the contextual auxiliary expressions of the sentences, confirm that the contextual expressions do not affect the consistency of the core semantics, and obtain a context-compatible sentence combination subset.

[0031] In practical implementation, semantic reasoning techniques, such as rule-based semantic reasoning or knowledge graph-based semantic reasoning, can be used to check for semantic conflicts in contextual auxiliary expressions. By analyzing the semantic information in the contextual expressions and comparing it with the core semantics, conflicts can be determined. After confirming that the contextual expressions do not affect the consistency of the core semantics, sentence combinations that meet the conditions are selected to obtain a context-compatible subset of sentence combinations.

[0032] Step S1235: Add a semantic consistency confirmation mark to each combination in the context-compatible sentence combination subset. The mark content includes the core field comparison result, the context compatibility confirmation result, and the semantic consistency judgment basis, to obtain a sentence combination set with semantic confirmation marks.

[0033] For example, the tag content could be: "Core field comparison result: Technical viewpoints, described objects, and implementation methods are all consistent; Context compatibility confirmation result: No conflict in contextual expressions; Semantic consistency judgment criteria: Core fields are completely overlapping and context is compatible." In practical implementation, data structures such as dictionaries or lists can be used to store the tag content. For each context-compatible sentence combination, the tag content is added to the corresponding record.

[0034] Step S1236: Sort and organize the set of sentence combinations with semantic confirmation tags according to the text length of the sentences, filter out the sentence combinations with complete core semantic expression, and finally generate a semantically consistent subset of sentence combinations.

[0035] For example, after sorting, some sentence combinations may be found to be semantically consistent but shorter in length, potentially lacking key information, while others are longer and contain a more complete core semantic expression. The latter should be selected. In practical implementation, a sorting algorithm can be used to rank the set of sentence combinations with semantic confirmation tags. Selecting sentence combinations with complete core semantic expressions can be based on evaluation criteria such as text length and semantic completeness. For example, a text length threshold can be set to filter out sentence combinations exceeding that threshold.

[0036] Step S124: Perform sentence structure analysis on each clause combination in the semantically consistent clause combination subset, identify structural elements such as the subject position, predicate type, and modifier order of the clause, compare the differences in structural elements between the two clauses, mark the clause combinations with different sentence structures, and obtain the clause combination subset with sentence structure differences.

[0037] For example, consider two semantically identical clauses: "The automatic control device realizes the intelligent adjustment function" and "The intelligent adjustment function is realized by the automatic control device." The subject positions differ; the subject of the first clause is "automatic control device," while the subject of the second clause is "intelligent adjustment function." In practice, sentence structure analysis can utilize syntactic analysis techniques, such as dependency parsing or constituent parsing. Dependency parsing analyzes the dependency relationships between words in a sentence, thereby determining the positions of structural elements such as the subject and predicate. Constituent parsing analyzes a sentence into different syntactic components, such as the subject, predicate, object, and modifiers, thus clarifying the sentence structure. To compare the differences in structural elements between two clauses, a contrastive algorithm can be used to compare each structural element one by one, identifying those with differences.

[0038] Step S125: Add a cross-text synonym association tag to each sentence combination in the subset of sentence combination combinations with sentence structure differences. The tag content includes the text source to which the sentence belongs, the semantic consistency confirmation result, and the sentence structure difference type, resulting in a set of sentence combination combinations with synonym tags. For example, the tag content can be "Text source: Patent document A, Patent document B; Semantic consistency confirmation result: Core semantics are consistent; Sentence structure difference type: Subject position is different". In specific implementation, data structures such as dictionaries or lists can be used to store the tag content. For each sentence combination with sentence structure differences, the tag content is added to the corresponding record.

[0039] Step S126: Group the sentence combination set with synonym tags according to the text source of the parallel corpus. Each group corresponds to the parallel corpus text from the same source. Finally, update and generate the tagged parallel corpus sentence combination set.

[0040] In practical implementation, grouping algorithms can be used to group the set of sentence combinations with synonym tags that differ in sentence structure. For example, dictionaries in Python can be used to implement grouping.

[0041] Step S130: Traverse each text segment in the text set of reference-related text partitions, trace the source of the reference marked in the text, sort out the reference transmission links between all texts, add source-tracing association tags to the starting text and ending text of each link, and obtain a set of tagged reference-related links.

[0042] For example, if one patent document cites another, and that cited document cites yet another, the entire citation chain can be identified by tracing the source of the citations and tracing the citation propagation path. Adding source-tracing association markers to the starting and ending texts of each link is to mark the citation source relationship between these texts, facilitating subsequent use and querying. The marker content can include information about the citation chain, the source of the starting and ending texts, etc. In practical implementation, a loop structure can be used to traverse the text, processing each segment. Tracing the source of citations can be done by parsing the citation annotation information in the text. For example, in patent documents, there are usually explicit citation annotation formats, such as "[Reference 1]", etc. By parsing this annotation information, the source of the citation can be found. Tracing the citation propagation path can be done using graph algorithms to construct a citation relationship graph, and then using graph traversal algorithms to find all citation propagation paths. Adding source-tracing association markers can be done using data structures, such as dictionaries or lists, adding the marker information to the corresponding text records.

[0043] Step S140: Perform full-text keyword co-occurrence statistics on all texts in the parallel corpus partitioned text set and the reference-related text partitioned text set, identify the same technical topic expression that appears simultaneously in different texts, add topic association tags to the corresponding texts of each co-occurring topic, and obtain a labeled technical topic co-occurrence text set.

[0044] In practical implementation, full-text keyword co-occurrence statistics can utilize keyword extraction techniques and statistical methods. First, keyword extraction algorithms, such as TF-IDF and TextRank, are used to extract keywords from the text. Then, the co-occurrence of keywords in different texts is statistically analyzed, using hash tables or databases to store the results. Identifying expressions of the same technical topic can be achieved through semantic similarity calculation methods. Semantic analysis is performed on the extracted keywords to identify keywords with similar or identical meanings, grouping them into the same technical topic. Adding topic-related tags can be done using data structures such as dictionaries or lists, adding the tag information to the corresponding text records.

[0045] Step S150: Integrate the set of sentence combinations with tags, the set of citation association links, and the set of texts co-occurring with technical topics, and remove text entries with duplicate tags so that each tag corresponds to only one text association relationship, to obtain a set of deduplicated multi-type association tags.

[0046] Removing duplicate text entries is necessary because different sets may contain duplicate tags for the same text association, leading to data redundancy and affecting subsequent processing and use. Ensuring each tag corresponds to only one unique text association improves data accuracy and effectiveness. For example, in a parallel corpus sentence combination set and a technical topic co-occurrence text set, the association between two texts may be tagged in multiple sets, requiring the removal of duplicate tags. In practice, the ensemble set can use a list or dictionary to store all tag information. Removing duplicate text entries can be implemented using hash tables or sets, comparing the text associations of the tags to remove duplicate entries.

[0047] Step S160: The deduplicated multi-type association tag set is formatted in a unified manner. The type, associated text position, and associated text content of each association tag are sorted according to fixed fields to generate a weak supervision signal set containing various implicit tag information.

[0048] For example, the type of associated tags, the location of associated text, and the content of associated text can be sorted in the order of "associated tag type - associated text location - associated text content".

[0049] Generating a set of weakly supervised signals containing various implicit labeling information involves summarizing the processed labeling information into a single set containing all implicit labeling information. This set will serve as the basis for subsequent processing. In practice, standardized formatting can be achieved using data structures such as dictionaries or lists, storing information for each associated label according to fixed fields. Sorting can be performed using sorting algorithms to rank the labeling information.

[0050] Step S200: Based on the weakly supervised signal set, extract synonym pairs from parallel corpora and semantic association pairs from cited texts. Perform cross-validation and deduplication on the extracted synonym pairs and semantic association pairs to obtain a text alignment benchmark set.

[0051] For example, step S200 may specifically include the following steps S210 to S260: Step S210: Read the weakly supervised signal set, filter out all entries with synonym mapping tags, extract the related sentence content and the source information of the parallel corpus of each entry, and pair the sentences with the same synonym association to obtain the set of synonym sentence pairs of the parallel corpus.

[0052] In practical implementation, the weakly supervised signal set can be read using file reading tools or database query tools, loading the data into memory for subsequent processing. Entries marked with synonym mappings can be filtered using conditional filtering statements, such as using a WHERE clause in a database query to filter entries marked with synonym mappings. Extracting the content of related clauses and their corresponding parallel corpus sources can be done using data extraction methods, such as extracting relevant field values ​​from database records. Pairing can be achieved by iterating through the filtered entries and pairing clauses with similar semantics.

[0053] Step S220: Read the weakly supervised signal set, filter out all entries with citation traceability tags and technical topic association tags, extract the associated text content and the source information of the citation associated text of each entry, and pair the text fragments with citation or topic association to obtain a set of citation associated text semantic association pairs.

[0054] For example, step S220 may specifically include the following steps S221 to S226: Step S221: Read the weak supervision signal set and extract all entries with reference tracing markers. Each entry contains the starting text, the ending text, and the reference link information. Extract the fragment content corresponding to the starting text and the ending text to obtain a set of reference tracing text fragment pairs.

[0055] In practical implementation, the set of weakly supervised signals can be read using file reading tools or database query tools, loading the data into memory for subsequent processing. Entries with reference tracing tags can be extracted using conditional filtering statements, such as using a WHERE clause in a database query to filter entries with reference tracing tags. The fragments corresponding to the starting and ending text can be extracted using text extraction methods to extract the corresponding text field values ​​from database records or files.

[0056] Step S222: Read the weak supervision signal set and extract all entries with technical topic association tags. Each entry contains co-occurring topic text fragments and a list of associated texts. Pair the co-occurring topic text fragments with the text fragments in each of the associated text lists to obtain a set of technical topic association text fragment pairs.

[0057] In practical implementation, reading the weakly supervised signal set and extracting entries can be done using methods similar to those in step S221. Extracting co-occurring topic text fragments and associated text lists can be done using data extraction methods to extract corresponding field values ​​from database records or files. Pairing can be done by looping through each entry, pairing the co-occurring topic text fragments with each text fragment in the associated text list.

[0058] Step S223: Traverse each fragment pair in the reference source text fragment pair set, check whether there is a direct reference statement annotation between the starting text fragment and the ending text fragment, confirm the authenticity of the reference relationship, remove fragment pairs without direct reference annotation, and obtain the verified reference source fragment pair set.

[0059] For example, a patent document that explicitly cites a technical solution from another document is a reliable citation. Removing fragments without direct citations is to eliminate unreliable citation relationships and improve data quality.

[0060] In practical implementation, iterating through fragment pairs can use a loop structure to check each fragment pair one by one. Checking for direct citations can use text matching methods to search for citation information in the starting and ending text fragments, such as searching for citations like "[Reference 1]". Fragment pairs without direct citations can be removed using list filtering methods, removing fragment pairs that do not meet the criteria from the collection.

[0061] Step S224: Traverse each fragment pair in the set of technical topic-related text fragment pairs, check whether the technical topic descriptions between the co-occurring topic text fragments and the related text fragments are completely consistent, confirm the rationality of the topic association, remove fragment pairs with topic description deviations, and obtain the verified set of technical topic-related fragment pairs.

[0062] For example, step S224 may specifically include the following steps S2241 to S2246: Step S2241: Read each fragment pair in the set of technical topic associated text fragment pairs, extract the complete content of co-occurring topic text fragments and associated text fragments, including the context descriptions before and after the fragments, to obtain a set of technical topic fragment pairs with context.

[0063] In practical implementation, reading fragment pairs can use a loop structure to iterate through each fragment pair in the collection. Extracting complete content and context descriptions can be done by extracting text fragments containing context from the original text based on the fragment's position within the original text. String manipulation techniques can be used to extract text containing context based on the start and end positions of the fragment.

[0064] Step S2242: For each pair of technical topic fragments with context, extract the core technical topic expression from the co-occurring topic text fragments, and decompose it into multiple key components such as technical field, technical object, and technical means to obtain a set of core technical topic elements.

[0065] In practical implementation, extracting core technical themes can be achieved using text extraction methods, extracting the core technical theme content from contextualized technical theme fragments. Decomposing the content into key components can utilize natural language processing techniques, such as named entity recognition and semantic role labeling. Named entity recognition identifies entities such as technical fields and technical objects, while semantic role labeling identifies semantic roles such as technical means.

[0066] Step S2243: Extract the core technical theme description from the related text fragments, obtain the related text technical theme element set according to the same decomposition method, compare the core technical theme element set with the related text technical theme element set item by item, record the comparison result of each element, and obtain the technical theme element comparison result set.

[0067] The purpose of comparing the core technology theme element set with the related text technology theme element set item by item is to check whether the technology theme elements in the two sets are consistent. For example, comparing whether elements such as technology field, technology object, and technology means are the same. Recording the comparison results of each element is for subsequent analysis and screening.

[0068] In practical implementation, extracting the core technical themes and decomposing elements from related text fragments can be done using methods similar to step S2242. Item-by-item comparison can be performed by iterating through the two element sets and comparing each element individually. The comparison results can be stored using a list or a dictionary; for example, a dictionary can be used to store the comparison results.

[0069] Step S2244: For fragment pairs in the technical topic element comparison result set where all key components completely overlap, further check whether there are any statements in the context description that contradict the technical topic, confirm that the context description does not affect the rationality of the topic association, and obtain a set of context-compatible technical topic fragment pairs.

[0070] In practical implementation, checking whether the contextual description contains statements that contradict the technical topic can utilize semantic reasoning techniques, such as rule-based semantic reasoning or knowledge graph-based semantic reasoning. By analyzing the semantic information in the contextual description and comparing it with the core technical topic, a conflict can be determined. After confirming that the contextual description does not affect the rationality of the topic association, the matching fragment pairs are selected to obtain a set of context-compatible technical topic fragment pairs.

[0071] Step S2245: For fragment pairs with element differences in the technical topic element comparison result set, they are determined to be topic expression deviations, and they are removed from the technical topic associated text fragment pair set to obtain the technical topic fragment pair set after deviation removal.

[0072] In the actual implementation, the set of technical topic element comparison results is traversed to find the fragment pairs with element differences, and then the list filtering method is used to remove these fragment pairs from the set of technical topic associated text fragment pairs.

[0073] Step S2246: Integrate the context-compatible set of technical topic fragment pairs with the set of technical topic fragment pairs after removing biases, add a topic association verification pass mark to each fragment pair, and finally generate a verified set of technical topic association fragment pairs.

[0074] In practical implementation, the collection can be merged using list merge operations, combining two collections into one. Adding a topic association verification pass mark can utilize a data structure, adding a mark field to each fragment pair and setting the mark value to "verification passed". For example, each fragment pair can be represented as a tuple or dictionary containing the text fragment and the verification mark: constructing a tuple in the form of ("co-occurrence_theme_text", "associated_text", "topic association verification pass mark"), or creating a dictionary like {"co-occurrence_theme_text": "co-occurrence_theme_text", "associated_text": "associated_text", "verification_mark": "topic association verification passed"} to more clearly express the relevant information of each fragment pair and its verified status.

[0075] Step S225: Integrate the verified set of citation source traceability fragment pairs with the verified set of technical topic association fragment pairs, remove fragment pairs with completely duplicate content, so that each fragment pair corresponds to only one unique association, and obtain the integrated set of citation association fragment pairs.

[0076] In practice, the consolidation of the collections can be done by merging lists, combining the verified collection of reference traceability fragment pairs with the verified collection of technical topic association fragment pairs. Duplicate entries can be removed by iterating through the merged collection and comparing the content of each fragment pair one by one. String matching can be used to determine if two fragment pairs have identical content; if they are identical, only one fragment pair is retained. For example, for two fragment pairs ("text fragment A1", "text fragment B1") and ("text fragment A1", "text fragment B1"), if their content is found to be identical during iteration, only one of them is retained.

[0077] Step S226: Add a semantic association type tag to each fragment pair in the integrated collection of reference association fragment pairs. The tag types are divided into reference source association and technical topic association, and finally generate a collection of reference association text semantic association pairs.

[0078] In practical implementation, the integrated set of reference-related fragment pairs can be traversed, and the association type of each fragment pair can be determined based on its source and verification information. If the fragment pair comes from the verified set of reference-originating fragment pairs, then a "reference-originating association" tag is added to it; if it comes from the verified set of technical topic-related fragment pairs, then a "technical topic-related association" tag is added.

[0079] For example, for a fragment pair ("Description of the technical solution in patent document X", "Description of the solution cited in patent document Y"), since it comes from the citation tracing verification process, the "citation tracing association" tag is added; for another fragment pair ("Introduction to the principle of artificial intelligence image recognition technology", "Application case of artificial intelligence image recognition technology in another document"), since it is obtained based on the technical topic association verification, the "technical topic association" tag is added.

[0080] Step S230: Compare the content of each clause pair in the set of synonymous clause pairs in the parallel corpus with each association pair in the set of semantic association pairs in the referenced text, identify the association pairs with completely overlapping content, record the association type and source information of the overlapping entries, and obtain the set of cross-type overlapping association pairs.

[0081] In practice, nested loops can be used to traverse the set of synonymous clause pairs and the set of semantically related text pairs in the parallel corpus. For each pair of synonymous clause pairs and each pair of semantically related text pairs in the parallel corpus, a string matching algorithm is used to compare whether their content is completely identical. If the content is completely identical, the association type of that pair in the set of synonymous clause pairs and the set of semantically related text pairs in the parallel corpus, as well as the text sources from which they respectively originate, are recorded.

[0082] For example, when a parallel corpus synonymous clause pair ("automatic control device", "automatic control equipment") is found to completely overlap with a reference-related text semantic association pair ("automatic control device", "automatic control equipment"), its association type in the parallel corpus synonymous clause pair set is recorded as "synonymous association", and its source is "patent document A and parallel corpus". Its association type in the reference-related text semantic association pair set is recorded as "technical topic association", and its source is "patent document B and reference-related text set".

[0083] Step S240: For each entry in the cross-type overlapping association pair set, verify whether its semantic reference in the parallel corpus and the referenced associated text is completely consistent. After confirming that there is no semantic deviation, retain the tag of one type of association and remove the entries with duplicate tags to obtain the deduplicated cross-type association pair set.

[0084] In implementing semantic referencing verification, semantic understanding and reasoning techniques can be employed. First, by combining the contextual information from parallel corpora and cited related texts, a deep semantic analysis of each overlapping association pair is conducted. Methods from natural language processing, such as semantic role labeling and event extraction, can be used to clarify the semantic role and function of each element in the association pair in different contexts. Then, the consistency of these semantic roles and functions in the parallel corpora and cited related texts is compared. For example, if an association pair ("intelligent sensor," "smart sensing device") represents a synonymous relationship in the parallel corpus but a technical topic association in the cited related text, it is necessary to analyze whether the technical concepts and application scenarios represented by "intelligent sensor" and "smart sensing device" are consistent in the two contexts.

[0085] If verification confirms there are no semantic discrepancies, select one type of association tag to retain. This selection can be based on actual needs or pre-defined priorities, such as prioritizing reference-based association tags, or choosing based on the universality of the association type. Then, remove entries with duplicate tags from the set of overlapping cross-type association pairs. List filtering or set operations can be used to implement this removal operation.

[0086] Step S250: Integrate the non-overlapping entries in the set of synonymous sentence pairs in parallel corpus, the non-overlapping entries in the set of semantic association pairs in referenced texts, and the set of cross-type association pairs after deduplication, so that each association pair appears only once, and obtain the complete set of integrated association pairs.

[0087] In practice, we can first list the non-overlapping entries in the parallel corpus synonymous sentence pair set, the non-overlapping entries in the citation-related text semantic association pair set, and the deduplicated cross-type association pair set. Then, we use a new set to store the integrated association pairs. During the process of adding association pairs to the new set, we check if each association pair already exists in the new set; if not, we add it; otherwise, we skip it. Hash tables or set data structures can be used to achieve fast lookup and deduplication operations.

[0088] Step S260: Classify the integrated set of association pairs according to the association type, add a unified alignment benchmark mark to each type of association pair, and generate a text alignment benchmark set containing synonym pairs and semantic association pairs.

[0089] In the specific implementation, the classification criteria for association types are first determined. Based on the tagging information of the association pairs, such as the tags added in previous steps ("synonymous associations," "reference-based associations," "technical topic associations," etc.), the integrated set of association pairs can be classified. Dictionary or list data structures can be used to store the classification results, where the key or index represents the association type, and the value represents the set of association pairs under that type.

[0090] For example, create a dictionary {"Synonym Association":[("Automatic Control Device", "Automatic Control Equipment")], "Citation Tracing Association":[("Technical Solution in Patent Document A", "Content Citing the Solution in Patent Document B")], "Technical Topic Association":[("Principles of Artificial Intelligence Image Recognition Technology", "Application of Artificial Intelligence Image Recognition Technology in Another Document")]} to store the categorized association pairs.

[0091] Then, add a uniform alignment benchmark tag to each type of association pair. The tag content can include information such as association type, importance level of association, and applicable alignment scenario. For example, add the tag "Synonym Alignment Benchmark - High Importance - Applicable to Lexical Substitution Alignment" to association pairs of synonym association type, and add the tag "Citation Alignment Benchmark - Medium Importance - Applicable to Document Citation Relationship Alignment" to association pairs of citation tracing association type, etc.

[0092] Step S300: Perform global association matching between the text alignment benchmark set and the target intellectual property text dataset. Through the mapping association derivation of the text alignment benchmark set, complete the automatic association and labeling of intellectual property entities and semantic relationships between entities in the text, and obtain the initial structured knowledge unit set. The initial structured knowledge unit set contains entity representations and relationship representations that have not undergone credibility verification.

[0093] For example, step S300 may specifically include the following steps S310-S360: Step S310: Read the text alignment reference set and the target intellectual property text dataset, divide the target intellectual property text dataset into multiple text blocks according to chapters, and retain the original chapter number, page number and paragraph mark of each text block to obtain the set of divided intellectual property text blocks.

[0094] In practical implementation, the first step is to use a data reading tool to read the text alignment benchmark set and the target intellectual property text dataset. For the target intellectual property text dataset, chapters can be divided based on information such as chapter titles and numbers in the text. Regular expressions or title recognition algorithms from natural language processing can be used to identify chapter titles, and then the text is split according to chapters. After splitting, corresponding chapter numbers, page numbers, and paragraph marks are added to each text block. Dictionary or list data structures can be used to store each text block and its related information. For example, a patent document containing multiple chapters can be split according to chapters to obtain multiple text blocks, each text block representing the content of a chapter, while recording the chapter number, page number, and paragraph mark.

[0095] Step S320: Traverse each association pair in the text alignment benchmark set, extract the core expression content in the association pair, use it as the search content and perform full-text matching with each text block in the segmented intellectual property text block set, identify the text blocks containing the core expression content, and obtain the matched intellectual property text block set.

[0096] The core statements are used as the search criteria, and a full-text match is performed between each text block in the segmented intellectual property text block set. This aims to find text blocks containing these core statements within the segmented intellectual property text blocks. Full-text matching can use string matching algorithms, such as the naive string matching algorithm or the KMP algorithm, to search for the existence of text fragments identical to the core statements within each text block.

[0097] After identifying text blocks containing core statements, these text blocks are collected to obtain a set of matching intellectual property text blocks. This set will contain all text blocks that match the core statements in the text alignment benchmark set, providing a scope for subsequent location of core statements and extraction of technical expression units within text blocks.

[0098] In the specific implementation, each association pair in the text alignment benchmark set is iterated through using a loop to extract the core expression content. Then, for each core expression content, each text block in the divided intellectual property text block set is iterated through using a loop to perform a full-text match. If a matching core expression content is found in a certain text block, that text block is added to the matching intellectual property text block set.

[0099] Step S330: For each matched intellectual property text block, locate the specific position of the core expression within the text block, extract the text content within a range before and after that position, identify the expression units with independent technical meaning, and obtain the set of technical expression units within the text block.

[0100] For example, step S330 may specifically include the following steps S331-S336: Step S331: Read each matched intellectual property text block, split the text block into multiple sentence units according to punctuation marks, and retain the original paragraph mark and sentence number in each sentence unit to obtain a set of sentence units after splitting the text block.

[0101] In practical implementation, string processing methods are used to split text blocks into multiple sentence units based on punctuation marks (such as periods, question marks, exclamation marks, etc.). Regular expressions or string splitting functions can be used for this. After splitting, paragraph tags and sentence numbers are added to each sentence unit. Dictionaries or list data structures can be used to store each sentence unit and its related information. For example, for a matching intellectual property text block, it can be split into multiple sentence units, each representing a complete sentence, while recording the paragraph number and its sequence number within the paragraph.

[0102] Step S332: Based on the core expression content in the text alignment benchmark set, locate the sentence unit containing the core expression content in the sentence unit set, mark the start and end positions of the sentence unit, and obtain the located core sentence unit set.

[0103] In this embodiment of the invention, based on the core expression content in the text alignment benchmark set, the sentence unit containing the core expression content is located in the sentence unit set. This is to find sentence units related to the core expression content from the split sentence unit set. String matching algorithms, such as full-text search or substring matching, can be used to check whether the core expression content is contained in each sentence unit. The start and end positions of the sentence unit are marked. The start and end position information can accurately determine the position of the core expression content in the sentence unit, providing an accurate range for subsequent extraction of extended text fragments surrounding the core sentence unit. Index values ​​can be used to represent the start and end positions.

[0104] In the implementation, each sentence unit in the sentence unit set is traversed, and a string matching method is used to check whether the sentence unit contains the core expression content in the text alignment benchmark set. If it does, the start and end positions of the sentence unit are recorded. A list or dictionary data structure can be used to store the located core sentence unit set, where the elements of the list or dictionary contain the content of the sentence unit and its start and end position information.

[0105] Step S333: Extract the content of the core sentence unit after positioning and the two sentence units before and after it, combine these sentence units into continuous text segments, retain the position mark of each sentence unit, and obtain the set of extended text segments around the core sentence unit.

[0106] In the specific implementation, for each core sentence unit in the located core sentence unit set, the content of the two sentence units before and after it is extracted based on its position in the sentence unit set. List indexing can be used to retrieve the corresponding sentence units. These sentence units are then combined into a continuous text segment, which can be achieved using string concatenation. Simultaneously, the positional markers of each sentence unit (such as paragraph number, sentence sequence number, etc.) are recorded, which can be stored using a dictionary or list data structure.

[0107] Step S334: Traverse each text fragment in the extended text fragment set, identify the expression content in the fragment that can independently express a complete technical concept, including technical name, technical method, technical effect, etc., mark the start and end positions of the expression content, and obtain an independent technical expression candidate set.

[0108] For example, step S334 may specifically include the following steps S3341-S3346: Step S3341: Read each text fragment in the extended text fragment set, split the text fragments according to words, and obtain a set of word sequences containing word, part of speech, and position information. Each word retains its original order in the text fragment.

[0109] In practical implementation, a word segmenter from a natural language processing toolkit is used to segment text fragments into words. For example, the Jieba word segmentation library can be used to segment Chinese text. After segmentation, a part-of-speech tagger is used to tag each word with its part of speech. Simultaneously, the position information of each word within the text fragment is recorded, which can be achieved by recording the starting index of the word. Finally, the words, part-of-speech tags, and position information are stored in a list or dictionary data structure to form a set of word sequences.

[0110] Step S3342: Traverse each word sequence in the word sequence set, identify words that can serve as the core of technical concepts, including technical terms, technical action verbs, etc., mark the positions of these core words, and obtain the core technical word tag set.

[0111] In the specific implementation, each word sequence in the word sequence set is traversed, and based on the word's part of speech and semantic information, it is determined whether it is a technical term or a technical action verb. A predefined dictionary of technical terms and technical action verbs can be used to assist in this determination. If it is a core technical term, its position information is recorded. A list or dictionary data structure can be used to store the set of core technical term tags, where the elements of the list or dictionary contain the core technical terms and their position information.

[0112] Step S3343: Taking each core technical term as the center, expand the term sequence forward and backward until the expanded term combination can form a complete technical concept expression. Record the start and end positions of the expanded term combination to obtain a set of candidate combinations of technical concept expressions.

[0113] In practical implementation, for each core technical term in the core technical term tagging set, the term sequence is expanded progressively forward and backward from its position. During the expansion process, based on the grammatical and semantic relationships between terms, it is determined whether the expanded term combinations can form a complete technical concept expression. Syntactic analysis and semantic understanding techniques can be used to assist in this determination. Once a complete technical concept expression is determined, its start and end positions are recorded. A list or dictionary data structure can be used to store the set of candidate combinations of technical concept expressions, where the elements of the list or dictionary contain the candidate combinations of technical concept expressions and their start and end position information.

[0114] Step S3344: For each candidate combination of technical concept expressions, check whether it can independently express a complete technical meaning, including whether it contains a clear technical object, technical action or technical attribute, confirm the logical integrity of the expression, and obtain a logically complete set of technical expression candidates.

[0115] In practical implementation, each candidate combination of technical concept expressions in the set is traversed and semantically analyzed. Semantic role labeling and event extraction techniques from natural language processing can be used to identify elements such as technical objects, technical actions, and technical attributes. Then, the logical relationships between these elements are checked to ensure they are reasonable and can independently express a complete technical meaning. If the conditions are met, the candidate combination of technical concept expressions is added to the set of logically complete technical expression candidates. A list or dictionary data structure can be used to store the set of logically complete technical expression candidates, where each element of the list or dictionary contains logically complete technical expression candidate combinations.

[0116] Step S3345: For each candidate in the logically complete candidate set of technical expressions, remove expressions containing redundant words and retain the most concise and semantically complete technical concept expressions to obtain a simplified candidate set of technical expressions.

[0117] In the specific implementation, each candidate in the logically complete set of technical expression candidates is traversed, and semantic analysis and lexical filtering are performed on it. Part-of-speech analysis and semantic understanding techniques from natural language processing can be used to determine the necessity of each word for the technical concept expression. If a word can be removed without affecting the expression of the technical concept, it is discarded. At the same time, it is ensured that the technical concept expression remains semantically complete after word removal. String manipulation methods can be used to implement word removal. Finally, the simplified technical concept expression candidate combinations are collected into the simplified technical expression candidate set.

[0118] Step S3346: Add start and end position markers to each candidate in the simplified technical description candidate set, organize them into description units with a unified format, and finally generate an independent technical description candidate set.

[0119] In the specific implementation, each candidate in the simplified technical statement candidate set is traversed, and its start and end positions are determined based on its lexical position information in the original text. This positional information is then added to the candidate. Next, according to the unified format requirements for statement units, the candidates are organized into a format containing fields such as statement content, start position, and end position. A dictionary or list data structure can be used to store the organized statement units, and finally, all such statement units are collected into an independent technical statement candidate set.

[0120] Step S335: For each independent technical expression candidate, verify its semantic independence in the extended text fragment, confirm that the expression content can be fully understood without relying on other text fragments, remove the expression content that requires context supplementation, and obtain a set of semantically independent technical expressions.

[0121] In practical implementation, each candidate in the independent technical statement candidate set is traversed and semantically analyzed. The statement content can be input into a trained semantic understanding model, which can analyze the completeness and logical relationships of the various elements in the statement. If the model determines that the statement requires contextual supplementation to be fully understood, it is removed from the independent technical statement candidate set. The final set of semantically independent technical statements will contain all semantically independent technical statement content.

[0122] Step S336: Sort the set of semantically independent technical expressions according to the length of the expression content, filter out the technical expression units with complete expressions, and finally generate the set of technical expression units in the text block.

[0123] In this embodiment of the invention, the set of semantically independent technical expressions is sorted according to the length of its content. The purpose of this sorting is to facilitate the subsequent selection of technical expression units with complete content. Generally, technical expressions with longer content may contain richer information and are more likely to be complete.

[0124] Complete technical description units are selected. A complete technical description unit should contain clearly defined technical objects, technical actions, and technical effects, comprehensively expressing a technical concept. Completeness can be determined based on the semantic structure and logical relationships of the technical description. Finally, a set of technical description units within the text block is generated, containing all the selected complete technical description units. In the specific implementation, a sorting algorithm is used to sort the semantically independent set of technical descriptions according to the length of their content. Then, the sorted set is traversed, and semantic analysis is performed on each technical description to determine its completeness. This can be done based on predefined rules for technical description completeness, such as checking whether it contains elements like technical objects, technical actions, and technical effects. If the description is complete, it is added to the set of technical description units within the text block.

[0125] Step S340: Based on the association relationships in the text alignment benchmark set, deduce the semantic associations between technical description units, bind technical description units that have synonym or reference associations, add semantic association tags to the bound units, and obtain a set of bound technical description units with tags.

[0126] In the specific implementation, each association pair in the text alignment benchmark set is traversed. For each association pair, a technical description unit containing the core description content of the association pair is searched in the technical description unit set. If found, the association type (synonym or reference) between these technical description units is determined, and they are bound together. Semantic association tags are added to the bound units, and dictionary or list data structures can be used to store the tag information. For example, for the bound technical description units ("high-precision detection technology of intelligent sensors", "high-precision detection technology of smart sensing devices"), the tag {"association_type":"synonymous association", "association_basis":"synonymous description pair (intelligent sensor, smart sensing device) in the text alignment benchmark set") is added}. Finally, all tagged bound technical description units are collected into the tagged bound technical description unit set.

[0127] Step S350: Organize each unit in the set of labeled binding technology description units into a structured form, clarifying the technical meaning of the unit, the associated unit information, and the position of the text block to which it belongs, forming a structured knowledge entry and obtaining an initial set of structured knowledge entries.

[0128] In this embodiment of the invention, each unit in the set of labeled binding technology expression units is structured and organized. This structuring involves organizing the information within the binding technology expression units in an orderly manner, giving them a clear structure and format. The technical meaning of the unit is clarified; the technical meaning is the core technical concept expressed by the binding technology expression unit. Related unit information is also clarified, including information about other units that are synonymous or referenced with this unit, as well as the type and basis of the association. For example, for the unit "high-precision detection technology for intelligent sensors," the related unit information could be the synonymous unit "high-precision detection technology for intelligent sensing devices," with the association type being "synonymous association," and the association basis being a synonymous expression pair in the text alignment reference set.

[0129] Clearly define the location of the associated text block. This location information helps accurately pinpoint the source of the technical description unit in the original text, such as its chapter number, page number, and paragraph mark. Create structured knowledge entries by organizing the clearly defined technical meaning, related unit information, and associated text block location according to a specific format. Obtain an initial set of structured knowledge entries. This set will contain all the structured knowledge entries, providing standardized data for subsequent removal of duplicate entries and obtaining the initial set of structured knowledge units.

[0130] In practical implementation, each unit in the set of tagged binding technology representation units is traversed, and its technical meaning, related unit information, and the position of its corresponding text block are analyzed. This information is then organized according to a predefined structured format to form knowledge entries. A list or dictionary data structure can be used to store the initial set of structured knowledge entries, where each element of the list or dictionary is a knowledge entry.

[0131] Step S360: Integrate all initial structured knowledge entries, remove duplicate knowledge entries, and ensure that each entry corresponds to only one unique technical expression unit and relationship, thus obtaining an initial structured knowledge unit set.

[0132] In the specific implementation, the initial set of all structured knowledge entries is first merged into a large set. Then, this set is traversed, and for each knowledge entry, it is checked whether it is a duplicate of other processed knowledge entries in the set. Duplicates can be determined by comparing key information such as the technical representation units, related unit information, and relationships of the knowledge entries. If a duplicate is found, it is removed from the set. The final set of initial structured knowledge units will contain all unique, integrated, and deduplicated knowledge entries, including entity and relationship representations that have not undergone credibility verification.

[0133] Step S400: Perform credibility normalization on the initial set of structured knowledge units, and retain entity and relation representations that meet the preset credibility requirements through semantic consistency verification and domain adaptability screening to obtain a standardized set of knowledge units.

[0134] For example, step S400 may specifically include the following steps S410-S460: Step S410: Read the initial set of structured knowledge units, split each knowledge unit into an entity representation part and a relation representation part, and assign them to the entity representation storage partition and the relation representation storage partition respectively. The representation in each partition retains the original knowledge unit association mark, and the split entity representation set and relation representation set are obtained.

[0135] In the specific implementation, each knowledge unit in the initial set of structured knowledge units is traversed, and based on its structure and information, it is split into entity representation and relation representation. This splitting can be achieved using string manipulation and information extraction techniques. Then, the split entity and relation representations are added to the entity representation storage partition and relation representation storage partition respectively, while retaining the original knowledge unit association tags. List or dictionary data structures can be used to store the representations in these two partitions.

[0136] Step S420: Traverse each entity representation in the entity representation set, compare its content with all other entity representations in the set, identify representation entries that represent the same entity but have different content, check whether there is a semantic conflict between these entries, remove entity representation entries with semantic conflicts, and obtain a semantically consistent entity representation set.

[0137] In practice, content comparison can employ semantic similarity calculation methods. For example, using a deep learning-based semantic similarity model, entity representations can be converted into vector representations, and the similarity between vectors can be calculated to determine whether the representations might refer to the same entity. Representation pairs with high similarity may require further semantic conflict checks.

[0138] Checking for semantic conflicts can combine domain knowledge and logical reasoning. Specialized dictionaries and standards in the field of intellectual property can be consulted to determine whether different expressions are consistent in technical concepts and logic. For example, relevant sensor technical standards can be consulted to determine whether "high-temperature tolerance" is a characteristic that such sensors should possess.

[0139] Step S430: Traverse each relation representation in the relation representation set, compare its content with all other relation representations in the set, identify the representation entries that express the relationship between the same entity but have different content, check whether there is a semantic contradiction between these entries, remove the relation representation entries with semantic contradictions, and obtain a set of semantically consistent relation representations.

[0140] After identifying entries that describe the same relationship between entities but differ in content, the focus is on checking for semantic contradictions between these entries. Semantic contradictions may manifest as conflicts in the description of the relationship between entities in terms of logic, direction, or degree. For example, one relationship might state "intelligent sensors control the operation of a data analysis system," while another might state "the data analysis system controls the operation of intelligent sensors," which presents a clear semantic contradiction.

[0141] Removing semantically contradictory relational entries ensures the semantic consistency of the relational representation set. Only semantically consistent relational representations can accurately reflect the real relationships between entities in the intellectual property field, providing reliable relational information for subsequent knowledge graph construction.

[0142] In the specific implementation of content comparison, methods similar to entity representation comparison can be used, such as calculating semantic similarity. By converting relational representations into vector representations, the similarity between vectors can be calculated to preliminarily determine whether they represent the same relation.

[0143] Detecting semantic contradictions requires a combination of domain knowledge and logical judgment. One can analyze the reasonableness of different relational expressions based on the technical principles and business logic of the intellectual property field. For example, based on the working principles of intelligent sensors and data analysis systems, one can determine whether the direction and control relationship of data transmission are logically sound.

[0144] Step S440: Compare the semantically consistent set of entity representations and the semantically consistent set of relation representations with the general representation specifications in the intellectual property field, identify the representation entries that conform to the specifications, and remove the representation content that does not conform to the general representation specifications to obtain the domain-adapted set of representations.

[0145] In the specific implementation of the comparison, a knowledge base of universally accepted expressions in the field of intellectual property can be established, containing standardized terminology, concepts, and relational descriptions. Each expression in the semantically consistent set of entity expressions and relational expressions is then matched against the content in the knowledge base. Methods such as string matching and semantic matching can be used for the comparison.

[0146] For entity representations, check whether the terminology used is accurately defined in the specification knowledge base and conforms to the domain's conceptual framework. For example, check whether the representation of "intelligent sensor" matches the classification and definition of sensors in the specification. For relationship representations, check whether the described relationships between entities conform to the domain's logic and rules. For example, check whether the description of "the connection relationship between intelligent sensors and the data analysis system" conforms to the logic of data transmission and processing.

[0147] Step S450: Rebind the domain-adapted entity representation set and relation representation set according to the original knowledge unit association tags, so that each entity representation corresponds to the correct relation representation, and obtain the bound standardized knowledge unit candidate set.

[0148] The original knowledge unit association tags recorded the original relationships between entity representations and relation representations, allowing for accurate mapping between them. For example, in the knowledge unit before splitting, the entity representation "intelligent sensor" was associated with the relation representation "intelligent sensor transmits data to data analysis system," and the association tags recorded this correspondence. Ensuring that each entity representation corresponds to the correct relation representation guarantees that the recombined knowledge unit accurately reflects the true state of entities and their relationships in the intellectual property field. Incorrect mapping between entity and relation representations leads to inaccurate information conveyed by the knowledge unit, affecting the subsequent construction and application of the knowledge graph.

[0149] In the specific implementation of binding, the set of entity representations and relationship representations adapted to the domain can be traversed, and the corresponding entity and relationship representation can be found based on the association marker. Data structures, such as dictionaries, can be used to store the association markers and their corresponding representation information. For each entity representation, the corresponding relationship representation is found through the association marker, and they are combined into a new knowledge unit.

[0150] Step S460: Perform final verification on the bound standardized knowledge unit candidate set to confirm that the entity representation and relation representation of each unit are semantically compatible and conform to the domain specifications. Remove units that fail the verification to obtain the standardized knowledge unit set.

[0151] In this embodiment of the invention, the final verification of the standardized knowledge unit candidate set after binding is to ensure that the quality of the knowledge units reaches the highest standard. Although semantic consistency verification and domain adaptability screening have been performed in the previous steps, some potential problems may still exist after rebinding, requiring a final check.

[0152] Ensure that the entity representation and relation representation of each unit are semantically compatible. Semantic compatibility means that the entity representation and relation representation match each other logically and semantically. For example, the entity representation "intelligent sensor" and the relation representation "intelligent sensor transmits data to the data analysis system" are semantically compatible because "intelligent sensor" has the ability to transmit data, and "transmits data to the data analysis system" is a relation description that conforms to its functional characteristics.

[0153] Each unit is verified to conform to domain specifications, which include terminology, concept definitions, and logical rules in the intellectual property field. For example, in the intellectual property field, there are strict specifications for the description of technical products and the definition of relationships, and knowledge units must follow these specifications. Units that fail verification are removed, as they may affect the accuracy and reliability of the entire knowledge graph. For example, if the entity representation and relation representation in a knowledge unit are semantically incompatible or do not conform to domain specifications, then this unit cannot be included in the standardized knowledge unit set.

[0154] In the final verification process, various methods can be employed. For semantic compatibility checks, semantic reasoning and logical analysis can be performed. For example, the reasonableness of relational expressions can be determined based on the entity's function and characteristics. For domain specification checks, the general expression specification knowledge base in the intellectual property field can be referenced again to check whether the terms, concepts, and relational descriptions in the knowledge units conform to the specifications.

[0155] Step S500: The standardized knowledge unit set is classified and organized into entities and relationships according to the hierarchical requirements of intellectual property retrieval, and a knowledge graph for intellectual property retrieval is constructed. The knowledge graph includes technical topic levels, entity association links and semantic mapping paths that are adapted to retrieval requirements.

[0156] For example, step S500 may specifically include the following steps S510-S560: Step S510: Read the standardized knowledge unit set, extract all entity representation content and relation representation content from it, and organize them into independent complete sets of entity representations and complete sets of relation representations respectively. Each representation retains the original knowledge unit association information, and the extracted complete sets of entity representations and complete sets of relation representations are obtained.

[0157] During extraction and organization, each knowledge unit in the standardized knowledge unit set can be traversed, and its entity representations and relation representations can be extracted separately. List or dictionary data structures can be used to store the complete set of entity representations and relation representations. For each representation, its associated information is added as an attribute to the representation record.

[0158] For example, for the knowledge unit {"entity":"intelligent sensor","relationship":"intelligent sensor transmits data to the data analysis system","association_info":"association number 1"} in the standardized knowledge unit set, "intelligent sensor" is extracted to the complete set of entity representations and the association information "association number 1" is added; "intelligent sensor transmits data to the data analysis system" is extracted to the complete set of relation representations and the association information "association number 1" is added.

[0159] Step S520: In accordance with the hierarchical requirements of intellectual property retrieval, classify the complete set of entity descriptions according to dimensions such as technical field, technical type, and application scenario. Each category corresponds to a node in the retrieval hierarchy. Add a hierarchy node marker to each category to obtain a hierarchical entity category set.

[0160] In implementing the classification, a classification rule base can be established, containing classification standards and rules based on dimensions such as technical field, technology type, and application scenario. Each entity representation in the complete set of entity representations is matched against the classification rule base, and then assigned to the corresponding category based on the rules it conforms to. For example, the entity representation "smart sensor," according to the classification rule base, belongs to the sensor technology type within the field of electronics technology, and has applications in industrial and home scenarios; therefore, it is assigned to the corresponding categories.

[0161] For example, step S520 may specifically include the following steps S521-S526: Step S521: Read the hierarchical requirement description of intellectual property retrieval, extract the dimension description content used for classification, break down each dimension into specific classification standard items, and obtain a set of retrieval hierarchical classification dimension standards. Each standard item contains a dimension name and classification rules.

[0162] In practical implementation, text processing techniques are first used to read the hierarchical requirement description for intellectual property retrieval. This requirement description can be stored in a text file, and its contents can be read into memory using a file reading tool. Then, natural language processing techniques, such as part-of-speech tagging and named entity recognition, are used to extract the dimensional descriptions.

[0163] For each dimension's content, further analyze its classification criteria. Refer to professional literature and standards in the field of intellectual property to determine the specific classification criteria for each dimension. For example, consult relevant standards in the field of electronic technology to determine the classification criteria for technology types within that field. Organize the classification criteria for each dimension into standard entries, with each entry containing a dimension name and classification rule. List or dictionary data structures can be used to store the set of classification dimension criteria for the retrieval hierarchy. For example, create a dictionary {"Dimension Name 1":["Classification Rule 1-1", "Classification Rule 1-2"], "Dimension Name 2":["Classification Rule 2-1", "Classification Rule 2-2"]} to store the standard entries.

[0164] Step S522: Read the complete set of entity representations, extract the core technical content of each entity representation, including the technical field, technical implementation type, and technical application scenario, to obtain the core classification information set of each entity representation.

[0165] The core technical content is extracted from each entity description. This core technical content reflects the essence and application of the entity's technology, including its technical domain, implementation type, and application scenario. The technical domain determines the entity's macro-level technological affiliation, such as electronics or mechanical technology; the implementation type describes the entity's implementation method, such as optical or acoustic sensors in sensor technology; and the application scenario reflects the entity's practical application, such as industrial production or medical testing. This yields a core classification information set for each entity description. This core classification information set organizes and stores the core technical content of each entity description, facilitating subsequent matching with the retrieval hierarchy classification dimension standard set.

[0166] In practical implementation, natural language processing techniques are used to analyze entity representations. Methods such as named entity recognition and semantic understanding can be employed to extract information such as the technology domain, technology implementation type, and technology application scenario from entity representations.

[0167] For example, for the entity description "Application of intelligent optical sensors in industrial automated production," named entity recognition can determine that "intelligent optical sensor" belongs to the sensor technology type under the field of electronic technology, "optical" is its technology implementation type, and "industrial automated production" is its application scenario. This information can be organized into a core classification set, such as {"Technology Field": "Electronic Technology", "Technology Implementation Type": "Optical Sensor", "Technology Application Scenario": "Industrial Automated Production"}.

[0168] Step S523: Match the core classification information set of each entity description with the retrieval hierarchical classification dimension standard set item by item, confirm each classification dimension node corresponding to the entity description, record the matching results, and obtain the entity description-hierarchical node matching result set.

[0169] In this embodiment of the invention, the core classification information set of each entity representation is matched item by item with the retrieval level classification dimension standard set in order to determine the specific position of each entity representation in the retrieval level. The retrieval level classification dimension standard set provides the classification criteria and rules, while the core classification information set contains the key classification information of the entity representation.

[0170] Identify each category dimension node corresponding to the entity description. Each category dimension node represents a category in the retrieval hierarchy. For example, in the technology field dimension, "electronic technology" is a node; in the technology type dimension, "sensor technology" is a node. Through matching, determine which nodes the entity description should be assigned to.

[0171] During the matching process, the core classification information set of each entity is traversed, and information such as the technology's domain, technology implementation type, and technology application scenario is compared with the classification rules in the standard set of retrieval level classification dimensions.

[0172] For example, for the entity description "intelligent optical sensor," its core classification information set is {"Technology Field": "Electronic Technology," "Technology Implementation Type": "Optical Sensor," "Technology Application Scenario": "Industrial Automation Production"}. Matching "Electronic Technology" with the classification rules of the technology field dimension determines its corresponding "Electronic Technology" node; matching "Optical Sensor" with the classification rules of the technology type dimension determines its corresponding "Optical Sensor" node under "Sensor Technology"; and matching "Industrial Automation Production" with the classification rules of the application scenario dimension determines its corresponding "Industrial Application" node.

[0173] The matching results are recorded to form an entity representation-hierarchical node matching result set. List or dictionary data structures can be used to store the matching results, for example, {"Entity Representation":"Intelligent Optical Sensor","Matching Node":["Electronic Technology","Sensor Technology - Optical Sensor","Industrial Applications"]}.

[0174] Step S524: For each entity representation in the entity representation-hierarchical node matching result set, verify whether its matched hierarchical nodes conform to the logical order of the retrieval requirements, confirm that the hierarchical relationship between nodes is clear, remove entity representations with chaotic hierarchical matching logic, and obtain a set of correctly matched entity representations.

[0175] It is essential to ensure that the hierarchical relationships between nodes are clear. A clear hierarchical relationship means that the superior-inferior relationships between nodes are explicit and there are no logical contradictions. For example, "sensor technology" should be a technology type node under the field of "electronic technology," and there should not be an incorrect hierarchical relationship where "sensor technology" is the parent node of "electronic technology."

[0176] Entity descriptions with confusing hierarchical matching logic should be removed, as these can negatively impact the accuracy and usability of the entire retrieval hierarchy. For example, if an entity description matches nodes with a confusing hierarchical relationship, it may prevent the entity from being accurately located and categorized during retrieval.

[0177] During the implementation and verification, the relationships between the hierarchical nodes matching each entity description are checked based on the hierarchical structure and logical rules of the retrieval requirements. A hierarchical relationship model can be established, which contains the correct hierarchical relationships between nodes of different classification dimensions. The nodes matching each entity description are then compared with the hierarchical relationship model.

[0178] If a matching entity description is found to have a disordered node hierarchy, such as reversed node order or logical contradictions between nodes, then that entity description will be removed from the entity description-hierarchical node matching result set. The final set of correctly matched entity descriptions will contain all entity descriptions whose matched hierarchical nodes conform to the logical order required for retrieval.

[0179] Step S525: Classify the set of correctly matched entity representations according to the corresponding hierarchical nodes. Each hierarchical node corresponds to an entity category. Assign a category name that is consistent with the retrieval level to each category to obtain a named hierarchical entity category set.

[0180] During classification, the set of correctly matched entity representations is traversed, and each entity representation is added to the corresponding entity category based on its corresponding hierarchical node. A dictionary data structure can be used to store entity categories, where the keys are hierarchical nodes and the values ​​are lists of entity representations corresponding to that node. The resulting named hierarchical entity category set will contain all classified and named entity categories, forming an ordered hierarchical structure.

[0181] Step S526: Add a hierarchical node marker to each category in the named hierarchical entity category set. The marker content includes the hierarchical position, classification criteria, and the number of entity descriptions to which it belongs, and finally generate a hierarchical entity category set.

[0182] In this embodiment of the invention, a hierarchical node marker is added to each category in the named hierarchical entity category set. The hierarchical node marker can provide detailed information about the category in the retrieval hierarchy. The marker content includes the hierarchical position, the classification criteria, and the number of entity descriptions to which it belongs.

[0183] The hierarchical position clearly indicates the category's specific location within the search hierarchy, such as its level and parent node. This helps users quickly locate the category during a search. The classification criteria explain the standards used to divide the category, such as technical definitions or application scope, which helps users understand the category's classification logic. The number of entity descriptions reflects the number of entities contained within the category, providing valuable insight into assessing its size and importance.

[0184] When adding tags, each category in the named hierarchical entity category set is traversed. The hierarchical position can be determined based on the position of the corresponding hierarchical node in the retrieval hierarchy. For example, if the node corresponding to the category is "Electronic Technology - Sensor Technology - Optical Sensors," its hierarchical position can be determined as the third level, with the parent node being "Sensor Technology." The classification criteria can be referenced from the classification rules of the node in the retrieval hierarchy classification dimension criteria set. For example, if the "Optical Sensors" category is divided according to the sensor's technical implementation method (optical principle), then the classification criteria can be recorded as "Sensor technology implemented according to optical principles."

[0185] To determine the number of entity descriptions, you can count the length of the list of entity descriptions contained within that category. For example, if there are 10 entity descriptions under the category "Electronics Technology - Sensor Technology - Optical Sensors," then the number of entity descriptions is marked as 10. These tags are added to each category, and the tagging information can be stored using a dictionary or list data structure. The resulting hierarchical set of entity categories will contain all the tagged entity categories, forming a complete and detailed retrieval hierarchy.

[0186] Step S530: For each category in the hierarchical entity category set, sort out the relationship descriptions between entity descriptions within the category, organize the relationship descriptions according to the association logic of hierarchical nodes, so that the relationship descriptions match the entity category hierarchy, and obtain the set of relationship descriptions organized within the category.

[0187] In this embodiment of the invention, for each category in the hierarchical entity category set, the relationships between entity representations within that category are analyzed. The hierarchical entity category set has categorized entity representations according to different hierarchical nodes, and each category contains entity representations with similar characteristics. Various relationships may exist between these entity representations, such as technical associations, application associations, etc., and it is necessary to analyze the descriptions of these relationships.

[0188] The relationship representations are organized according to the association logic of hierarchical nodes. The association logic of hierarchical nodes defines the rules for the relationships between entities at different levels and within the same level. For example, entity representations under the same technology type node may have relationships such as technology competition or cooperation; entity representations under different technology type nodes may have relationships such as technology reference or application expansion.

[0189] Matching relational expressions to entity category hierarchy means that the relational expressions should conform to the position and characteristics of the category in the retrieval hierarchy. For example, within the category "Electronic Technology - Sensor Technology - Optical Sensors", the relational expressions should be related to the technical characteristics and application scenarios of optical sensors, and conform to the hierarchical positioning of this category in the field of electronic technology.

[0190] In the specific implementation, each category in the hierarchical entity category set is traversed. For each category, relational representations related to entity representations within that category are selected from the complete set of relational representations. This selection can be based on the association information in the entity representation-hierarchical node matching result set. Then, the selected relational representations are organized according to the association logic of the hierarchical nodes. A relational organization model can be established, which includes relational rules and organization methods under different hierarchical nodes. The relational representations are then classified and sorted according to this model.

[0191] For example, for the category "Electronic Technology - Sensor Technology - Optical Sensors", we filter out relational expressions related to the entity description of optical sensors, such as "Performance Comparison of Intelligent Optical Sensors and Digital Optical Sensors" and "Application Expansion of Optical Sensors in Industrial Inspection". Based on the category's position and association logic in the hierarchical structure, we organize these relational expressions, such as classifying performance comparison relationships into the technology competition relationship category and application expansion relationships into the application association relationship category.

[0192] Step S540: Among different levels of entity categories, sort out the relationship descriptions between entity descriptions across categories, organize these relationship descriptions according to the hierarchical association logic of categories, construct the association links between entities across categories, and obtain the set of relationship descriptions after cross-category organization.

[0193] In this embodiment of the invention, the relationships between entity representations across different hierarchical entity categories are analyzed. The hierarchical entity category set categorizes entity representations according to different hierarchical nodes. Various relationships may exist between categories at different levels, such as intersections between technical fields, or associations between technology types and application scenarios. It is necessary to identify these relationships between entity representations across categories from the complete set of relationship representations.

[0194] These relationships are organized according to a hierarchical logic of categories, which defines the rules governing the relationships between different levels of categories. For example, a technology field is a macro-level classification, a technology type is a sub-classification within a technology field, and an application scenario is the actual application of a technology type. Therefore, there are clear hierarchical relationships between different technology fields and technology types, and between technology types and application scenarios.

[0195] Constructing cross-category entity relationship links clearly demonstrates the connections between entities at different levels, helping users understand the technical links and application extensions between different categories during retrieval. In implementation, the combination relationships between entity categories at different levels are first determined. This can be done by traversing the hierarchical entity category set to find all possible cross-category combinations. For each cross-category combination, relationship expressions related to entity expressions within these categories are selected from the complete set of relationship expressions. Then, the selected relationship expressions are organized according to the hierarchical relationship logic of the categories. A cross-category relationship organization model can be established, containing the relationship rules and organization methods between different levels of categories. The relationship expressions are then classified and sorted according to this model.

[0196] Step S550: Integrate the hierarchical entity category set, the set of relational descriptions organized within categories, and the set of relational descriptions organized across categories. Set the display order of entities and relations according to the retrieval logic to form a hierarchical knowledge framework, thus obtaining a hierarchical knowledge framework set.

[0197] In practical implementation, the hierarchical entity category set, the set of relational representations organized within categories, and the set of relational representations organized across categories are first merged. A dictionary or list data structure can be used to store the merged information, where each hierarchical node corresponds to an entity category and its associated relational representation. Then, the merged information is sorted according to the retrieval logic. A sorting rule can be defined, ordering by technical field, technical type, and application scenario, while within each category, further sorting is done based on factors such as entity importance and relational relevance.

[0198] Step S560: Optimize the path of the hierarchical knowledge framework set so that any entity representation can be linked to other related entities through relation representations, and finally construct a knowledge graph for intellectual property retrieval.

[0199] A knowledge graph should possess strong connectivity and relevance, enabling connections between any entity representation and other related entities via relational representations. For example, when a user retrieves an entity representation, the knowledge graph should be able to quickly find other related entity representations through relational representations, whether they are entities within the same category or entities from different categories. Graph algorithms can be used for path optimization to analyze entities and relationships within a hierarchical knowledge framework. Entity representations can be viewed as nodes in a graph, and relational representations as edges, thus constructing a knowledge graph model.

[0200] Graph algorithms, such as breadth-first search and depth-first search, are used to check the connectivity between nodes. If it is found that the paths between some nodes are too long or do not exist, the paths can be optimized by adding or adjusting the relational descriptions. For example, if the relationship between "intelligent optical sensor" and "medical image diagnostic system" is not obvious, the relationship between them can be strengthened by adding a relational description such as "research on the application of intelligent optical sensor in medical image diagnostics".

[0201] Furthermore, weights can be assigned to relational representations, assigning different weights to each edge based on the importance and relevance of the relationship. During retrieval, associated paths can be prioritized based on these weights, improving the accuracy and efficiency of the search. Ultimately, through path optimization, a knowledge graph for intellectual property retrieval is constructed. This knowledge graph will possess a well-defined hierarchical structure, connectivity, and relevance, meeting various user needs in intellectual property retrieval and helping users quickly and accurately obtain the intellectual property information they require.

[0202] Figure 2 This is a schematic diagram of a hardware entity of a knowledge graph construction device provided in an embodiment of the present invention, such as... Figure 2 As shown, the hardware entity of the knowledge graph construction device 1000 includes a processor 1001 and a memory 1002, wherein the memory 1002 stores a computer program that can run on the processor 1001, and the processor 1001 executes the program to implement the steps in the method of any of the above embodiments.

[0203] The memory 1002 stores computer programs that can run on the processor. The memory 1002 is configured to store instructions and applications that can be executed by the processor 1001. It can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) in the processor 1001 and the various modules in the knowledge graph construction apparatus 1000. It can be implemented by flash memory or random access memory (RAM).

[0204] When the processor 1001 executes the program, it implements the steps of the knowledge graph construction method for intellectual property retrieval described above. The processor 1001 typically controls the overall operation of the knowledge graph construction apparatus 1000.

[0205] This invention provides a computer storage medium storing one or more programs that can be executed by one or more processors to implement the steps of the knowledge graph construction method for intellectual property retrieval as described in any of the above embodiments.

[0206] The aforementioned computer storage media / memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic random access memory (FRAM), flash memory, magnetic surface memory, optical disc, or compact disc read-only memory (CD-ROM), etc.; or it can be various terminals that include one or any combination of the above-mentioned memories, such as mobile phones, computers, tablet devices, personal digital assistants, etc.

[0207] The above description is merely an embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for constructing a knowledge graph for intellectual property retrieval, characterized in that, The method includes: Read the target intellectual property text dataset and its built-in parallel corpus and reference-related texts, and mine the implicit synonym mapping clues, reference tracing clues and technical topic association clues between the texts. Organize them to obtain a weak supervision signal set, which contains various implicit marker information used to indicate the semantic association of the texts. Based on the weakly supervised signal set, synonym pairs in parallel corpora and semantic association pairs in referenced texts are extracted. The extracted synonym pairs and semantic association pairs are cross-validated and deduplicated to obtain a text alignment benchmark set. The text alignment benchmark set is globally associated and matched with the target intellectual property text dataset. The mapping association derivation of the text alignment benchmark set completes the automatic association and labeling of intellectual property entities and semantic relationships between entities in the text, resulting in an initial structured knowledge unit set. The initial structured knowledge unit set contains entity representations and relationship representations that have not undergone credibility verification. The initial set of structured knowledge units is subjected to credibility normalization. Through semantic consistency verification and domain adaptability screening, entity and relation representations that meet the preset credibility requirements are retained to obtain a standardized set of knowledge units. The standardized knowledge unit set is organized into entities and relationships according to the hierarchical requirements of intellectual property retrieval, thereby constructing a knowledge graph for intellectual property retrieval. The knowledge graph includes technical topic levels, entity association links, and semantic mapping paths that are adapted to retrieval requirements.

2. The method as described in claim 1, characterized in that, The process involves reading the target intellectual property text dataset and its inherent parallel corpus and citation-related texts, mining implicit synonym mapping clues, citation tracing clues, and technical topic association clues between texts, and organizing them into a weakly supervised signal set, including: Read the target intellectual property text dataset and its built-in parallel corpus and referenced text. According to the original paragraph division mark of the text, classify the parallel corpus and referenced text into independent storage partitions. The text in each partition retains the original page number, paragraph number and reference mark information, and obtain the partitioned storage of parallel corpus partition text set and referenced text partition text set. Traverse each text segment in the parallel corpus partitioned text set, compare it sentence by sentence with all other texts in the set, identify sentence combinations that have the same core meaning but different sentence structures, add cross-text association tags to each sentence combination, and obtain a set of labeled parallel corpus sentence combinations; Traverse each text segment in the text set of reference-related text partitions, trace the source of the reference marked in the text, sort out the reference passing links between all texts, add source-tracing association tags to the starting text and ending text of each link, and obtain a set of tagged reference-related links; Full-text keyword co-occurrence statistics are performed on all texts in the parallel corpus partitioned text set and the citation-related text partitioned text set to identify the same technical topic expression that appears simultaneously in different texts. Topic association tags are added to the corresponding texts of each co-occurring topic to obtain a labeled technical topic co-occurrence text set. By integrating the set of sentence combinations with tags in parallel corpora, the set of citation association links, and the set of texts co-occurring with technical topics, and removing text entries with duplicate tags, so that each tag corresponds to only one text association relationship, we obtain a set of deduplicated multi-type association tags. The deduplicated set of multi-type association tags is formatted in a unified manner, and the type, position, and content of each association tag are sorted according to fixed fields to generate a set of weak supervision signals containing various implicit tag information.

3. The method as described in claim 2, characterized in that, The process involves traversing each text segment in the parallel corpus partitioned text set, comparing it sentence-by-sentence with all other texts in the set, identifying sentence combinations that have identical core meanings but differ in sentence structure, and adding cross-text association tags to each sentence combination to obtain a set of tagged parallel corpus sentence combinations, including: Read the labeled set of parallel corpus sentence combinations, extract the content of the two sentences in each sentence combination and the source information of the parallel corpus text to which they belong, split the sentence content into a continuous sequence of individual words, retain the original word order and part-of-speech information of the words, and obtain a set of sentence word sequences; For each set of sentence word sequences, perform word matching on each of the two sentence word sequences, mark the completely overlapping words and their positions in the sequence, count the proportion of overlapping words to the total number of words in the two sentences, record the continuous distribution of the matched words, and obtain the sentence word matching detail set. For sentence combinations where the continuous distribution of matched words in the sentence word matching detail set meets the requirements, the complete semantic expression of the two sentences is restored. The technical meaning, applicable scenarios, and limited scope expressed by the two sentences are compared to confirm that the core expression is completely consistent, thus obtaining a semantically consistent subset of sentence combinations. Sentence structure analysis is performed on each sentence combination in the semantically consistent sentence combination subset to identify the structural elements of the sentences, compare the differences in the structural elements of two sentences, mark the sentence combinations with different sentence structures, and obtain the sentence combination subset with sentence structure differences. Add cross-text synonym association tags to each sentence combination in the sentence difference sentence combination subset. The tag content includes the text source of the sentence, the semantic consistency confirmation result and the sentence difference type, resulting in a sentence difference sentence combination set with synonym tags. The set of sentence segments with synonym tags is grouped and organized according to the text source of the parallel corpus. Each group corresponds to the parallel corpus text from the same source. Finally, the set of parallel corpus sentence segments with tags is updated and generated.

4. The method as described in claim 3, characterized in that, The process involves combining clauses whose consecutive distribution of matched words in the clause word matching detail set meets the requirements, restoring the complete semantic expression of the two clauses, comparing their expressed technical meanings, applicable scenarios, and limited scopes, confirming that their core expressions are completely consistent, and obtaining a semantically consistent subset of clause combinations, including: Read the sentence combinations that meet the requirements for the continuous distribution ratio of matching words in the sentence word matching details set, extract the complete text content of the two sentences in each combination, including the contextual auxiliary expressions before and after the sentences, and obtain the sentence content set with context. Semantically restore the content of each clause with context, sort out the core information expressed by the clause, and organize the core information into structured semantic description items according to a fixed logical order to obtain a set of structured semantic descriptions of the clauses; For each sentence combination, the structured semantic description items of the two sentences are compared item by item to check whether the content of each core field is completely overlapping. The comparison results of each field are recorded to obtain a set of semantic field comparison results. For sentence combinations where all core fields in the semantic field comparison result set are completely overlapping, further check whether there are semantic conflicts in the contextual auxiliary expressions of the sentences, confirm that the contextual expressions do not affect the consistency of the core semantics, and obtain a context-compatible subset of sentence combinations. Add a semantic consistency confirmation tag to each combination in the context-compatible sentence combination subset. The tag content includes the core field comparison result, the context compatibility confirmation result, and the semantic consistency judgment basis, resulting in a sentence combination set with semantic confirmation tags. The set of sentence combinations with semantic confirmation tags is sorted and organized according to the text length of the sentences, and the sentence combinations with complete core semantic expression are selected to finally generate a semantically consistent subset of sentence combinations.

5. The method as described in claim 1, characterized in that, Based on the weakly supervised signal set, synonym pairs are extracted from parallel corpora and semantic association pairs from cited texts. The extracted synonym pairs and semantic association pairs are cross-validated and deduplicated to obtain a text alignment benchmark set, including: Read the weakly supervised signal set, filter out all entries with synonym mapping tags, extract the related clause content and the source information of the parallel corpus of each entry, and pair up the clauses with the same synonym association to obtain the set of synonym clause pairs of the parallel corpus. Read the weakly supervised signal set, filter out all entries with citation traceability tags and technical topic association tags, extract the associated text content and the source information of the citation associated text of each entry, and pair the text fragments with citation or topic association to obtain a set of citation associated text semantic association pairs. The content of each clause pair in the set of synonymous clause pairs in the parallel corpus is compared with that of each association pair in the set of semantic association pairs in the referenced text. The association pairs with completely overlapping content are identified, and the association type and source information of the overlapping entries are recorded to obtain a set of cross-type overlapping association pairs. For each entry in the set of cross-type overlapping association pairs, verify whether its semantic reference in the parallel corpus and the referenced text is completely consistent. After confirming that there is no semantic deviation, retain the tag of one type of association and remove the entries with duplicate tags to obtain the deduplicated set of cross-type association pairs. The non-overlapping entries in the set of synonymous sentence pairs in parallel corpora, the non-overlapping entries in the set of semantic association pairs in referenced texts, and the deduplicated cross-type association pairs are integrated to ensure that each association pair appears only once, resulting in the integrated set of association pairs. The integrated set of association pairs is categorized according to association type, and a unified alignment benchmark is added to each type of association pair to generate a text alignment benchmark set containing synonym pairs and semantic association pairs.

6. The method as described in claim 5, characterized in that, The process involves reading the weakly supervised signal set, filtering out all entries with citation traceability markers and technical topic association markers, extracting the associated text content and source information of each entry's citation association text, and pairing text fragments with citation or topic association to obtain a set of citation association text semantic association pairs, including: Read the weakly supervised signal set and extract all entries with reference tracing tags. Each entry contains the starting text, the ending text, and the reference link information. Extract the fragment content corresponding to the starting text and the ending text to obtain a set of reference tracing text fragment pairs. Read the weak supervision signal set and extract all entries with technical topic association tags. Each entry contains co-occurring topic text fragments and a list of associated texts. Combine the co-occurring topic text fragments with the text fragments in each of the associated text lists to obtain a set of technical topic association text fragment pairs. Iterate through each fragment pair in the set of reference source text fragment pairs, check whether there is a direct reference statement annotation between the starting text fragment and the ending text fragment, confirm the authenticity of the reference relationship, remove fragment pairs without direct reference annotation, and obtain the verified set of reference source fragment pairs; Iterate through each fragment pair in the set of technical topic-related text fragment pairs, check whether the technical topic descriptions between the co-occurring topic text fragments and the related text fragments are completely consistent, confirm the rationality of the topic association, remove fragment pairs with topic description deviations, and obtain the verified set of technical topic-related fragment pairs. The integrated and verified set of citation source fragment pairs and the verified set of technical topic-related fragment pairs are combined. Fragment pairs with completely duplicate content are removed, so that each fragment pair corresponds to only one unique relationship, resulting in the integrated set of citation-related fragment pairs. To integrate the reference-related fragments, a semantic association type tag is added to each fragment pair in the collection. The tag types are divided into reference source association and technical topic association, ultimately generating a collection of reference-related text semantic association pairs. Specifically, the process of traversing each fragment pair in the set of related technical topic text fragments involves checking whether the technical topic descriptions between co-occurring topic text fragments and related text fragments are completely consistent, confirming the rationality of the topic association, and removing fragment pairs with topic description deviations. This results in a verified set of related technical topic fragment pairs, including: Read each fragment pair in the set of text fragment pairs associated with technical topics, extract the complete content of co-occurring topic text fragments and associated text fragments, including the context descriptions before and after the fragments, to obtain a set of technical topic fragment pairs with context; For each pair of technical topic fragments with context, the core technical topic expression in the co-occurring topic text fragments is extracted, and it is broken down into multiple key components to obtain a set of core technical topic elements. Extract the core technical theme statements from the related text fragments, and obtain the related text technical theme element set by the same decomposition method. Compare the core technical theme element set with the related text technical theme element set item by item, record the comparison result of each element, and obtain the technical theme element comparison result set. For fragment pairs in the technical topic element comparison result set where all key components completely overlap, further check whether there are any statements in the context description that contradict the technical topic, confirm that the context description does not affect the rationality of the topic association, and obtain a set of context-compatible technical topic fragment pairs; For fragment pairs in the technical topic element comparison result set that have element differences, they are judged as topic expression deviations and removed from the technical topic associated text fragment pair set to obtain the technical topic fragment pair set after deviation removal; The set of context-compatible technical topic fragment pairs is integrated with the set of technical topic fragment pairs after bias removal. A topic association verification pass mark is added to each fragment pair, and finally a verified set of technical topic associated fragment pairs is generated.

7. The method as described in claim 1, characterized in that, The process involves globally associating and matching the text alignment benchmark set with the target intellectual property text dataset. Through mapping and association derivation using the text alignment benchmark set, the automatic association and labeling of intellectual property entities and semantic relationships between entities in the text are completed, resulting in an initial set of structured knowledge units, including: Read the text alignment reference set and the target intellectual property text dataset, divide the target intellectual property text dataset into multiple text blocks according to chapters, and retain the original chapter number, page number and paragraph mark of each text block to obtain the divided intellectual property text block set; Traverse each association pair in the text alignment benchmark set, extract the core expression content in the association pair, use it as the search content and perform full-text matching with each text block in the segmented intellectual property text block set to identify the text block containing the core expression content, and obtain the matching intellectual property text block set. For each matched intellectual property text block, the specific location of the core expression content within the text block is located, and the text content within a range before and after that location is extracted. The expression units with independent technical meaning are identified, and the set of technical expression units within the text block is obtained. Based on the association relationships in the text alignment benchmark set, the semantic associations between technical expression units are derived. Technical expression units with synonym or reference associations are bound together, and semantic association tags are added to the bound units to obtain a set of bound technical expression units with tags. Each unit in the set of labeled binding technology description units is organized in a structured manner to clarify the technical meaning of the unit, the information of related units, and the position of the text block to which it belongs, forming a structured knowledge item and obtaining an initial set of structured knowledge items; By integrating all the initial set of structured knowledge entries and removing duplicate entries, each entry corresponds to a unique technical expression unit and its relationship, resulting in the initial set of structured knowledge units.

8. The method as described in claim 7, characterized in that, For each matched intellectual property text block, the specific location of the core expression content within the text block is located, and the text content within a range before and after that location is extracted. Expression units with independent technical meaning are identified within these extraction units, resulting in a set of technical expression units within the text block, including: Read each matched intellectual property text block, split the text block into multiple sentence units according to punctuation marks, and retain the original paragraph marks and sentence sequence numbers in each sentence unit to obtain a set of sentence units after splitting the text block; Based on the core expression content in the text alignment benchmark set, the sentence unit containing the core expression content is located in the sentence unit set, and the start and end positions of the sentence unit are marked to obtain the located core sentence unit set. Extract the content of the core sentence unit after positioning and the two sentence units before and after it, combine these sentence units into continuous text segments, retain the position mark of each sentence unit, and obtain the set of extended text segments around the core sentence unit; Traverse each text fragment in the extended text fragment set, identify the expression content in the fragment that can independently express a complete technical concept, mark the start and end positions of the expression content, and obtain a candidate set of independent technical expressions; For each independent technical expression candidate, its semantic independence in the extended text fragment is verified to confirm that the expression can be fully understood without relying on other text fragments. Expressions that require contextual supplementation are removed to obtain a set of semantically independent technical expressions. The set of semantically independent technical expressions is sorted according to the length of the expression content, and the technical expression units with complete expressions are selected to finally generate the set of technical expression units within the text block.

9. A knowledge graph construction apparatus, comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the steps of the method according to any one of claims 1 to 8.