Semantic retrieval method and device based on patent document database
By pre-processing and extracting tags from the user-input search text, and combining this with similarity calculations from a patent document database, the problem of low semantic search accuracy in existing technologies is solved, resulting in more accurate patent search results and improved utilization of patent information.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2026-04-07
AI Technical Summary
Existing semantic retrieval methods for patents suffer from low accuracy and poor performance. They fail to effectively utilize the classification and association of different keywords in the text, resulting in poor customer understanding of the relevance of the search results.
By pre-processing the search text input by the user, extracting domain tags and technical descriptions, and converting them into tag word vectors, the patent number with the highest similarity is determined in the patent document database, and a scoring dictionary is constructed. The patent number with the highest total score is then returned as the search result.
It achieves more accurate semantic retrieval, improves retrieval precision, provides retrieval results that better meet user needs, and enhances the utilization rate and value of patent information.
Smart Images

Figure CN121807992A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates to the technical field of data processing, and particularly relates to a semantic retrieval method and device based on a patent file database, a storage medium and an electronic device. BACKGROUND
[0002] Currently, the information of the technical field, the technical feature, the technical problem and the technical effect in the patent is of high value to the customer.
[0003] The existing patent semantic retrieval method described in the platform and the literature is generally a pure black box method, and the key information in the input text is not displayed to the customer. The customer simply inputs the text and then performs background retrieval, and then a certain number of related patents are displayed according to the retrieval results. However, the customer has poor understanding of the original text and the relevance of the retrieval results. In addition, in the current retrieval method, the retrieval is mainly based on the keywords in the original text, and the different keywords in the text are not classified and associated. Therefore, part of the semantic content is not fully understood and utilized, resulting in low semantic retrieval accuracy and poor use effect. SUMMARY
[0004] The purpose of the embodiments of the present disclosure is to provide a semantic retrieval method and device based on a patent file database, a storage medium and an electronic device, so as to solve the problem of low semantic retrieval accuracy and poor use effect in the prior art.
[0005] The embodiments of the present disclosure adopt the following technical solutions: a semantic retrieval method based on a patent file database, comprising: performing preset processing on a retrieval text input by a user to obtain preset labels of the retrieval text, the preset labels at least including field labels and / or technical description combinations; converting each preset label into a label word vector, and taking each label word vector as a target word vector in turn, determining m patent numbers with the highest similarity to the target word vector in a patent file database, and determining a score of the patent number according to the similarity between the patent number and the target word vector; constructing a score dictionary of each patent number, determining a total score of each patent number based on the score dictionary; and feeding back the k patent numbers with the highest total scores to the user as the retrieval result of the retrieval text.
[0006] The embodiment of the present disclosure further provides a semantic retrieval device based on a patent file database, comprising: a preprocessing module configured to perform preset processing on a retrieval text input by a user to obtain preset labels of the retrieval text, wherein the preset labels at least include field labels and / or technical description combinations; a retrieval module configured to convert each of the preset labels into a label word vector, and sequentially use each of the label word vectors as a target word vector to determine m patent numbers with the highest similarity to the target word vector in a patent file database, and determine scores of the patent numbers according to the similarity between the patent numbers and the target word vector; a score statistics module configured to construct a score dictionary of each of the patent numbers, and determine total scores of each of the patent numbers based on the score dictionary; and a feedback module configured to feed back k patent numbers with the highest total scores as retrieval results of the retrieval text to the user.
[0007] The embodiment of the present disclosure further provides a storage medium storing a computer program, wherein the computer program is executed by a processor to implement the steps of the semantic retrieval method based on the patent file database.
[0008] The embodiment of the present disclosure further provides an electronic device comprising at least a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the semantic retrieval method based on the patent file database when executing the computer program stored in the memory.
[0009] The embodiment of the present disclosure has the beneficial effects that: the retrieval text input by the user is identified and extracted based on the field labels and the technical description combinations, the retrieval text is structured and split, the most critical data corresponding to the main content module of the patent in the retrieval text is extracted for subsequent retrieval, and the retrieval process based on the patent file database is performed through more fine-grained and more accurate retrieval information, the retrieval results most meeting the user's demand are obtained by combining the scores of the retrieval results under different preset labels, and the purpose of improving the retrieval precision is achieved. BRIEF DESCRIPTION OF DRAWINGS
[0010] In order to more clearly illustrate the technical solutions in the one or more embodiments of the present disclosure or the prior art, the drawings needed to be used in the embodiment or prior art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments described in the present disclosure, and other drawings can be obtained by those skilled in the art without any creative effort on the basis of these drawings.
[0011] Figure 1 The flowchart of the semantic retrieval method based on the patent file database in the first embodiment of the present disclosure;
[0012] Figure 2Fig. 1 is a schematic diagram of an architecture of a patent document database in a first embodiment of the present disclosure;
[0013] Figure 3 Fig. 2 is a schematic diagram of a structure of a semantic retrieval device based on a patent document database in a second embodiment of the present disclosure. DETAILED DESCRIPTION
[0014] In order to make the person skilled in the art better understand the technical solutions in one or more embodiments of the present specification, the technical solutions in one or more embodiments of the present specification will be described clearly and completely in conjunction with the drawings in one or more embodiments of the present specification. Obviously, the described embodiments are only a part of the embodiments of the present specification, not all the embodiments. Based on one or more embodiments of the present specification, all other embodiments obtained by those skilled in the art without creative labor should belong to the protection scope of the present document.
[0015] The patent semantic retrieval method described in the existing platform and literature is generally a pure black box method, which does not show the key information in the input text to the customer. The customer simply inputs the text and then performs background retrieval, and then displays a certain number of related patents according to the retrieval results. However, the customer has poor understanding of the original text and the relevance of the retrieval results. In addition, in the current retrieval method, the retrieval is mainly based on the keywords in the original text, and the different keywords in the text are not classified and associated. Therefore, part of the semantic content is not fully understood and utilized, resulting in low semantic retrieval accuracy and poor use effect.
[0016] In order to solve the above problems, the first embodiment of the present disclosure provides a semantic retrieval method based on a patent document database, the flow chart of which is as shown in Figure 1 The steps S10 to S40 include at least:
[0017] S10, performing preset processing on the retrieval text input by the user to obtain preset labels of the retrieval text.
[0018] When a user performs semantic retrieval, the user inputs a relevant retrieval text in a search box. The retrieval text can include technical features, technical problems solved by the technical features, technical effects achieved by the technical features, and application fields, and the retrieval text also includes redundant information and interference words used in the process of text description. After receiving the retrieval text input by the user, the retrieval engine first performs preset processing on the retrieval text to extract preset labels of the retrieval text. The preset labels in the embodiment include at least field labels and / or technical description combinations. The field labels are mainly used to indicate technical field information corresponding to the retrieval text, and the technical description combinations are mainly used to represent descriptions of technical features, technical problems, and technical effects in the retrieval text, such as technical feature subjects and association relationships between the subjects, technical problem subjects and description information of the problem subjects, technical effect subjects and description information of the effect subjects, and the like. Through extraction of the preset labels of the retrieval text, the retrieval text is structurally split, more refined field information and description information corresponding to main content of a patent file can be extracted from the retrieval text, the influence of redundant information is removed in subsequent retrieval, and the accuracy of semantic retrieval is improved.
[0019] In some embodiments, after obtaining the preset labels of the retrieval text, all the preset labels can be displayed to the user, so that the user can have a deeper understanding of the input retrieval text, and the user can conveniently filter and use the final retrieval result.
[0020] In some embodiments, the process of performing preset processing on the retrieval text input by the user to obtain the preset labels of the retrieval text mainly includes: performing first named entity recognition on the retrieval text by using a first named entity recognition model to identify field labels in the retrieval text. Specifically, the first named entity recognition model is a model pre-trained for entity recognition of field information in a text. The model can extract field labels by using specific keywords or key subjects in the retrieval text. In actual use, there are cases where field labels cannot be extracted from the retrieval text, for example, the retrieval text content is less, the user's expected use scenario is not clear, or there is an entity with a specific function.
[0021] In some embodiments, the process of performing preset processing on the retrieval text input by the user to obtain the preset labels of the retrieval text can also include: performing second named entity recognition on the retrieval text by using a second named entity recognition model to identify subject labels and description labels in the retrieval text; and performing punctuation-based feature sentence division on the retrieval text to combine all the subject labels and / or all the description labels belonging to the same feature sentence to obtain a technical description combination.
[0022] Specifically, the second named entity recognition model is a pre-trained model for entity recognition of subject information and description information in the text. The subject information can include subjects for representing features, subjects for representing technical problems, and subjects for representing technical effects. The description information includes feature relationships between feature subjects, specific descriptions of technical problem subjects and technical effect subjects, etc. Corresponding to the search text, it can have multiple different subject labels and description labels. These labels are actually a simplified content extraction of the technical description content of the search text, which realizes the filtering of redundant information and interference words.
[0023] After extracting the labels, the association and combination between the subject labels and the description labels need to be performed to clarify the real content disclosed in the search text and represent it through a simplified and accurate label combination. Generally, there is an association between the subjects described in the same sentence. Therefore, the embodiment takes the position of the punctuation mark as the basis to divide the feature sentences of the search text, and takes the feature sentence as the basic unit for determining the association between the labels to determine the technical description combination.
[0024] For the divided feature sentences, the sentence start index and the sentence end index corresponding to each feature sentence are determined, which can be the position information of the character marking the start of the feature sentence and the position information of the character marking the end of the feature sentence. Corresponding to the subject label and the description label, the corresponding word start index and word end index need to be determined when determining the feature sentence to which they belong. Then, according to the word start index and the word end index, all subject labels and / or description labels located within the sentence start index and the sentence end index of each feature sentence are determined, which is equivalent to determining a set of subject labels and description labels having an association relationship. Arranging the subject labels and description labels belonging to the same feature sentence according to the ascending order of the word start index, a technical description combination is obtained, which corresponds to recording all key subjects and specific relationships between key subjects in the feature sentence, achieving the purpose of simplifying the technical solution. In specific implementation, for each feature sentence, the subject labels and / or description labels with the word start index greater than or equal to the sentence start index and the word end index less than or equal to the sentence end index are determined as the subject labels and / or description labels located within the sentence start index and the sentence end index of each feature sentence.
[0025] S20, convert each preset label into a label word vector, and take each label word vector as a target word vector in turn to determine m patent numbers with the highest similarity to the target word vector in the patent file database, and determine the score of the patent number according to the similarity between the patent number and the target word vector.
[0026] After step S10, one or more preset labels of the search text can be extracted, i.e., all the preset labels mentioned above can be searched in the patent file database respectively. When searching, each preset label is converted into a corresponding label word vector, and each label word vector is sequentially taken as a target word vector to determine the m most similar patent numbers in the patent file database. The m patent numbers are the search results of the preset label currently taken as the target word vector in the database, and the similarity between the patent numbers and the target word vector is the score of the patent numbers under the corresponding preset label. In actual implementation, m patent numbers and their scores are output for each preset label, so there may be m times the number of preset labels and their scores as the output of this step.
[0027] Specifically, corresponding to the combination of field labels and technical descriptions, the patent file database has different index libraries to record the association between all patent numbers and fields and technical descriptions, i.e., the field index library records the corresponding relationship between all patent numbers and their corresponding field indexes, and the technical description index library records the corresponding relationship between all patent numbers and their technical description indexes. Based on the identification and extraction results of the preset labels in step S10, different index libraries are used for searching when the preset labels include different types.
[0028] In some embodiments, when the preset labels include field labels, the field labels are first converted into corresponding field word vectors, and then the distance between each field index and the field word vector is determined in the field index library. Based on the calculation results of the distance between all field indexes and the field word vector, the m field indexes closest to the field word vector are determined, and the distance between each field index in the m field indexes and the field word vector is normalized to obtain the score of the m field indexes. Finally, the m field indexes respectively corresponding to the patent numbers are determined in the field index library, and the score of the field index is taken as the score of the corresponding patent number under the current field label.
[0029] In some embodiments, when the preset labels include technical description labels, the technical description labels are first converted into corresponding technical description word vectors, and then the distance between each technical description index and the technical description word vector is determined in the technical description index library. Based on the calculation results of the distance between all technical description indexes and the technical description word vector, the m technical description indexes closest to the technical description word vector are determined, and the distance between each technical description index in the m technical description indexes and the technical description word vector is normalized to obtain the score of the m technical description indexes. Finally, the m technical description indexes respectively corresponding to the patent numbers are determined in the technical description index library, and the score of the technical description index is taken as the score of the corresponding patent number under the current technical description label.
[0030] Generally, there is only one extracted field tag corresponding to the search text, which represents the field for the most desired application or search, but the technical description tag can reflect the contents of technical features, technical problems and technical effects, and there can be multiple technical subjects and technical descriptions corresponding to each aspect, so the number of technical description tags is usually not unique. In the process of determining the corresponding patent number and score, the determination of the corresponding patent number and score needs to be performed for each technical description tag. There can be the same patent or different patents between the m patent numbers corresponding to different technical description tags.
[0031] S30, constructing a score dictionary of each patent number, and determining the total score of each patent number based on the score dictionary.
[0032] The score dictionary is constructed based on the patent number, which records the scores of each patent under different preset tags. It can be constructed in the following form: {‘patent number’: [{‘technical field’: []}, ‘technical description’: []]}, wherein the brackets after ‘technical field’ and ‘technical description’ are used to record the score of the current patent number under the combination of field tag and technical description. Especially the technical description score is recorded as the sum of the scores of the patent number under different technical description combinations. Corresponding to the calculation process of the total score, first, calculate the product of the sum of the scores of the above patent number under different technical description combinations and the preset technical weight. The role of the technical weight is to balance the quantity difference between the technical field score and the technical description score. Since the technical description score is usually the sum of the scores of the patent number under multiple technical description combinations, its specific value will be much larger than the technical field score, so it is adjusted by the technical weight. Then add the product to the technical field score of the patent number, and the sum is recorded as the total score of the patent number. It should be noted that the specific value of the above technical weight can be set based on actual demand and historical experience, and the embodiment does not make specific limitation.
[0033] S40, feeding back the k patent numbers with the highest total scores to the user as the search results of the search text.
[0034] After calculating the total scores of all patent numbers, the patent file search results based on the extraction results of the preset tags of the search text can be obtained. The k patent numbers with the highest scores are selected as the search results and fed back to the user. The scheme of the above k patent numbers corresponds to the technical content most similar to the actual technical content represented by the search text.
[0035] In actual implementation, the total scores of all patent numbers can be normalized and sorted, the total scores are limited between score_low and score_high, and then the first k patent numbers are returned.
[0036] The embodiment extracts the recognition of the search text input by the user based on the combination of the field label and the technical description, realizes the structured splitting of the search text, extracts the most critical data in the search text corresponding to the main content module of the patent for subsequent search, realizes the search based on the patent file database, and realizes the search based on the search information with finer granularity and higher accuracy. Combining the score of the search result under different preset labels, the most suitable search result for the user demand is obtained, and the purpose of improving the search precision is achieved.
[0037] In the actual implementation process, before the user performs semantic search, the construction process of the patent file database can also be included. The patent file database also needs to be realized based on information with higher fine degree and better correlation such as field label and technical description combination during the construction process, and cooperates with the semantic search method provided in the embodiment to realize more accurate search result feedback.
[0038] For a patent file, it mainly includes the following parts: technical field part, technical feature part, technical problem part, and technical effect part. Among them, the technical field part mainly refers to the part in the specification for indicating the actual technical field to which the patent file belongs. The technical field can be described by a specific field information, or can be described by an entity with specific function contained in the patent file, such as a method entity or a device entity. The technical feature part mainly refers to the part for describing the actual technical solution of the patent file, which can correspond to the content of the claims or the invention content part in the specification. The technical problem part mainly refers to the technical problem actually solved by the patent file, which is usually described in the background technology part of the specification. The technical effect part mainly refers to the beneficial effect achieved by the technical solution, which is usually in the latter half of the invention content.
[0039] In some embodiments, the database of patent documents can be first constructed based on technical fields. For example, the technical field part of the patent document is subjected to named entity recognition by a first named entity recognition model, where there is a description of the field information and the specific device or method involved in the patent document. The corresponding field label and subject label can be identified by the NER model, which corresponds to the technical field information and the specific device or method of the patent, respectively. In actual application process, there are some patent documents that may not have a description of the field information in the technical field part, so the corresponding field label of this part of the patent document may not be identified, but the technical field part will record the subject of the disclosed scheme, so the subject label of each patent document should be identified. Based on the named entity recognition result, all patent documents are divided into first type patent documents and second type patent documents, wherein the first type patent documents are the patent documents with identified field labels and subject labels, and the second type patent documents are the patent documents with only identified subject labels.
[0040] According to the extracted field and corresponding specific device / method information of all first subject labels that identify the field label and the subject label, a mapping dictionary is constructed. After the mapping dictionary for all first type patent documents is constructed, the corresponding field label of each second subject label in the second type patent document can be determined based on the mapping dictionary, so as to realize the allocation of specific field information for each patent document. Specifically, based on the constructed mapping dictionary, for each second subject label, the first subject label most similar to the second subject label can be determined from the mapping dictionary. Then, based on the similarity of the subject labels, the corresponding field labels should also have similarity, so at least one first field label corresponding to the first subject label most similar to the second subject label in the mapping dictionary is assigned to the second subject label as at least one second field label corresponding to the second subject label. After the above processing is performed on all second type patent documents, the field information applicable to each patent document can be determined.
[0041] After the field label of each patent document is determined, the corresponding relationship between the patent number of each patent document and the corresponding first field label is established, and a patent document database based on technical fields is formed. In actual establishment process, the patent number can be the application number or the publication number of the patent document, and the corresponding relationship between each patent number and the field can be saved in the form of a matrix.
[0042] In some embodiments, the patent document database can also be constructed based on the technical description. For example, the technical description part of the patent document is subjected to named entity recognition by a second named entity recognition model, where there are description subjects and specific description contents of the patent document. The corresponding subject label and description label can be identified by the NER model, corresponding to the description subjects and specific description contents in the patent document. In actual implementation, the subject label and the description label correspond to different representations of different content parts of the patent document. Specifically, in the case of a technical description part being a claim, the subject label is used to represent the feature subjects in the claim, and the description label is used to represent the feature relationships between the feature subjects; in the case of a technical description part being a background technology part, the subject label is used to represent the technical problem subjects in the background technology part, and the description information is used to represent the problem description of the technical problem subjects; in the case of a technical description part being a beneficial effect part, the subject label is used to represent the technical effect subjects in the beneficial effect part, and the description information is used to represent the effect description of the technical effect subjects.
[0043] After the labels are extracted, the association and combination between the subject labels and the description labels need to be performed to clearly express the disclosed true content of the patent document and to express it through a concise and accurate label combination. Generally, there is an association between the subjects described in the same sentence, so this embodiment takes the position of the punctuation mark as the basis to divide the feature sentences in the technical description part, and takes the feature sentence as the basic unit for determining the association between the labels to determine the technical description combination. For the divided feature sentences, the sentence start index and the sentence end index corresponding to each feature sentence are determined, which can be the position information of the character marking the beginning of the feature sentence and the position information of the character marking the end of the feature sentence. Corresponding to the subject label and the description label, the corresponding word start index and word end index need to be determined when determining the feature sentence to which they belong. The indexes can be determined when performing entity recognition, or after the sentence-related indexes are determined.
[0044] Subsequently, according to the word start index and the word end index, all subject labels and / or description labels located within the sentence start index and the sentence end index of each feature sentence are determined, which is equivalent to determining a group of subject labels and description labels having an association relationship. Arranging the subject labels and the description labels belonging to the same feature sentence according to the ascending order of the word start index, a group of technical description combinations can be obtained, which correspond to recording all key subjects and specific relationships between the key subjects in the feature sentence, achieving the purpose of simplifying the technical solution.
[0045] After traversing all the feature sentences, all the technical description combinations of the patent file are obtained, and the correspondence between the patent number of each patent file and its corresponding all technical description combinations is established to form a patent file database based on technical description. After all the patent files are determined to correspond to all their technical description combinations, the patent number can be the application number or the publication number of the patent file during actual establishment, and the correspondence between each patent number and technical description combination can be saved in the form of a matrix.
[0046] After the above-mentioned patent file database based on technical field and the patent file database based on technical description are both constructed, the field label and the technical description combination corresponding to each patent number can be converted into a word vector, and the faiss library is further converted into a corresponding index library to speed up the subsequent retrieval task. Finally, the correspondence between the patent number and the field index and the technical description index is converted to form a field index library and a technical description index library, and the overall construction of the complete patent file database is completed.
[0047] In some embodiments, when actually constructing the patent file database based on technical description, the technical features, technical problems and technical effects can also be split, and the patent file data can be constructed based on the above three parts, i.e., the patent file database based on technical features, the patent file database based on technical problems and the patent file database based on technical effects are formed. The construction process of the above three databases can be directly performed by using the construction method of the patent file database based on technical description, which will not be repeated here. At the same time, the construction and conversion of the feature index library, the problem index library and the effect index library are performed, and when performing semantic retrieval, the technical description combination can be determined based on the technical features, technical problems and technical effects, and the patent number retrieval determination based on the feature index library, the problem index library and the effect index library is performed. At this time, the corresponding patent file database is formed as shown in Figure 2 At this time, the calculation formula for determining the total score of each patent can be:
[0048] Total score = technical field score * (q1 * technical feature score + q2 * technical problem score + q3 * technical effect score), q1, q2 and q3 are the trained weights, the technical feature score is the comprehensive score of the current patent under each different technical feature, and the technical problem score and the technical effect score are the same.
[0049] Based on the above, this disclosure combines technical field information, technical feature information, technical problem information, and technical effect information as a foundational retrieval database, thereby more comprehensively and accurately representing the key information of the input text. During semantic retrieval, key information from multiple modules is integrated, making the retrieval results more complete and better representing the full key information of the input text. Simultaneously, combined retrieval through weighting and combination methods makes the retrieval results more reasonable and accurate, more comprehensively utilizing the data information in the patent, improving the value and utilization rate of patent information, and enabling its application to patent texts with different fields and technical features, demonstrating strong versatility and applicability.
[0050] Based on the same inventive concept, the second embodiment of this disclosure provides a semantic retrieval device based on a patent document database, the structural schematic diagram of which is shown below. Figure 3 As shown, the system mainly includes a preprocessing module 10, which performs preprocessing on the user-input search text to obtain pre-defined tags for the search text. The pre-defined tags include at least a combination of domain tags and / or technical descriptions. A retrieval module 20 is used to convert each pre-defined tag into a tag word vector, and uses each tag word vector as a target word vector to determine the m patent numbers with the highest similarity to the target word vectors in the patent document database. The score of each patent number is determined based on the similarity between the patent number and the target word vector. A score statistics module 30 is used to construct a score dictionary for each patent number and determine the total score for each patent number based on the score dictionary. A feedback module 40 is used to provide the user with the k patent numbers with the highest total scores as the search results for the search text.
[0051] For a detailed description of the above modules and their corresponding functions, please refer to the content disclosed in the first embodiment of this disclosure. This embodiment will not repeat the description again.
[0052] Based on the same inventive concept, the third embodiment of this disclosure provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of the semantic retrieval method based on a patent document database described in the first embodiment of this disclosure.
[0053] Based on the same inventive concept, the fourth embodiment of this disclosure provides an electronic device, including at least a memory and a processor. The memory stores a computer program, and the processor, when executing the computer program in the memory, implements the steps of the semantic retrieval method based on a patent document database described in the first embodiment of this disclosure.
[0054] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this disclosure, and are not intended to limit them. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this disclosure.
Claims
1. A semantic retrieval method based on a patent document database, characterized in that, include: The search text input by the user is processed in a preset manner to obtain preset tags for the search text. The preset tags include at least a combination of domain tags and / or technical descriptions. Each preset label is converted into a label word vector, and each label word vector is used as a target word vector in turn. The m patent numbers with the highest similarity to the target word vector are determined in the patent document database. The score of the patent number is determined based on the similarity between the patent number and the target word vector. Construct a score dictionary for each of the patent numbers, and determine the total score for each patent number based on the score dictionary; The k patent numbers with the highest total scores are used as the search results and fed back to the user.
2. The semantic retrieval method according to claim 1, characterized in that, After performing preset processing on the user-input search text to obtain preset tags for the search text, the method further includes: Display all the preset labels to the user.
3. The semantic retrieval method according to claim 1, characterized in that, The step of performing preset processing on the user-input search text to obtain preset tags for the search text includes: The first named entity recognition model is used to perform first named entity recognition on the search text in order to identify the domain tags in the search text.
4. The semantic retrieval method according to claim 1, characterized in that, The step of performing preset processing on the user-input search text to obtain preset tags for the search text includes: The search text is subjected to second named entity recognition using a second named entity recognition model to identify the subject tags and description tags in the search text. The retrieved text is segmented into feature sentences based on punctuation marks. All the main tags and / or all the description tags belonging to the same feature sentence are combined to obtain a technical description combination.
5. The semantic retrieval method according to claim 1, characterized in that, When the preset tags include the domain tags, the step of converting each preset tag into a corresponding tag word vector, and using each tag word vector as a target word vector in sequence, determining the m patent numbers with the highest similarity to the target word vectors in the patent document database, and determining the score of the patent number based on the similarity between the patent number and the target word vectors, includes: The domain labels are converted into corresponding domain word vectors, and the distance between each domain index and the domain word vector is determined in the domain index library; Determine the m domain indices that are closest to the domain word vector, and normalize the distance between each of the m domain indices and the domain word vector to obtain the score of the domain index; Determine the patent number corresponding to each of the m domain indices, and use the score of the domain index as the score of the patent number under the domain label.
6. The semantic retrieval method according to claim 5, characterized in that, When the preset tags include the combination of technical descriptions, the step of converting each preset tag into a corresponding tag word vector, and using each tag word vector as a target word vector in sequence, determining the m patent numbers with the highest similarity to the target word vectors in the patent document database, and determining the score of the patent number based on the similarity between the patent number and the target word vectors, includes: The technical descriptions are combined and converted into corresponding technical description word vectors. The distance between each technical description index and the technical description word vectors is determined in the technical description index library. Determine the m technical description indices that are closest to the technical description word vector, and normalize the distance between each of the m technical description indices and the technical description word vector to obtain the score of the technical description index; Determine the patent number corresponding to each of the m technical description indices, and use the score of the technical description index as the score of the patent number under the technical description combination.
7. The semantic retrieval method according to claim 6, characterized in that, The process of constructing a score dictionary for each patent number and determining the total score for each patent number based on the score dictionary includes: Based on the patent number and the score of the patent number under the field label and / or the score under the combination of technical descriptions, a score dictionary is constructed, wherein the score of the patent number under the combination of technical descriptions is the sum of the scores of the patent number under different combinations of technical descriptions; Calculate the product between the sum and the technology weight, and sum the product with the score of the patent number under the field label to obtain the total score of the patent number.
8. A semantic retrieval device based on a patent document database, characterized in that, include: The preprocessing module is used to perform pre-processing on the search text input by the user to obtain pre-defined tags for the search text. The pre-defined tags include at least a combination of domain tags and / or technical descriptions. The retrieval module is used to convert each of the preset tags into tag word vectors, and use each of the tag word vectors as target word vectors in turn to determine the m patent numbers with the highest similarity to the target word vectors in the patent document database, and determine the score of the patent number based on the similarity between the patent number and the target word vector; The scoring statistics module is used to construct a scoring dictionary for each patent number and determine the total score for each patent number based on the scoring dictionary. The feedback module is used to provide the user with the k patent numbers with the highest total scores as the search results of the search text.
9. A storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the semantic retrieval method based on the patent document database as described in any one of claims 1 to 7.
10. An electronic device, comprising at least a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program on the memory, it implements the steps of the semantic retrieval method based on the patent document database as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Patent retrieval method, storage medium and device
CN112836010A
Semantic search method and device based on technical keyword extraction, equipment and medium
CN116303968A
System and method for matching expertise
US20080183759A1
Extracting fine-grained topics from text content
US20230161964A1