Document retrieval method and apparatus, storage medium, and electronic device

By preset processing and tag extraction of user input text, combined with the similarity and score calculation of the patent file database, the problem of low semantic search accuracy in the prior art is solved, and more accurate feedback on patent search results is achieved.

WO2025140165A1PCT designated stage expired Publication Date: 2025-07-03BEIJING AUGUST MELON TECHNOLOGY CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/141699
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-12-24
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In the existing patent semantic search methods, the semantic search accuracy is low, the usage effect is poor, and the different keywords in the text are not fully utilized for classification and association, resulting in poor customer understanding of the relevance of search results.

Method used

By presetting the search text input by the user, the domain tag and technical description combination is extracted, and converted into a tag word vector, the patent number with the highest similarity is determined in the patent file database, and a scoring dictionary is constructed, and the patent number with the highest total score is feedback as the search result.

Benefits of technology

It realizes finer granularity and more accurate patent search, improves semantic search accuracy, meets user needs, and provides search results that are more in line with actual content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024141699_03072025_PF_FP_ABST
    Figure CN2024141699_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present disclosure are a document retrieval method and apparatus, a storage medium, and an electronic device. The method comprises: performing preset processing on a retrieval text input by a user, so as to obtain preset labels of the retrieval text; correspondingly converting each preset label into a label word vector, and by using each label word vector as a target word vector sequentially, determining from a patent document database m patent numbers having the highest similarity to the target word vector, and determining a score of each patent number on the basis of the similarity between the patent number and the target word vector; constructing a score dictionary for each patent number, and determining a total score of each patent number; and using k patent numbers having the highest total scores as retrieval results of the retrieval text and feeding back the retrieval results to the user. The present disclosure performs structured splitting on the retrieval text, so that in a retrieval process based on the patent document database, a retrieval result that best meets user's needs is obtained by means of more fine-grained and more accurate retrieval information in combination with scores of retrieval results under different preset labels.
Need to check novelty before this filing date? Find Prior Art

Description

File retrieval method, device, storage medium and electronic equipment

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on December 29, 2023, with application number 202311873439.0 and invention name “A semantic retrieval method and device based on patent document database”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present disclosure relates to the field of data processing technology, and in particular to a file retrieval method, device, storage medium, and electronic device. Background Art

[0003] Currently, the information that customers are concerned about and that has the highest information value in patents includes the technical field information, technical feature information, technical problem information and technical effect information.

[0004] Patent semantic retrieval methods described in existing platforms and literature are generally pure black-box approaches, failing to expose key information from the input text to the client. Clients simply enter text, which then triggers a backend search, displaying a certain number of relevant patents based on the search results. However, this approach provides little insight into the original text and a poor understanding of the relevance of the search results. Furthermore, current retrieval methods primarily rely on keywords within the original text, without categorizing and associating the different keywords within the text. Consequently, some semantic content is not fully understood and utilized, leading to low semantic retrieval accuracy and poor performance. Summary of the Invention

[0005] The purpose of the embodiments of the present disclosure is to provide a file retrieval method, device, storage medium and electronic device to solve the problems of low semantic retrieval accuracy and poor usage effect in the prior art.

[0006] The embodiment of the present disclosure adopts the following technical solution: a document retrieval method, comprising: performing preset processing on a search text input by a user to obtain a preset label of the search text, the preset label including at least a field label and / or a technical description combination; converting each of the preset labels into a label word vector, and using each of the label word vectors as a target word vector in turn, determining m patent numbers with the highest similarity to the target word vector in a patent document database, and determining the score of the patent number based on the similarity between the patent number and the target word vector, wherein m is a natural number greater than 1; constructing a score dictionary for each of the patent numbers, and determining the total score of each patent number based on the score dictionary; and feeding back the k patent numbers with the highest total scores as retrieval results of the search text to the user, wherein k is a natural number less than or equal to m.

[0007] The embodiment of the present disclosure also provides a semantic retrieval device based on a patent document database, including: a preprocessing module, used to perform preset processing on a search text input by a user to obtain a preset label of the search text, wherein the preset label includes at least a field label and / or a technical description combination; a retrieval module, used to convert each of the preset labels into a label word vector, and use each of the label word vectors as a target word vector in turn, determine the m patent numbers with the highest similarity to the target word vector in the patent document database, and determine the score of the patent number based on the similarity between the patent number and the target word vector, wherein m is a natural number greater than 1; a score statistics module, used to construct a score dictionary for each of the patent numbers, and determine the total score of each patent number based on the score dictionary; a feedback module, used to feed back the k patent numbers with the highest total scores as the retrieval results of the search text to the user, wherein k is a natural number less than or equal to m.

[0008] The embodiment of the present disclosure further provides a storage medium storing a computer program, which implements the steps of the above-mentioned file retrieval method when executed by a processor.

[0009] An embodiment of the present disclosure further provides an electronic device, comprising at least a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above-mentioned file retrieval method when executing the computer program on the memory.

[0010] The beneficial effects of the disclosed embodiments are: identifying and extracting the search text input by the user based on the combination of domain labels and technical descriptions, realizing structured splitting of the search text, extracting the most critical data in the search text corresponding to the main content module of the patent for subsequent search, and realizing, in the search process based on the patent document database, obtaining the search results that best meet the user's needs through more fine-grained and more accurate search information and combining the scores of the search results under different preset labels, thereby achieving the purpose of improving the search accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate one or more embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this specification. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0012] FIG1 is a flow chart of a file retrieval method according to a first embodiment of the present disclosure;

[0013] FIG2 is a schematic diagram of the architecture of the patent document database in the first embodiment of the present disclosure;

[0014] FIG3 is a schematic structural diagram of a file retrieval device in a second embodiment of the present disclosure. DETAILED DESCRIPTION

[0015] In order to enable those skilled in the art to better understand the technical solutions in one or more embodiments of this specification, the technical solutions in one or more embodiments of this specification will be clearly and completely described below in conjunction with the drawings in one or more embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on one or more embodiments of this specification, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this document.

[0016] Patent semantic retrieval methods described in existing platforms and literature are generally pure black-box approaches, failing to expose key information from the input text to the client. Clients simply enter text, which then triggers a backend search, displaying a certain number of relevant patents based on the search results. However, this approach provides little insight into the original text and a poor understanding of the relevance of the search results. Furthermore, current retrieval methods primarily rely on keywords within the original text, without categorizing and associating the different keywords within the text. Consequently, some semantic content is not fully understood and utilized, leading to low semantic retrieval accuracy and poor performance.

[0017] In order to solve the above problem, the first embodiment of the present disclosure provides a file retrieval method, the flow chart of which is shown in FIG1 , and at least includes steps S10 to S40:

[0018] S10, performing preset processing on the search text input by the user to obtain a preset label of the search text.

[0019] When a user performs a semantic search, he or she enters the relevant search text in the search box. The search text may include the technical features that the user wishes to search, the technical problems solved by the technical features, the technical effects achieved, and the technical fields of application. At the same time, the search text also contains redundant information and interference words used in the text description process. After receiving the search text input by the user, the search engine first performs a preset process on the search text and extracts the preset tags of the search text. The preset tags in this embodiment include at least a field tag and / or a technical description combination, wherein the field tag is mainly used to explain the technical field information corresponding to the search text, and the technical description combination is mainly used to characterize the description of the technical features, technical problems, and technical effects in the search text, such as the technical feature subject and the relationship between the subjects, the technical problem subject and the description information of the problem subject, the technical effect subject and the description information of the effect subject, etc. By extracting the preset tags from the search text, a structured splitting of the search text is achieved, and more refined field information and description information corresponding to the main content of the patent document can be extracted from the search text. The influence of redundant information is removed during subsequent searches, thereby improving the accuracy of semantic retrieval.

[0020] In some embodiments, after obtaining the preset tags of the search text, all preset tags can be displayed to the user to facilitate the user to have a deeper understanding of the search text entered by the user and to facilitate the user to filter and use the final search results.

[0021] In some embodiments, the process of pre-processing the search text input by the user to obtain a preset label for the search text mainly includes: performing a first named entity recognition on the search text through a first named entity recognition model to identify the domain label in the search text. Specifically, the first named entity recognition model is a pre-trained model for entity recognition of domain information in the text, which can extract domain labels through specific keywords or key subjects in the search text. In actual use, there are also cases where domain labels cannot be extracted from the search text, such as when the search text content is small, the user's expected usage scenario is not clear, or entities with specific functions, etc.

[0022] In some embodiments, the process of performing preset processing on the search text input by the user to obtain preset labels for the search text may also include: performing second named entity recognition on the search text through a second named entity recognition model to identify the subject labels and description labels in the search text; and dividing the search text into feature sentences based on punctuation marks, combining all subject labels and / or all description labels belonging to the same feature sentence to obtain a technical description combination.

[0023] Specifically, the second named entity recognition model is a pre-trained model for entity recognition based on subject information and descriptive information in the text. Subject information can include entities representing features, technical issues, and technical effects. Descriptive information includes the feature relationships between feature entities, and detailed descriptions of technical issue entities and technical effect entities. The search text may have multiple different subject and description labels. These labels effectively represent a streamlined content extraction of the technical description of the search text, filtering out redundant information and noise words.

[0024] After extracting tags, it is necessary to associate and combine subject tags with description tags to clarify the actual content disclosed in the search text and represent it through a concise and accurate tag combination. Generally, the subjects described in the same sentence are necessarily associated with each other. Therefore, this embodiment divides the search text into feature sentences based on the position of punctuation marks, and uses feature sentences as the basic unit for determining the association between tags to determine the technical description combination.

[0025] For the feature sentences after division, the sentence start index and sentence end index are determined for each feature sentence, which can be the position information of the character used to mark the start of the feature sentence and the position information of the character used to mark the end of the feature sentence. Corresponding to the subject label and description label, when determining the feature sentence to which they belong, it is also necessary to determine the corresponding word start index and word end index. Then, based on the word start index and word end index, all subject labels and / or description labels within the sentence start index and sentence end index of each feature sentence are determined, which is equivalent to determining a group of subject labels and description labels with an associated relationship. The above-mentioned subject labels and description labels belonging to the same feature sentence are arranged in ascending order of the word start index to obtain a group of technical description combinations. The technical description combination records all the key subjects in the feature sentence and the specific relationship between the key subjects, thereby achieving the purpose of streamlining the technical solution. In the specific implementation, each feature sentence is traversed, and the main tags and / or description tags whose word start index is greater than or equal to the sentence start index and whose word end index is less than or equal to the sentence end index are determined as the main tags and / or description tags located within the sentence start index and sentence end index of each feature sentence.

[0026] S20, convert each preset label into a label word vector, and use each label word vector as the target word vector in turn, determine the m patent numbers with the highest similarity to the target word vector in the patent document database, and determine the score of the patent number based on the similarity between the patent number and the target word vector.

[0027] After executing step S10, one or more preset tags of the search text can be extracted, and then searches can be performed in the patent document database based on all the above preset tags. When searching, each preset tag is converted into a corresponding tag word vector, and each tag word vector is used as a target word vector in turn to determine the m most similar patent numbers in the patent document database. The above m patent numbers are the search results of the preset tag currently used as the target word vector in the database, and the similarity between the patent number and the target word vector is the score of the patent number under the corresponding preset tag. In actual implementation, m patent numbers and their scores are output corresponding to each preset tag, so there may actually be m times the number of patent numbers and their scores as the output of this step. In this embodiment, m is a natural number greater than 1.

[0028] Specifically, corresponding to the combination of domain labels and technical descriptions, the patent document database has different index libraries to record the associations between all patent numbers and domains and technical descriptions. That is, the domain index library records the correspondence between all patent numbers and their corresponding domain indexes, and the technical description index library records the correspondence between all patent numbers and their corresponding technical description indexes. Based on the identification and extraction results of the preset labels in step S10, when the preset labels include different types, different index libraries are used for retrieval.

[0029] In some embodiments, when the preset label includes a domain label, the domain label is first converted into a corresponding domain word vector, and then the distance between each domain index and the domain word vector is determined in the domain index library; based on the calculation results of the distances between all domain indexes and the domain word vectors, the m domain indexes closest to the domain word vector are determined, and the distance between each of the above m domain indexes and the domain word vector is normalized to obtain the scores of the m domain indexes; finally, the patent numbers corresponding to the m domain indexes are determined in the domain index library, and the scores of the domain indexes are used as the scores of their corresponding patent numbers under the current domain label.

[0030] In some embodiments, when the preset label includes a technical description label, the technical description label is first converted into a corresponding technical description word vector, and then the distance between each technical description index and the technical description word vector is determined in the technical description index library; based on the calculation results of the distances between all technical description indexes and the technical description word vectors, the m technical description indexes closest to the technical description word vector are determined, and the distance between each technical description index and the technical description word vector in the above m technical description indexes is normalized to obtain the scores of the m technical description indexes; finally, the patent numbers corresponding to the m technical description indexes are determined in the technical description index library, and the scores of the technical description indexes are used as the scores of their corresponding patent numbers under the current technical description label.

[0031] Normally, there is only one field label extracted from a search text, which represents the field most desired for application or search. However, the technical description label can reflect the three aspects of technical features, technical problems and technical effects, and there may be multiple technical subjects and technical descriptions corresponding to each aspect. Therefore, the number of technical description labels is usually not unique. In the process of determining the corresponding patent number and calculating the score, it is necessary to determine the corresponding patent number and score for each technical description label. The m patent numbers corresponding to different technical description labels may contain the same patent or different patents.

[0032] S30, constructing a score dictionary for each patent number, and determining a total score for each patent number based on the score dictionary.

[0033] The score dictionary is constructed based on the patent number, which records the scores of each patent under different preset labels and can be constructed in the following form: {'patent number': [{'technical field': []}, 'technical description': []]}, wherein the brackets after 'technical field' and 'technical description' are used to record the score value of the current patent number under the combination of field label and technical description, especially the technical description score, which is recorded according to the sum of the scores of the patent number under different technical description combinations. Corresponding to the calculation process of the total score, first calculate the product between the sum of the scores of the above-mentioned patent number under different technical description combinations and the preset technical weight. The role of the technical weight is to balance the quantitative difference between the technical field score and the technical description score. Since the technical description score is usually the sum of the scores of the patent number under multiple technical description combinations, its specific value is much larger than the technical field score, so it is adjusted by the technical weight; then the product is added to the technical field score of the patent number, and the resulting sum is recorded as the total score of the patent number. It should be noted that the specific numerical value of the above-mentioned technical weight can be set based on actual needs and historical experience, and this embodiment does not impose specific restrictions.

[0034] S40: Feedback the k patent numbers with the highest total scores to the user as the search results of the search text.

[0035] After calculating the total score of all patent numbers, the patent document search results based on the preset label extraction results of the search text are obtained. The k patent numbers with the highest scores are selected and fed back to the user as the search results. The solutions corresponding to these k patent numbers are the most similar to the technical content actually expressed in the search text. In this embodiment, k is a natural number less than or equal to m.

[0036] In actual implementation, the total scores of all patent numbers can be normalized and sorted, the total scores can be limited to between score_low and score_high, and then the first k patent numbers can be intercepted and returned.

[0037] This embodiment identifies and extracts the search text input by the user based on the combination of domain labels and technical descriptions, realizes the structured splitting of the search text, extracts the most critical data in the search text that corresponds to the main content module of the patent for subsequent search, and realizes that in the search process based on the patent document database, through more fine-grained and more accurate search information, combined with the scores of the search results under different preset labels, the search results that best meet the user's needs are obtained, thereby achieving the purpose of improving the search accuracy.

[0038] In the actual implementation process, before the user performs semantic search, it can also include the process of building a database of patent documents. During the construction process of the patent document database, it is also necessary to implement more refined and better relevant information based on field labels, technical description combinations, etc., and cooperate with the semantic retrieval method provided in this embodiment to achieve more accurate search result feedback.

[0039] A patent document mainly includes the following parts: technical field part, technical features part, technical problem part, and technical effect part. Among them, the technical field part mainly refers to the technical field to which the patent document actually belongs in the specification. The technical field can be described by a specific field information or by an entity with specific functions actually contained in the patent document, such as a method entity or an apparatus entity; the technical features part mainly refers to the part used to describe the actual technical solution of the patent document, such as the content of the claims or the invention content part in the specification; the technical problem part mainly refers to the technical problem that the patent document actually aims to solve, which is usually described in the background technology part of the specification; the technical effect part mainly refers to the beneficial effects achieved by the technical solution, which is usually in the second half of the invention content.

[0040] In some embodiments, a database of patent documents can be first constructed based on the technical field. For example, the technical field part of the patent document is subjected to named entity recognition through the first named entity recognition model, in which there is a description of the field information of the patent document and the specific device or method involved. The corresponding field label and subject label can be identified through the NER model, which respectively correspond to the technical field information of the patent and the subject information such as the specific device or method. In actual application, there are some patent documents that may not describe the field information in the technical field part, so these patent documents may not be able to identify the corresponding field label, but the technical field part will definitely record the subject of the disclosed solution. Therefore, each patent document should be able to identify the subject label. Based on the named entity recognition results, all patent documents are divided into the first category of patent documents and the second category of patent documents, wherein the first category of patent documents are patent documents that identify the field label and the subject label, and the second category of patent documents are patent documents that only identify the subject label.

[0041] Corresponding to all the first subject tags that have been identified as domain tags and subject tags, a mapping dictionary is constructed based on the extracted domains and corresponding specific device / method information. After the mapping dictionary for all first-category patent documents is constructed, the corresponding domain tags for each second subject tag in the second-category patent document can be determined based on the mapping dictionary, so that specific domain information can be assigned to each patent document. Specifically, based on the mapping dictionary that has been constructed, corresponding to each second subject tag, the first subject tag that is most similar to the second subject tag can be determined from the mapping dictionary. Based on the similarity between the two subject tags, the corresponding domain tags should also be similar. Therefore, at least one first domain tag corresponding to the first subject tag that is most similar to the second subject tag in the mapping dictionary is assigned to the second subject tag as at least one second domain tag corresponding to the second subject tag. After the above processing is performed on all second-category patent documents, the domain information that can be applied to each patent document can be determined.

[0042] After all patent documents have their corresponding field labels determined, a correspondence can be established between each patent document's patent number and its corresponding first field label, forming a patent document database based on technical fields. In the actual establishment process, the patent number can be the patent document's application number or publication number, and the correspondence between each patent number and field can be stored in the form of a matrix.

[0043] In some embodiments, a patent document database can also be constructed based on the technical description. For example, the technical description part in the patent document is subjected to named entity recognition through the second named entity recognition model, in which there is a description subject and specific description content of the patent document. The corresponding subject label and description label can be identified through the NER model, which respectively correspond to the description subject and specific description content appearing in the patent document. In actual implementation, the subject label and description label have different representational meanings for different content parts of the patent document. Specifically, in the case where the technical description part is the claims, the subject label is used to represent the feature subject in the claims, and the description label is used to represent the feature relationship between the feature subjects; in the case where the technical description part is the background technology part, the subject label is used to represent the technical problem subject in the background technology part, and the description information is used to represent the problem description of the technical problem subject; in the case where the technical description part is the beneficial effect part, the subject label is used to represent the technical effect subject in the beneficial effect part, and the description information is used to represent the effect description of the technical effect subject.

[0044] After the labels are extracted, it is necessary to make an association combination between the subject label and the description label to clarify the true content disclosed by the patent document and to express it through a concise and accurate label combination. Under normal circumstances, there must be an association between the subjects described in the same sentence. Therefore, this embodiment divides the technical description part into feature sentences based on the position of punctuation marks, and uses the feature sentences as the basic unit for determining the association between labels to determine the technical description combination. For the divided feature sentences, the sentence start index and the sentence end index are determined for each feature sentence, which can be the position information of the character that marks the beginning of the feature sentence and the position information of the character that marks the end of the feature sentence. Corresponding to the subject label and the description label, when determining the feature sentence to which it belongs, the corresponding word start index and word end index are also needed. The index can be performed when performing entity recognition or after the sentence-related index is determined.

[0045] Then, based on the word start index and the word end index, all subject tags and / or description tags within the sentence start index and sentence end index of each feature sentence are determined, which is equivalent to determining a group of subject tags and description tags with an associated relationship. The above-mentioned subject tags and description tags belonging to the same feature sentence are arranged in ascending order of the word start index to obtain a group of technical description combinations. The technical description combination corresponds to recording all the key subjects in the feature sentence and the specific relationship between the key subjects, thereby achieving the purpose of streamlining the technical solution.

[0046] After traversing all feature statements, all technical description combinations of the patent document are obtained. A correspondence between the patent number of each patent document and all corresponding technical description combinations can be established, forming a patent document database based on technical descriptions. After all technical description combinations corresponding to all patent documents are determined, in the actual establishment process, the patent number can be the application number or publication number of the patent document, and the correspondence between each patent number and technical description combination can be stored in the form of a matrix.

[0047] After the above-mentioned database of patent documents based on technical fields and the database of patent documents based on technical descriptions are constructed, the field labels and technical description combinations corresponding to each patent number can be converted into word vectors, and further converted into corresponding index libraries using the faiss library to accelerate subsequent retrieval tasks. Finally, it is converted into a correspondence between the patent number and the field index and the technical description index to form a field index library and a technical description index library, completing the overall construction of the patent document database.

[0048] In some embodiments, when actually constructing a patent document database based on technical descriptions, the three parts of technical features, technical problems and technical effects can also be split, and the patent document data can be constructed based on the above three parts, that is, a patent document database based on technical features, a patent document database based on technical problems and a patent document database based on technical effects can be formed accordingly. The construction process of the above three databases can all be directly carried out using the construction method of the patent document database based on technical descriptions, and will not be repeated here. At the same time, the construction and transformation of the feature index library, the problem index library and the effect index library are carried out accordingly, and when performing semantic retrieval, the technical description combination can also be determined based on the three parts of technical features, technical problems and technical effects in the process of determining the technical description combination, and the patent number search and determination based on the feature index library, the problem index library and the effect index library can be carried out accordingly. The corresponding patent document database formed at this time is shown in Figure 2. At this time, when determining the total score of each patent, the calculation formula can be:

[0049] Total score = technical field score * (q1 * technical feature score + q2 * technical problem score + q3 * technical effect score), where q1, q2, and q3 correspond to trained weights. The technical feature score is the combined score of the current patent under each different technical feature. The same applies to the technical problem score and technical effect score.

[0050] Based on the above content, the embodiment of the present disclosure combines technical field information, technical feature information, technical problem information and technical effect information as the basic retrieval database, so as to more comprehensively and accurately represent the key information of the input text. In the process of semantic retrieval, the key information of multiple modules is integrated to make the retrieval results more complete and better represent the complete key information of the input text. At the same time, combined retrieval is performed through weights and combinations, making the retrieval results more reasonable and accurate, and making more comprehensive use of the data information in the patent, thereby improving the value and utilization rate of the patent information, and can be applied to patent texts in different fields and technical features, with strong versatility and applicability.

[0051] Based on the same inventive concept, the second embodiment of the present disclosure provides a document retrieval device, the structural diagram of which is shown in Figure 3, mainly including a preprocessing module 10, which is used to perform preset processing on the search text input by the user to obtain a preset label of the search text, and the preset label at least includes a field label and / or a technical description combination; a retrieval module 20, which is used to convert each of the preset labels into a label word vector, and use each of the label word vectors as a target word vector in turn, determine the m patent numbers with the highest similarity to the target word vector in the patent document database, and determine the score of the patent number according to the similarity between the patent number and the target word vector, wherein m is a natural number greater than 1; a score statistics module 30, which is used to construct a score dictionary for each of the patent numbers, and determine the total score of each patent number based on the score dictionary; and a feedback module 40, which is used to feed back the k patent numbers with the highest total scores as the retrieval results of the search text to the user, wherein k is a natural number less than or equal to m.

[0052] For a detailed description of the above modules and their corresponding functions, please refer to the contents disclosed in the first embodiment of the present disclosure, which will not be repeated in this embodiment.

[0053] Based on the same inventive concept, a third embodiment of the present disclosure provides a storage medium storing a computer program, which, when executed by a processor, implements the steps of the file retrieval method described in the first embodiment of the present disclosure.

[0054] Based on the same inventive concept, the fourth embodiment of the present disclosure provides an electronic device, comprising at least a memory and a processor, wherein a computer program is stored on the memory, and the processor implements the steps of the file retrieval method described in the first embodiment of the present disclosure when executing the computer program on the memory.

[0055] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure, rather than to limit them. Although the present disclosure has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present disclosure.

Claims

1. A file retrieval method, characterized in that Including: Performing preset processing on the retrieval text input by the user to obtain a preset label of the retrieval text, where the preset label at least includes a field label and / or a technical description combination; Converting each of the preset labels into a label word vector, and using each of the label word vectors as a target word vector in turn to determine the m patent numbers with the highest similarity to the target word vector in the patent document database, and determining the score of the patent number according to the similarity between the patent number and the target word vector, where m is a natural number greater than 1; Constructing a score dictionary for each of the patent numbers, and determining the total score of each patent number based on the score dictionary; and Feeding back the k patent numbers with the highest total scores to the user as the retrieval result of the retrieval text, where k is a natural number less than or equal to m.

2. The document retrieval method according to claim 1, characterized in that After performing preset processing on the retrieval text input by the user to obtain a preset label of the retrieval text, the method further includes: Displaying all the preset labels to the user.

3. The document retrieval method according to claim 1, wherein Performing preset processing on the retrieval text input by the user to obtain a preset label of the retrieval text includes: Performing first named entity recognition on the retrieval text through a first named entity recognition model to identify the field label in the retrieval text.

4. The document retrieval method according to claim 1, characterized in that, Performing preset processing on the retrieval text input by the user to obtain a preset label of the retrieval text includes: Performing second named entity recognition on the retrieval text through a second named entity recognition model to identify the subject label and the description label in the retrieval text; and Performing punctuation-based characteristic sentence division on the retrieval text, and combining all the subject labels and / or all the description labels belonging to the same characteristic sentence to obtain a technical description combination.

5. The document retrieval method according to claim 1, wherein When the preset label includes the field label, converting each of the preset labels into a label word vector and using each of the label word vectors as a target word vector in turn to determine the m patent numbers with the highest similarity to the target word vector in the patent document database and determining the score of the patent number according to the similarity between the patent number and the target word vector includes: Converting the field label into a corresponding field word vector, and determining the distance between each field index and the field word vector in the field index library; Determining the m field indexes with the closest distance to the field word vector, and normalizing the distance between each of the m field indexes and the field word vector to obtain the score of the field index; and Determining the patent numbers corresponding to the m field indexes, and using the score of the field index as the score of the patent number under the field label.

6. The document retrieval method according to claim 5, characterized in that, When the preset label includes the technical description combination, converting each of the preset labels into a label word vector and using each of the label word vectors as a target word vector in turn to determine the m patent numbers with the highest similarity to the target word vector in the patent document database and determining the score of the patent number according to the similarity between the patent number and the target word vector includes: Convert the technical description combination into a corresponding technical description word vector, and determine the distance between each technical description index and the technical description word vector in the technical description index library; Determine the m technical description indexes that are closest to the technical description word vector, and normalize the distance between each of the m technical description indexes and the technical description word vector to obtain the score of the technical description index; and Determine the patent numbers corresponding to the m technical description indexes, and use the score of the technical description index as the score of the patent number under the technical description combination.

7. The document retrieval method according to claim 6, wherein Constructing a score dictionary for each of the patent numbers and determining the total score of each patent number based on the score dictionary includes: Construct a score dictionary based on the patent number and the score of the patent number under the field label and / or under the technical description combination, wherein the score of the patent number under the technical description combination is the sum of the scores of the patent number under different technical description combinations; and Calculate the product of the sum value and the technical weight, and sum the product and the score of the patent number under the field label to obtain the total score of the patent number.

8. A file retrieval device, characterized in that Includes: A preprocessing module for performing preset processing on the retrieved text input by the user to obtain a preset label of the retrieved text, where the preset label includes at least a field label and / or a technical description combination; A retrieval module for converting each of the preset labels into a label word vector, and using each of the label word vectors as a target word vector in turn to determine the m patent numbers with the highest similarity to the target word vector in the patent document database, and determining the score of the patent number according to the similarity between the patent number and the target word vector, where m is a natural number greater than 1; A score statistics module for constructing a score dictionary for each of the patent numbers and determining the total score of each patent number based on the score dictionary; and A feedback module for using the k patent numbers with the highest total scores as the retrieval results of the retrieved text and feedbacking them to the user, where k is a natural number less than or equal to m.

9. A storage medium stores a computer program, characterized in that, When the computer program is executed by a processor, it implements the file retrieval method according to any one of claims 1 to 7.

10. An electronic device, at least comprising a memory and a processor, wherein a computer program is stored on the memory, characterized in that, When the processor executes the computer program on the memory, it implements the file retrieval method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Retrieval method and device based on semantic analysis and keyword recognition

    CN112507109A

  • Text retrieval method, device, storage medium and server

    CN114090799A

  • Text retrieval method and device, equipment and storage medium

    CN114416954A

  • Method and system for generating patent summary information based on semantic understanding

    CN116842173A