Patent document retrieval method, device, and readable storage medium
By segmenting and screening the patent text, the patent document guidance information is determined for searching, and the inaccurate results caused by abstract search in the prior art is solved, and more accurate patent document retrieval is achieved.
Patent Information
- Application Number
- CN202510251965.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-05
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-03-05
AI Technical Summary
The existing patent document search method is based on abstracts, resulting in inaccurate results.
By performing word segmentation on the patent text, a phrase set is obtained, and divided into first and second weight phrase sets according to the adjacency relationship between the phrases, and filtering is performed separately to determine the target phrases, thereby determining the patent document guidance information for searching.
It improves the accuracy of the search and can more comprehensively reflect the technical solutions of patent documents.
Smart Images

Figure CN119739682B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the computer field, and in particular to a patent document retrieval method, device, and readable storage medium. Background Art
[0002] When searching, the existing patent document retrieval method only compares the keywords entered by the user and the relationship between the keywords with the headings and abstracts of the patent document in order to quickly obtain search results, and determines whether their contents meet the search conditions entered by the user, and then returns the patent documents that meet the search conditions entered by the user.
[0003] However, since the abstract of a patent document often cannot fully reflect the entire technical solution of the patent document, the results obtained by the existing patent document retrieval method based on the abstract of the patent document are not accurate enough. Summary of the invention
[0004] In view of this, the embodiments of the present application provide a patent document retrieval method, apparatus and terminal device, which can solve the problem in the prior art that the results obtained by searching based on the abstracts of patent documents are not accurate enough.
[0005] A first aspect of an embodiment of the present application provides a patent document retrieval method, comprising:
[0006] Obtain patent search information input by the user;
[0007] Retrieving a target patent document from the patent document database according to a matching relationship between the patent search information and preset patent document guide information corresponding to each patent document in the patent document database;
[0008] The patent document guide information is obtained by the following steps:
[0009] Perform word segmentation on the patent text to obtain a set of phrases;
[0010] Dividing the phrase set according to the adjacency relationship between the phrases in the phrase set to obtain a first weighted phrase set and a second weighted phrase set;
[0011] Based on the first information corresponding to each phrase in the first weighted phrase set and the second information corresponding to each phrase in the second weighted phrase set, the first weighted phrase set and the second weighted phrase set are respectively screened to obtain a first target phrase and a second target phrase;
[0012] Patent document guidance information is determined according to the first target phrase and the second target phrase.
[0013] A second aspect of the embodiment of the present application provides a patent document retrieval device, comprising:
[0014] An input information acquisition module is used to acquire patent search information input by a user;
[0015] a patent document retrieval module, configured to retrieve a target patent document from the patent document database according to a matching relationship between the patent retrieval information and preset patent document guide information corresponding to each patent document in the patent document database; and
[0016] A patent document guidance information generation module is used to perform word segmentation processing on a patent text to obtain a phrase set; divide the phrase set according to the adjacency relationship between each phrase in the phrase set to obtain a first weighted phrase set and a second weighted phrase set; based on the first information corresponding to each phrase in the first weighted phrase set and the second information corresponding to each phrase in the second weighted phrase set, the first weighted phrase set and the second weighted phrase set are respectively screened to obtain a first target phrase and a second target phrase; and determine the patent document guidance information according to the first target phrase and the second target phrase.
[0017] The third aspect of an embodiment of the present application provides a computer device, wherein the terminal device includes a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, the steps of the patent document retrieval method as described in any one of the first aspects above are implemented.
[0018] The fourth aspect of the embodiments of the present application provides a computer-readable storage medium, comprising: storing a computer program, characterized in that when the computer program is executed by a processor, the steps of the patent document retrieval method as described in any one of the first aspects above are implemented.
[0019] Compared with the prior art, the embodiments of the present invention have the following beneficial effects:
[0020] By performing word segmentation on the text of the patent document, a phrase set is obtained. According to the connection relationship between the phrases in the phrase set, the phrases in the phrase set are divided into different weighted phrase sets, and keyword selection is performed on the phrases in these weighted sets respectively, so as to obtain target phrases that can summarize the full text of the specification, and the patent document guidance information is determined based on these target phrases. Then, when performing a search, the search is performed based on the patent document guidance information, so as to improve the accuracy of the search. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0022] Figure 1 It is an application environment diagram of a patent search method provided by an embodiment of the present application;
[0023] Figure 2 It is a flowchart of a patent search method provided in an embodiment of the present application;
[0024] Figure 3 It is a flowchart of a method for determining patent document guidance information provided by an embodiment of the present application;
[0025] Figure 4 It is a flowchart of a method for dividing a first weighted phrase set and a second weighted phrase set provided in an embodiment of the present application;
[0026] Figure 5 It is a flowchart of a method for determining adjacency weights provided in an embodiment of the present application;
[0027] Figure 6 It is a flowchart of a method for obtaining first and second information provided by an embodiment of the present application;
[0028] Figure 7 is a flowchart of a method for determining a first adjustment coefficient provided in an embodiment of the present application;
[0029] Figure 8 is a flowchart of a method for determining second information provided in an embodiment of the present application;
[0030] Fig. 9 is a flow chart of a method for determining a second adjustment coefficient provided in an embodiment of the present application;
[0031] Fig.10 It is a schematic diagram of the structure of a patent document retrieval device provided in an embodiment of the present application;
[0032] Fig.11 It is a schematic diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0033] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0034] In order to illustrate the technical solution described in this application, a specific embodiment is provided below for illustration.
[0035] Figure 1 An application environment diagram of a patent document retrieval method provided by an embodiment of the present invention, such as Figure 1 As shown, in the application environment, a terminal 110 and a computer device 120 are included.
[0036] The computer device 120 may be an independent physical server or terminal, or a server cluster composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud servers, cloud databases, cloud storage, and CDN.
[0037] The terminal 110 device may be a station (STAION, ST) in a WLAN, a personal digital assistant (PDA) device, a handheld device with wireless communication function, a computing device or other processing device connected to a wireless modem, a vehicle-mounted device, a vehicle networking terminal, a computer, a laptop computer, a handheld communication device, a handheld computing device, a satellite wireless device, a customer premises equipment (CPE) and / or other devices for communicating on a wireless system and a next-generation communication system, for example, a mobile terminal in a 5G network or a mobile terminal in a future-evolved public land mobile network (PLMN) network, etc. The terminal 110 and the computer device 120 may be connected via a network, and the present invention is not limited thereto.
[0038] like Figure 2 As shown, in one embodiment, a patent document retrieval method is proposed. This embodiment mainly uses the method applied to the above-mentioned computer device 120 as an example. The patent document retrieval method includes:
[0039] Step 202: Obtaining patent search information input by the user;
[0040] Step 204: Retrieve a target patent document from the patent document database based on the patent search information and the matching relationship between the preset patent document guide information corresponding to each patent document in the patent document database.
[0041] The patent search information input by the user may be just some keywords or logical expressions, which is not limited in this application. After obtaining the user's search information, preferably, the user's search information is matched with the subject matter, abstract and patent document guide information of each patent document in the patent document database, and the successfully matched patent documents are returned to the terminal 110.
[0042] like Figure 3 As shown, the patent document guidance information is obtained through the following steps:
[0043] Step 302: Perform word segmentation on the patent text to obtain a phrase set.
[0044] Among them, since in this embodiment, the patent document guide information and the subject matter and abstract of the patent document are matched with the user's search information together, the patent text in this embodiment is the description part of the patent document.
[0045] Preferably, step 302 includes:
[0046] The patent text is processed by sentence segmentation, word segmentation, and stop word filtering to obtain a phrase set.
[0047] Among them, since it is necessary to determine whether each phrase appears in the same sentence in the subsequent processing, the patent document is firstly processed by sentence segmentation, so as to determine whether each phrase appears in the same sentence later. Then each sentence is processed by word segmentation, and the stop words in the sentence are deleted to prevent the stop words without semantics from affecting the subsequent processing. It should be noted that in the patent document, there will be "first spring" or similar descriptions. At this time, "first spring" needs to be identified as a phrase, and it cannot be identified as "first" or "spring". After the above processing, a phrase set can be obtained for subsequent processing.
[0048] Step 304: Divide the phrase set according to the adjacency relationship between the phrases in the phrase set to obtain a first weighted phrase set and a second weighted phrase set.
[0049] Among them, if a phrase is in the same sentence as another phrase, and the number of phrases between the phrase and the other phrase is less than k-1, then the two phrases are adjacent. For example: the patent text states that "by setting a disassembly assembly composed of a placement slot, a clamping rod, and an elastic member to cooperate with a fixed ring, multiple sets of climbing poles can be conveniently disassembled and installed." After preprocessing the above sentence, two subsets are obtained: [Setting a placement slot, a clamping rod, an elastic member, a disassembly assembly, and a fixed ring to cooperate with each other] and [Convenient disassembly and installation of multiple sets of climbing poles]. At this time, k=2, then corresponding to the "mutual" in the first subset, it is adjacent to the "assembly", "fixed ring", and "cooperation" in the first set, but not adjacent to the "convenient" in the second subset, so the two do not belong to the same sentence. According to the connection relationship between each phrase in the phrase set, the connection weight between each phrase can be obtained, and then according to the adjacency weight of each phrase, each phrase is divided into a first weight phrase set and a second weight phrase set, so that the two phrase sets can be processed separately later.
[0050] Step 306: Based on the first information corresponding to each phrase in the first weighted phrase set and the second information corresponding to each phrase in the second weighted phrase set, the first weighted phrase set and the second weighted phrase set are respectively screened to obtain a first target phrase and a second target phrase.
[0051] Among them, the first information corresponding to each phrase in the first weighted phrase set and the second information corresponding to each phrase in the second weighted phrase set are the weights corresponding to each phrase in the two phrase sets, respectively. This phrase weight can be obtained by a keyword extraction method based on weight / score in the prior art, such as TF (Term-Frequency) algorithm, YAKE (YetAnother Keyword Extractor) algorithm, etc. Then, the phrases in the two phrase sets whose corresponding weights are higher than the preset threshold are respectively determined as target phrases, or the phrases with the highest weights in the two phrase sets are respectively determined as target phrases. Here, this application does not limit the target phrase selection method. In addition, since a phrase may appear multiple times in the first weighted set, if the phrase is selected as the first target phrase, the phrase can only appear once in the first target phrase, so that there are no repeated phrases in the first target phrase. Similarly, there are no repeated phrases in the second target phrase.
[0052] Step 308: Determine patent document guidance information according to the first target phrase and the second target phrase.
[0053] Among them, the present application does not limit the method of determining the patent document guidance information according to the first and second target phrases. The first and second target phrases can be placed in the same set, sorted according to the size of the first information of the first target phrase and the second information of the second target phrase, that is, sorted according to the weight of the first and second target phrases, and then several phrases with the highest weights are selected as the patent document guidance information. Alternatively, several phrases are selected from the first target phrase and the second target phrase according to a certain ratio to form the patent document guidance information together.
[0054] Preferably, several phrases with the highest first information in the first weight phrase set are determined as first target phrases, and several phrases with the highest second information in the second weight phrase set are determined as second target phrases, and the number of first target phrases and second target phrases is the same, and then the first target phrases and the second target phrases directly constitute the patent document guidance information.
[0055] In one embodiment, Figure 4 As shown, step 304 includes:
[0056] Step 402: Determine the adjacency weight corresponding to each phrase in the phrase set according to the adjacency relationship between each phrase in the phrase set.
[0057] Among them, the method of determining the adjacency weight corresponding to each phrase according to the adjacency relationship of each phrase in the phrase set can be found in Figure 5 The adjacency weight of each word group is determined according to the adjacency relationship of each word group in the word group set, so that each word group can be divided according to the adjacency weight of each word group later.
[0058] Step 404: Based on the adjacency weights of each phrase in the phrase set and a preset first weight threshold and a second weight threshold, the phrases whose adjacency weights are higher than the first weight threshold are divided into a first weight phrase set, and the phrases whose adjacency weights are lower than the second weight threshold are divided into a second weight phrase set.
[0059] Among them, the preset first weight threshold and the second weight threshold can be set dynamically, for example: the first weight threshold is the average of the adjacency weights of each phrase in the phrase set multiplied by 1.2, and the second weight threshold is the average of the adjacency weights of each phrase in the phrase set multiplied by 0.8, or it can be a constant. The present application does not limit the specific numerical values of the first and second weight thresholds, and technicians in this field can obtain them based on their own experience. It should be noted that when dividing each phrase in the phrase set into the first and second weight phrase sets, it is necessary to divide the sentence information of each phrase in the phrase set into the first and second weight phrase sets together, so that when processing each phrase in the first and second weight phrase sets, it is possible to determine whether each phrase belongs to the same sentence. If a phrase appears frequently in a patent text, its adjacency weight is also high. In the description of a patent document, phrases that appear frequently are generally the names of various components, while phrases that appear less frequently are generally phrases that describe the technical effects. Therefore, through the first and second weight thresholds, each phrase in the phrase set is divided into a high-weight technical feature set and a low-weight technical effect set, and then keyword selection is performed on each phrase in the technical feature set and the technical effect set respectively, which can screen out phrases that describe the technical effects with low frequency but more important words. Moreover, by dividing each phrase in the phrase set by two weights, some unnecessary phrases can be preliminarily filtered out to avoid interference with subsequent processing.
[0060] In one embodiment, Figure 5 As shown:
[0061] Step 402 includes:
[0062] Step 502: construct a phrase graph model according to the adjacency relationship between each phrase in the phrase set and a preset co-occurrence window length.
[0063] Among them, each node in the phrase graph model is each phrase in the phrase set. It should be noted that this implementation records whether each phrase belongs to the same sentence by putting phrases belonging to the same sentence into the same subset. Therefore, a phrase may appear multiple times in the phrase set, but the word only appears once in the phrase graph model. The connection relationship between each node of the phrase graph model is determined by the adjacency relationship between each phrase and the preset co-occurrence window length k. For example, it is recorded in the patent document that "by setting a disassembly and assembly component composed of a placement slot, a clamping rod, and an elastic part to cooperate with a fixed ring, multiple sets of climbing poles can be conveniently disassembled and installed." After preprocessing the above sentence, the two subsets [setting a placement slot, a clamping rod, an elastic part, a disassembly and assembly component, and a fixed ring to cooperate with each other] and [convenient disassembly and installation of multiple sets of climbing poles] are obtained. Then, the following node exists in the phrase graph model [setting a placement slot, a clamping rod, an elastic part, a disassembly and assembly component, and a fixed ring to cooperate with each other to facilitate the disassembly and installation of multiple sets of climbing poles]. Assuming that the co-occurrence window length at this time is 2, then the phrases connected to the phrase "mutually" in the first subset are "component", "fixed ring", and "cooperation" respectively. Therefore, in the phrase graph model, there is an undirected edge between any two nodes among the nodes "mutually", "component", "fixed ring", and "cooperation". It should be noted that since a phrase can appear in different sentences, it has different adjacency relationships in different sentences, but it has only one corresponding node in the phrase graph model, so the adjacency relationship of the phrase in different sentences will be reflected in the corresponding nodes in the phrase graph model. For example, the phrase "spring" is adjacent to "support plate" in one sentence and to "bracket" in another sentence. Therefore, in the phrase graph model, there is an undirected edge between the node "spring" and the node "support plate", and there is also an undirected edge between the node "spring" and the node "between". After performing the above processing on each phrase in the phrase set, the phrase graph model can be obtained.
[0064] Step 504: Obtain the initial weight of each phrase node in the phrase graph model.
[0065] The initial weight of each phrase node may be obtained based on the experience of those skilled in the art, and the present application does not impose any limitation thereto. Preferably, the initial weight of each phrase node is 1.
[0066] Step 506: Determine the connection weight of each phrase node in the phrase graph model according to the initial weight of each phrase node in the phrase graph model and the connection relationship between each phrase node.
[0067] The connection weight of each phrase node can be iteratively calculated using the following formula:
[0068]
[0069] In the formula, Representative Node The connection weight of , a is the damping coefficient, and its value is between 0 and 1. The specific value can be determined by the experience of technicians in this field. , Represents a node in the phrase graph model, Represents the node is the set of incoming edges of the endpoint, Represents the node is the set of outgoing edges from the starting point, Representative Node With Node The initial weight of each node is substituted into the above formula for iterative calculation until the difference between the weights of the same node in the previous and next iterations is lower than the preset threshold, thereby obtaining the connection weight of each node.
[0070] Step 508: Determine the adjacency weight corresponding to each phrase in the phrase set according to the connection weight of each phrase node in the phrase graph model.
[0071] The connection weight of each node in the phrase graph model is set to the adjacency weight corresponding to each phrase in the phrase set. It should be noted that when a node in the phrase graph model has multiple corresponding phrases in the phrase set, the adjacency weights of the multiple phrases are all the connection weight of the node.
[0072] In one embodiment, Figure 6 As shown, before step 306, the following steps are included:
[0073] Step 602: Determine the adjacency weight of each phrase in the first weighted phrase set according to the adjacency relationship of each phrase in the first weighted phrase set.
[0074] The method of determining the adjacency weight of each phrase according to the adjacency relationship of each phrase in the first weight set and Figure 5 The method described in the corresponding embodiment is similar and will not be repeated here.
[0075] Step 604: Determine first information of each phrase in the first weighted phrase set according to the adjacency weight of each phrase in the first weighted phrase set and a preset first adjustment coefficient.
[0076] Among them, the preset first adjustment coefficient can be used to scale the adjacent weights of each phrase in the first weight phrase set, so as to obtain the first information of each phrase in the first weight phrase set, and the range of the first information is similar to the range of the second information of each phrase in the second weight phrase set, so that each phrase in the first weight phrase set and each phrase in the second weight phrase set can be put together for comparison later. The preset first adjustment coefficient can also be used to adjust the adjacent weights of each phrase in the first weight phrase set according to specific rules, so as to facilitate the screening of more suitable keywords from the first weight phrase set. Preferably, the first adjustment coefficient of each phrase in the first weight phrase set can be determined based on the drawing information of the patent document, so as to facilitate the screening of more important technical feature parts therefrom.
[0077] Step 606: Obtain the position information of each phrase in the second weighted phrase set in the corresponding patent text.
[0078] Among them, the position information of each phrase in the second weighted phrase set in the corresponding patent text includes: the paragraph information of each phrase in the second weighted phrase set in the patent text, the number of times each phrase in the second weighted phrase set appears in the patent text, etc.
[0079] Step 608: Determine the second information of each phrase in the second weighted phrase set based on the position information of each phrase in the second weighted phrase set in the patent text and a preset second adjustment coefficient.
[0080] Among them, the preset second adjustment coefficient can be used to adjust the range of the second information of each phrase in the second weighted phrase set so that the range of the second information is similar to the range of the first information of each phrase in the first weighted phrase set, so as to compare each phrase in the first weighted phrase set and each phrase in the second weighted phrase set together. The preset second adjustment coefficient can also be used to adjust each phrase in the second weighted phrase set according to a specific rule to control the size of the second information of each phrase. Preferably, the second adjustment coefficient of each phrase in the second weighted phrase set is determined according to the part of speech of each phrase in the second weighted phrase set.
[0081] In one embodiment, Figure 7 As shown, before step 604, the following steps are included:
[0082] Step 702: Based on the distance between the position of each component in the patent document drawing and the corresponding center point of the drawing, a preset weight database is queried to obtain the weight coefficient corresponding to each component in the patent document drawing.
[0083] Among them, according to the positions pointed to by the corresponding indicator lines in the drawings of the patent document, and the corresponding relationship between the numbers of the drawings and the names of the components in the description of the drawings in the patent text, the positions of the components in the drawings of the patent document are determined, the center positions of the components are calculated according to the positions of the components in the drawings of the patent document, the distance between the center position of the components and the center point of the corresponding drawings is calculated, and then the weight coefficient of the components in the corresponding drawings is determined according to the distance between the center position of the components and the center point of the corresponding drawings. The specific corresponding relationship between the distance from the center position of each component to the center point of the corresponding drawings and the weight coefficient of each component can be obtained based on the experience of technicians in this field, and this application is not limited here.
[0084] Step 704: For each component in the patent document drawing, if the component has multiple corresponding positions in the corresponding patent document drawing, the preset weight coefficients corresponding to the positions are averaged.
[0085] When a component appears in multiple drawings in a patent document, it may have multiple corresponding weight coefficients. At this time, all corresponding weight coefficients of the component are averaged, and the average is determined as the weight coefficient of the component, so that the component has only one weight coefficient.
[0086] Step 706: Determine a first adjustment coefficient for each phrase in the first weight phrase set based on the weight coefficients corresponding to each component in the drawings of the patent document.
[0087] The weight coefficients of the components in the drawings of the patent document are set as the first adjustment coefficients of the phrases corresponding to the names of the components in the first weight phrase set. Since some phrases in the first weight phrase set are not the names of components, these phrases do not have corresponding weight coefficients and first adjustment coefficients, and the first adjustment coefficients of these phrases are 0. At this time, the calculation formula for the first information of each phrase in the first weight set is as follows:
[0088]
[0089] In the above formula, represents the first information of the i-th phrase in the first weight set, represents the first adjustment coefficient of the i-th phrase in the first weight set, Represents the adjacency weight of the i-th word group in the first weight set.
[0090] In one embodiment, Figure 8 As shown, step 608 includes:
[0091] Step 802: Determine the weighted word frequency of each word group in the second weighted word group set according to the word frequency of each word group in the second weighted word group set and a preset second adjustment coefficient.
[0092] The calculation formula of the weighted word frequency of each phrase is as follows:
[0093]
[0094] In the above formula, represents the weighted word frequency of the i-th word group in the second weight set, represents the second adjustment coefficient of the i-th phrase in the second weight set, represents the number of times the i-th phrase in the second weight set appears in the patent document, Represents the total number of occurrences of all phrases in the patent document.
[0095] Step 804: Determine the reverse paragraph frequency of each phrase in the second weighted phrase set according to the total number of paragraphs in the patent text and the number of paragraphs corresponding to each phrase in the second weighted phrase set in the patent text.
[0096] Among them, the calculation formula for the reverse paragraph of each phrase is as follows:
[0097]
[0098] In the above formula, represents the inverse paragraph frequency of the i-th phrase in the second weight set, represents the total number of paragraphs in the patent text, Represents the number of paragraphs in the patent text that contain the i-th phrase in the second weight set. It should be noted that a phrase in the second weight set may appear multiple times in a paragraph, but it will only be recorded once. For example, a phrase in the second weight set is "spring", which appears twice in the second paragraph of the patent text and once in the third paragraph of the patent text, and does not appear in other areas of the patent text. Then the corresponding number of the phrase is is 2.
[0099] Step 806: Determine the second information of each phrase in the second weighted phrase set according to the weighted word frequency of each phrase in the second weighted phrase set and the corresponding inverse paragraph frequency.
[0100] The calculation formula of the second information of each phrase is:
[0101]
[0102] In the above formula, The second information representing the i-th phrase in the second weighted phrase set, where k are all the phrases participating in the operation in the second weight set, that is, each phrase in the second weight set.
[0103] After calculating the second information of each phrase in the second weighted phrase set, the phrases in the second weighted phrase set can be sorted according to the magnitude of the second information, and several second target phrases with the highest second information can be selected therefrom. It should be noted that there may be multiple identical phrases in the second weighted phrase set, and the second information corresponding to multiple identical phrases is also the same. If such a phrase is selected as a second target phrase, it will only be selected once, so that there are no duplicate phrases in the second target phrases.
[0104] Such as Fig. 9 As shown, in one embodiment, before step 608, it includes:
[0105] Step 902: Query a preset coefficient database based on the part of speech of each phrase in the second weighted phrase set to obtain the second adjustment coefficient corresponding to each phrase in the second weighted phrase.
[0106] Among them, the coefficients corresponding to each part of speech can be obtained according to the experience of those skilled in the art, and the present application does not make specific limitations here. Preferably, if the part of speech of a phrase is a verb, its corresponding coefficient is 0.2; if the part of speech of a phrase is a noun, its corresponding coefficient is 0.17; if the part of speech of a phrase is an adjective, its corresponding coefficient is 0.15.
[0107] The second adjustment coefficient corresponding to each phrase in the phrase.
[0108] For each phrase in the second weighted phrase set, if the phrase has multiple corresponding parts of speech, the maximum candidate coefficient among the preset candidate coefficients corresponding to each part of speech is set as the second adjustment coefficient of the phrase.
[0109] Among them, since a phrase may appear in different positions in the patent text and has different parts of speech at different positions, resulting in it having multiple parts of speech, the maximum candidate coefficient among the preset candidate coefficients corresponding to the multiple parts of speech of the phrase is set as the second adjustment coefficient of the phrase. For example: There is a phrase "lifting" in the second weighted phrase. Its part of speech at position 1 is a verb, and the corresponding candidate coefficient is 0.2; its part of speech at position 2 is a noun, and the corresponding candidate coefficient is 0.17. Then the second adjustment coefficient of this phrase is 0.2.
[0110] In one embodiment, a patent retrieval device is provided, such as Fig.10 As shown, the patent retrieval device includes:
[0111] An input information acquisition module 1010 is used to acquire patent search information input by a user; and
[0112] A patent document retrieval module 1020 is used to retrieve a target patent document from the patent document database according to a matching relationship between the patent retrieval information and preset patent document guide information corresponding to each patent document in the patent document database; and
[0113] The patent document guidance information generation module 1030 is used to perform word segmentation processing on the patent text to obtain a phrase set; divide the phrase set according to the adjacency relationship between each phrase in the phrase set to obtain a first weighted phrase set and a second weighted phrase set; based on the first information corresponding to each phrase in the first weighted phrase set and the second information corresponding to each phrase in the second weighted phrase set, the first weighted phrase set and the second weighted phrase set are respectively screened to obtain a first target phrase and a second target phrase; determine the patent document guidance information according to the first target phrase and the second target phrase.
[0114] Among them, the process of the patent search device module realizing its respective functions can be specifically referred to the description of the aforementioned embodiment, which will not be repeated here.
[0115] In one embodiment, a computer device is provided, such as Fig.11 As shown, the computer device includes a memory and a processor, the memory stores a computer program that can be run on the processor, and the processor performs the following steps when executing the computer program:
[0116] Obtain patent search information input by the user;
[0117] Retrieving a target patent document from the patent document database according to a matching relationship between the patent search information and preset patent document guide information corresponding to each patent document in the patent document database;
[0118] The patent document guide information is obtained by the following steps:
[0119] Perform word segmentation on the patent text to obtain a set of phrases;
[0120] Dividing the phrase set according to the adjacency relationship between the phrases in the phrase set to obtain a first weighted phrase set and a second weighted phrase set;
[0121] Based on the first information corresponding to each phrase in the first weighted phrase set and the second information corresponding to each phrase in the second weighted phrase set, the first weighted phrase set and the second weighted phrase set are respectively screened to obtain a first target phrase and a second target phrase;
[0122] Patent document guidance information is determined according to the first target phrase and the second target phrase.
[0123] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the processor performs the following steps:
[0124] Obtain patent search information input by the user;
[0125] Retrieving a target patent document from the patent document database according to a matching relationship between the patent search information and preset patent document guide information corresponding to each patent document in the patent document database;
[0126] The patent document guide information is obtained by the following steps:
[0127] Perform word segmentation on the patent text to obtain a set of phrases;
[0128] Dividing the phrase set according to the adjacency relationship between the phrases in the phrase set to obtain a first weighted phrase set and a second weighted phrase set;
[0129] Based on the first information corresponding to each phrase in the first weighted phrase set and the second information corresponding to each phrase in the second weighted phrase set, the first weighted phrase set and the second weighted phrase set are respectively screened to obtain a first target phrase and a second target phrase;
[0130] Patent document guidance information is determined according to the first target phrase and the second target phrase.
[0131] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0132] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.
[0133] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0134] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0135] In addition, in the description of the present specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish descriptions and should not be understood as indicating or implying relative importance. It should also be understood that although the terms "first", "second", etc. are used in the text to describe various elements in some embodiments of the present application, these elements should not be limited by these terms. These terms are only used to distinguish one element from another element.
[0136] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0137] The present application implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by the processor, the steps of each of the above-mentioned method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc.
[0138] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0139] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0140] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0141] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.
Claims
1. A patent document retrieval method, characterized in that: include: Obtain patent search information input by the user; Retrieving a target patent document from the patent document database according to a matching relationship between the patent search information and preset patent document guide information corresponding to each patent document in the patent document database; The patent document guide information is obtained by the following steps: Perform word segmentation on the patent text to obtain a set of phrases; Dividing the phrase set according to the adjacency relationship between the phrases in the phrase set to obtain a first weighted phrase set and a second weighted phrase set; Based on the first information corresponding to each phrase in the first weighted phrase set and the second information corresponding to each phrase in the second weighted phrase set, the first weighted phrase set and the second weighted phrase set are respectively screened to obtain a first target phrase and a second target phrase; Determining patent document guidance information according to the first target phrase and the second target phrase; The step of dividing the phrase set according to the adjacency relationship between the phrases in the phrase set to obtain a first weighted phrase set and a second weighted phrase set includes: Determine the adjacency weight corresponding to each phrase in the phrase set according to the adjacency relationship between each phrase in the phrase set; According to the adjacency weights of the respective phrases in the phrase set and a preset first weight threshold and a second weight threshold, the phrases whose adjacency weights are higher than the first weight threshold are divided into a first weight phrase set, and the phrases whose adjacency weights are lower than the second weight threshold are divided into a second weight phrase set; The step of determining the adjacency weight corresponding to each phrase in the phrase set according to the adjacency relationship between each phrase in the phrase set includes: According to the adjacency relationship between each phrase in the phrase set and the preset co-occurrence window length, a phrase graph model is constructed; Obtaining the initial weight of each phrase node in the phrase graph model; Determining the connection weight of each phrase node in the phrase graph model according to the initial weight of each phrase node in the phrase graph model and the connection relationship between each phrase node; According to the connection weight of each phrase node in the phrase graph model, the adjacency weight corresponding to each phrase in the phrase set is determined.
2. The patent document retrieval method according to claim 1, characterized in that: The first weighted phrase set and the second weighted phrase set are respectively screened based on the first information corresponding to each phrase in the first weighted phrase set and the second information corresponding to each phrase in the second weighted phrase set to obtain the first target phrase and the second target phrase, including: Determining the adjacency weight of each phrase in the first weighted phrase set according to the adjacency relationship of each phrase in the first weighted phrase set; Determining first information of each phrase in the first weighted phrase set according to the adjacency weight of each phrase in the first weighted phrase set and a preset first adjustment coefficient; Obtaining position information of each phrase in the second weighted phrase set in the corresponding patent text; The second information of each phrase in the second weighted phrase set is determined according to the position information of each phrase in the patent text and a preset second adjustment coefficient.
3. The patent document retrieval method according to claim 2, characterized in that: Before determining the first information of each phrase in the first weighted phrase set according to the adjacent weight of each phrase in the first weighted phrase set and the preset first adjustment coefficient, the method includes: Based on the distance between the position of each component in the drawings of the patent document and the center point of the corresponding drawings, a preset weight database is queried to obtain the weight coefficient corresponding to each component in the drawings of the patent document; For each component in the drawings of the patent document, if the component has multiple corresponding positions in the corresponding drawings of the patent document, the preset weight coefficients corresponding to the positions are averaged; According to the weight coefficients corresponding to the various components in the drawings of the patent document, the first adjustment coefficients of the various phrases in the first weighted phrase set are determined.
4. The patent document retrieval method according to claim 2, characterized in that: The determining of the second information of each phrase in the second weighted phrase set according to the position of each phrase in the patent text and the preset second adjustment coefficient includes: Determining the weighted word frequency of each word group in the second weighted word group set according to the word frequency of each word group in the second weighted word group set and a preset second adjustment coefficient; Determine the reverse paragraph frequency of each phrase in the second weighted phrase set according to the total number of paragraphs in the patent text and the number of paragraphs corresponding to each phrase in the second weighted phrase set in the patent text; The second information of each phrase in the second weighted phrase set is determined according to the weighted word frequency of each phrase in the second weighted phrase set and the corresponding inverse paragraph frequency.
5. The patent document retrieval method according to claim 2, characterized in that: Before determining the second information of each phrase in the second weighted phrase set according to the position of each phrase in the second weighted phrase set in the patent text and the preset second adjustment coefficient, the method includes: Based on the part of speech of each phrase in the second weighted phrase set, a preset coefficient database is searched to obtain a second adjustment coefficient corresponding to each phrase in the second weighted phrase set; For each phrase in the second weighted phrase set, if the phrase has multiple corresponding parts of speech, the largest candidate coefficient among the preset candidate coefficients corresponding to the parts of speech is set as the second adjustment coefficient of the phrase.
6. A patent document retrieval device, characterized in that: include: An input information acquisition module is used to acquire patent search information input by a user; A patent document retrieval module, used to retrieve a target patent document from the patent document database according to a matching relationship between the patent retrieval information and preset patent document guide information corresponding to each patent document in the patent document database; as well as, A patent document guidance information generation module is used to perform word segmentation processing on the patent text to obtain a phrase set; divide the phrase set according to the adjacency relationship between each phrase in the phrase set to obtain a first weighted phrase set and a second weighted phrase set; based on the first information corresponding to each phrase in the first weighted phrase set and the second information corresponding to each phrase in the second weighted phrase set, respectively screen the first weighted phrase set and the second weighted phrase set to obtain a first target phrase and a second target phrase; determine the patent document guidance information according to the first target phrase and the second target phrase; The step of dividing the phrase set according to the adjacency relationship between the phrases in the phrase set to obtain a first weighted phrase set and a second weighted phrase set includes: Determine the adjacency weight corresponding to each phrase in the phrase set according to the adjacency relationship between each phrase in the phrase set; According to the adjacency weights of the respective phrases in the phrase set and a preset first weight threshold and a second weight threshold, the phrases whose adjacency weights are higher than the first weight threshold are divided into a first weight phrase set, and the phrases whose adjacency weights are lower than the second weight threshold are divided into a second weight phrase set; The step of determining the adjacency weight corresponding to each phrase in the phrase set according to the adjacency relationship between each phrase in the phrase set includes: According to the adjacency relationship between each phrase in the phrase set and the preset co-occurrence window length, a phrase graph model is constructed; Obtaining the initial weight of each phrase node in the phrase graph model; Determining the connection weight of each phrase node in the phrase graph model according to the initial weight of each phrase node in the phrase graph model and the connection relationship between each phrase node; According to the connection weight of each phrase node in the phrase graph model, the adjacency weight corresponding to each phrase in the phrase set is determined.
7. A computer device, characterized in that: The computer device comprises a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor implements the steps of the method according to any one of claims 1 to 5 when executing the computer program.
8. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Searching method and device based on data analysis and terminal
CN110083681A
Text retrieval method, device, storage medium and server
CN114090799A