Information retrieval method and related apparatus
Patent Information
- Application Number
- PCT/CN2026/084253
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-03-19
- Filing Date
- 2026-03-18
- Publication Date
- 2026-09-24
Smart Images

Figure CN2026084253_24092026_PF_FP_ABST
Abstract
Description
Information retrieval methods and related devices
[0001] Cross-references to related applications
[0002] This application claims priority to Chinese patent application filed on March 19, 2025, application number 202510329888.1, entitled "Information Retrieval Method and Related Apparatus", the entire contents of which are incorporated herein by reference. Technical Field
[0003] The embodiments described in this application relate to the field of computer technology, and in particular to an information retrieval method and related apparatus. Background Technology
[0004] Social development relies on technological progress, which generates a vast number of technological solutions. Those who possess these solutions desire legal protection or want their implementation unaffected by existing protected technologies. Therefore, prior art searches are necessary to generate analytical reports characterizing the differences between the proposed solution and existing technologies, thereby determining the innovativeness of the solution. However, current search methods still suffer from low information retrieval efficiency. Summary of the Invention
[0005] In view of this, multiple embodiments of this application aim to provide an information retrieval method and related apparatus that can improve retrieval efficiency and save information retrieval resource consumption on the server.
[0006] One embodiment of this application provides an information retrieval method, the method comprising: acquiring, through a client's visual interface, user-inputted information for characterizing a technical solution; generating an analysis report that meets expected analysis objectives; wherein the analysis report is generated based on the information to be retrieved and prior art documents retrieved based on a retrieval query, and the analysis report is at least used to characterize the feature differences between the information to be retrieved and the prior art documents, the retrieval query is generated based on the information to be retrieved, and the expected analysis objectives are used to characterize the analysis direction provided by the analysis report that the user expects; and sending the analysis report to the client to instruct the client to present and modify the feature differences in the analysis report through the visual interface.
[0007] In some possible implementations, the method further includes: receiving a model instruction submitted by a user; wherein the model instruction is a prompt instruction for instructing a large model to generate a search query; responding to the model instruction and generating a search query based on the information to be searched; and sending the search query to the client to instruct the client to display the search query through the visual interface.
[0008] In some possible implementations, the step of responding to the model instruction and generating a search query based on the information to be retrieved includes: responding to the model instruction by extracting technical features from the information to be retrieved; sending the technical features to a client to instruct the client to display the technical features through a visual interface; wherein the technical features are determined based on a pre-trained semantic parsing model and a named entity recognition algorithm; receiving user-confirmed technical features from the client to extract search elements from the user-confirmed technical features; wherein the search elements and technical features are used to characterize the information to be retrieved at different granularities; sending the search elements to the client to instruct the client to display and provide feedback on the user-confirmed search elements; and generating a search query based on the user-confirmed search elements.
[0009] In some possible implementations, generating a search expression based on the user-confirmed search elements includes: obtaining supplementary elements corresponding to the search elements, wherein the supplementary elements include at least one of the following: classification number, material keywords, functional descriptive terms, process parameters, equipment requirements, and quality control indicators; generating a search expression based on a first relationship between the search elements and the supplementary elements, and a second relationship between each of the search elements; wherein the first relationship includes any one of a superior relationship, a subordinate relationship, and a synonym relationship, and the second relationship includes a parallel relationship.
[0010] In some possible implementations, the prior art document retrieval process includes: performing a retrieval process based on the retrieval query to retrieve prior art documents related to the information to be retrieved; wherein the retrieval process includes at least one of the following: semantic retrieval process, classification number retrieval process, keyword retrieval process, and family citation retrieval process; sending the prior art documents to the client to instruct the client to visualize the real-time status of the prior art document retrieval process through a dynamic progress bar and result preview component.
[0011] In some possible implementations, the search query includes a semantic search query; when the search process includes the semantic search process, the prior art documents include a first set of prior art documents; the process of determining the first set of prior art documents includes: constructing a semantic search query based on the semantic features of the information to be searched; and recalling the first set of prior art documents based on the semantic search query.
[0012] In some possible implementations, the search query includes a classification number search query; if the search process also includes the classification number search process, the prior art documents also include a second set of prior art documents; the process of determining the second set of prior art documents includes: constructing the classification number search query based on the classification number statistics of the first set of prior art documents; and recalling the second set of prior art documents based on the classification number search query.
[0013] In some possible implementations, the search query includes a keyword search query; when the search process includes the keyword search process, the prior art documents include a third set of prior art documents; the process of determining the third set of prior art documents includes: generating the keyword search query based on the search elements in the information to be searched, and recalling a set of candidate prior art documents based on the keyword search query; determining keywords corresponding to the search elements from the set of candidate prior art documents based on the feature comparison results of each document in the set of candidate prior art documents with the information to be searched; wherein, the keywords include at least one of the synonyms, related terms, hypernyms, and hyponyms of the search elements; iterating the keyword search query based on the keywords to update the set of candidate prior art documents until the number of iterations meets the expected number, and determining the set of candidate prior art documents obtained from the last iteration as the third set of prior art documents.
[0014] In some possible implementations, the method further includes: determining the frequency of occurrence of keywords corresponding to each search element in the candidate prior art document set based on the feature comparison results between each document in the candidate prior art document set and the information to be retrieved; determining the weight corresponding to each search element based on the frequency of occurrence; and updating the weight in the keyword search expression.
[0015] In some possible implementations, the search query includes a family citation search query; when the search process includes the family citation search process, the prior art documents include a fourth set of prior art documents; the process of determining the fourth set of prior art documents includes: generating the family citation search query based on the associated files of the prior art documents retrieved in at least one of the semantic search process, the classification number search process, and the keyword search process; wherein, the associated files are at least one of the family files and citation files of the retrieved prior art documents; and recalling the fourth set of prior art documents according to the family citation search query.
[0016] In some possible implementations, the expected analysis objective includes at least one of FTO analysis, invalidity analysis, novelty search analysis, and infringement analysis. Generating an analysis report that meets the expected analysis objective includes: when the expected analysis objective is FTO analysis, obtaining prior art patent maintenance status information and territorial scope information to determine the relevance between the prior art documents and the technical solution, and generating an analysis report; when the expected analysis objective is invalidity analysis, using a feature matrix matching algorithm, selecting at least one target document with novelty or inventiveness challenge potential from the prior art documents, and constructing a model based on the argument chain to logically map the technical features in the target documents to the technical features in the information to be retrieved, and generating an analysis report; when the expected analysis objective is novelty search analysis, performing feature comparison between the information to be retrieved and the prior art documents to determine the distinguishing technical features between the information to be retrieved and the prior art documents, and evaluating the inventiveness of the distinguishing technical features, and generating an analysis report; when the expected analysis objective is infringement analysis, using a combination of semantic parsing and neural network models to determine infringement between the information to be retrieved and the prior art documents, and generating an analysis report.
[0017] In some possible implementations, the analysis report includes a technical feature comparison table between the technical solution and the prior art documents; the process of generating the technical feature comparison table includes: extracting corresponding technical features from the prior art documents and the information to be retrieved to form technical feature pairs; and generating the technical feature comparison table based on the technical feature pairs.
[0018] In some possible implementations, the step of extracting corresponding technical features from the prior art documents and the information to be retrieved to form technical feature pairs includes: selecting prior art documents that meet the expected relevance from the prior art documents based on the relevance between the prior art documents and the information to be retrieved; and extracting corresponding technical features from the prior art documents that meet the expected relevance and the information to be retrieved to form technical feature pairs.
[0019] In some possible implementations, the method further includes: receiving a file operation instruction sent by a client for a specified file; wherein the file operation instruction includes a file addition instruction and a file deletion instruction; the specified file includes at least one file among existing technical documents retrieved based on the retrieval query and other files imported by the user; in response to the file addition instruction, adding the specified file to existing technical documents used to form the technical feature pair; or, in response to the file deletion instruction, deleting the specified file from existing technical documents used to form the technical feature pair.
[0020] In some possible implementations, after receiving user-inputted information representing a technical solution via a client, the method further includes: if the amount of data in the information to be retrieved is greater than or equal to a preset amount of data, extracting technical features from the extracted results after semantic extraction of the information to be retrieved; or, if the amount of data in the information to be retrieved is less than the preset amount of data, extracting technical features from the information to be retrieved; extracting search elements from the technical features, and generating a search expression based on the search elements; wherein the search elements and technical features are used to represent the information to be retrieved at different granularities; comparing the existing technical documents recalled based on the search expression with the information to be retrieved; based on the feature comparison results, selecting multiple existing technical documents with the highest relevance to the information to be retrieved as target documents from the existing technical documents; wherein the analysis report is generated based on the target documents and the information to be retrieved.
[0021] One embodiment of this application also provides an information retrieval device, the device comprising: an information acquisition module, configured to acquire, through a client's visual interface, user-inputted information representing a technical solution to be retrieved; an analysis report display module, configured to generate an analysis report that meets the expected analysis objectives; wherein the analysis report is generated based on the information to be retrieved and prior art documents retrieved based on a retrieval query, and the analysis report is at least used to represent the feature differences between the information to be retrieved and the prior art documents, the retrieval query is generated based on the information to be retrieved, and the expected analysis objectives are used to represent the analysis direction provided by the analysis report that the user expects; and sending the analysis report to the client to instruct the client to present and modify the feature differences in the analysis report through the visual interface.
[0022] In some possible implementations, the apparatus further includes: a search query display module, configured to receive a model instruction submitted by a user; wherein the model instruction is a prompt instruction for instructing a large model to generate a search query; in response to the model instruction, generate a search query based on the information to be searched; and send the search query to the client to instruct the client to display the search query through the visual interface.
[0023] One embodiment of this application also provides a computer device, the computer device including a memory and a processor, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to implement the method as described above.
[0024] One embodiment of this application also provides a computer-readable storage medium storing at least one computer program that, when executed by a processor, can implement the method described above.
[0025] One embodiment of this application also provides a computer program product for implementing the method as described above.
[0026] In several embodiments provided in this application, user-inputted information representing a technical solution can be obtained through a client's visual interface; an analysis report conforming to the expected analysis objectives can be generated; wherein, the analysis report is generated based on the information to be retrieved and prior art documents retrieved based on a retrieval query, and the analysis report is at least used to characterize the feature differences between the information to be retrieved and the prior art documents, the retrieval query is generated based on the information to be retrieved, and the expected analysis objectives are used to characterize the analysis direction expected by the user in the analysis report; the analysis report is sent to the client to instruct the client to present and modify the feature differences in the analysis report through the visual interface. Since the retrieval information is based on user input, and prior art documents related to the information to be retrieved are retrieved using a retrieval query; and the analysis report can be summarized and formed based on the association between the information to be retrieved and the retrieved prior art documents, the manual writing of analysis work can be reduced. In summary, implementing the embodiments of this application can achieve automatic generation of retrieval-based analysis reports, reduce the workload of manually writing analysis reports, and improve information retrieval efficiency. Attached Figure Description
[0027] Figure 1 is a schematic diagram of a system architecture for implementing an information retrieval method according to one embodiment of this application.
[0028] Figure 2 is a flowchart of an information retrieval method provided in one embodiment of this application.
[0029] Figure 3 is a flowchart of an information retrieval method provided in another embodiment of this application.
[0030] Figure 4 is a schematic diagram of an information input interface provided in one embodiment of this application.
[0031] Figure 5 shows the technical features of one embodiment of this application.
[0032] Figure 6 shows a search element display interface provided in one embodiment of this application.
[0033] Figure 7 is a schematic diagram of an information retrieval device provided in one embodiment of this application.
[0034] Figure 8 is a schematic diagram of the information retrieval device provided in another embodiment of this application.
[0035] Figure 9 is a schematic diagram of a computer device provided in one embodiment of this application. Detailed Implementation
[0036] The information to be retrieved in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments.
[0037] In the description of the embodiments of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly specified.
[0038] Technological development drives social progress. To accelerate this development, relevant laws and regulations protect technological innovation. A series of intellectual achievements generated during technological development can be solidified through patents and other means. Within the scope of patent protection, technical solutions are effectively protected. For the holder of a technical solution, whether implementing the solution themselves, protecting it through patents, conducting infringement checks on existing patents, or requesting invalidation of others' patents, a search of existing technologies is necessary. This search aims to identify one or more existing technologies highly relevant to the technical solution. Based on the differences between these existing technologies and the technical solution, conclusions such as novelty search results, infringement determinations, and invalidation evidence conclusions are derived. These conclusions are presented in the form of a search report, demonstrating the differences between the existing technologies and the technical solution, and the conclusions logically derived from these differences.
[0039] A search query is a logical statement that follows specified query rules to retrieve existing technologies that match the query from a prior art database. Search queries are typically constructed by relevant personnel combining their professional knowledge and understanding of the technical solutions. For example, TACD_ALL:(("anti-glare" AND "glass") OR ("nano" AND "touchscreen") OR ("display")). The quality of the search query directly determines the relevance of the retrieved existing technologies to the technical solutions, and the quality of the search query depends on the search experience of the relevant personnel. TACD_ALL is a search field identifier, indicating a cross-field search (i.e., a search across all fields) within the main paragraphs of patent documents, such as the title (T), abstract (A), claims (C), and description (D). Specifically, relevant personnel need to manually break down the technical solution into several technical features and construct complex logical expressions as search queries using professional search syntax. In the field of information retrieval, relevant information retrieval methods mainly rely on manually constructed keyword matching rules or simple semantic analysis based on statistical models. However, when dealing with complex technical solutions involving multiple fields, special technical expressions, or implicit technical features, relevant information retrieval methods often face problems such as excessively long search query construction time and incomplete semantic coverage. Especially in professional scenarios such as patent novelty searches and infringement analysis, manually constructing search queries requires relevant personnel to have a deep technical background and search experience, which leads to long search times and low search efficiency.
[0040] Therefore, it is necessary to provide an information retrieval method that can acquire user-inputted information representing a technical solution and then provide an analysis report; wherein the analysis report is generated based on the information to be retrieved and prior art documents recalled based on the retrieval query. Implementing the embodiments of this application can achieve automatic generation of retrieval-based analysis reports, reduce manual writing of analysis reports, and improve information retrieval efficiency.
[0041] Please refer to Figure 1. In several embodiments provided in this application, the information retrieval method can be applied to an information retrieval device. The information retrieval device can be an electronic device with certain computing power and network access capabilities. This electronic device can be a desktop computer, laptop computer, tablet computer, or a server. This electronic device can connect to the server via a network. The server can be a distributed server, including multiple processors, memory, network communication modules, etc., working together to achieve various functions. Alternatively, the server can also be a server cluster formed by several servers, possessing higher computing and data processing capabilities. With the development of science and technology, the server can also be implemented using new forms of technology, such as a new type of "server" based on quantum computing. Of course, in some embodiments, the information retrieval device can also be a program module running in an electronic device.
[0042] Specifically, the electronic device includes a processor, a memory, a display module, and a network access module for network communication. The processor acquires user-inputted information representing a technical solution and stores it in the memory. The network access module reads the information from the memory and sends it to a server. The server generates a search query based on the information. The network access module receives the search query returned by the server and stores it in the memory. The display module presents the search query in the memory to the user. Furthermore, the server generates an analysis report based on the information and prior art documents retrieved using the search query. The network access module receives the analysis report returned by the server and stores it in the memory. The display module presents the analysis report in the memory to the user.
[0043] In some implementations, the processor in the electronic device can also be used to generate a search query based on the information to be retrieved, and to generate an analysis report based on the information to be retrieved and prior art documents retrieved based on the search query.
[0044] Please refer to Figure 2. Embodiments of this application also provide an information retrieval method, which includes:
[0045] Step S110: Obtain the search information input by the user to characterize the technical solution through the client's visual interface.
[0046] Step S120: Generate an analysis report that meets the expected analysis objectives; wherein the analysis report is generated based on the information to be retrieved and existing technical documents retrieved based on the retrieval formula, and the analysis report is at least used to characterize the feature differences between the information to be retrieved and the existing technical documents, the retrieval formula is generated based on the information to be retrieved, and the expected analysis objectives are used to characterize the analysis direction that the user expects the analysis report to provide.
[0047] Step S130: Send the analysis report to the client to instruct the client to present and modify the feature differences in the analysis report through the visualization interface.
[0048] In this embodiment, user-inputted information representing a technical solution can be retrieved, and an analysis report can be provided. This analysis report is generated based on the retrieval information and prior art documents retrieved using the search query. Implementing this embodiment enables the automatic generation of search-based analysis reports, reducing the workload of manually writing analysis reports and improving information retrieval efficiency. Furthermore, since it reduces the need for personnel with specialized knowledge to complete the retrieval process, it also lowers the threshold and difficulty of information retrieval.
[0049] Optionally, the information retrieval method may further include: extracting corresponding technical features from the existing technical documents and the information to be retrieved to form technical feature pairs; and generating a technical feature comparison table based on the technical feature pairs.
[0050] The specific functions and effects of the information retrieval method implemented in this embodiment can be explained by referring to other embodiments of this application, and will not be repeated here.
[0051] In some embodiments, the information retrieval method may further include: displaying a search query generated based on the information to be retrieved. Please refer to Figure 3. One embodiment of this application provides an information retrieval method. The information retrieval method can be applied to an information retrieval device. The information retrieval method may include the following steps.
[0052] Step S115: Receive a model instruction submitted by the user; wherein the model instruction is a prompt instruction for instructing the large model to generate a search query; respond to the model instruction and generate a search query based on the information to be searched; send the search query to the client to instruct the client to display the search query through the visualization interface.
[0053] In some implementations, step S115 is performed between steps S110 and S120 to generate and display the search query for user confirmation. Alternatively, step S115 can be an optional iterative step triggered after step S130 to regenerate the search query based on user feedback and update the analysis report, etc.
[0054] In this embodiment, based on the needs of protecting technical solutions and avoiding infringement, the information retrieval device can provide an information input interface as shown in Figure 4. The information input interface includes an input area for information to be retrieved. The information retrieval device can obtain the information to be retrieved input by the user from the input area. The information to be retrieved is used to represent information for which prior art needs to be retrieved. Optionally, the information retrieval device can also obtain the information to be retrieved in other ways, such as receiving the information to be retrieved transmitted by the user via Bluetooth, or downloading files from a cloud server as the information to be retrieved based on user instructions. The information to be retrieved can represent a technical solution and can include, but is not limited to, text, images, videos, files, links, etc. The input method of the information to be retrieved includes, but is not limited to, typing, importing, pasting, etc. For example, the information to be retrieved may include: "A smart blackboard, including a blackboard frame, with three anti-glare black transparent glass panels (left, middle, and right) arranged inside the blackboard frame. The three anti-glare black transparent glass panels form a blackboard plane. A nano-touch film is attached to the back of the middle anti-glare black transparent glass panel, and a display screen is installed on the back of the nano-touch film."
[0055] In some implementations, the input area for the information to be retrieved shown in Figure 4 also provides a "search criteria limitation control," a "model selection control," and a "start" control. Furthermore, the input area for the information to be retrieved may also support at least one of the following functions: editing functions for the information to be retrieved and highlighting functions for sentence errors in the information to be retrieved. Sentence errors refer to any of the following: grammatical errors, grammatical mistakes, or word errors.
[0056] The search criteria control allows users to specify arbitrary conditions such as publication date, application date, application type, and field of application. When the search criteria control is triggered, a trigger result is determined, which serves as a filtering condition for the retrieved prior art documents, ensuring that the retrieved prior art documents meet the trigger result. For example, if the trigger result includes "application date earlier than 2025-1-1," then the retrieved prior art documents are filtered accordingly, ensuring that all prior art documents displayed to the user meet the condition of "application date earlier than 2025-1-1." The model selection control provides multiple callable model identifiers for the user to choose from. When the model selection control is triggered, the corresponding model is called based on the model identifier corresponding to the model selection control to perform the search query generation operation. When the start control is triggered, the search query generation operation is performed based on the trigger status of each control in the input area of the information to be searched and the information to be searched.
[0057] Furthermore, the information input interface shown in Figure 4 may also include "Content Example 1", "Content Example 2", ..., where the content examples are samples of the information to be retrieved for the user's reference. In addition, the information input interface shown in Figure 4 may also integrate other functions to facilitate the user in providing the information to be retrieved or to facilitate the user in defining the search requirements, such as a grammar check function for the information to be retrieved, etc. This embodiment of the application does not limit this aspect.
[0058] In this embodiment, after obtaining the information to be retrieved, a pre-trained retrieval model can be invoked to generate a large model based on the information to be retrieved and the prompts used to instruct the large model to generate the retrieval formula. This triggers the large model to generate and display the retrieval formula. The retrieval formula is a logical statement that associates logical features through logical relationship identifiers (e.g., AND, OR, etc.), and includes at least one of the following logical features: features in the information to be retrieved, and other features related to the content to be retrieved (e.g., synonymous features, near-synonymous features, etc.). In some embodiments, the type of retrieval formula generated based on the information to be retrieved can be one or more, such as semantic retrieval formulas, classification number retrieval formulas, keyword retrieval formulas, etc., and the number of retrieval formulas corresponding to each type can be one or more.
[0059] In this embodiment, with a defined search query, prior art related to the content to be searched can be retrieved from one or more prior art databases based on the search query. Based on the prior art, the information to be searched, and a prompt instructing a large model to generate an analysis report, the large model for generating the analysis report is invoked. This triggers the large model to perform feature comparison between the prior art and the information to be searched, and then generate a corresponding analysis report. Specifically, the feature comparison method involves: refining the information to be searched so that the amount of data in the refined result is less than the amount of data in the information to be searched; performing feature identification on the refined result to obtain the technical features to be compared; or directly performing feature identification on the information to be searched to obtain the technical features to be compared; comparing these technical features with the corresponding features of existing technical documents to obtain comparison results that reflect multiple sets of technical features (one set of technical features includes a technical feature from the refined result and a technical feature from an existing technical document); and combining the comparison results to obtain a comprehensive comparison result. The comprehensive comparison result can represent the overall relevance between the information to be searched and the existing technical documents. The analysis report can at least be used to characterize the feature differences between the information to be searched and the existing technical documents. In some implementations, the analysis report may also include at least one of the following: the results of extracting the inventive concept of the information to be retrieved, the content of the prior art, the evaluation of innovativeness obtained based on the feature differences analysis between the information to be retrieved and the prior art, and the modification suggestions for the information to be retrieved based on the evaluation.
[0060] In this embodiment, user-inputted information representing a technical solution can be obtained, and a search query generated based on the search information can be displayed, thereby providing an analysis report. The analysis report is generated based on the search information and prior art documents recalled based on the search query. Implementing this embodiment enables automated generation of search queries and analysis reports based on those queries, reducing manual construction of search queries and writing of analysis reports, thus improving information retrieval efficiency. Furthermore, since it reduces the need for personnel with specialized knowledge to complete the retrieval process, it also lowers the threshold and difficulty of information retrieval.
[0061] In some implementations, the information retrieval device can provide an analysis report that meets the expected analysis objectives; wherein the expected analysis objectives include at least one of the following: FTO (Freedom to Operate) analysis, invalidity analysis, novelty search analysis, and infringement analysis.
[0062] In this embodiment, to enhance the analytical directions offered by the analysis report and further increase interaction frequency, the information retrieval device provides FTO analysis, invalidity analysis, novelty search analysis, and infringement analysis. The information retrieval device can dynamically configure the analysis report generation strategy based on the user-selected expected analysis target. After retrieving relevant prior art documents based on retrieval, it can parse the user's input instructions (e.g., FTO analysis, invalidity analysis, etc.) through the analysis target selection control and trigger the corresponding analysis engine call process. The expected analysis target can be used to characterize the analytical directions the user expects the analysis report to provide.
[0063] When the intended analysis objective is FTO (Freedom to Take) analysis, the information retrieval device can obtain parameters such as the patent maintenance status and territorial scope of prior art based on a legal status database interface. Based on these parameters, it can clarify the correlation between prior art and technical solutions, and evaluate whether the implementation of the technical solution will be limited by prior art. When the intended analysis objective is invalidation analysis, the information retrieval device can use feature matrix matching algorithms or other related algorithms to screen one or more prior art documents with novelty or inventive step potential from the recalled prior art documents. Then, based on an argument chain construction model, it can logically map the technical features in the prior art documents to the technical features in the information to be retrieved, and generate invalidation grounds accordingly. When the intended analysis objective is novelty search analysis, it can perform feature comparison between the information to be retrieved and prior art documents, thereby clarifying the distinguishing technical features between the information to be retrieved and the prior art documents. Based on relevant laws and regulations, it can provide suggestive evaluations of the inventiveness of the distinguishing technical features and other patent layout suggestions. When the intended analysis objective is infringement analysis, the information retrieval device can use a combination of semantic parsing and neural network models to determine infringement between the information to be retrieved and prior art documents, generating an analysis report that includes at least fields such as literal infringement and equivalent infringement.
[0064] In some implementations, the information retrieval device can also dynamically update the corresponding analysis report when it detects a user's editing operation on the information to be retrieved. The user can edit the information to be retrieved at any stage of the information retrieval process; this application embodiment does not limit this.
[0065] In some implementations, the information retrieval device can display technical features extracted from the information to be retrieved; extract search elements from the technical features confirmed by the user and display the search elements; wherein the search elements and technical features are used to represent the information to be retrieved at different granularities; and generate a search expression based on the search elements confirmed by the user. The step of generating a search expression based on the information to be retrieved in response to the model instruction includes: extracting technical features from the information to be retrieved in response to the model instruction; sending the technical features to a client to instruct the client to display the technical features through a visual interface; wherein the technical features are determined based on a pre-trained semantic parsing model and a named entity recognition algorithm; receiving the user-confirmed technical features from the client to extract search elements from the user-confirmed technical features; wherein the search elements and technical features are used to represent the information to be retrieved at different granularities; sending the search elements to the client to instruct the client to display and provide feedback on the user-confirmed search elements; and generating a search expression based on the user-confirmed search elements.
[0066] In this embodiment, to enhance interactivity and improve the accuracy of query generation, the information retrieval device can determine technical features based on a pre-trained semantic parsing model and named entity recognition algorithm, and display these features through a visual interface for user confirmation. The visual interface can display technical features in the form of options, lists, etc., and provides editing functions, allowing users to edit one or more technical features. Optionally, when multiple technical features need to be edited, they can be selected in advance before triggering the editing function, enabling unified editing of multiple technical features and reducing the need to trigger the editing function separately for each feature, thus saving user operations.
[0067] Furthermore, the technical features confirmed by the user are used to extract search elements. After the search elements are confirmed by the user, a search query containing one or more search elements can be constructed. Since the search elements used to generate the search query are confirmed by the user, it can help improve the relevance between the existing technologies retrieved by the search query and the information to be searched.
[0068] In some implementations, the information retrieval device can also directly extract search elements from the information to be retrieved and display them for user confirmation, thereby forming a search query based on the user-confirmed search elements. Compared to extracting technical features first, then obtaining user confirmation, and finally extracting search elements, directly extracting search elements from the information to be retrieved for user confirmation improves retrieval efficiency and saves user operations. Granularity refers to the degree of refinement of the information to be retrieved; the higher the granularity, the higher the degree of refinement. Since technical features and search elements represent the information to be retrieved at different granularities, search elements are finer than technical features. Part of the characters in a technical feature (e.g., "three anti-glare black transparent glass pieces form a blackboard plane") (e.g., "anti-glare") constitute search elements. Based on finer-grained search elements, more existing technical documents related to the information to be retrieved can be recalled. Therefore, extracting technical features first, then obtaining user confirmation, and finally extracting search elements improves the accuracy of search element extraction and thus improves the recall quality of the search query.
[0069] In one scenario, technical features can be displayed in the same style. In another scenario, there is a priority relationship between technical features to characterize their importance, and they can be displayed in different styles to differentiate their importance. In some implementations, the information retrieval device can provide a technical feature display interface as shown in Figure 5. After retrieving the technical features, it displays each technical feature (e.g., "three anti-glare black transparent glass panels are provided on the left, middle, and right sides within the blackboard frame", "three anti-glare black transparent glass panels form a blackboard plane", "a nano-touch film (300) is attached to the back of the middle anti-glare black transparent glass", "a display screen (400) is installed on the back of the nano-touch film (300)") and corresponding confirmation and denial controls in the interface shown in Figure 5. As an example, in Figure 5, the confirmation control is represented as "is a technical feature", and the denial control is represented as "is not a technical feature". When the confirmation control is detected to be triggered, it can be determined that the technical feature corresponding to the confirmation control is a technical feature confirmed by the user.
[0070] In addition, optionally, before the user triggers the confirmation or denial control, the technical feature can be pre-determined. The result of the pre-determination is used to indicate whether the technical feature is a genuine technical feature or not. Based on the result of the pre-determination, the corresponding control can be highlighted when displaying Figure 5 for the user's reference. The user can adjust the result of the pre-determination. For example, the technical feature highlighted corresponding to the denial control can be modified to the technical feature highlighted corresponding to the confirmation control.
[0071] Furthermore, the information retrieval device can provide a search element display interface as shown in Figure 6. The search elements obtained after extracting search elements from the technical features confirmed by the user (such as "anti-glare black transparent glass," "nano-touch film," and "display screen") can be displayed in Figure 6. In addition, confirmation and denial controls corresponding to each search element can be displayed. As an example, in Figure 6, the confirmation control is represented as "is a search element," and the denial control is represented as "is not a search element." When a confirmation control is triggered, it can be determined that the search element corresponding to the confirmation control is a search element confirmed by the user.
[0072] Furthermore, the information retrieval device can generate a search query based on the search elements confirmed by the user. Specifically, the search query can be obtained by connecting the search elements with logical identifiers based on the relationship between the search elements (such as parallel relationships).
[0073] In some implementations, generating a search query based on user-confirmed search elements may also include: obtaining information such as classification numbers, material keywords, functional descriptors, process parameters, equipment requirements, and quality control indicators corresponding to the search elements as supplementary elements; and constructing the search query based on the relationships between the search elements and supplementary elements (e.g., superior relationships, subordinate relationships, synonym relationships, etc.) and the relationships between the search elements themselves. The information retrieval device provides a search query editing function for users to modify or supplement the search query.
[0074] In some implementations, the method of generating a search query based on user-confirmed search elements may further include: extracting key information corresponding to key fields from the information to be searched, and generating a search query based on the key information and user-confirmed search elements. The key fields include at least one of the following: technical field, technical topic, technical problem, technical means, technical efficacy, technical application, technical problem phrase, technical efficacy phrase, technical application phrase, and technical summary.
[0075] For example, the extracted key information can be represented as follows:
[0076] In some implementations, the information retrieval device can also stream the retrieval process to recall prior art documents related to the information to be retrieved; wherein the retrieval process includes at least one of the following: semantic retrieval process, classification number retrieval process, keyword retrieval process, and family citation retrieval process. The information retrieval device can also send the prior art documents to the client to instruct the client to visualize the real-time status of the prior art document retrieval process through a dynamic progress bar and result preview component.
[0077] In this embodiment, the retrieval process involves a large-scale logical reasoning process. To improve the problem of long user waiting time, the information retrieval device can also display the retrieval process in a streaming manner, that is, display each retrieval process sequentially. The display order of the retrieval processes and the logical calling relationship between each retrieval process can be adjusted according to actual needs, and this application embodiment does not limit this. Specifically, the information retrieval device can use a streaming data processing framework to visualize the real-time status of the semantic retrieval process, classification number retrieval process, keyword retrieval process, and family citation retrieval process through dynamic progress bars and result preview components.
[0078] As an optional embodiment, the semantic retrieval process includes: the information retrieval device performing vector space mapping on the existing technical document database, combining the cosine similarity algorithm to calculate the semantic relevance between the technical solution and the existing technical documents in real time, and dynamically adding documents with a matching degree exceeding a preset threshold (e.g., 0.75) as existing technical documents related to the information to be retrieved.
[0079] As an optional embodiment, the classification number retrieval process includes: the information retrieval device maps the information to be retrieved to the corresponding classification number range (e.g., G06F3 / 041) based on a pre-trained classification prediction model, and recalls prior art documents related to the information to be retrieved through the classification number index interface of the patent database.
[0080] As an optional embodiment, the keyword retrieval process includes: the information retrieval device performing weighted matching of keywords in the search query based on an inverted index algorithm combined with a position weight algorithm, in order to recall existing technical documents related to the information to be retrieved.
[0081] As an optional embodiment, the family citation retrieval process includes: the information retrieval device calling the patent family database interface to obtain family prior art documents and / or cited prior art documents, and recalling prior art documents related to the information to be retrieved. Optionally, the information retrieval device may also display a citation / family relationship graph when the family citation graph display control is triggered, to reflect the citation relationships between prior art documents, the influence of prior art documents, etc.
[0082] Furthermore, the semantic retrieval process, classification number retrieval process, keyword retrieval process, and family citation retrieval process each include a semantic search query, a classification number retrieval query, a keyword retrieval query, and a family citation retrieval query, respectively. In one optional implementation, these search queries are displayed in the corresponding retrieval process, and an editing function is provided for editing any of the search queries. When the editing function is triggered, the retrieved prior art documents are updated in real time based on the edited search query. When there is a calling relationship between any two retrieval processes, the prior art documents involved in other related retrieval processes can also be updated in real time based on the edited search query.
[0083] In some embodiments, the search query includes a semantic search query, and the semantic search process includes at least a first set of prior art documents; the information retrieval device may also construct a semantic search query based on the semantic features of the information to be retrieved; and recall the first set of prior art documents based on the semantic search query.
[0084] In this embodiment, to enrich the retrieval dimensions and improve the accuracy of the retrieval results, the information retrieval device can semantically encode the information to be retrieved based on a pre-trained BERT-wwm model or other related models, obtaining feature vectors conforming to a specified dimension (e.g., 768 dimensions) as semantic features. After representing these semantic features as semantic retrieval expressions, it can recall a first set of prior art documents with a semantic similarity higher than a preset similarity (e.g., 0.82) that are related to these semantic features. The BERT-wwm model is a variant of a pre-trained language model that employs a "whole word masking" strategy to mask complete terms during the pre-training stage to improve the ability to model the overall semantics of Chinese (or other languages) words, thus typically performing better in downstream natural language processing tasks. The semantic retrieval expression (e.g., "SEMANTIC:2cbe5207ae1638bba06030c50f9dd02a") is used to represent semantic features in a specified form (e.g., text form, string form, etc.), and the semantic similarity between each first prior art document in the first set of prior art documents and the information to be retrieved is higher than the preset similarity. Semantic similarity can be represented by methods such as cosine distance and Euclidean distance.
[0085] Furthermore, it is understood that, optionally, the identification information of each first prior art document in the first prior art document set (such as title, publication number, application number, IPC (International Patent Classification) classification number, CPC (Cooperative Patent Classification) classification number, etc.) can be displayed in a streaming manner, and the similarity (e.g., 97.50%) and feature matching status (e.g., feature matching: 2 / 2) between each first prior art document and the information to be searched can also be displayed. Additionally, optionally, the identification information of the first prior art documents can be displayed in a triggerable form such as a link. When the identification information is triggered, the user is redirected to a document details display interface, which displays all the data in the first prior art documents.
[0086] In some embodiments, the search query includes a classification number search query, and the classification number search process includes at least a second set of prior art documents; the information retrieval device may also construct the classification number search query based on the classification number statistics of the first set of prior art documents; and recall the second set of prior art documents based on the classification number search query.
[0087] In this embodiment, to enrich the search dimensions and improve the accuracy of the search results, the information retrieval device can also statistically analyze the classification numbers of the first set of prior art documents and construct a classification number search expression accordingly to recall the second prior art documents. The classification numbers include, but are not limited to, IPC classification numbers and CPC classification numbers. Specifically, after obtaining the classification numbers corresponding to each first prior art document in the first set of prior art documents, the first prior art documents can be classified based on the classification numbers to obtain file clusters corresponding to different classification numbers. The N file clusters with the largest data volume can then be used as the basis for constructing the search expression, and a classification number search expression composed of classification numbers can be constructed accordingly. The classification numbers in this search expression are taken from the classification numbers corresponding to the N file clusters, where N is a positive integer. It should be noted that a first prior art document can correspond to multiple IPC classification numbers and / or multiple CPC classification numbers, a file cluster can correspond to a specific IPC classification number or CPC classification number, and a first prior art document can exist in one or more file clusters. For example, the classification number search expression can be represented as IPC_FACET:(G06F3 / 041OR G09B5 / 02OR bB43L1). The prior art documents retrieved based on the above classification number search are considered as second prior art documents, thus forming a second prior art document set.
[0088] In some implementations, the information retrieval device may also determine a classification number corresponding to the retrieval element in the information to be retrieved, and then construct a classification number retrieval formula based on the classification number.
[0089] In some embodiments, the search query includes a keyword search query, and the keyword search process includes at least a third set of prior art documents. The information retrieval device can also generate the keyword search query based on the search elements in the information to be retrieved, and recall a set of candidate prior art documents according to the keyword search query. Based on the feature comparison results of each document in the candidate prior art document set with the information to be retrieved, keywords corresponding to the search elements are determined from the candidate prior art document set. The keywords include at least one of the synonyms, related words, hypernyms, and hyponyms of the search elements. The keyword search query is iterated based on the keywords to update the set of candidate prior art documents until the number of iterations meets the expected number, and the final set of candidate prior art documents obtained from the iteration is determined as the third set of prior art documents.
[0090] In this embodiment, to enrich the retrieval dimensions and improve the accuracy of the retrieval results, the information retrieval device can also construct keyword search expressions related to the retrieval elements, thereby recalling existing technical documents related to the keywords as third existing technical documents to enrich the retrieval results and improve their accuracy. Specifically, the information retrieval device can combine retrieval elements into keyword search expressions based on the search expression construction syntax. Since the retrieval elements are extracted from the information to be retrieved, there may be retrieval limitations. To reduce missed detections and false detections, optionally, three elements of the information to be retrieved can be determined based on a semantic recognition model trained on the target language. These three elements include the technical problem, technical efficacy, and technical means. Then, these three elements and the retrieval elements are combined to generate a keyword search expression. For example, TACD_ALL:(("Anti-glare" AND "Light" AND "Black" AND "Transparent" AND "Glass") OR ("Nano" AND "Touch Film") OR ("Display Screen")).
[0091] Prior art documents retrieved based on the aforementioned keyword retrieval are used as candidate prior art documents. Furthermore, to improve the comprehensiveness of the retrieved prior art documents, the information retrieval device can perform feature comparison between the candidate prior art documents and the information to be retrieved / user-confirmed technical features / user-confirmed search elements, in order to extract one or more features related to the information to be retrieved from the candidate prior art documents, which serve as the basis for generating the feature comparison results.
[0092] In some implementations, the feature comparison results between existing technical documents and the information to be retrieved can be used as input to a large model to trigger the model to generate importance scores for each feature in the feature comparison results. Based on this, the semantic relevance between the information to be retrieved and existing technical documents, as well as the aforementioned importance, can be combined to generate the public disclosure of existing technical documents relative to the information to be retrieved. Based on this, the recalled existing technical documents can be filtered to obtain the N existing technical documents with the highest public disclosure and display them in the user interface. This saves display resources, where N is a positive integer.
[0093] After obtaining the feature comparison results corresponding to each candidate prior art document, keywords corresponding to the search element can be determined from the set of candidate prior art documents. These keywords are taken from the features of prior art documents in the feature comparison results. Keywords refer to features related to the search element. Keywords can be synonyms, related words, hypernyms, hyponyms, or diverse expressions of the same concept. This application embodiment does not limit this.
[0094] Updating keywords to the keyword search expression can improve the recall scope of the keyword search expression. For example, the keyword search expression after the keyword update can be: TACD_ALL:((("anti-glare" OR "anti-reflective" OR "anti-glare") AND ("light" OR "light" OR "light source") AND ("black" OR "dark" OR "dark color") AND ("perspective" OR "transparent" OR "transparent") AND ("glass" OR "glass plate" OR "glass material")) OR (("nano" OR "micron" OR "ultra-micro") AND ("touch film" OR "touch film" OR "touch layer")) OR (("display screen" OR "monitor" OR "LCD screen" OR "display panel"))).
[0095] In some implementations, to mitigate the problem of excessive irrelevant prior art documents being recalled after the updated keyword search query, recall criteria can be limited when performing recall based on the updated keyword search query, thereby reducing the number of documents recalled. Recall criteria can be used to limit the conditions that the recalled prior art documents must meet, such as technical problems, technical effects, and technical fields.
[0096] By iteratively executing the aforementioned steps based on the new keyword search formula, multiple iterations of the keyword search formula can be achieved until the number of iterations meets the expected number (e.g., 4). The candidate prior art documents recalled based on the keyword search formula obtained in the last iteration are then used as the third prior art document. The number of iterations can be set according to actual needs, and this embodiment does not limit this. Regarding the expected number of iterations, it is assumed that there are iteration number N1 and iteration number N2 (iteration number N2 = iteration number N1 + 1). If the recall result of iteration number N2 does not contain a new target prior art document compared to the recall result of iteration number N1, then iteration number N1 can be considered the expected number of iterations. Here, N1 and N2 are both positive integers, and the relevance between the target prior art document and the information to be retrieved is higher than the relevance between any prior art document in the recall result of iteration number N1 and the information to be retrieved.
[0097] As an example, taking iteration number = 3, the keyword search expression "TACD_ALL:(("anti-glare" AND "light" AND "black" AND "transparent" AND "glass") OR ("nano" AND "touch film") OR ("display"))" generated based on the search elements can recall A1 candidate prior art documents. Based on the feature comparison results of A1 candidate prior art documents, the keyword search expression can be iterated as "ACD_ALL:((("anti-glare" OR "anti-reflective" OR "anti-glare") AND ("light" OR "light" OR "light source") AND ("black" OR "dark" OR "dark color") AND ("perspective" OR "transparent" OR "transparent") AND ("glass" OR "glass plate" OR "glass material")) OR (("nano" OR "micrometer" OR "ultra-micro") AND ("touch film" OR "touch film" OR "touch layer")) OR (("display screen" OR "monitor" OR "LCD screen" OR "display panel")))", thereby recalling A2 candidate prior art documents. Based on the feature comparison results of A2 candidate prior art documents, the keyword search expression can be iterated as "PROBLEM_SUM:(("anti-glare" AND "light" AND "black" AND "transparent" AND "glass") OR ("nano" AND "touch film") OR ("display screen")) OR BENEFIT_SUM:(("anti-glare" AND "light" AND "black" AND "transparent" AND "glass") OR ("nano" AND "touch film") OR ("display screen"))". On this basis, recall conditions including technical problems and technical effects are specified, so that the recalled A3 candidate prior art documents meet the specified technical problems and technical effects. Here, ACD_ALL is a search field identifier, typically indicating that the abstract, claims, and description are merged into a single "all" search field for full-text or multi-domain matching; PROBLEM_SUM is a search field identifier, typically representing a summary / abstract field of technical problems or problem statements in the document, used for searching for phrases related to technical problems. Based on the feature comparison results of A3 candidate prior art documents, the keyword search formula can be updated in the same way as the previous iteration to obtain a new keyword search formula, and then A4 candidate prior art documents can be recalled as the third prior art document.
[0098] In some implementations, the information retrieval device may further determine the frequency of occurrence of keywords corresponding to each retrieval element in the candidate prior art document set based on the feature comparison results between each document in the candidate prior art document set and the information to be retrieved; determine the weight corresponding to each retrieval element based on the frequency of occurrence; and update the weight to the keyword retrieval formula.
[0099] In this embodiment, to improve the accuracy of the recalled prior art documents, the information retrieval device can also determine the frequency of occurrence of keywords corresponding to each retrieval element in the candidate prior art document set based on the feature comparison results of each document in the candidate prior art document set with the information to be retrieved. Based on the frequency of occurrence of each retrieval element, a pre-trained large model can be called to generate a score corresponding to that retrieval element as a weight. The relationship between weight and frequency of occurrence is positively correlated, which can be linear or non-linear; this embodiment does not limit this. The weight represents the importance of the corresponding retrieval element. Updating the weights corresponding to each retrieval element to the keyword search expression can recall prior art documents that conform to the weight distribution. Prior art documents recalled in this way have a higher relevance to the information to be retrieved.
[0100] In some implementations, updating the weights corresponding to each search element to the keyword search expression can be done by updating the weights corresponding to the search elements to the positions of the corresponding search elements and related keywords in the keyword search expression, so as to serve as scaling factors and thus characterize the importance of the corresponding search elements and related keywords.
[0101] In some embodiments, the search query includes a family citation search query; the family citation search process includes at least a fourth set of prior art documents; the information retrieval device can also generate the family citation search query based on the associated files of prior art documents retrieved in any of the semantic search process, classification number search process, or keyword search process; wherein, the associated files are at least one of the family files and citation files of the retrieved prior art documents; and the fourth set of prior art documents is recalled according to the family citation search query.
[0102] In this embodiment, in order to enrich the search dimensions and improve the accuracy of the search results, the information retrieval device can also perform a family citation query on the prior art documents recalled in any of the semantic search process, classification number search process, and keyword search process to clarify the family documents and / or citation documents corresponding to these prior art documents. Then, the family documents and / or citation documents can be generated into a family citation search expression, and the prior art documents recalled in this way can be used as the fourth prior art document.
[0103] For example, a citation search query can be expressed as "CITEORCITEDBY:(CN11836****A OR CN10840****B OR WO202206****A1 OR CN21431****U OR WO202117****A1 OR CN21395****U OR CN21369****U OR CN11308****A OR CN21344****U OR CN11276****A OR CN21314****U OR CN21296****U OR CN21236****U OR CN11212****A OR CN10851****B OR CN21169****U)". Here, CITEORCITEDBY is a search field identifier, typically used to retrieve documents cited by other documents or documents that cite other documents (i.e., a "cited / cited" search).
[0104] In some implementations, the analysis report includes a technical feature comparison table; the information retrieval device can also extract corresponding technical features from the prior art documents and the information to be retrieved to form technical feature pairs; and generate the technical feature comparison table based on the technical feature pairs.
[0105] In this embodiment, to improve the readability of the analysis report, the information retrieval device can also provide a technical feature comparison table in the analysis report. This technical feature comparison table is used to show the feature comparison details between the information to be retrieved and various existing technical documents. The technical feature comparison table includes corresponding features, which are respectively derived from the information to be retrieved and existing technical documents. The technical feature comparison table also includes relevant information of the existing technical documents (e.g., applicant, publication date, title, patent number, etc.). Furthermore, it should be noted that existing technical documents can be patents, journals, papers, etc.
[0106] For example, a technical feature comparison table can be represented as follows:
[0107] In some implementations, the information retrieval device may, based on the relevance between the prior art documents and the information to be retrieved, select prior art documents that meet the expected relevance from the prior art documents; and extract corresponding technical features from the prior art documents that meet the expected relevance and the information to be retrieved to form a technical feature pair.
[0108] In this embodiment, to reduce the computational resources consumed in the feature comparison process, the information retrieval device can screen existing technical documents that meet the expected relevance (e.g., 80%) and extract corresponding technical features from the existing technical documents and the information to be retrieved as technical feature pairs (e.g., [“A nano-touch film (300) is attached to the back of the anti-glare black transparent glass in the middle”, “A nano-touch film 300 is attached to the back of the anti-glare black transparent glass 200 in the middle”]). Multiple technical feature pairs are used as the basis for generating the technical feature comparison table.
[0109] In some embodiments, the information retrieval device may also provide file addition and file deletion functions; wherein the file addition function is used to supplement existing technology documents that need to participate in feature comparison, and the file deletion function is used to delete existing technology documents that need to participate in feature comparison. In other words, the information retrieval device may also receive file operation instructions sent by a client for a specified file; wherein the file operation instructions include file addition instructions and file deletion instructions; the specified file includes at least one file among existing technology documents retrieved based on the retrieval query and other files imported by the user; in response to the file addition instruction, the specified file is added to the existing technology documents used to form the technical feature pair; or, in response to the file deletion instruction, the specified file is deleted from the existing technology documents used to form the technical feature pair.
[0110] In this embodiment, to enhance interactivity, the information retrieval device provides file addition and file deletion functions. When the file addition function is triggered, a specified file can be identified as a prior art file requiring feature comparison. The specified file can be a prior art file based on retrieval or other files imported by the user. When the file deletion function is triggered, the specified file can be deleted from the prior art files, reducing the need for feature comparison on the specified file and saving computational resources.
[0111] One embodiment of this application also provides an information retrieval method. After receiving user-inputted information representing a technical solution through a client, the method includes: when the data volume of the information to be retrieved is greater than or equal to a preset data volume, extracting technical features from the extracted results after semantic extraction of the information to be retrieved; when the data volume of the information to be retrieved is less than the preset data volume, extracting technical features from the information to be retrieved; extracting retrieval elements from the technical features and generating a retrieval formula based on the retrieval elements; wherein the retrieval elements and technical features are used to represent the information to be retrieved at different granularities; comparing the features of existing technical documents recalled based on the retrieval formula with the information to be retrieved; based on the feature comparison results, selecting multiple existing technical documents with the highest relevance to the information to be retrieved as target documents from the existing technical documents; and providing an analysis report; wherein the analysis report is generated based on the target documents and the information to be retrieved.
[0112] In some implementations, by performing semantic extraction on the information to be retrieved and submitting only the extracted results or user-confirmed search elements to the large model used to generate search queries or analysis reports, the length of text submitted to the large model and the frequency of calls can be reduced, thereby saving token consumption and lowering computational / usage costs. Furthermore, the confirmation and editing controls provided by the client allow users (such as professionals) to freely modify the analysis report or search elements based on their judgment. This manual correction can, to some extent, replace or reduce repetitive generation steps in the large model, improving the accuracy and resource utilization efficiency of the overall retrieval and analysis process. The aforementioned token can refer to the basic text unit (such as a word, subword, or character fragment) used in language processing and retrieval models. The aforementioned token can serve as the smallest semantic or statistical unit for model input and quantitative processing.
[0113] For example, when the amount of data to be retrieved is large, semantic extraction (e.g., extractive summarization or key sentence extraction algorithms) can be performed locally or on the server side to generate a concise summary or a structured set of search elements. The number of characters or tokens in the summary / search elements is significantly less than the original text. When subsequently used to call a large model to generate search queries or analysis reports, only the summary / search elements are used as input, thereby reducing the number of tokens required for each call. On the other hand, the technical features, search elements, or analysis conclusions initially generated by the model can be displayed to users (e.g., search / patent experts) on a visual interface. Users can directly edit, confirm, or annotate these entries on the client side. Uploading only the minimum information containing "Change Summary" or "User Confirmation Identifier" can trigger necessary recalculation or model calls, helping to reduce repeated full-text uploads or repeated full-model inference, thus further saving token consumption and response time. Based on the above mechanism, users (such as professionals) can freely correct the results based on their judgment (e.g., modify feature extraction results, add or delete candidate documents, or adjust search elements). The modifications made will be used to locally update the analysis results or as input for subsequent models, thereby reducing the consumption of computing resources while maintaining the quality of the results.
[0114] The specific functions and effects of the information retrieval method implemented in this embodiment can be explained by referring to other embodiments of this application, and will not be repeated here.
[0115] Please refer to Figure 7. This application also provides an information retrieval device. The information retrieval device may include: an information acquisition module 101, used to acquire user-inputted information representing a technical solution through a client's visual interface; a retrieval expression display module 102, used to receive a model instruction submitted by the user; wherein the model instruction is a prompt instruction for instructing a large model to generate a retrieval expression; responding to the model instruction, generating a retrieval expression based on the information to be retrieved; sending the retrieval expression to the client to instruct the client to display the retrieval expression through the visual interface; and an analysis report display module 103, used to generate an analysis report that meets the expected analysis objectives; wherein the analysis report is generated based on the information to be retrieved and existing technical documents retrieved based on the retrieval expression, and the analysis report is used to represent the feature differences between the information to be retrieved and existing technical documents, the retrieval expression is generated based on the information to be retrieved, and the expected analysis objectives are used to represent the analysis direction provided by the user's expected analysis report; sending the analysis report to the client to instruct the client to present and modify the feature differences in the analysis report through the visual interface.
[0116] In this embodiment, the specific functions and effects of the information retrieval device can be explained by referring to other embodiments of this application, and will not be repeated here.
[0117] Please refer to Figure 8. One embodiment of this application also provides an information retrieval device, comprising: an information acquisition module 101, configured to acquire user-inputted information representing a technical solution through a client's visual interface; an analysis report display module 103, configured to generate an analysis report that meets the expected analysis objectives; wherein the analysis report is generated based on the information to be retrieved and prior art documents retrieved based on a retrieval query, and the analysis report is used to represent the feature differences between the information to be retrieved and the prior art documents, the retrieval query is generated based on the information to be retrieved, and the expected analysis objectives are used to represent the analysis direction expected by the user from the analysis report; the analysis report is sent to the client to instruct the client to present and modify the feature differences in the analysis report through the visual interface.
[0118] In this embodiment, the specific functions and effects of the information retrieval device can be explained by referring to other embodiments of this application, and will not be repeated here.
[0119] In some implementations, the aforementioned information acquisition module, retrieval display module, and analysis report display module can be implemented as software modules running on a general-purpose computing device. This computing device includes a processor (such as a CPU (Central Processing Unit) or other general-purpose processor), memory (volatile memory for storing program instructions and temporary data, and non-volatile memory for long-term storage), an input interface (e.g., keyboard, mouse, touchscreen, file import interface, or Bluetooth), a display module (for user interface rendering), and a network access module / transceiver (Ethernet port or other communication link for communicating with a remote server or database). The information acquisition module, retrieval display module, and analysis report display module operate by the processor loading and executing program instructions stored in the memory. For example, the information acquisition module is responsible for receiving and preprocessing data from the input interface or network; the retrieval display module is responsible for rendering program-generated retrieval and interactive controls on the display module; and the analysis report display module is responsible for storing the analysis report generated by the server or local processor in non-volatile memory and displaying it to the user on the display module.
[0120] In some implementations, the information acquisition module, retrieval display module, and analysis report display module described above can be partially or entirely implemented by dedicated hardware, such as Application Specific Integrated Circuit (ASIC), Field Programmable Gate Array (FPGA), or other programmable logic circuits, or implemented in a combination of hardware and software. The physical implementation also includes a bus or network interface for inter-module communication, a network transceiver for interfacing with external patent / document databases, and local / remote storage units for storing copies of retrieved documents and feature vectors. The software and hardware implementations described above can be arbitrarily combined or deployed in a distributed manner (e.g., the client is responsible for information acquisition and display, while the server is responsible for retrieval-based document retrieval and analysis report generation), and each of the aforementioned information acquisition module, retrieval display module, and analysis report display module can be implemented by any one or more of the aforementioned physical components.
[0121] In some implementations, the aforementioned information acquisition module, retrieval display module, and analysis report display module refer to logical entities implemented by specific physical structures or combinations thereof to achieve their respective functions. These physical structures include: one or more processors (e.g., general-purpose processors, microprocessors, digital signal processors, etc.) and their loaded and executed program instructions (stored in volatile or non-volatile memory), memory (e.g., flash memory), input / output interfaces (e.g., keyboard, mouse, touchscreen, file import interface, or Bluetooth, etc.), display modules, network access modules or transceivers (e.g., Ethernet ports), buses or network interfaces for inter-module communication, and application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices. The aforementioned physical structures or any combination thereof can serve as the corresponding physical implementation structures for achieving the corresponding functions of each module; alternatively, the module function can be implemented by storing the corresponding program instructions in a computer-readable storage medium and loading and executing them with a processor.
[0122] For example, the information acquisition module can be implemented by a processor loading and executing program instructions stored in memory, and receiving user input or external information to be retrieved through an input interface or network transceiver; the retrieval display module can be implemented by the aforementioned processor in conjunction with a display module and memory, used to render the program-generated retrieval queries and related interactive controls on the user interface; the analysis report display module can be implemented by a local processor or a processor on a remote server, and the analysis results can be stored in a local or remote non-volatile storage unit (e.g., memory) and displayed to the user through the display module. Furthermore, if implemented with dedicated hardware, the above modules can be partially or entirely implemented by ASICs, FPGAs, or other programmable logic circuits, and the implementation of each module can be centralized (all implemented on a single device / server), distributed (client / server collaborative implementation), or a hybrid deployment of any two (e.g., the client is responsible for information acquisition and display, and the server is responsible for retrieval-based literature retrieval and analysis report generation). The reference numerals shown in the accompanying drawings are merely examples to illustrate the mapping relationship between modules and specific entities; the above-mentioned correspondence is an exemplary disclosure and should not be construed as a limitation on the implementation of the above-mentioned modules. Any scheme based on an equivalent structure or the combination described above is also included within the scope of this specification.
[0123] Please refer to Figure 9. This application also provides a computer device, comprising: a memory and a processor, wherein the memory stores at least one computer program, which is loaded and executed by the processor to implement the information retrieval method as described above.
[0124] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, causes the processor to implement the information retrieval method as described above.
[0125] This application also provides a computer program product containing instructions that, when executed by a processor, implements the information retrieval method as described above.
[0126] The user information or user account information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, etc.) involved in various embodiments of this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws and regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0127] It is understood that the specific examples in this document are only intended to help those skilled in the art better understand the implementation methods of this application, and are not intended to limit the scope of this application.
[0128] It is understood that in the various embodiments of this application, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0129] It is understood that the various implementation methods described in this application can be implemented individually or in combination, and the implementation methods in this application are not limited in this respect.
[0130] Unless otherwise stated, all technical and scientific terms used in the embodiments of this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The term "and / or" as used in this application includes any and all combinations of one or more of the associated listed items. The singular forms "a," "the," and "the" as used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.
[0131] It is understood that the processor in the embodiments of this application can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method embodiments can be completed by the integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly implemented by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.
[0132] It is understood that the memory in the embodiments of this application may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Specifically, non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM). It should be noted that the memory in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0133] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the information to be retrieved. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0134] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the aforementioned method implementations, and will not be repeated here.
[0135] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0136] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.
[0137] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0138] If the aforementioned function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the information to be retrieved in this application, essentially or in terms of its contribution to the prior art, or a portion of the information to be retrieved, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0139] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. An information retrieval method, wherein, The method includes: The client's visual interface allows for the acquisition of user-inputted information to characterize the technical solution. An analysis report that meets the expected analysis objectives is generated; wherein the analysis report is generated based on the information to be retrieved and existing technical documents retrieved based on the retrieval formula, and the analysis report is at least used to characterize the feature differences between the information to be retrieved and the existing technical documents, the retrieval formula is generated based on the information to be retrieved, and the expected analysis objectives are used to characterize the analysis direction that the user expects the analysis report to provide; The analysis report is sent to the client to instruct the client to present and modify the feature differences in the analysis report through the visualization interface.
2. The method according to claim 1, wherein, The method further includes: Receive model instruction submitted by the user; wherein the model instruction is a prompt instruction used to instruct the large model to generate a search query; In response to the model instruction, a search query is generated based on the information to be searched; The search query is sent to the client to instruct the client to display the search query through the visual interface.
3. The method according to claim 2, wherein, The response to the model instruction, based on the information to be retrieved, generates a search query, including: In response to the model instruction, technical features are extracted from the information to be retrieved; The technical features are sent to the client to instruct the client to display the technical features through a visual interface; wherein the technical features are determined based on a pre-trained semantic parsing model and a named entity recognition algorithm; The system receives user-confirmed technical features from the client and extracts search elements from these features; wherein the search elements and technical features are used to characterize the information to be retrieved at different granularities. Send the search elements to the client to instruct the client to display and provide feedback on the search elements confirmed by the user; A search query is generated based on the search elements confirmed by the user.
4. The method according to claim 3, wherein, The step of generating a search query based on the user-confirmed search elements includes: Obtain supplementary elements corresponding to the search elements, wherein the supplementary elements include at least one of the following: classification number, material keywords, functional descriptive terms, process parameters, equipment requirements, and quality control indicators; A search expression is generated based on the first relationship between the search element and the supplementary element, and the second relationship between each of the search elements; wherein the first relationship includes any one of the following: a superior relationship, a subordinate relationship, and a synonym relationship, and the second relationship includes a parallel relationship.
5. The method according to claim 1, wherein, The recall process for the prior art documents includes: The retrieval process is executed based on the retrieval query to recall existing technical documents related to the information to be retrieved; wherein the retrieval process includes at least one of the following: semantic retrieval process, classification number retrieval process, keyword retrieval process, and family citation retrieval process; The existing technical documents are sent to the client to instruct the client to visualize the real-time status of the existing technical document retrieval process through a dynamic progress bar and result preview component.
6. The method according to claim 5, wherein, The search query includes a semantic search query; when the search process includes the semantic search process, the prior art documents include a first set of prior art documents; the process of determining the first set of prior art documents includes: Based on the semantic features of the information to be retrieved, a semantic retrieval formula is constructed; Based on the semantic search query, the first set of prior art documents is recalled.
7. The method according to claim 6, wherein, The search query includes a classification number search query; if the search process also includes the classification number search process, the prior art documents also include a second set of prior art documents; the process for determining the second set of prior art documents includes: Based on the statistical results of the classification numbers of the first set of prior art documents, the classification number retrieval formula is constructed; Based on the classification number search formula, the second set of prior art documents is recalled.
8. The method according to claim 5, wherein, The search query includes a keyword search query; when the search process includes the keyword search process, the prior art documents include a third set of prior art documents; the process of determining the third set of prior art documents includes: The keyword search expression is generated based on the search elements in the information to be searched, and a set of candidate prior art documents is recalled based on the keyword search expression; Based on the feature comparison results of each document in the candidate prior art document set with the information to be retrieved, keywords corresponding to the retrieval element are determined from the candidate prior art document set; wherein, the keywords include at least one of the synonyms, related words, hypernyms, and hyponyms of the retrieval element; The keyword retrieval formula is iterated based on the keywords to update the candidate prior art document set until the number of iterations meets the expected number. The candidate prior art document set obtained from the last iteration is then determined as the third prior art document set.
9. The method according to claim 8, wherein, The method further includes: Based on the feature comparison results between each document in the candidate prior art document set and the information to be retrieved, the frequency of occurrence of the keywords corresponding to each retrieval element in the candidate prior art document set is determined. The weight of each search element is determined based on its frequency of occurrence. Update the weights in the keyword search expression.
10. The method according to claim 5, wherein, The search query includes a family citation search query; when the search process includes the family citation search process, the prior art documents include a fourth set of prior art documents; the process of determining the fourth set of prior art documents includes: Based on the associated files of existing technical documents retrieved in at least one of the semantic retrieval process, the classification number retrieval process, and the keyword retrieval process, the same family citation retrieval formula is generated; wherein, the associated files are at least one of the same family files and citation files of the retrieved existing technical documents; Based on the aforementioned citation search formula, the fourth set of prior art documents is recalled.
11. The method according to claim 1, wherein, The expected analysis objectives include at least one of FTO analysis, invalidity analysis, novelty search analysis, and infringement analysis. Generating an analysis report that meets the expected analysis objectives includes: When the expected analysis objective is FTO analysis, information on the patent maintenance status and territorial scope of the prior art is obtained to determine the correlation between the prior art documents and the technical solution, and an analysis report is generated. If the expected analysis target is invalid, at least one target document with novelty or inventiveness challenge potential is selected from the existing technical documents based on the feature matrix matching algorithm. A model is constructed based on the argument chain to logically map the technical features in the target document with the technical features in the information to be retrieved, and an analysis report is generated. When the expected analysis objective is novelty search analysis, feature comparison is performed on the information to be retrieved and the existing technical documents to determine the distinguishing technical features between the information to be retrieved and the existing technical documents, and the inventiveness of the distinguishing technical features is evaluated to generate an analysis report; When the intended analysis objective is infringement analysis, an infringement determination is made on the information to be retrieved and the existing technical documents based on a combination of semantic parsing and neural network models, and an analysis report is generated.
12. The method according to claim 1, wherein, The analysis report includes a comparison table of technical features between the technical solution and the existing technical documents; The process of generating the technical feature comparison table includes: Extract the corresponding technical features from the existing technical documents and the information to be retrieved to form technical feature pairs; Based on the aforementioned technical feature pairs, a technical feature comparison table is generated.
13. The method according to claim 12, wherein, The step of extracting corresponding technical features from the existing technical documents and the information to be retrieved to form technical feature pairs includes: Based on the relevance between the existing technical documents and the information to be retrieved, existing technical documents that meet the expected relevance are selected from the existing technical documents; Extract corresponding technical features from existing technical documents that meet the expected relevance and the information to be retrieved to form technical feature pairs.
14. The method according to claim 12, wherein, The method further includes: The system receives file operation instructions sent by the client for a specified file; wherein the file operation instructions include file addition instructions and file deletion instructions; the specified file includes at least one file among existing technical documents retrieved based on the search query and other files imported by the user; In response to the file addition instruction, the specified file is added to the prior art documents used to form the pair of technical features; or, In response to the file deletion command, the specified file is deleted from the prior art documents that make up the pair of technical features.
15. The method according to claim 1, wherein, After receiving the retrieval information for characterizing the technical solution from the user input via the client, the method further includes: If the amount of data in the information to be retrieved is greater than or equal to a preset amount of data, the technical features in the extracted results are extracted after semantic extraction of the information to be retrieved; or, if the amount of data in the information to be retrieved is less than the preset amount of data, the technical features in the information to be retrieved are extracted. Extract retrieval elements from technical features and generate a retrieval formula based on the retrieval elements; wherein, the retrieval elements and technical features are used to characterize the information to be retrieved at different granularities; The existing technical documents retrieved based on the search query are compared with the information to be retrieved based on their features; Based on the feature comparison results, multiple existing technology documents with the highest relevance to the information to be retrieved are selected as target documents from the existing technology documents; wherein, the analysis report is generated based on the target documents and the information to be retrieved.
16. An information retrieval device, wherein, The device includes: The information acquisition module is used to acquire the information to be retrieved by the user, which is used to characterize the technical solution, through the client's visual interface; An analysis report display module is used to generate an analysis report that meets the expected analysis objectives. The analysis report is generated based on the information to be retrieved and existing technical documents retrieved using a retrieval query. The analysis report at least characterizes the feature differences between the information to be retrieved and the existing technical documents. The retrieval query is generated based on the information to be retrieved. The expected analysis objectives characterize the analysis direction the user expects the analysis report to provide. The analysis report is sent to the client to instruct the client to present and modify the feature differences in the analysis report through the visualization interface.
17. The apparatus according to claim 16, wherein, The device further includes: The retrieval display module is used to receive model instruction commands submitted by users; wherein, the model instruction command is a prompt command used to instruct the large model to generate a retrieval query; in response to the model instruction command, the module generates a retrieval query based on the information to be retrieved; and sends the retrieval query to the client to instruct the client to display the retrieval query through the visualization interface.
18. A computer device, wherein, The computer device includes a memory and a processor, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to implement the method as described in any one of claims 1 to 15.
19. A computer-readable storage medium, wherein, The computer-readable storage medium stores at least one computer program, which, when executed by a processor, is capable of implementing the method as described in any one of claims 1 to 15.
20. A computer program product, wherein, The computer program product is used to implement the method as described in any one of claims 1 to 15.