Methods, systems and computer-readable media for actively monitoring patent infringement
Fine-tuning LLMs with litigation data allows for automated and efficient identification of patent infringement, addressing the inefficiencies of manual methods and enhancing the accuracy and speed of infringement detection.
Patent Information
- Application Number
- PCT/CA2025/050102
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-30
- Filing Date
- 2025-01-24
- Publication Date
- 2025-07-31
AI Technical Summary
Current methods for identifying patent infringement are manual, time-consuming, and costly, lacking efficient and automated solutions for discovering potential infringers.
Utilizing Large Language Models (LLMs) fine-tuned with a litigation dataset to automatically identify products or services that may infringe patents, and vice versa, through vector database generation and machine learning model refinement.
Enables efficient and accurate identification of potential patent infringement, reducing manual effort and costs, and allowing for real-time monitoring and adjustment of patent claims.
Smart Images

Figure CA2025050102_31072025_PF_FP_ABST
Abstract
Description
TITLE: METHODS, SYSTEMS AND COMPUTER-READABLE MEDIA FOR ACTIVELY MONITORING PATENT INFRINGEMENTCROSS-REFERENCE TO RELATED PATENT APPLICATIONS
[0001] This application claims the benefit of United States Provisional Patent Application No. 63 / 625,934 filed January 27, 2024; United States Provisional Patent Application No. 63 / 689,269 filed August 30, 2024; United States United States Provisional Patent Application No. 63 / 689,571 filed August 30, 2024 and United States Provisional Patent Application No. 63 / 689,551 filed August 30, 2024. The entire contents of United States Provisional Patent Application No. 63 / 625,934; United States Provisional Patent Application No. 63 / 689,269; United States United States Provisional Patent Application No. 63 / 689,571 and United States Provisional Patent Application No. 63 / 689,551 is hereby incorporated by reference.FIELD
[0002] The present embodiments relate generally to systems for analyzing patent claims and for actively monitoring patent infringement, methods of operating thereof and related computer-readable media.BACKGROUND
[0003] An issued patent offers a patentee enforceable rights to exclude others from making, using, selling, or importing products and services that fall within the bounds of the claims of the issued patent.
[0004] The discovery of potential infringers of the claims of a patent, however, has conventionally been performed based on market surveillance including private investigations of competitors and the analysis of commercially available products and services. This can be done by searching online marketplaces, product catalogs, and patent databases to identify products that may be in violation of their patents.
[0005] Such monitoring and discovery of potential infringers is complex, costly, and time consuming since it is driven by manual search methodologies. There remains an unmet need for efficient and automated solutions for the discovery of alleged infringers.SUMMARY
[0006] The present embodiments are directed to systems, methods and computer-readable media for automatically identifying potential patent infringers, performing patent analysis using Large Language Models (LLMs) and ranking results products or services that may infringe a patent and / or ranking patents that may be infringed by products or services. The use of such LLMs introduces technical challenges. Namely, in order to aim to provide efficient and accurate results, the use of the LLM systems may be fine-tuned. To perform fine tuning, a dataset including pairs of patent claims and litigation outcomes is illustrative since it identifies ground truth and improves the ability of the LLM system to provide accurate results when prompted.
[0007] In a first aspect, there is provided a computer-implemented method for fine- tuning a machine learning model. The method involves: providing, in a memory, a litigation dataset comprising a plurality of historical patent litigation records, each patent litigation record comprising at least one patent claim, and a litigation outcome corresponding to each of the at least one patent claim; receiving, at a processor in communication with the memory, a product or service content item associated with a candidate patent litigation record in the plurality of historical patent litigation records; updating, at the processor, the candidate patent litigation record based on the product or service content item; and generating, at the processor, a vector database based on the litigation dataset, the vector database used to fine-tune a machine learning model.
[0008] In some embodiments, the method further involves: receiving, at a network device in communication with the processor, a user submitted product or service content item associated with a product or service in the candidate patent litigation record.
[0009] In some embodiments, the product or service content item comprises a user-submitted URL content item corresponding to the product or service in the candidate patent litigation record and the method further comprises: sending, using a network device in communication with the processor, a content request, the content request based on a URL content item; receiving, using the network device, a content response based on the content request; performing web scraping on the content response to identify at least one image content item and at least one text content item; and updating the candidate patent litigation record by storing the at least one image content item and the at least one text content item in association with the candidate patent litigation record.
[0010] In some embodiments, the content item comprises a user-submitted text content item corresponding to the product or service in the candidate patent litigation record, and the method further comprises: storing the user-submitted text content item inassociation with the candidate patent litigation record.
[0011] In some embodiments, the method further comprises: sending, using the network device, a captioning request to an LLM-system, the captioning request comprising the at least one image content item; receiving, using the network device, a captioning response from the LLM-system based on the captioning request, the captioning response comprising at least one text-based caption based on the at least one image content item; and updating, at the processor, the candidate patent litigation record based on the at least one text-based caption.
[0012] In some embodiments, the method further comprises: determining, at the processor, a hash value corresponding to the content response; updating, at the processor, the candidate patent litigation record with the hash value; and subsequently sending, using a network device in communication with the processor, a second content request based on the URL content item; receiving, using the network device, a second content response based on the second content request; determining, at the processor, that a second hash value corresponding to the second content response is different from the hash value corresponding to the content response; and updating, at the processor, the candidate patent litigation record based on the second content response.
[0013] In some embodiments, the generating the vector database comprises generating a vector embedding for the product or service content item of the candidate patent litigation record.
[0014] In some embodiments, each patent litigation record in the dataset comprises at least one selected from the group of: a corporate identification and an estimated damages quantum.
[0015] In some embodiments, the method further comprises generating, at the processor, a prediction prompt for an LLM-system based on the vector database; sending, using a network device in communication with the processor to the LLM-system, the prediction prompt; and receiving, using the network device, a prompt response comprising at least two search elements.
[0016] In some embodiments, the method further comprises generating, at the processor, at least two machine learning models based on the vector database, the at least two machine learning models comprising a ranking model and a classification model; and updating, at the processor, each of the at least two search elements based on the at least two machine learning models, each updated search element comprisinga rank and a classification.
[0017] In accordance with another aspect, there is provided a system for fine- tuning a machine learning model. The system includes a memory, comprising a litigation dataset comprising a plurality of historical patent litigation records, each patent litigation record comprising at least one patent claim, and a litigation outcome corresponding to each of the at least one patent claim; a processor in communication with the memory, configured to: receive a product or service content item corresponding to a candidate patent litigation record in the plurality of historical patent litigation records; update the candidate patent litigation record based on the product or service content item; and generate a vector database based on the litigation dataset, the vector database used to fine-tune a machine learning model.
[0018] In some embodiments, the system includes a network device in communication with the processor, the network device configured to receive a user submitted product or service content item associated with a product or service in the candidate patent litigation record.
[0019] In some embodiments, the product or service content item comprises a user-submitted URL content item corresponding to the product or service in the candidate patent litigation record, and the network device is further configured to: send a content request, the content request based on a URL content item; receive a content response based on the content request; and the processor is further configured to: perform web scraping on the content response to identify at least one image content item and at least one text content item; and update the candidate patent litigation record by storing the at least one image content item and the at least one text content item in association with the candidate patent litigation record.
[0020] In some embodiments, the content item comprises a user-submitted text content item corresponding to the product or service in the candidate patent litigation record, and the processor is further configured to: store the user-submitted text content item in association with the candidate patent litigation record.
[0021] In some embodiments, the network device is further configured to: send a captioning request to an LLM-system, the captioning request comprising the at least one image content item; receive a captioning response from the LLM-system based on the captioning request, the captioning response comprising at least one text-based caption based on the at least one image content item; and the processor is further configured to:update the candidate patent litigation record in the memory based on the at least one text-based caption.
[0022] In some embodiments, the processor is further configured to: determine a hash value corresponding to the content response; update the candidate patent litigation record in the memory with the hash value; determine that a second hash value corresponding to a second content response is different from the hash value corresponding to the content response; and update the candidate patent litigation record based on the second content response; and the network device is further configured to: subsequently send a second content request based on the URL content item; receive the second content response based on the second content request.
[0023] In some embodiments, the generating the vector database comprises generating a vector embedding for each patent litigation record in the plurality of historical patent litigation records.
[0024] In some embodiments, each patent litigation record in the dataset comprises at least one selected from the group of: a corporate identification and an estimated damages quantum.
[0025] In some embodiments, the processor is further configured to: generate a prediction prompt for an LLM-system based on the vector database; and a network device in communication with the processor is configured to: send to the LLM-system, the prediction prompt; and receive a prompt response comprising at least two search elements.
[0026] In some embodiments, the processor is configured to: generate at least two machine learning models based on the vector database, the at least two machine learning models comprising a ranking model and a classification model; and update each of the at least two search elements based on the at least two machine learning models, each updated search element comprising a rank and a classification.
[0027] In another aspect, there is provided a computer-implemented method for providing updated search results by fine-tuning of machine learning models using a feedback dataset. The method involves: receiving, at a network device from a user device, a search query comprising at least one of a patent document identifier or a product or service identifier; transmitting, from the network device to the user device, a search response user interface comprising a search identifier and at least one search result item; receiving, at the network device from the user device, a feedback itemcorresponding to a candidate search item in the at least one search result item; storing, in a memory, a feedback item record in a feedback dataset, the feedback item record stored in association with the search identifier, the feedback item record comprising the feedback item, at least one of a patent claim associated with the patent document identifier or a product or service description, and the search result item; and transmitting, from the network device to the user device, at least one updated search result item.
[0028] In some embodiments, the method further comprises: extracting, at a processor in communication with the memory, patent information associated with the patent document identifier; and generating, at the processor, an initial analysis prompt based on the patent information.
[0029] In some embodiments, the method further comprises extracting, at a processor in communication with the memory, product or service information associated with the product or service description; and generating, at the processor, an initial analysis prompt based on the product information.
[0030] In some embodiments, the patent information associated with the patent document identifier comprises at least one selected from the group of: claim text; drawing images; caption text for the drawing images; description text; abstract text; forward or backward citations, each citation comprising a patent document identifier; chart images; caption text for the chart images; formulae text; inventor information; assignee information; classification information; prosecution file wrapper text; maintenance payment information; and litigation history text.
[0031] In some embodiments, the product or service information associated with the product or service identifier comprises at least one selected from the group of: a text snippet, a URL, an icon, an image, and a citation.
[0032] In some embodiments, the method comprises: transmitting, from the network device to an LLM system, an initial analysis request comprising the initial analysis prompt; receiving, from the network device from the LLM system, an initial analysis response based on the initial analysis request, the initial analysis response comprising at least one remote device search query; transmitting, from the network device to a remote device, the at least one remote device search query; receiving, from the network device to the remote device, at least one remote device search response; and transmitting, from the network device to the user device, the at least one search result item based on the at least one remote device search response.
[0033] In some embodiments, the method comprises: generating, at the processor, a subsequent analysis prompt based on the initial analysis prompt and the feedback item record; transmitting, from the network device to an LLM system, a subsequent analysis request comprising the subsequent analysis prompt; receiving, from the network device from the LLM system, a subsequent analysis response based on the subsequent analysis request, the subsequent analysis response comprising at least one subsequent remote device search query; transmitting, from the network device to the remote device search provider, the at least one subsequent remote device search query; receiving, from the network device to the remote device search provider, at least one subsequent remote device search response; and transmitting, from the network device to the user device, at least one subsequent search result item based on the at least one subsequent remote device search response.
[0034] In some embodiments, the feedback item comprises a rating, the rating corresponding to a feedback control selected by the user.
[0035] In some embodiments, the feedback control comprises at least one selected from the group of: two or more button controls, a radio button control, a slider control, a star selector control.
[0036] In some embodiments, the method involves providing, in a memory, the feedback item dataset comprising a plurality of historical feedback records, each feedback item record stored in association with a corresponding search identifier, the feedback item record comprising a historical feedback item, at least one of a patent claim associated with a corresponding patent document identifier or a product or service description, and a corresponding search result item.
[0037] In some embodiments, the method involves generating, at the processor, at least two machine learning models based on the feedback item dataset, the at least two machine learning models comprising a ranking model and a classification model; and updating, at the processor, each of the at least one search result item based on the at least two machine learning models, each of the at least one search result item comprising a rank and a classification.
[0038] In accordance with another aspect, there is provided a system for providing updated search results by fine-tuning machine learning models using a feedback dataset. The system includes: a network device configured to: receiving from a user device, a search query comprising at least one of a patent document identifier, a product identifieror service identifier; transmitting to the user device, a search response user interface comprising a search identifier and at least one search result item; and receiving from the user device, a feedback item corresponding to a candidate search item in the at least one search result item; a memory; a processor configured to: store in the memory, a feedback item record in a feedback dataset, the feedback item record stored in association with the search identifier, the feedback item record comprising the feedback item, at least one of a patent claim associated with the patent document identifier or a product or service description, and the search result item; and transmitting, using the network device to the user device, at least one updated search result item based on the feedback item record.
[0039] In some embodiments, the processor is configured to: extract patent information associated with the patent document identifier; and generate an initial analysis prompt based on the patent information.
[0040] In some embodiments, the processor is configured to: extract product information associated with the product or service description; and generate an initial analysis prompt based on the product or service information.
[0041] In some embodiments, the patent information associated with the patent document identifier comprises at least one selected from the group of: claim text; drawing images; caption text for the drawing images; description text; abstract text; forward or backward citations, each citation comprising a patent document identifier; chart images; caption text for the chart images; formulae text; inventor information; assignee information; classification information; prosecution file wrapper text; maintenance payment information; and litigation history text.
[0042] In some embodiments, the product or service information associated with the patent document identifier comprises at least one selected from the group of: a text snippet, a URL, an icon, an image, and a citation.
[0043] In some embodiments, the processor is further configured to: transmit using the network device to an LLM system, an initial analysis request comprising the initial analysis prompt; receive from the network device from the LLM system, an initial analysis response based on the initial analysis request, the initial analysis response comprising at least one remote device search query; transmit from the network device to a remote device search provider, the at least one remote device search query; receive from the network device to the remote device search provider, at least one remote device searchresponse; and transmit from the network device to the user device, the at least one search result item based on the at least one remote device search response.
[0044] In some embodiments, the processor is further configured to: generate a subsequent analysis prompt based on the initial analysis prompt and the feedback item record; transmit from the network device to an LLM system, a subsequent analysis request comprising the subsequent analysis prompt; receive from the network device from the LLM system, a subsequent analysis response based on the subsequent analysis request, the subsequent analysis response comprising at least one subsequent remote device search query; transmit from the network device to the remote device search provider, the at least one subsequent remote device search query; receive from the network device to the remote device search provider, at least one subsequent remote device search response; and transmit from the network device to the user device, at least one subsequent search result item based on the at least one subsequent remote device search response.
[0045] In some embodiments, the feedback item comprises a rating, the rating corresponding to a feedback control selected by the user.
[0046] In some embodiments, the feedback control comprises at least one selected from the group of: two or more button controls, a radio button control, a slider control, a star selector control.
[0047] In some embodiments, the feedback item dataset further comprises a plurality of historical feedback records, each feedback item record stored in association with a corresponding search identifier, the feedback item record comprising a historical feedback item, at least one of a patent claim associated with a corresponding patent document identifier or a product or service description, and a corresponding search result item.
[0048] In some embodiments, the processor is further configured to: generate at least two machine learning models based on the feedback dataset, the at least two machine learning models comprising a ranking model and a classification model; and update each of the at least one search result item based on the at least two machine learning models, each of the at least one search result item comprising a rank and a classification.
[0049] In accordance with another aspect, there is provided a method of ranking product-patent or service-patent evaluation search results. The method involvesoperating at least one processor to: receive a plurality of search results relevant to a search input, the search input corresponding to a product description or a service description, or a patent claim; for each search result, compare a vector embedding of the search result and a vector embedding of the search input to obtain a relevancy score for the search result; apply, by the processor, one or more ranking refinement models to the search results to obtain refined relevancy scores; and display, the search results according to the refined relevancy scores.
[0050] In accordance with another aspect, there is provided a method of ranking product-patent or service-patent evaluation search results. The method involves operating at least one processor to: receive a plurality of search results corresponding to products or services relevant to a patent claim associated with a patent document; calculate a vector embedding of the patent claim; for each search result, calculate a vector embedding of the search result and comparing the vector embedding of the claim and the vector embedding of the search result to obtain a relevancy score for the search result; apply one or more ranking refinement models to the search results to obtain refined relevancy scores; and display the search results according to the refined relevancy scores.
[0051] In some embodiments, the one or more ranking refinement models are selected according to one or more of: a subject matter of the patent document and classifications of the patent document.
[0052] In some embodiments, applying the one or more ranking refinement models comprises applying two or more ranking refinement models, and the method further comprises operating the at least one processor to determine a weighting for each of the two or more ranking refinement models according to one or more of: a predetermined weighting, a subject matter of the patent document, classifications of the patent document and a confidence of the ranking refinement model.
[0053] In some embodiments, at least one of the one or more ranking refinement models is a large language model and applying the one or more ranking refinement models comprises operating the at least one processor to generate a guidance prompt for the at least one ranking refinement model based one or more of: a subject matter of the patent document, classifications of the patent document, search parameters, feedback data.
[0054] In some embodiments, applying the one or more ranking refinementmodels to obtain refined relevancy scores comprises operating the at least one processor to: for each of the one or more ranking refinement models, for each search result, obtain a model relevancy score; and determine the refined relevancy score based on the relevancy score and the model relevancy score.
[0055] In some embodiments, the plurality of results are URLs of product or service websites and the method further comprises operating the at least one processor to parse each of the URLs to obtain a corresponding product or service description.
[0056] In some embodiments, the method further comprises operating the at least one processor to identify the K most relevant search results based on the relevancy scores; and apply the one or more ranking refinement models to the K most relevant search results.
[0057] In some embodiments, applying the one or more ranking refinement models to the search results comprises operating the at least one processor to input a text of the patent claim and a text of a product or service description corresponding to the search result into the one or more ranking refinement models, and the one or more ranking refinement models are configured to determine a measure of similarity between the text of the claim and the text of the product or service description corresponding to the search result.
[0058] In some embodiments, calculating the vector embedding of the claim comprises operating the at least one processor to: identify one or more portions of the patent corresponding to the claim; and calculate the vector embedding of the patent claim and the one or more portions of the patent document.
[0059] In some embodiments, the one or more portions comprise one or more of: a description, a drawing, an equation, a table and a formula.
[0060] In accordance with another aspect, there is provided a method of ranking product-patent or service-patent evaluation search results. The method comprises operating at least one processor to: receive a description of a product or a service; convert the description of the product or the description of the service into a vector embedding of the product or service; retrieve, from a database, a plurality of vector embeddings of search results corresponding to patent claims; for each search result, compare the vector embedding of the search result and the vector embedding of the product or service description to obtain a relevancy score for the search result; apply one or more ranking refinement models to the search results to obtain refined relevancyscores; and display by the search results according to the refined relevancy scores.
[0061] In accordance with another aspect there is provided a system for ranking product-patent evaluation or service-patent evaluation search results, the system comprising at least one processor operable to: receive a plurality of search results corresponding to products or services relevant to a patent claim associated with a patent document; calculate a vector embedding of the patent claim; for each search result, calculate a vector embedding of the search result and comparing the vector embedding of the claim and the vector embedding of the search result to obtain a relevancy score for the search result; apply one or more ranking refinement models to the search results to obtain refined relevancy scores; and display the search results according to the refined relevancy scores.
[0062] In some embodiments, the at least one processor is operable to select the one or more ranking refinement models according to one or more of: a subject matter of the patent document and classifications of the patent document.
[0063] In some embodiments, applying the one or more ranking refinement models comprises applying two or more ranking refinement models, and the at least one processor is further operable to determine a weighting for each of the two or more ranking refinement models according to one or more of: a predetermined weighting, a subject matter of the patent document, classifications of the patent document and a confidence of the ranking refinement model.
[0064] In some embodiments, at least one of the one or more ranking refinement models is a large language model and applying the one or more ranking refinement models comprises generating a guidance prompt for the at least one ranking refinement model based one or more of: a subject matter of the patent document, classifications of the patent document, search parameters, feedback data.
[0065] In some embodiments, applying the one or more ranking refinement models to obtain refined relevancy scores comprises operating the at least one processor to: for each of the one or more ranking refinement models, for each search result, obtain a model relevancy score; and determine the refined relevancy score based on the relevancy score and the model relevancy score.
[0066] In some embodiments, the plurality of results are URLs of product or service websites and the at least one processor is operable to parse each of the URLs to obtain a corresponding product or service description.
[0067] In some embodiments, the at least one processor is further operable to: identify the K most relevant search results based on the relevancy scores; and apply the one or more ranking refinement models to the K most relevant search results.
[0068] In some embodiments, applying the one or more ranking refinement models to the search results comprises operating the at least one processor to input a text of the claim and a text of a product or service description corresponding to the search result into the one or more ranking refinement models, and the one or more ranking refinement models are configured to determine a measure of similarity between the text of the claim and the text of the product or service description corresponding to the search result.
[0069] In some embodiments, calculating the vector embedding of the claim comprises identifying one or more portions of the patent document corresponding to the patent claim; and calculating the vector embedding of the claim and the one or more portions of the patent document corresponding to the patent claim.
[0070] In some embodiments, the one or more portions comprise one or more of: a description, a drawing, an equation, a table and a formula.
[0071] In accordance with another aspect, there is provided a computer-readable medium having instructions stored thereon that when executed, cause at least one processor to: receive a plurality of search results corresponding to products or services relevant to a patent claim associated with a patent document; calculate a vector embedding of the claim for each search result, calculate a vector embedding of the search result and comparing the vector embedding of the claim and the vector embedding of the search result to obtain a relevancy score for the search result; apply one or more ranking refinement models to the search results to obtain refined relevancy scores; and display the search results according to the refined relevancy scores.
[0072] In some embodiments, the instructions, when executed, further cause the at least one processor to perform steps of the methods described herein.
[0073] In accordance with another aspect, there is provided a method for identifying products or services relevant to a target patent claim, the method comprising operating at least one processor to: receive a patent document identifier associated with the target patent claim from a computing device in communication with the at least one processor via a network; obtain the target patent claim; retrieve, from a database in communication with the at least one processor via the network, patent documentdetails associated with the target patent claim, using the patent document identifier; generate a set of queries based on the retrieved patent document details; search remote databases using the set of queries to identify products or services satisfying the query set; score the identified products or services to provide a relevancy score for each identified product or service; and generate a ranked list of the products or services having a relevancy score greater than a predetermined threshold; and display the generated ranked list.
[0074] In some embodiments, obtaining the target patent claim comprises operating the at least one processor to: receive a free-form text of the target patent claim from the computing device; or retrieve, from the database, the target patent claim from a patent document associated with the patent document identifier.
[0075] In some embodiments, the method further comprises operating the at least one processor to: when no identified product or service has a relevancy score greater than the predetermined threshold, display an indication indicating a lack of relevant results.
[0076] In some embodiments, the method further comprises operating the at least one processor to: receive a digital image comprising the patent number of the target patent; and extract the patent number from the digital image using optical character recognition.
[0077] In some embodiments, the method further comprises operating the at least one processor to: access a database comprising patent information to retrieve the patent details associated with the target patent.
[0078] In some embodiments, the method further comprises operating the at least one processor to: analyze the retrieved patent details to identify terms and phrases relevant to the target patent.
[0079] In some embodiments, the remote databases comprise one or more of: databases of online marketplaces, databases comprising product catalogs, service offering catalogs, and patent databases.
[0080] In some embodiments, the method further comprises operating the at least one processor to: score the identified products and services using a machine learning algorithm.
[0081] In some embodiments, the method further comprises operating the at leastone processor to: display the generated ranked list as a table or a graph; display at least one interactable graphical user interface element, the interactable graphical user interface element causing the at least one processor to filter or sort the ranked list according to a selectable criterion.
[0082] In some embodiments, the method further comprises operating the at least one processor to: determine one or more recommendations.
[0083] In accordance with another aspect, there is provided, a system for identifying products or services relevant to a target patent claim, the system comprising at least one processor implementing: an input module for receiving a patent document identifier associated with the target patent claim from a computing device in communication with the at least one processor; a patent document detail fetching module for retrieving patent document details associated with the target patent claim using the patent document identifier; a query generation module for generating a set of queries based on the retrieved patent document details; a search module for searching remote databases using the set of queries to identify products or services satisfying the query set; a scoring module for scoring the identified products or services to provide a relevancy score for each identified product or service; a ranking module for generating a ranked list of the products or services having a relevancy score greater than predetermined threshold; and a feedback module for receiving feedback on the ranked list of the products and services from the computing device.
[0084] In some embodiments, the input module is further configured to receive a free-form text of the target patent claim from the computing device.
[0085] In some embodiments, the patent document detail fetching module is configured to retrieve the target patent claim from a patent document associated with the patent document identifier.
[0086] In some embodiments, the patent detail-fetching module comprises a communication interface for connecting to a remote database comprising patent information and the patent detail fetching module is configured to retrieve the patent document details associated with the target patent claim from the remote database.
[0087] In some embodiments, the query-generation module comprises a natural language processing component for analyzing the retrieved patent document details to identify terms and phrases relevant to the target patent claim.
[0088] In some embodiments, the search module comprises a web crawlingcomponent for online searching at least one of: databases comprising online marketplaces, databases comprising product catalogs, service offering catalogs and patent databases.
[0089] In some embodiments, the scoring module comprises a machine learning component that uses a pre-trained model to score the identified products and services.
[0090] In some embodiments, the ranking module comprises a visualization component configured to display the generated ranked list as a table or a graph and display at least one interactable graphical user interface element, the interactable graphical user interface element causing the at least one processor to filter or sort the ranked list according to a selectable criterion.
[0091] In accordance with another aspect, there is provided a non-transitory computer-readable medium having a set of executable instructions stored thereon that, when executed by at least one processor, cause the at least one processor to: receive a patent document identifier associated with a target patent claim from a computing device in communication with the at least one processor via a network; obtain the target patent claim; retrieve, from a database, patent document details associated with the target patent claim, using the patent document identifier; generate a set of queries based on the retrieved patent document details; search remote databases using the set of queries to identify products or services satisfying the query set; score the identified products or services to provide a relevancy score for each identified product or service; and generate a ranked list of the products or services having a relevancy score greater than a predetermined threshold; and display the generated ranked list.
[0092] In some embodiments, the set of executable instructions further cause the at least one processor to: receive a free-form text of the target patent claim from the computing device; or retrieve, from the database, the target patent claim from a patent document associated with the patent document identifier.
[0093] In some embodiments, the set of executable instructions, when executed by the at least one processor, further cause the at least one processor to receive a digital image comprising the patent document identifier associated with the target patent claim and extract the patent document identifier from the digital image using optical character recognition.
[0094] In some embodiments, the set of executable instructions, when executed by the at least one processor, further cause the at least one processor to access adatabase comprising patent information to retrieve the patent document details associated with the target patent claim.
[0095] In some embodiments, the set of executable instructions, when executed by the at least one processor, further cause the at least one processor to analyze the retrieved patent document details to identify terms and phrases relevant to the target patent claim.
[0096] In some embodiments, the set of executable instructions, when executed by the at least one processor, further cause the at least one processor to score the identified products and services using a machine learning algorithm.
[0097] In some embodiments, the set of executable instructions, when executed by the at least one processor, further cause the at least one processor to: display the generated ranked list as a table or a graph; and display at least one interactable graphical user interface element, the interactable graphical user interface element causing the at least one processor to filter or sort the ranked list according to a selectable criterion.
[0098] In accordance with another aspect, there is provided a method for identifying patents relevant to a product or a service, the method comprising operating at least one processor to: receive a description of the product or service from a computing device in communication with the at least one processor; extract description details from the received description; generate a set of search queries based on the extracted description details; search at least one database for patents satisfying the set of search queries; score the patents satisfying the set of search queries to obtain a relevancy score for each patent; generate a ranked list of the patents having a relevancy score greater than a predetermined threshold.
[0099] In some embodiments, the method further comprises operating the at least one processor to: extract information including: product features or service features, technical specifications, or method parameters from the received description.
[0100] In accordance with another aspect, there is provided a system for identifying patents relevant to a product or a service, the system comprising operating at least one processor implementing: an input module for receiving a description of the product or the service from a computing device in communication with the at least one processor; a detail extraction module for extracting description details from the received description; a query generation module for generating a set of queries based on the extracted description details; a search module for searching at least one database forpatents satisfying the set of search queries; a scoring module for scoring the patents satisfying the set of search queries to obtain a relevancy score for each patent; a list generation module for generating a ranked list of the patents having a relevancy score greater than predetermined threshold.
[0101] In some embodiments, the detail extraction module comprises means for extracting information including: product features or service features, technical specifications, or method parameters from the received description.
[0102] In accordance with another aspect, there is provided a non-transitory computer-readable medium having a set of executable instructions stored thereon that, when executed by at least one processor, cause the at least one processor to: receive a description of the product or service from a computing device in communication with the at least one processor; extract description details from the received description; generate a set of search queries based on the extracted description details; search at least one database for patents satisfying the set of search queries; score the patents satisfying the set of search queries to obtain a relevancy score for each patent; generate a ranked list of the patents having a relevancy score greater than a predetermined threshold.
[0103] In some embodiments, the set of executable instructions, when executed by the at least one processor, further cause the at least one processor to extract information including: product features or service features, technical specifications, or method parameters from the received description.
[0104] In accordance with another aspect, there is provided a computer-readable having instructions stored thereon that when executed cause at least one processor to perform any of the methods described herein.DRAWINGS
[0105] An illustrative embodiment of the present invention will now be described in detail with reference to the diagrams, in which:
[0106] FIG. 1 shows a patent analytics and infringement monitoring system in communication with external components, in accordance with one or more embodiments;
[0107] FIG. 2 shows a device diagram for a server of the patent analytics and infringement monitoring system of FIG. 1 , in accordance with one or more embodiments;
[0108] FIG. 3 shows a method diagram of a method of operating the patentanalytics and infringement monitoring system in accordance with one or more embodiments.
[0109] FIG. 4 shows an input / output diagram in accordance with one or more embodiments;
[0110] FIG. 5 shows another method diagram of a method of operating the patent analytics and infringement monitoring system in accordance with one or more embodiments;
[0111] FIG. 6 shows a flowchart of an example method for fine-tuning a machinelearning model, in accordance with one or more embodiments;
[0112] FIG. 7 shows another method diagram of operating the patent analytics and infringement monitoring system in in accordance with one or more embodiments;
[0113] FIG. 8 shows another method diagram of a method of operating the patent analytics and infringement monitoring system in accordance with one or more embodiments;
[0114] FIG. 9 shows a data diagram in accordance with one or more embodiments;
[0115] FIG. 10A shows a screenshot of graphical user interface (GUI) illustrating an example embodiment of an application of a system for patent analytics and for monitoring patent infringement;
[0116] FIG. 10B shows another screenshot of graphical user interface (GUI) illustrating an example embodiment of an application of a system for patent analytics and for monitoring patent infringement;
[0117] FIG. 11 shows another screenshot of a GUI illustrating an example embodiment of an application of a system for patent analytics and for monitoring patent infringement;
[0118] FIG. 12 shows a flowchart of an example method for providing updated search results by fine-tuning of machine learning models using a feedback dataset, in accordance with one or more embodiments;
[0119] FIG. 13 shows another data diagram in accordance with one or more embodiments;
[0120] FIG. 14 shows a flowchart of an example method for identifying products or services that may infringe a patent, in accordance with one or more embodiments;
[0121] FIG. 15 shows a flowchart of an example method for identifying one or more patents that may be infringed by a product or service, in accordance with one or more embodiments;
[0122] FIG. 16 shows a flowchart of an example method of ranking products or services that may infringe a patent, in accordance with one or more embodiments;
[0123] FIG. 17 is a schematic representation of another method of operating the patent analytics and infringement monitoring system to identify products or services that may infringe a patent, in accordance with one or more embodiments;
[0124] FIG. 18 is a flowchart of an example method for identifying a list of products potentially infringing a target patent, in accordance with one or more embodiments;
[0125] FIG. 19A is a schematic representation of functionalities of a patent analytics and infringement monitoring system, in accordance with one or more embodiments;
[0126] FIG. 19B is a schematic representation of functionalities of a patent analytics and infringement monitoring system , in accordance with one or more embodiments;
[0127] FIG. 19C is a schematic representation of functionalities of a patent analytics and infringement monitoring system , in accordance with one or more embodiments;
[0128] FIG. 19D is a schematic representation of functionalities of a patent analytics and infringement monitoring system, in accordance with one or more embodiments;
[0129] FIG. 19E is a schematic representation of a patent analytics and infringement monitoring system, in accordance with one or more embodiments;
[0130] FIG. 20 shows another screenshot of a GUI illustrating an example embodiment of an application of a system for patent analytics and for monitoring patent infringement;
[0131] FIG. 21 shows another screenshot of a GUI illustrating an example embodiment of an application of a system for patent analytics and for monitoring patent infringement;
[0132] FIG. 22 is a schematic representation of another method for operating thepatent analytics and infringement monitoring system to identify patents or patent applications that may be infringed by products or services, in accordance with one or more embodiments;
[0133] FIG. 23 is a schematic diagram showing components of the patent analytics and infringement monitoring system; and
[0134] FIG. 24 is a schematic diagram showing a training model for a machine learning framework, in accordance with one or more embodiments.DESCRIPTION OF VARIOUS EMBODIMENTS
[0135] Various embodiments will now be described below to provide an example of the claimed subject matter. No example described below limits any claimed subject matter and any claimed subject matter may cover embodiments such as systems or methods that differ from those described below.
[0136] Furthermore, it will be appreciated that for simplicity and clarity of illustration, where considered appropriate, reference numerals may be repeated among the figures to indicate corresponding or analogous elements. In addition, numerous specific details are set forth in order to provide a thorough understanding of the examples described herein. However, it will be understood by those of ordinary skill in the art that the examples described herein may be practiced without these specific details. In other instances, well-known methods, procedures, and components have not been described in detail so as not to obscure the examples described herein. Also, the description is not to be considered as limiting the scope of the examples described herein.
[0137] It should also be noted that, as used herein, the wording “and / or” is intended to represent an inclusive-or. That is, “X and / or Y” is intended to mean X or Y or both, for example. As a further example, “X, Y, and / or Z” is intended to mean X or Y or Z or any combination thereof.
[0138] It should be noted that terms of degree such as "substantially", "about" and "approximately" as used herein mean a reasonable amount of deviation of the modified term such that the end result is not significantly changed. These terms of degree may also be construed as including a deviation of the modified term if this deviation would not negate the meaning of the term it modifies.
[0139] Furthermore, the recitation of numerical ranges by endpoints herein includes all numbers and fractions subsumed within that range (e.g., 1 to 5 includes 1 ,1 .5, 2, 2.75, 3, 3.90, 4, and 5). It is also to be understood that all numbers and fractions thereof are presumed to be modified by the term "about" which means a variation of up to a certain amount of the number to which reference is being made if the end result is not significantly changed.
[0140] Some elements herein may be identified by a part number, which is composed of a base number followed by an alphabetical or subscript-numerical suffix (e.g., 112a, or 1121 ). Multiple elements herein may be identified by part numbers that share a base number in common and that differ by their suffixes (e.g., 1121 , 1122, and 1123). All elements with a common base number may be referred to collectively or generically using the base number without a suffix (e.g., 112).
[0141] The example systems and methods described herein may be implemented in hardware or software, or a combination of both. In some cases, the examples described herein may be implemented, at least in part, by using one or more computer programs, executing on one or more programmable devices comprising at least one processing element, a data storage element (including volatile and non-volatile memory and / or storage elements), and at least one communication interface. These devices may also have at least one input device (e.g., a keyboard, a mouse, a touchscreen, and the like), and at least one output device (e.g., a display screen, a printer, a wireless radio, and the like) depending on the nature of the device. For example, and without limitation, the programmable devices (referred to below as computing devices) may be a server, network appliance, embedded device, computer expansion module, a personal computer, laptop, personal data assistant, cellular telephone, smart-phone device, tablet computer, a wireless device or any other computing device capable of being configured to carry out the methods described herein.
[0142] In some examples, the communication interface may be a network communication interface. In examples in which elements are combined, the communication interface may be a software communication interface, such as those for inter-process communication (IPC). In still other examples, there may be a combination of communication interfaces implemented as hardware, software, and a combination thereof.
[0143] Program code may be applied to input data to perform the functions described herein and to generate output information. The output information is applied to one or more output devices, in known fashion.
[0144] Each program may be implemented in a high-level procedural, declarative, functional or object-oriented programming and / or scripting language, or both, to communicate with a computer system. However, the programs may be implemented in assembly or machine language, if desired. In any case, the language may be a compiled or interpreted language. Each such computer program may be stored on a storage media or a device (e.g., ROM, magnetic disk, optical disc) readable by a general or special purpose programmable computer, for configuring and operating the computer when the storage media or device is read by the computer to perform the procedures described herein. Examples of the system may also be considered to be implemented as a non- transitory computer-readable storage medium, configured with a computer program, where the storage medium so configured causes a computer to operate in a specific and predefined manner to perform the functions described herein.
[0145] Furthermore, the example system, processes and methods are capable of being distributed in a computer program product comprising a computer readable medium that bears computer usable instructions for one or more processors. The medium may be provided in various forms, including one or more diskettes, compact disks, tapes, chips, wireline transmissions, satellite transmissions, internet transmission or downloads, magnetic and electronic storage media, digital and analog signals, and the like. The computer useable instructions may also be in various forms, including compiled and non-compiled code.
[0146] Various examples of systems, methods and computer programs products are described herein. Modifications and variations may be made to these examples without departing from the scope of the invention, which is limited only by the appended claims. Also, in the various user interfaces illustrated in the figures, it will be understood that the illustrated user interface text and controls are provided as examples only and are not meant to be limiting. Other suitable user interface elements may be used with alternative implementations of the systems and methods described herein.
[0147] Conventionally, patent infringement is discovered through market surveillance, including through monitoring the commercial activities of competitors, or by monitoring the Web through tools such as search engine alerts. These techniques, however, are time-consuming and costly, and contribute to the majority of patents being unenforced. Further, monitoring the web often requires specialized knowledge of search terms to be used and requires patent owners to manually create alerts or other monitoring tools.
[0148] Similarly, assessing whether a product or service is likely to infringe a patent is often costly and time-consuming, since the process typically requires specialized patent databases, expert knowledge, and a manual assessment of claims.
[0149] The disclosed systems and methods can identify products or services that may infringe claims of a patent and patents that may be infringed by a product or service and patent applications that may be relevant to a product or a service (e.g., patent applications containing claims, that if issued, could be infringed by a product or service. The disclosed systems and methods can be used for actively monitoring for patent infringement.
[0150] At least some of the embodiments described herein can identify products or services that may be relevant to a claim being drafted. For example, at least some of the embodiments described herein can enable a user, such as a patent practitioner to evaluate a draft claim drafted and / or amended (e.g., a mock claim) by the patent practitioner and identify products or services that would infringe the proposed, if the proposed claim was a granted patent claim to assist the patent practitioner in drafting or amending the claim. In this manner, the language of a claim may be adjusted prior to the filing of an application (e.g., a continuation application, a divisional application, a continuation-in-part application) or during its prosecution in order to take into account potential infringement targets.
[0151] At least some of the embodiments described herein can analyze a patent to identify its validity. For example, at least some of the embodiments described herein can enable a user, such as a patent practitioner, a patent owner, or a person seeking to invalidate a patent to evaluate whether the validity of a given patent may be challenged due to products or services that existed prior to the priority date of the patent.
[0152] The embodiments described herein can involve fine-tuning a large language model (LLM) used for identifying products or services that may infringe claims of a patent using a litigation dataset.
[0153] As used herein, the term “product” includes any subject matter that involves physical elements, including but not limited to, apparatuses, devices, systems, formulations that may be patented and is not limited to physical, software or any other type of products. The term “service” includes any subject matter that involves including methods, processes and that may be patented.
[0154] As used herein, the term “patent”, unless otherwise indicated, include utilitymodels.
[0155] As will be described, at least some of the embodiments described herein can determine a relevancy score that is indicative of the similarity between claims of a patent and a product or service, as described by a product or service description. The relevancy score can enable patent owners, product manufacturers or service providers to assess the risk of infringement. The disclosed systems and methods allow patent owners to automatically identify products that may infringe their patents and allow product manufacturers to automatically identify patents that may be infringed by their products, reducing the need for manual searches.
[0156] At least one of the embodiments described herein can determine a relevancy score that is indicative of a risk of a patent being invalidated.
[0157] At least some of the embodiments described herein can involve fine-tuning the LLM using a feedback dataset so that improved search results can be obtained.
[0158] Reference is first made to FIG. 1 , which illustrates an example block diagram 100 of a patent analytics and infringement monitoring system 108 in communication with an external data storage 102 and a computing device 106 via a network 104. Although only one computing device 106 is shown in FIG. 1 , the patent analytics and infringement monitoring system 108 may be in communication with a greater number of computing devices 106. The patent analytics and infringement monitoring system 108 communicate with the computing device(s) 106 over a wide geographic area via the network 104. While the patent analytics and infringement monitoring system 108 and the computing device 106 are shown as separate components, in some cases, the patent analytics and infringement monitoring system 108 or one or more components of patent infringement monitoring system 108 may be implemented within the computing device 106.
[0159] The patent analytics and infringement monitoring system 108 includes a storage component 110, a processor 112, and a communication component 114. The patent analytics and infringement monitoring system 108 can be implemented with more than one computer server distributed over a wide geographic area and connected via the network 104. The storage component 110, the processor 112 and the communication component 114 may be combined into a fewer number of components or may be separated into further components. The patent analytics and infringement monitoring system 108 can be configured to identify products or services that may infringe patent ofinterest or identify patents that may be infringed by a product or service, depending on the implementation of patent infringement monitoring system 108.
[0160] The patent analytics and infringement monitoring system 108 can be configured to rank the products or services, or patents, according to their relevance, as will be described in further detail below. For example, the patent analytics and infringement monitoring system 108 can implement a subsystem for ranking product / services-patent evaluation search results. The patent analytics and infringement monitoring system 108 can also include an LLM fine-tuning subsystem that can fine-tune one or more LLM subsystems for identifying products or services that may infringe a patent or interest.
[0161] The LLM subsystem(s) may use a pre-trained model or fine-tuned model (e.g. GPT-4 or similar) to process inputs from the computing device 106 via server. The LLM subsystem may process the inputs from the computing device 106, collect patent data and analyze it using the LLM subsystem(s), generate one or more search queries (for example, Google searches, requests to web servers, or requests to user remote devices). The results of said searches may be used to generate search results that may be provided in user interfaces (see e.g. FIG. 11 ) for display on the computing device 106. Detailed prompts and predefined crafted prompts provided by a server may enable the LLM subsystem(s) to adapt to the volume and type of input data without making any modifications to its base model. For example, prompts may be created to translate patent applications in non-English languages, to caption patent application drawings orformulae in patent documents, to parse the structure of different national patent documents, etc. The LLM subsystem(s) may include one or more Large Language Models (LLMs).
[0162] The prompts sent to the LLM subsystem(s) may include additional information identified in a vector database in the storage component 110 of the patent analytics and infringement monitoring system 108 prior to the prompt being executed against an LLM subsystem as described herein.
[0163] The LLM subsystem(s) may include a first party LLM trained and / or finetuned based on the litigation dataset (see e.g. FIG. 9), the feedback dataset, and other patent data or product or service catalogs or any other dataset as are known.
[0164] The one LLM subsystem(s) may include one or more of the following (it is noted that the following are non-limiting examples, and that other models may be used):GPT-2 made by OpenAI®GPT-3 made by OpenAI®GPT-3.5 Turbo made by OpenAI®GPT-4 made by OpenAI®GPT-4 Turbo made by OpenAI®GPT-4o made by OpenAI®Claude 1 made by Anthropic®Claude 2 made by Anthropic®Claude 3 made by Anthropic®LaMDA (Language Models for Dialog Applications) made by Google®LLaMA (Large Language Model Meta Al) made by Meta ®PaLM 2 (Pathways Language Model 2) Google®Llama 2 made by Meta®SDXL made by Stability Al®A BART trained model, i.e. , a model trained using a denoising autoencoderfor pretraining sequence-to-sequence models;ALBERT, A Lite BERT for Self-supervised Learning of Language Representations.
[0165] The processor 112 can be implemented with any suitable processor, controller, digital signal processor, graphics processing unit, application specific integrated circuits (ASICs), and / or field programmable gate arrays (FPGAs) that can provide sufficient processing power for the configuration, purposes and requirements of patent infringement monitoring system 108. The processor 112 can include more than one processor with each processor being configured to perform different dedicated tasks. For example, the processor 112 can include a one or more processors for fine-tuning the LLM subsystem, one or more processors for identifying products or services that may infringe a target patent claim (e.g., a claim of a patent of interest, a draft claim of a patent application) and / or one or more processors for ranking products or services identified as potentially infringing a target patent claim.
[0166] The communication component 114 can include any interface that enables the patent analytics and infringement monitoring system 108 to communicate with various devices and other systems. For example, the communication component 114 canreceive inputs (e.g., patent numbers, patent publication numbers, patent application numbers, utility model numbers, claims of interest including draft claims, URL, product or service description) from the computing device 106 and store the inputs in the storage component 110 or external data storage 102. The processor 112 can then process the inputs according to the methods described herein.
[0167] The communication component 114 can include at least one of a serial port, a parallel port or a USB port, in some embodiments. The communication component 114 may also include an interface to component via one or more of an Internet, Local Area Network (LAN), Ethernet, Firewire, modem, fiber, or digital subscriber line connection. Various combinations of these elements may be incorporated within the communication component 114. For example, the communication component 114 may receive input from various input devices, such as a mouse, a keyboard, a touch screen, a thumbwheel, a trackpad, a track-ball, a card-reader, voice recognition software and the like depending on the requirements and implementation of the patent analytics and infringement monitoring system 108.
[0168] The storage component 110 can include RAM, ROM, one or more hard drives, one or more flash drives or some other suitable data storage elements such as disk drives. The storage component 110 can include one or more databases for storing information extracted from patents, vector embeddings of patent claims, vector embeddings of product or service descriptions or product / service website contents, predefined search queries, search results, litigation datasets, user feedback information, a feedback dataset, previous product-patent pairs or service-patent pairs evaluated, models for extracting information from patents and / or product / service websites, models for generating search queries, search parameters or filters and ranking refinement models. The litigation dataset can be saved according to the method in FIG. 6. The database(s) may store a vector database populated with document embeddings based on the litigation dataset as described herein. The storage component 110 can include databases including a Structured Query Language (SQL) such as PostgreSQL® or MySQL® or a not only SQL (NoSQL) database such as MongoDB™. The vector database(s) can include a Pinecone database.
[0169] The external data storage 102 can store data similar to that of the storage component 110. The external data storage 102 can, in some embodiments, be used to store data that is less frequently used and / or older data. In some embodiments, the external data storage 102 can be a third-party data storage that stores data that can beanalyzed by the patent analytics and infringement monitoring system 108. For example, the external data storage 102 can store patents or patent data derived from patents (e.g., subject matter, industry, classifications). The external data storage 102, for example, can be the patent database of a national patent office, such as, for example, the United States Patent and Trademarks Office or commercially available databases, for example, by Questel™, LexisNexis™, ReedTech™, Google Patents™, etc.
[0170] The external data storage 102 can include data servers that store data uploaded on the internet. The data stored in the external data storage 102 can be retrieved by the computing device 106 and / or patent analytics and infringement monitoring system 108 via the network 104. For example, the patent analytics and infringement monitoring system 108 can retrieve data via remote systems 120 using search engines (e.g., Google®), retrieve websites locatable using search engines and / or data otherwise available from publicly available websites. Although only one external data storage 102 is shown, it will be understood that the external data storage can include multiple data storages that can be distributed over a wide geographic area. In at least one embodiment, the external data storage 102 is implemented as a cloud storage.
[0171] The remote systems 120 may include web servers and may be reachable using HTTP requests, or other network requests as known. Alternatively, the remote systems 120 may be located at a client venue (not shown). This may include a private network controlled by the user for example, the end-user of the patent analytics and infringement monitoring system 108. The user at computing device 106 may specify one or more remote systems 120 to be searched when a user submits a search request to the server. The remote system(s) 120 may include web servers and may be reachable using HTTP requests, or may be accessible using a connector to a proprietary client storage system. The remote system (s) 120 may be created by the user and may describe evidence of third-party products or services. The remote system(s) 120 may include public websites, websites for news publications, websites for product monographs, etc. The remote system(s) 120 can include one or more remote devices.
[0172] The computing device 106 can include any device capable of communicating with other devices through a network such as the network 104 through a wireless connection. The computing device 106 can include a processor and memory, and may be an electronic tablet device, a personal computer, workstation, server, portable computer, mobile device, personal digital assistant, laptop, smart phone, WAP phone, an interactive television, video display terminals, gaming consoles, and portableelectronic devices or any combination of these. The computing device 106 can be a user device that can transmit patent numbers of patents of interest, claims of interest, and / or product or service descriptions and search parameters and receive ranked search results.
[0173] The computing device 106 can be a user device used by an end-user such as a patent owner, a lawyer, a patent practitioner, an engineer, a product manufacturer, a product designer, or any other intended end-user of the patent analytics and infringement monitoring system 108 , i.e., an end-user of software application for evaluating patent analytics, including but not limited to patent portfolio review and analysis of potential infringing products or services, or infringed patents. The computing device 106 can be used to access a software application (not shown) running on a server (not shown) over network 104, which may be provided by a downloadable application on user device 102 that connects to the server using an API, or using a web-based application in a browser on computing device 106. The one or more computing devices 106 may be any two-way communication device with capabilities to communicate with other devices. A computing device 106 may be, for example, a mobile device such as mobile devices running the Google® Android® operating system or Apple® iOS® operating system. A computing device 106 may also be, for example, a personal computer operating the Windows® or MacOS® operating system.
[0174] The application provided to the end-user either via an application at computing device 106 or via a web interface from the server may include a user interface that accepts inputs (e.g., patent numbers, claims, product descriptions, service descriptions, other search parameters) in various forms, such as text or image-based inputs.
[0175] The software application running on the one or more computing devices 106 may communicate with the server using an Application Programming Interface (API) endpoint, and may send various inputs and other requests for processing by the patent analytics and infringement monitoring system 108.
[0176] The software application running on the one or more computing devices 106 may display one or more user interfaces on a display device of the user device computing device 106, including, but not limited to, the graphical user interface shown in FIGS. 10 and 11 . A browser may be used at the computing device 106 to access the web application running on the server.
[0177] The network 104 may be any network or network components capable of carrying data including the Internet, Ethernet, fiber optics, satellite, mobile, wireless (e.g. Wi-Fi, WiMAX), SS7 signaling network, fixed line, local area network (LAN), wide area network (WAN), a direct point-to-point connection, mobile data networks (e.g., Universal Mobile Telecommunications System (UMTS), 3GPP Long-Term Evolution Advanced (LTE Advanced), Worldwide Interoperability for Microwave Access (WiMAX), etc.) and others, including any combination of these.
[0178] Referring next to FIG. 2, there is shown a device diagram 200 of a server of the patent analytics and infringement monitoring system 108 of FIG. 1.
[0179] The network unit 204 can include wired or wireless connection capabilities. The network unit 204 can include a radio that communicates using standards such as IEEE 802.11 a, 802.11 b, 802.11g, or 802.11 n. The network unit 204 can be used by the server 200 to communicate with other devices or computers.
[0180] Network unit 204 may communicate with a network, such as network 104 (see FIG. 1).
[0181] The display 206 may be an LED or LCD based display and may be a touch sensitive user input device that supports gestures.
[0182] The processor unit 208 controls the operation of the server 200. The processor unit 208 can be any suitable processor, controller or digital signal processor that can provide sufficient processing power depending on the configuration, purposes and requirements of the server 200 as is known by those skilled in the art. For example, the processor unit 208 may be a high-performance general processor. In alternative embodiments, the processor unit 208 can include more than one processor with each processor being configured to perform different dedicated tasks. The processor unit 208 may include a standard processor, such as an Intel® processor or an AMD® processor. The processor unit 208 can be the processor 112 of FIG. 1 .
[0183] The memory unit 210 comprises software code for implementing an operating system 220, programs 222, databases 224, web scraper 226, LLM unit 228, user interface engine 230 and web / API unit 232.
[0184] The memory unit 210 can include RAM, ROM, one or more hard drives, one or more flash drives or some other suitable data storage elements such as disk drives, etc. The memory unit 210 is used to store an operating system 220 and programs 222 as is commonly known by those skilled in the art.
[0185] The I / O unit 212 can include at least one of a mouse, a keyboard, a touch screen, a thumbwheel, a trackpad, a trackball, a card-reader, an audio source, a microphone, voice recognition software and the like again depending on the particular implementation of the server 200. In some cases, some of these components can be integrated with one another.
[0186] The power unit 216 can be any suitable power source that provides power to the server 200 such as a power adaptor or a rechargeable battery pack depending on the implementation of the server 200 as is known by those skilled in the art.
[0187] The operating system 220 may provide various basic operational processes for the server 200. For example, the operating system 220 may be a server operating system such as Ubuntu® Linux, Microsoft® Windows Server® operating system, or another operating system.
[0188] The programs 222 include various user programs. They may include several hosted applications delivering services to users over the network, for example, a patent portfolio management system, or patent analysis system.
[0189] In one or more embodiments, the programs 222 may provide a patent analytics and infringement monitoring platform that is web-based, or client-server-based application via Web / API Unit 232 that provides for the patent analysis and infringement monitoring in conjunction with an application at a user device.
[0190] The databases 224 may be at least one of a relational database or a vector database. The databases 224 can be stored in the storage component 110 of FIG. 1. There may be multiple different databases 224. The database 224 may run on the same server as the other components of the patent searching system, or alternatively, may run on a different server in network communication with the other components of the patent searching system. The relational database may include a litigation dataset or / and a feedback dataset. The litigation dataset (for example, see FIG. 9) may include productpatent or service-patent associations, including content associated with a product or service subject to patent litigation. The content in relational database 224 may include content collected by web scraper 226. The content in the vector database 224 may include document embeddings associated with the content in the litigation dataset or the feedback dataset.
[0191] The databases 224 may index or store content ingested from remote system(s) 120 or retrieved from the external database 102, which may include a pluralityof user content including documents associated with a products or services.
[0192] While the databases 224 is shown as operating on a single server 200, it is understood that the database may be located separately on another server or system in network communication with the server 200. The databases 224 may store the user content from the remote system(s) 120 in an indexed format.
[0193] The product / service content items associated with litigation records in the litigation dataset in database 224 may be stored using a collection timestamp by the web scraper 226. The web scraper 226 may, on user request or as a function of a periodic polling of remote systems may make a request to a remote system 120 to capture the content (including both text content and other media content) from the remote system and store it in the database 224.
[0194] The product / service content items in database 224 may include text data, media data such as image data, and screenshot data associated with a browser rendering of a website at a remote system. The content entries may include language translations of text data performed by the LLM subsystem.
[0195] The product / service content items in database 224 may include content created by the LLM Unit 228 when an LLM response is received from an LLM subsystem of the patent analytics and infringement monitoring system 108. For example, where the product / service content item includes an image content item, the image may be transmitted to the LLM subsystem and a text-based caption may be received and stored. Additionally, where the product / service content item includes a text content item, a text translation may be produced by the LLM subsystem based on the language of the claims or based upon a user language preference.
[0196] The web scraper 226 may be used to populate candidate litigation records in the litigation dataset, in order to capture content for the litigation dataset in association with patent litigation. The collected content may be described in FIG. 9. The collected data may be used to generate a vector database, a relational database, a non-relational database, etc. as described herein, and for measuring the performance of the patent search results.
[0197] The web scraper 226 may be a software system for connecting to remote system 120. The web scraper 226 may be triggered on a regular schedule, or on request of a user using a user interface described herein. The web scraper 226 may further be triggered automatically by a user submitting updated product / service content items to thefine-tuning subsystem based on a candidate patent litigation record in the litigation dataset. The web scraper 226 may be a headless browser-based scraper, such as a scraper based on Chromium.
[0198] The web scraper 226 may further be used to scrape content when the user submits a search query. This may include accessing the search result items identified at a remote device or by a search engine (see e.g., 802 in FIG. 8), collecting the associated content items, and then storing them.
[0199] The web scraper 226 may be used to transmit a content request to a remote system 120, and store the results in the database 224. The content request may be configured based on a user submitted configuration information database 224, including for example, one or more URLs associated with the candidate litigation record. The one or more URLs may include a URL that links to a website, and one or more third-party URLs that are associated with other publications associated with a product or service that is associated with the candidate litigation record. The web scraper 226 may generate a hash value and store it in association with product / service content items. Subsequently, the web scraper 226 may compare a subsequent hash value of the content response with the prior acquired hash to identify if the content has subsequently been updated. The web scraper 226 may capture static content, image or other media content, JavaScript content, cascading style sheet content, or other web assets. The web scraper 226 may further store a screenshot of one or more documents.
[0200] The web scraper 226 may perform post-scraping processing. Upon scraping, processing may include categorizing the content data (e.g., text, images, videos), filtering irrelevant content, translating content, and captioning image or video content, etc. The web scraper 226 may similarly break the content down into chunks. For example, a large content item may be broken down into a plurality of smaller content items.
[0201] Images and other media collected by the web scraper 226 are stored in the database 224 may include captioning by the LLM system, enabling the media to be searchable and used when generating the vector database for fine-tuning. Similarly, images may be processed using an OCR program to store the text used within them. For example, PDFs may be scraped for their content and may have embedding data created in association with their content.
[0202] The LLM Unit 228 provides integration with the LLM subsystem. This mayinclude a variety of LLMs such as those described above, with reference to FIG.1
[0203] This submitted search query information may be received by the LLM Unit 228 and may be used to produce a prompt for the LLM subsystem. There may be a static portion of the prompt, and a dynamic portion of the prompt.
[0204] The patent document identified by the user in their search query submission may be requested, including claims, drawings, descriptions, abstracts, forward and backward citations, charts, formulae, inventor information, assignee information, classification information, prosecution file wrapper, maintenance payment information, litigation history etc. The drawings may be submitted using the LLM Unit 228 for captioning by the LLM subsystem so that there is text-based information that summarizes the drawings. This information may be stored in a structured format in the database 224 in association with a search query record for the user’s search.
[0205] The LLM Unit 228 may generate a prompt based on the search query submitted by the user as described in FIG. 8. The prompt may include a static and a dynamic portion. The dynamic portion may be generated using the vector database, and used to fine-tune the LLM subsystem (e.g. based on the vector database generated based on FIG. 6).
[0206] The LLM Unit 228 may connect to different LLM subsystems as described above with reference to FIG. 1 .
[0207] The user interface engine 230 is configured to generate interfaces for users to perform patent searches (including, but not limited to patentability, invalidity, freedom- to-operate (FTO, infringement, and others), to review search results including providing search result feedback, to manage configuration of the patent analytics and infringement monitoring system 108, etc. The various interfaces generated by the user interface engine 230 may be transmitted to a user device by virtue of the Web / API Unit 232 and the network unit 204.
[0208] The processor unit 208 can also execute a user interface (Ul) creation engine 230 that is used to generate various Uls for delivery via a web application provided by the Web / API Unit 232, some examples of which are shown and described herein, such as interfaces shown in FIGS. 10 and 11 .
[0209] The Web / API Unit 232 may be a web-based application or Application Programming Interface (API) such as a REST (Representational State Transfer) API. The API may communicate in a format such as XML, JSON, or other interchange format.
[0210] The Web / API Unit 232 may receive a user request for a patent search or other functions of the system. The Web / API Unit 232 may apply methods herein to provide the search results in a user interface as provided herein (for example, FIG. 11).
[0211] Referring next to FIG. 3, there is shown a high-level method diagram 300 for servicing a user search query in accordance with one or more embodiments. The details of the inputs, search universe, and outputs of each of the different types of searches is summarized in FIG. 4.
[0212] At 302, a user provides the patent document identifier (e.g., patent number, patent application publication number) for searching to the patent analytics and infringement monitoring system 108. For example, the patent document identifier can be provided by the user via the computing device 106. This could include the user entering the patent document identifier of a patent of interest into a graphical user interface such as FIGS. 10-10B. In some embodiments, the user provides more than one patent document identifier. For example, the user may wish to identify products or services relevant to multiple patent documents. In such cases, the patent analytics and infringement monitoring system 108 can receive multiple patent document identifiers and conduct separate searches for each patent document and / or patent claim. In some embodiments, patent analytics and infringement monitoring system 108 can receive a selection of one or more claims of the patent document associated with the patent document identifier. For example, the user may wish to only assess specific claims of a patent document.
[0213] In some embodiments, in addition to the patent document identifier, the user can input the text of a claim of the patent of interest, or the text of a claim of an application of interest. In some embodiments, the user can input the text of a mock claim. For example, a patent attorney or other patent professional may wish to use the patent analytics and infringement monitoring system 108 as a claim drafting and enforcement strategy tool, in order to verify the content and scope of a proposed claim as against target products or services identified by the system 108. In this manner, the language of a claim may be adjusted prior to the filing of an application or during its prosecution in order to take into account potential infringement targets. The submission of the patent document identifier triggers data collection surrounding the patent document identified. This information can include additional information associated with the patent document, images, prosecution file history, etc. The submitted images in the search query may be captioned as described herein, and the captioned text may be used as additionalinformation searched in conjunction with the patent document identifier. In some embodiments, the patent analytics and infringement monitoring system 108 can provide an assignee search tool that enables a user to search for patent documents assigned to a particular assignee. In such embodiments, the patent analytics and infringement monitoring system 108 can display a list of patent documents assigned to a particular assignee and receive a selection of one or more patent documents.
[0214] At 304, search results are provided to the user based on the query. For example, if the user requests an infringement search, then the search results will include web pages or other publications identified as featuring or describing products or services that potentially infringe the patent of interest. This may include references available on the public internet, or optionally at a remote system 120 at a client venue.
[0215] In the case that a validity search is requested, the search results output may include a listing of publications that are known to have been published prior to the priority date of the user submitted patent document. For example, when a validity search is requested, the patent analytics and infringement monitoring system 108 can determine the priority date of the patent of interest and analyze search results published prior to the priority date of the patent of interest. When an infringement search is requested, the patent analytics and infringement monitoring system 108 can determine the application date or the issue date, depending on the implementation of the patent analytics and infringement monitoring system 108 and analyze search results published after the relevant date.
[0216] The search results provided at 304 may be explained to the user using a request to the LLM subsystem via the LLM unit, as will be explained in further detail with reference to FIGS. 11 and 16.
[0217] At 306, a user reviewing the search results may provide user feedback in association with a particular search result in the list. This may be accomplished for example using a thumbs-up and thumbs-down feedback button as described herein.
[0218] At 308, the user may set a periodic trigger that may monitor an infringement search. This may trigger email notifications or application notifications for newly identified search results that may infringe.
[0219] Referring next to FIG. 5, there is shown another high-level method diagram 500 of a method of operating the patent analytics and infringement monitoring system 108, in accordance with one or more embodiments.
[0220] At 502, litigation dataset construction, and vector database creation occurs.This is described in further detail in FIG. 6.
[0221] At 504, the vector database is used to fine-tune the operation of the LLM subsystem by generating dynamic portions of prompt requests to the LLM subsystem. The dynamic portions of the prompt may provide examples for the LLM subsystem that narrow or focus the LLM subsystem appropriately to provide improved output. This may improve the operation of the LLM subsystem in order to service a user search query 508 for a patent search. This is described in further detail in FIG. 7.
[0222] At 506, the search results are generated and provided to the user. The user may provide feedback to the LLM subsystem, which may in turn be stored in the database. This is described further in FIG. 8.
[0223] Referring next to FIG. 6, there is shown another method diagram 600 for fine-tuning a machine learning model in accordance with one or more embodiments. At least some of the steps of method 600 can be performed by a processor of the patent analytics and infringement monitoring system 108, for example, processor 112 or a processor of processor unit 208.
[0224] At 602 a litigation dataset is provided in a memory for example the memory unit 210, the dataset comprising a plurality of historical patent litigation records, each patent litigation record comprising at least one patent claim, and a litigation outcome corresponding to each of the at least one patent claim.
[0225] At 604, a product / service content item associated with a candidate patent litigation record in the plurality of historical patent litigation records is received at the processor in communication with the memory.
[0226] At 606, the candidate patent litigation record is updated at the processor based on the product / service content item.
[0227] At 608, a vector database based on the litigation dataset is generated at the processor, the vector database used to fine-tune a machine learning model.
[0228] Optionally, in some embodiments, at 610 a prediction prompt is generated at the processor for an LLM-subsystem based on the vector database.
[0229] Optionally, in some embodiments, at 612, the prediction prompt is sent using a network device in communication with the processor to the LLM subsystem.
[0230] Optionally, in some embodiments, at 614, a prompt response that includesat least one search element is received using the network device.
[0231] Optionally, in some embodiments, a prompt response including at least two search elements is received using the network device.
[0232] In some embodiments, the method 600 involves receiving, at a network device in communication with the processor, a user submitted product / service content item associated with a product or service in the candidate patent litigation record.
[0233] In some embodiments, the content item includes a user-submitted URL content item corresponding to the product or service in the candidate patent litigation record and method 600 further involves sending, using a network device in communication with the processor, a content request, receiving, using the network device, a content response based on the content request, performing web scraping on the content response to identify at least one image content item and at least one text content item, and updating the candidate patent litigation record by storing the at least one image content item and the at least one text content item in association with the candidate patent litigation record. The content request can be based on a URL content item.
[0234] In some embodiments, the content item includes a user-submitted text content item corresponding to the product or service in the candidate patent litigation record, and method 600 includes storing the user-submitted text content item in association with the candidate patent litigation record.
[0235] In some embodiments, method 600 includes sending, using the network device, a captioning request to an LLM-subsystem, receiving, using the network device, a captioning response from the LLM-subsystem based on the captioning request and updating, at the processor, the candidate patent litigation record based on the at least one text-based caption. The captioning request can include the at least one image content item and the captioning response includes at least one text-based caption based on the at least one image content item.
[0236] In some embodiments, method 600 further includes determining, at the processor, a hash value corresponding to the content response, updating, at the processor, the candidate patent litigation record with the hash value; and subsequently sending, using a network device in communication with the processor, a second content request based on the URL content item; receiving, using the network device, a second content response based on the second content request; determining, at the processor,that a second hash value corresponding to the second content response is different from the hash value corresponding to the content response; and updating, at the processor, the candidate patent litigation record based on the second content response.
[0237] In some embodiments, the vector database is generated by generating an embedding for the product / service content item of the candidate patent litigation record.
[0238] In some embodiments, generating the vector database involves generating a plurality of vector embeddings for a corresponding plurality of chunks of the product / service content item of the candidate patent litigation record.
[0239] In some embodiments, the vector embedding is one text-embedding-3- small / large, text-embedding-ada-002, BAAI / llm-embedder and / or BAAI / bge-large-en.
[0240] In some embodiments, each patent litigation record in the dataset includes a corporate identification, an estimated damages quantum or any combination thereof.
[0241] In some embodiments, method 600 involves generating, at the processor, at least two machine learning models based on the vector database and updating, at the processor, each of the at least two search elements based on the at least two machine learning models. The at least two machine learning models can include a ranking model and a classification model, and each updated search element can include a rank and a classification.
[0242] In some embodiments, the vector database may be Pinecone™.
[0243] Referring next to FIG. 7, there is shown another method diagram 700 in accordance with one or more embodiments.
[0244] At 702, the user of the patent analytics and infringement monitoring system 108 may enter a patent document identifier into a user interface such as the one in FIGS. 10A-10B. The identifier may be a patent publication number, an issued patent number, a design patent registration number, a patent application number, etc. The patent document identifier may be for the US, Canada, Europe, or other patent jurisdictions as known.
[0245] The user submission of the patent document identifier in FIGS. 10A-10B may correspond to a search query submission in the system 108, i.e. the user wants to identify patent analytics, provide an invalidity search, or search for possible infringing products or services of a patent document or a mock claim. Once submitted, the user may select one or more patent claims associated with the patent document identifiedusing the patent document identifier. The user may also separately provide specific information associated with the search query, such as a company name (such as a competitor), additional keywords, additional descriptive terms, geographical information, etc.
[0246] The user of the patent analytics and infringement monitoring system 108 may enter a patent document identifier into a user interface such as the one in FIGS. 10A-10B. The identifier may be a patent publication number, an issued patent number, a design patent registration number, a patent application number, etc. The patent document identifier may be for the US, Canada, Europe, or other patent jurisdictions as known.
[0247] The user submission of the patent document identifier in FIGS. 10A-10B may correspond to a search query submission in the system 100, i.e. the user wants to identify patent analytics, provide an invalidity search, or search for possible infringing products or services of a patent document. Once submitted, the user may select one or more patent claims associated with the patent document identified using the patent document identifier. The user may also separately provide specific information associated with the search query, such as a company name (such as a competitor), additional keywords, additional descriptive terms, geographical information, etc.
[0248] Alternatively, rather than selecting one or more patent claims, in some embodiments, the user can input the text of one or more mock claims (e.g., draft claims, draft amended claims) as free-form text. The one or more mock claims may be assigned names by the user.
[0249] The user may select what type of search, e.g., validity or infringement searching.
[0250] At 704, analysis of the user-submitted patent document may be performed. This can include using the patent document including the drawings, the specifications, the one or more claims. As noted, other information may be used to build an analysis prompt for the LLM Unit 228 to send to the LLM subsystem of the patent analytics and infringement monitoring system 108. This could include incorporating patent data related to the abstract, forward and backward citations, charts, formulae, inventor information, classification information, prosecution file wrapper, assignment, citations, classes, maintenance fee information, remaining patent term, potential users of the system, potential industries, keywords, etc. The analysis prompt may include a static portion anda dynamic portion. The dynamic portion may include patent data including the claim text of the patent. The static portion may be for example “you are a patent attorney, summarize the above text in 10 words, provide the answer in 4 programmatic JSON file in following structure {"summary": "summarized example"}”. As noted above, the drawings may have a textual caption identified by the LLM Unit 228 using the LLM subsystem so that there is text-based information that summarizes the drawings. This information may be stored in a structured format in the database 224 in associate with a search query record for the user’s search.
[0251] The fine-tuning of the LLM-subsystem may involve submitting matching contents of the vector database during the generation of the analysis prompt. The matching contents identified in the vector database may include records from the litigation dataset. The matching may rely on similarity matching within the litigation dataset, and the results may include litigation records that are similar to the user submitted query. These matching records may be included in the analysis prompt that is sent to the LLM subsystem as examples.
[0252] Citation data may be further evaluated and incorporated into the analysis prompt. This could include separate patent data associated with the cited documents.
[0253] The prompt may provide for keywords to be extracted by the LLM subsystem.
[0254] At 706, a prompt is generated by the LLM Unit 228 for the LLM subsystem. This includes sending a request to the LLM subsystem using the analysis prompt from 704 to generate queries for searching. Two approaches may be used. The first may be a heuristic approach that builds search queries using specific rules including extracted keywords. The search queries may be created for use with Google Search, Bing Search, or may be created to run against the remote devices at the user venue.
[0255] The second approach includes LLM subsystem-generated search queries for searching based on the instructions in the analysis prompt.
[0256] Referring next to FIG. 8, there is shown another method diagram 800 of using the patent analytics and infringement monitoring system 108, in accordance with one or more embodiments. At least some of the steps of method 800 may be performed by a processor of the patent analytics and infringement monitoring system 108, for example, processor 112 or a processor of the processing unit 208.
[0257] At 802, the search queries generated at 706 are received and executedagainst a search engine (either a public one like Google, or alternatively, against a private search index accessible on a remote device at a client venue).
[0258] The results from the search may be accessed by the web scraper 226 (e.g. , FIG. 2), and the contents of each search result item may be collected and stored in the database. The image content of each URL associated with the search items may be captioned by the LLM Unit 228 as described herein, or translated as described herein.
[0259] The results from the search may be filtered to remove advertising, noncommercial links, blacklisted domains, user-chosen sites or categories of links.
[0260] At 804, ranking of the search results may be performed using a machine learning model. The ranking can include, for each search result item, using a machine learning model with input from the patent claim and supporting information, product / service content items, identifying and ranking within the set of search results to enable.
[0261] At 806, classification of the search results may be performed using a machine learning model. The classifying step may classify each search result into different entities or categories.
[0262] At 808, the search results from classifying step 806 and the ranking step 804 are collected and combined.
[0263] At 810, different actions may be offered to the user on the search results page, including feedback buttons.
[0264] At 812, users may select or click the feedback buttons, for example, thumbs-up and thumbs-down buttons in the search results or a classification of the search results (e.g., high relevance, medium relevance, low relevance) . This selection sends a request back to the server with the associated feedback and a search result identified. The request, including the feedback, the analysis prompt, the patent claims, and the particular selected search result may be stored in the feedback dataset in the database.
[0265] Referring next to FIG. 9, there is shown a data diagram 900 in accordance with one or more embodiments. FIG. 9 shows the litigation dataset. FIG. 9 may be a ground-truth dataset for patent litigation outcomes.
[0266] The litigation dataset may include both positive (where infringement was found by a decision maker such as a judge or jury) and negative examples (whereinfringement was not found by a decision maker such as a judge or jury).
[0267] Column 902 may be provided in the litigation dataset, including the year of a patent litigation record.
[0268] Column 904 may be a docket citation or case title of the court of the patent litigation record.
[0269] Column 906 may be a case title or style of cause including the party names of the patent litigation record.
[0270] Column 908 may be the plaintiff party of the patent litigation record.
[0271] Column 910 may be the defendant party of the of the patent litigation record.
[0272] Column 912 may identify whether the patent litigation record includes a finding of contributory infringement.
[0273] Column 914 may associate the patent litigation record with a patent document identifier.
[0274] Columns 916, 920 and 924 may be the first, second, and third product / service content items (in this case, three URL content items).
[0275] Column 918, 922 and 926 may be corresponding sources of the product / service content items 916, 920 and 924 respectively.
[0276] Other fields in the litigation dataset can include a quantum of damages, patent claims, and individual decision result for infringement of each of the corresponding claims. Separately, data and content from Markman hearings may be provided as well.
[0277] Referring next to FIG. 10A, there is shown a user interface 1000A in accordance with one or more embodiments.
[0278] A user may access the user interface 1000A either using a browser to connect to the server 106, or through an application running on a user device.
[0279] The user enters a patent document identifier in the text input 1002 (for example, US patent US1234567B2), and then may submit the search query by hitting enter or by clicking the search button. Alternatively, referring to FIG. 10B, which shows a user interface 1000B in accordance with one or more embodiments, the patent document identifier may correspond to a patent application publication number and the user may enter the text of a mock claim.
[0280] Referring next to FIG. 11 , there is shown another screenshot of a graphical user interface diagram 1100 of a system for monitoring patent infringement, in accordance with one or more embodiments.
[0281] Upon submitting the search query in FIG.S 10A, 10B, the patent search system may return search results 1102a-1102f. For example, the search query US1234567B2 may return search results 1102a through 1102f as well as others. The search results may be websites, including public websites, or websites hosted on a remote device at a client venue.
[0282] Each search result 1102a-1102f may be presented with feedback buttons or controls that can allow a user to rate the results. For example, an upvote (e.g., thumbs- up) 1106 and a downvote (e.g., thumbs-down) 1104 button may be displayed. The use may click on, or select these buttons in order to generate a request back to the server 106. The feedback request may be stored in the feedback dataset in the database. The ratings can be used to train one or more of the ranking refinement models, as will be explained further below. FIG. 21 shows an alternative embodiment of a graphical user interface of the patent analytics and infringement monitoring system 108, showing feedback buttons.
[0283] Referring next to FIG. 12, there is shown another a flowchart of a method 1200 for providing updated search results by fine-tuning of machine learning models using a feedback dataset, in accordance with one or more embodiments. The method 1200 may be performed by a processor of the patent analytics and infringement monitoring system 108, for example, processor 112 or a processor of the processing unit 208.
[0284] At 1202, a search query comprising at least one of a patent document identifier or a product or service identifier is received at a network device from a user device.
[0285] At 1204, a search response user interface comprising a search identifier and at least one search result item is transmitted from the network device to the user device.
[0286] At 1206, a feedback item corresponding to a candidate search item in the at least one search result item is received at the network device from the user device.
[0287] At 1208, a feedback item record in a feedback dataset is stored in a memory, the feedback item record stored in association with the search identifier, thefeedback item record comprising the feedback item, at least one of a patent claim associated with the patent document identifier or a product or service description, and the search result item.
[0288] At 1210, at least one update search result item is transmitted from the network device to the user device.
[0289] In some embodiments, method 1200 can involve extracting, at the processor patent information associated with the patent document identifier; and generating, at the processor, an initial analysis prompt based on the patent information.
[0290] In some embodiments, method 1200 can involve extracting, at the processor, product or service information associated with the product or service description; and generating, at the processor, an initial analysis prompt based on the product or service information.
[0291] In some embodiments, the patent information associated with the patent document identifier includes: claim text, drawing images, caption text for the drawing images, description text, abstract text, forward or backward citations, each citation including a patent document identifier, chart images, caption text for the chart images, formulae text, inventor information; assignee information, classification information, prosecution file wrapper text, maintenance payment information, or litigation history text or any combination thereof.
[0292] In some embodiments, the product or service information associated with the product or service identifier may include a text snippet, a URL, an icon, an image, a citation or any combination thereof.
[0293] In some embodiments, method 1200 involves transmitting, from the network device to an LLM subsystem, an initial analysis request comprising the initial analysis prompt; receiving, from the network device from the LLM subsystem, an initial analysis response based on the initial analysis request, transmitting, from the network device to a remote device, the at least one remote device search query; receiving, from the network device to the remote device, at least one remote device search response; and transmitting, from the network device to the user device, the at least one search result item based on the at least one remote device search response. The initial analysis response includes at least one remote device search query.
[0294] In some embodiments, the method 1200 may further include: generating, at the processor, a subsequent analysis prompt based on the initial analysis prompt andthe feedback item record, transmitting, from the network device to an LLM system, a subsequent analysis request comprising the subsequent analysis prompt, receiving, from the network device from the LLM subsystem, a subsequent analysis response based on the subsequent analysis request, transmitting, from the network device to the remote device search provider, the at least one subsequent remote device search query, receiving, from the network device to the remote device search provider, at least one subsequent remote device search response, and transmitting, from the network device to the user device, at least one subsequent search result item based on the at least one subsequent remote device search response. The subsequent analysis response can include at least one subsequent remote device search query.
[0295] In some embodiments, the feedback includes a rating corresponding to a feedback control selected by the user.
[0296] In some embodiments, the feedback control includes two or more button controls, a radio button control, a slider control or a star selector control.
[0297] In some embodiments, method 1200 includes: providing, in a memory such as memory unit 210, the feedback item dataset comprising a plurality of historical feedback records. Each feedback item record can be stored association with a corresponding search identifier and the feedback item record can include a historical feedback item, at least one of a patent claim associated with a corresponding patent document identifier or a product or service description, and a corresponding search result item.
[0298] In some embodiments, the method 1200 may further include: generating, at the processor, at least two machine learning models based on the feedback dataset and updating, at the processor, each of the at least one search result item based on the at least two machine learning models. The at least two machine learning models can include a ranking model and a classification model and each of the at least one search result item can include a rank and a classification.
[0299] Referring next to FIG. 13, there is shown another data diagram 1300 in accordance with one or more embodiments. The feedback requests may be stored in the feedback dataset that is stored in the database. Each row or record may identify a single feedback item, including a search identifier field 1302 (which may be a hash or other identifier associated with a particular search submitted by a user), claim index 1304 (which may correspond to a particular claim of a patent document identified by a user ina search query), a feedback identifier 1306 (which may uniquely identify the feedback that was submitted), and the rating provided by the user in the search (for example, 1 may correspond to a thumbs up, and -1 may correspond to a thumbs down feedback control submission). Various ratings may be stored in association with the different types of controls, for example, 0 to 5 may be stored for a star rating feedback control showing between 0 and 5 star buttons.
[0300] Referring next to FIG. 14, shown therein is a flowchart of a method 1400 for identifying products or services that may infringe a patent claim. As explained above, the disclosed systems and methods can identify products or services that may infringe patent claims, including mock patent claims, and determine a relevancy score that is indicative of the similarity between patent claims and a product or service, as described by the product or service description. The method 1400 can be performed by a processor of the patent analytics and infringement monitoring system 108, such as processor 112 or a processor of the processing unit 208. In at least one embodiment, method 1400 can be performed to identify products or services that may invalidate a patent. As explained above, the patent analytics and infringement monitoring system 108 can be used to assess whether a patent may be challenged for invalidity.
[0301] At 1410, the processor 112 receives a patent document identifier. The patent document identifier can be a patent number. The patent number can for example, be inputted by a user via a graphical user interface of the computing device 106 and transmitted to the processor 112 via the network 104. In at least one embodiment, the processor 112 receives an application number or an application publication number instead of a patent number. The processor 112 can retrieve the patent document from a database such as the external data storage 102.
[0302] The processor 112 can obtain a target patent claim. In at least one embodiment, as explained above, the processor 112 receives the text of a claim in freeform text. For example, the processor 112 can be configured to receive a natural language input. The text of the claim can be the text of a claim being drafted by a patent practitioner and that the patent practitioner seeks to evaluate, or the text of a claim of an issued patent or pending patent application. Alternatively, the processor 112 can retrieve one or more claims from the patent document identified by the patent document identifier.
[0303] In at least one embodiment, the processor 112 is configured to receive one or more search parameters from the computing device 106. The search parameters caninclude but are not limited to, a domain, a jurisdiction, a type of file, keywords, competitor names, a date range, an assignee and / or a selected patent claim. For example, the processor 112 can receive a selection of a patent claim of a patent document. For example, the user can specify a specific claim when inputting the patent identifier or provide the text of a claim.
[0304] At 1420, the processor 112 parses the contents of the patent document associated with the patent document identifier to extract information about the patent document. The processor 112 can extract the text, the formulas and the drawings of the patent document and extract bibliographic information about the patent document (e.g., classification, citations, assignee). Based on the contents of the patent document, the processor 112 can determine potential applications of a product or service targeted by the patent claim, product / service industries and / or potential users of a product / service targeted by the patent and / or the processor 112 can extract keywords and entities from the claims and / or the description, extract a textual representation of the unique features and / or concepts of the patent. The processor 112 can generate a claim summary of the claims of a patent document when the patent document contains more than one selected claim.
[0305] At 1430, the processor 112 generates a set of search queries for locating products or services that may infringe the patent claim. The processor 112 can generate a set of search queries for each claim of the patent or for each independent claim of the patent. The processor 112 can generate the set of search queries based on the information extracted at 1420 and in some cases, the search parameters received at 210. The processor 112 can employ a trained machine-learning model such as a large language model (LLM) to generate search queries. In at least one embodiment, the processor generates a guidance prompt for the LLM.
[0306] Alternatively, or in addition thereto, the processor 112 can generate the search queries according to predefined search query rules, using the information extracted at 1420 and in some cases, the search parameters received at 1410. For example, predefined search query rules can be defined for various keywords, industries, competitors, product or service or method names if applicable, and / or images. The storage component 110 and / orthe external data storage 102 can store predefined search query rules relating to the information extracted.
[0307] At 1440, the processor 112 searches the Web for websites containingproduct or service descriptions, using the set of search queries generated at 1430. In at least one embodiment, prior to searching the Web, the processor 112 can modify the format of the search queries according to the search engine used. The processor 112 can execute the searches in parallel, using the set of search queries.
[0308] In at least one embodiment, when searching the Web, the processor 112 can apply one or more filters. For example, the processor 112 can apply the search parameters received at 1410 and / or exclude web domains that have been determined to be unlikely to yield useful results. Alternatively, the processor 112 can filter the search results to exclude search results that have been determined to be unlikely to pose a risk of infringement. For example, the processor 112 can exclude non-commercial, academic websites and / or predetermined domains.
[0309] At 1450, the processor 112 scores and ranks the search results located at 1440, as will be described in further detail with reference to FIG. 16.
[0310] Referring next to FIG.15, shown therein is a flowchart of a method 1500 for identifying one or more patents that may be infringed by a product or service. The method 1500 can be performed by the processor 112 of the patent analytics and infringement monitoring system 108. The method 1500 can mirror the method 1400.
[0311] At 1510, the processor 112 receives a product or service description. Similar to 1410, the product or service description can be inputted by a user via a graphical user interface of the computing device 106 and transmitted to the processor 112 via the network 104. The processor 112 can parse the product or service description and extract information from the product or service description, such as keywords, industries, competitors, product / service names and / or images.
[0312] At 1520, the processor 112 searches a database of patents, such as the external data storage 102 using the product or service description received at 1510. For example, the processor 112 can search the database using the extracted information. In at least one embodiment, the database of patents is a database of vector embeddings of claims, as will be explained in further detail with reference to FIG. 16.
[0313] At 1530, the processor 112 scores and rank the results located, as will be described in further detail with reference to FIG. 16.
[0314] Referring next to FIG. 16, shown therein is a flowchart of a method 1600 for ranking search results obtained by a product / service-patent evaluation subsystem of the patent analytics and infringement monitoring system 108. The method 1600 can beimplemented by a processor, such as the processor 112 of the patent analytics and infringement monitoring system 108 or by a processor of a separate system for ranking search results that receives search results located by the patent analytics and infringement monitoring system 108. Method 1600 can be implemented at step 1450 of method 1400 or step 1530 of method 1500.
[0315] At 1610, the processor receives the search results. The search results can include patents, patent claims, URLs of product or service websites, or contents of product or service websites, depending on the implementation of the patent analytics and infringement monitoring system 108. In embodiments where URLs of product or service websites are received, the processor can scrape the contents of the product or service websites. Similarly, in embodiments where patents are received, the processor can parse the patents to extract the claims. Alternatively, in some embodiments, the processor can search a database of vector embeddings of claims, for example storage component 110 or databases 224. In such embodiments, prior to method 1600, the processor can parse patent documents stored in a database such as the external data storage 120 and transform each patent claim of the patent documents stored in the database into its vector embedding.
[0316] At 1620, the processor compares each located product / service-patent pair. In at least one embodiment, the processor determines a relevancy score for each located product / service-patent pair. For example, in embodiments where the product / service- patent evaluation system 108 is used to locate products or services that may infringe a patent, the processor determines a relevancy score for each of the located products or services. The relevancy score can correspond to a measure of similarity between the located products or services and a claim of the patent of interest or a mock claim, and can be indicative of a risk of infringement. As another example, in embodiments where the patent analytics and infringement monitoring system 108 is used to locate patents that may be infringed by a product or service, the processor determines a relevancy score for each claim of the patents located. The relevancy scores can be used by the processor to rank the search results. For example, a relevancy score approaching 1 can correspond to a highly relevant search result while a relevancy score approaching 0 can correspond to a search result that is unlikely to be relevant. The processor can rank the search results according to the relevancy scores. In embodiments where the product / service-patent evaluation system 108 is used to locate patents that may be infringed by a product or service and the patents have multiple claims, the processor can rank each patentaccording to the relevancy score of the claim having the highest relevancy score.
[0317] The processor can determine the relevancy score by comparing vector embeddings. The processor can transform the claims of each of the patents and the product / service information derived from the product or service websites into vectors that represent the semantics of the claims and the contents of the product or service websites. The processor can use any vector embedding model known to those skilled in the art, including, but not limited to text-embedding-3-small / large, text-embedding-ada- 002, BAAI / llm-embedder, and BAAI / bge-large-en.
[0318] The processor can embed each independent claim of a patent document into a vector representation. Alternatively, the processor can embed each independent claim and its associated dependent claims into a vector representation. In at least one embodiment, the processor embeds each independent claim, its dependent claims and portions of the description and optionally, equations, chemical formulations, sequence listings, tables and / or figures corresponding to the claims as a vector. For example, the processor can process the patent, including the description, equations, chemical formulations, sequence listing tables (if applicable), and figures of the patent to identify portions of the patent that correspond to the claims so that the vector representation of the claims can include information for interpreting the language of the claims. The processor can employ one or more models for converting non-textual information into textual data, for example visual transformer models and other large language models and for identifying portions of the patent that correspond to the claims.
[0319] The vector embeddings are then stored in a database such as the storage component 110 or databases 224, where each index represents a different product or service or patent claim, depending on the implementation of the patent analytics and infringement monitoring system 108. In embodiments where the patent analytics and infringement monitoring 108 identifies products or services that may infringe a claim of a patent or that may be relevant to a mock claim, the processor compares the vector embedding corresponding to the claim with the vector embeddings of each of the products or services, stored in the database. In embodiments where the patent analytics and infringement monitoring 108 identifies patent claims that may be infringed by a product or service, the processor compares the vector embedding of the product or service description with the vector embeddings of each of the patent claims, stored in the database. To compare the vectors, the processor can compute the distance between the vectors. For example, the processor can compute the cosine distance, the Euclidiandistance, the Hamming distance, or any other measure of distance between vectors.
[0320] In at least one embodiment, the processor identifies the K most relevant results, based on the relevancy scores, where K is a predetermined number.
[0321] At 1630, the processor applies one or more ranking refinement models to obtain refined relevancy scores. In embodiments, where the processor identifies the K best results, the processor can apply the one or more ranking refinement to the K best results identified at 1620. By only applying the ranking refinement models to a subset of the search results, the processor can avoid using processing resources to evaluate search results that are unlikely to be relevant. At least some of the ranking refinement models can be models that require more computing resources than vector comparison and accordingly, by applying the ranking refinement models to search results that have been pre-ranked through vector comparison, the processor can save processing resources when compared to applying the ranking refinement models without first preranking the search results. By refining the relevancy scores of the search results, the accuracy of the ranking of the search results can be improved.
[0322] In at least one embodiment, the one or more ranking refinement models are machine-learning models that are trained to evaluate similarities between product or service information and patent claims. For example, one or more ranking refinement model can be a text classification machine-learning model.
[0323] In at least one embodiment, at least one of the ranking refinement models is a transformer model, such as, but not limited to, a BERT classifier, which are particularly effective for tasks that require understanding the relationship between two pieces of text, such as text entailment or paraphrase identification. The BERT classifier can be fine-tuned on a dataset containing text-summary pairs or claim-description pairs and trained to train to score the pairs based on their similarity. Alternatively, or in addition thereto, at least one of the ranking refinement models is an embedding model, such as a BGE model, which are well-suited for retrieval tasks and text matching. An embedding model can generate embeddings for product or service descriptions and patent claims and then rank the descriptions based on their similarity to the claim embeddings. The ranking refinement models can determine a qualitative or quantitative measure of similarity between the text of a product or service description (e.g., an inputted product or service description, a product or service description extracted from a website) and the text of a claim. In at least one embodiment, the ranking refinement models can determinethe measure of similarity based additionally on the information extracted from the patent, for example, drawings, portions of the description, formulas, equations and / or tables where applicable. In at least one embodiment, the drawings can be converted to a textual representation and the ranking refinement models can compare the textual representation of the drawings with the text of the product or service description. Alternatively, the ranking refinement models can be models that can compare drawings with images of a product or service (e.g., an image on a website).
[0324] In at least one embodiment, at least one of the ranking refinement models is a LLM that determines a measure of similarity between product or service information and patent claims based on the patent claims, the product or service information and guidance prompts generated by the processor, that provide guidelines for the ranking refinement model. For example, the guidance prompts can configure the ranking refinement model to consider particular keywords, based on the subject matter of the patent, as determined by the patent analytics and infringement monitoring system 108. As another example, the guidance prompts can configure the ranking refinement model to favor search results associated with a particular industry or associated with particular companies. In at least one embodiment, the processor can generate the guidance prompts based in part on the search parameters received at 1410. In at least one embodiment, the guidance prompts can include feedback received from the user, as will be explained in further detail below.
[0325] In at least one embodiment, the processor selects the one or more ranking refinement models to be applied based on one or more factors, including but not limited to, the subject matter of the patent or of the mock claim as determined by the patent analytics and infringement monitoring 108, the patent classifications of the patent and the confidence level of the ranking refinement models.
[0326] In at least one embodiment, the processor applies two or more ranking refinement models to the search results, according to weights determined by pre-defined rules, for example based on the subject matter of the patent as determined by the patent analytics and infringement monitoring system 108 and the patent classifications of the patent. For example, it may be determined that a first ranking refinement model is more accurate than a second ranking refinement model for determining similarities between product or service information and patent claims when the subject matter relates to a software. In such a case, the first ranking refinement model may be assigned a higher weight than the second ranking refinement model. In at least one embodiment, theweights are predetermined. For example, prior to applying the one or more refinement models, the processor can assign weights to the ranking refinement models and compare results obtained by applying the ranking refinement models to a test set against a ground truth and iteratively modifying the weights.
[0327] In at least one embodiment, the processor determines a refined relevancy score. The refined relevancy score can correspond to a score determined by the one or more ranking refinement models. In embodiments where more than one ranking refinement model is applied, the refined relevancy score can correspond to an average or a weighted average of the scores determined by the ranking refinement models.
[0328] At 1640, the processor displays the ranked results on a graphical user interface. For example, as shown in FIG. 11 , which shows an example graphical user interface 1100, the located products 1102a-1102f that may infringe the patent can be displayed, along with the relevancy score 1140. The score 1140 can correspond to the score determined at 1620 or the refined relevancy score determined at 1630. The ranked results can be stored in a database, such as the storage component 110, databases 224 or the external data storage 102. Although FIG. 11 shows located products 1102a-1102f, it will be understood that a similar graphical user interface may be used to display located services that may infringe a patent. When more than one patent claim is analyzed by the patent analytics and infringement monitoring system 108, the results associated with each patent claim can be displayed as separate tabs (not shown) and the graphical user interface can include an interactable graphical element that enables a user to toggle between the tabs.
[0329] In at least one embodiment, the patent analytics and infringement monitoring system 108 can display a reason for the ranking and / or for the search results. For example, the GUI can include an interactable element that can enable the user to obtain an explanation of the patent-product / service evaluation subsystem’s ranking of the results and / or an explanation explaining the reasons for the patent analytics and infringement monitoring system 108 identification of the results as relevant. The explanation can include the similarities between the patent claims and the product or service, references to specific claim elements, a website reference of the product or service, an image on the website, a description on the website, a specific webpage of a website, etc. In at least one embodiment, the patent analytics and infringement monitoring system 108 can generate an intelligence report that can be viewed and optionally, downloaded by the end-user of the patent analytics and infringementmonitoring system 108.
[0330] The intelligence report can include information about a potentially infringing product or service (e.g., a description of the product or service, technical specifications of the product or technical details about the service, images of the product or images illustrating the service), information about the manufacturer and / or retailer (e.g., address, individuals associated with the business, links to Linkedln® profiles of individauls associated with the business, Securities Exchange Commission filings, business size, financial information, public relations extracts and summaries, links to products or service offerings, litigation data if applicable and available indicating whether a company appeared as a defendant or plaintiff, company attorneys, competitors, information about competitors, etc.). In at least one embodiment, the intelligence report can be compiled and prepared by an external system, for example, an artificial intelligence (Al) system that can search publicly available sources (e.g., the Web) and retrieve information about the product / service and / or the manufacturer / service provider and that can prepare a report based on the information retrieved. For example, the external system can be GPT Researcher. The patent analytics and infringement monitoring system 108 can transmit, to the external system, the name of the product / service and / or the name of the manufacturer or retailer selling the product and the external system can retrieve information about the product / service and / or the manufacturer / retailer / service provider.
[0331] Although in the example shown in FIG. 11 , the relevancy score is shown as a score between 0 and 1 , it will be understood that the relevancy score can be presented in other manners. For example, the relevancy score can be expressed as a percentage. As other non-limiting examples, the relevancy score can be presented as a level of relevancy on a scale or as a color gradient.
[0332] In at least one embodiment, as shown in FIG. 11 , the results displayed can be rated by a user via feedback buttons 1104-1106. The ratings can be used to improve the search results. For example, the ratings can be used to train one or more of the ranking refinement models. For example, the processor can generate the guidance prompts based in part on the ratings. As another example, the feedback can be used to generate the feedback dataset used for fine-tuning the LLM subsystem, as described above.
[0333] Reference is next made to FIG. 17, which illustrates another method of identifying products or services that may infringe a target patent claim. The target patentclaim can be an issued claim, or a patent application claim, including a claim being drafted or amended by a patent practitioner and that could be infringed if the patent claim was to issue. The method involves receiving, from a computing device associated with a user, patent document identifier (a patent number, a patent application number, patent publication number, or utility model number); obtaining the target patent claim, fetching (retrieving) predetermined patent document details associated with the target patent claim, using the patent document identifier; generating a set of queries based on the fetched (retrieved) patent document details; using the set of queries to search commercially relevant resources for identifying products or services each associated with the query set; scoring the search results to provide a relevancy score; and if the relevancy score is over a predetermined value, generating a ranked list of the products, methods or services for the user; else providing feedback to the computing device on the lack of relevant scores.
[0334] The patent document details fetched can be, for example, those illustrated in FIG. 19B namely, but not limited to, inventor(s) name(s), class, industry-specific terms, cited documents, and the like. In an exemplary implementation, the user can input a patent number using an input device in communication with the computing device 106. These can be, for example any computer input devices that allow a user to input data into a computer. For example, a keyboard having a set of keys used to input text and commands into a computer, a mouse used to point, click, and select items on the screen, a touchpad used to point, click, and select items on the screen. In general, use of the term “input device” is intended to include all possible types of devices and mechanisms for inputting information to computer system. In some embodiments, the patent document identifier can be received by the patent analytics and infringement monitoring system 108, as a digital image of the patent document and the patent analytics and infringement monitoring system can use techniques such as optical character recognition (OCR) to extract the patent document identifier from the digital image of the document.
[0335] Fetching the data (see e.g., FIG. 19A) whether related to the patents, patent applications, patent publications or utility models, or related to products or services involves executing a fetch request. The term “fetch request” comprises any selection received from the computing device 106 and can be associated with any action taken by the user that results in the computing device 106 performing a selection.
[0336] An example of a pseudo code for such a fetch request is provided below:# import necessary libraries import requests import json# function to fetch patent details def fetch_patent_details(patent_number):# construct the URL for the USPTO API url = "https: / / patent- api.uspto.gov / patent / application?patent_number={}".format(patent_number)# send a GET request to the USPTO API response = requests.get(url)# check if the request was successful if response. status_code == 200:# parse the JSON response data = json. Ioads(response. text)# extract the patent details title = data["title"] abstract = data["abstract"] assignee = data["assignee"] inventor = data [“first name inventor”] patent_type = data["patent_type"]# print the patent details print("Title:", title) print("Abstract:", abstract) print("Assignee:", assignee) print(“first name inventor:”, first name inventor) print("Type:", patent_type) else:# print error message printf'Error:", response. status_code)# example usage of the function fetch_patent_details("US 1234567")
[0337] The structure of the Json response may vary depending on the API endpoint used, and the token (or other authentication methods) used. Upon receiving the patent documentidentifier, the patent analytics and infringement monitoring system 108 retrieves the corresponding patent document (patent application publication, patent publication) details from databases 1702, e.g. , by opening a connection to the databases 1702 and using the patent document identifier to retrieve the information, following which the connection is terminated and the predetermined patent document details are retrieved. The databases may be remote databases such as databases of the external data storage 102. In an exemplary embodiment, depending on the input data, namely patent number, or application number or a publication number, the retrieved data may beadjusted based on a menu displayed on a graphical user interface of the patent analytics and infringement monitoring system 108 and provided to the user.
[0338] Generating the set of queries (see e.g., FIG. 19B), can include sorting the retrieved data for common phrases whether related to the patents, patent applications, patent publications or utility models, or related to products or services, and assembling the relevant search terms or expressions or formulae.
[0339] An exemplary implementation of a pseudo code for generating the search query(s) is provided below:# Input: patent details# Output: set of queries def generate_queries(patent_details):# Extract key terms and phrases from patent details key_terms = extract_key_terms(patent_details)# Generate queries based on key terms queries = generate_queries_from_terms(key_terms) return queries
[0340] The above pseudo code is also an exemplary implementation of a code within a query-generation module. The set of executable instructions takes in the patent document identifier details as input, extracts key terms and phrases from the patent document details, and uses those key terms to generate queries. The extract_key_terms function can use different techniques such as natural language processing, text mining or machine learning to extract the key terms from the patent (or application / publication or utility model) details. The generate_queries_from_terms function can use the extracted key terms and phrases to generate the queries. Depending on the specific requirements of the system, the code may need to be adjusted to match the specific requirements of the module.
[0341] Following the generation of the set of parameters and creating the query set (see e.g., FIG. 19B), the method further involves searching free and / or commercial resources online. For example, the method can involve searching remote systems 120 and / or databases stored in the external data storage 102. An example of a pseudocode used to affect the search based on the query terms generated is provided below:# import necessary libraries import requests from bs4 import BeautifulSoup# function to search online sources def search_online_sources(title, abstract, assignee):# create a list to store the search results search_results = []# create a list of online sources to search sources = ["website1.com", "website2.com", "website3.com"]# loop through each online source for source in sources:# construct the URL for the search url = "https: / / {} / search?q={}+{}+{}".format(source, title, abstract, assignee)# send a GET request to the source response = requests. get(url)# parse the HTML response soup = BeautifulSoup(response.text, "htm I. parser")# extract the search results results = soup.find_all("div", class_="search-result")# loop through each result for result in results:# extract the product name, description, and link product_name = result.find("h3").text product_description = result.find("p").text productjink = result.find("a")["href"]# add the result to the search results list search_results.append({"name": product_name,"description": product_description,"link": productjink})# return the search results return search_results# example usage of the function patent_details = fetch_patent_details("US1234567") search_results = search_online_sources(patent_details["title"], patent_details["abstract"], patent_details["assignee"]) print(search_results)
[0342] In an exemplary implementation, the pathway to search, the structure of the HTML, and the parsing may vary from one website to another requiring adapting the code. Also, specific API calls may be necessary for some websites, and this pseudocode is focused on web scraping of free web pages. Entering commercial databases may require stored authentication protocols that can be stored in the system’s modules. In an exemplary implementation, the step of searching commercially relevant resources comprises searching online marketplaces, product catalogs, and patent databases.
[0343] In an exemplary implementation, the method involves ranking the retrieved results based on relevancy score assigned to the search results (see e.g., FIG. 19C). An example of a pseudocode to generate the relevancy score is provided (see also, FIGS. 19C-19D), hereinbelow:# Input: search results, patent details# Output: relevancy score def score_results(search_results, patent_details):# Initialize relevancy score relevancy_score = 0# Iterate through search results for result in search_results:# Compare result to patent details score = compare_result o_patent(result, patent_details)# Add score to relevancy score relevancy_score += score# Normalize relevancy score relevancy_score = relevancy_score I len(search_results) return relevancy_score
[0344] The above pseudocode provides an example of how the code associated with the scoring module used in the systems disclosed could work. It takes in the search results and the patent document details as input and iterates through each search result, comparing it to the patent document details using the compare_result_to_patent (etc.) function. The function can use different techniques such as natural language processing, text mining, or machine learning to determine the similarity or relevancy between the search result and the patent details. In certain exemplary implementations, the score is then added to the total relevancy score, and divided by the number of search results to get a normalized score. Depending on the specific requirements of the patent analytics and infringement monitoring system 108, in certain other exemplary implementations, the code is adjusted to match the specific requirements.
[0345] The method also involves providing feedback on the search results, by for example, generating a ranked list of products or services and the like that potentially would infringe the input patent document (e.g., patent, patent application, patent publication, or utility patent) number(s). The ranked list can be provided on a graphical user interface of the patent analytics and infringement monitoring system. For example, as shown in FIGS. 11 and 20, the results can be shown as a ranked list with relevancy scores. The ranked list can be provided in other formats, such as, but not limited to, a table or graph, and can include interactable graphical user interface elements that can enable the user to filter or sort the list by various user defined criteria.
[0346] Reference is next made to FIG. 22, which illustrates another method of operating the patent analytics and infringement monitoring system 108 to identify patents that may be infringed by products or services. The method involves: receiving, from a computing device 106 associated with a user, a description of a product or service; extracting predetermined details from the inputted description; generating a set of queries based on the extracted details; searching at least one database for patents or patent applications associated with the query set; scoring the search results to provide a relevancy score; and if the relevancy score is over a predetermined value, generating aranked list of the patents or patent applications to the user. In an exemplary implementation, the step of detail extraction step includes extracting information such as product or service features, technical specifications, or method parameters / steps.
[0347] The patent analytics and infringement monitoring system 108 can include modules that can carry out the method steps. For example, the system can include an input module for receiving a patent document identifier from a computing device 106, a patent detail fetching module for fetching predetermined patent details; a query generation module for generating a set of queries based on the fetched patent details; a search module for using the set of queries to search commercially relevant resources for products, methods or services associated with the query set; a scoring module for scoring the search results to provide a relevancy score; a ranking module for generating a ranked list of the products, methods or services for the user if the relevancy score is over a predetermined value; and a feedback module.
[0348] As another example, the system can include an input module receiving a description of a product or service; a detail extraction module for extracting predetermined details from the inputted description; a query generation module for generating a set of queries based on the extracted details; a search module for searching at least one database for patents or patent applications associated with the query set; a scoring module for scoring the search results to provide a relevancy score; and a list generation module for generating a ranked list of the patents or patent applications to the user if the relevancy score is over a predetermined value. The computing device can receive an input (e.g., description of a product or service) from a user, via a graphical user interface of the computing device.
[0349] In at least some embodiments, the patent analytics and infringement monitoring system 108 can use at least some of the same modules and databases for identifying patents that may infringe a product or service and for identifying products or services that may infringe claims of a patent.
[0350] In at least one embodiment, the patent detail-fetching module includes a communication interface for connecting to a remote database of patent information, for example, databases of the external data storage 102, the query-generation module includes a natural language processing component for analyzing the fetched patent document details to identify key terms and phrases relevant to the patent document associated with the patent document identifier, the search module includes a webcrawling component for online searching of at least one of: marketplaces, product catalogs, service offering catalogs, and patent databases, the scoring module includes a machine learning component that uses a pre-trained model as described herein to determine the relevancy of the search results to the patent document identifier inputted, and the ranking module includes a visualization component that generates a user-friendly interface to present the list of products, methods or services and that includes elements with which a user can interact to filter or sort the list by various selectable criteria, as well as a feedback module that includes an output device (e.g., printer, screen, removable memory device and the like). These components can be used to identify products or services that may infringe a patent and / or to identify patents that may be infringed by a product or service.
[0351] As illustrated in FIG. 24, the patent analytics and infringement monitoring system 108 can include a module that enable the patent analytics and infringement monitoring system 108 to receive feedback. For example, the patent analytics and infringement monitoring system 108 can prompt the user to provide feedback to the patent analytics and infringement monitoring system 108 on the relevancy score, which can enable the patent analytics and infringement monitoring system 108 to modify, for example, the patent terms extracted, the type of product / service data inputted, the weighting and the calculation of the relevancy score. These can be done using machine learning models trained on the search results. The models can be used with a plurality of users and implemented regardless of the inputted search parameters.
[0352] In at least one embodiment, the computing device can include a user interface module for providing a user interface for inputting the description of the product, service, apparatus, composition, or method. The patent analytics and infringement monitoring system can include a detail extraction module that includes means for extracting information such as, for example, product or service features, technical specifications, or method parameters can be used in the generation of a ranked list of patent(s), patent application(s) patent publication(s) utility model(s), or design patent(s) relevant to products or services and the like that are potentially infringing the patent.
[0353] In the context of the disclosure, the term “module”, as used herein, means, but is not limited to, a software or hardware component, such as a Field Programmable Gate-Array (FPGA) or Application-Specific Integrated Circuit (ASIC), which performs certain tasks. A module may advantageously be configured to reside on the addressable storage medium and configured to execute on one or more processors. Thus, a modulemay include, by way of example, components, such as software components, object- oriented software components, class components and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, and variables. The functionality provided for in the components and modules may be combined into fewer components and modules or further separated into additional components and modules.
[0354] For example, the module is a self-contained piece of code that performs a specific task or set of tasks. Typically, the code is a separate file that can be imported into other code to be used as a reusable component. The code module can have its own variables, functions, and classes, in certain exemplary implementations, “module” can be a "building block" that can be combined with other modules to create a larger program. In certain exemplary implementations, the modules are also used to encapsulate functionality and data, so that it can be used without exposing its internal details or implementation.
[0355] In an exemplary implementation, provided herein is an article of manufacture for generating a list of products or services potentially infringing a target patent comprising a non-transitory memory device storing thereon a computer-readable media with a set of executable instructions embodied thereon configured, when executed by the at least one computer to: receive a patent number; fetch predetermined patent details; generate a set of queries based on the fetched patent details; using the set of queries to search commercially relevant resources for product, service, apparatus, composition, or method each associated with the query set score the search results to provide a relevancy score; and, if the relevancy score is over a predetermined value, generate a ranked list of the products, methods or services for the user.
[0356] In an exemplary implementation, the set of executable instructions further comprise program code for performing optical character recognition (OCR) on a digital image of the patent document identifier received, as well as program code for connecting to a remote database of patent information and retrieving information specific to the input patent document identifier, program code for analyzing the fetched patent document details to identify key terms and phrases relevant to the patent and generating a set of queries, program code for using a pre-trained machine learning model to determine the relevancy of the search results to the input patent number and generating a ranked list of product, service, apparatus, composition, or method for the user, and program codefor presenting the list of product, service, apparatus, composition, or method in a user- friendly format, such as a table or graph, and allowing the user to filter or sort the list by various selectable criteria, such as date, price, or popularity.
[0357] In another exemplary implementation, provided herein is an article of manufacture for generating a list of patents potentially reading on a product, a service, an apparatus, a composition, or a method, comprising a non-transitory memory device, storing thereon a computer-readable media (CRM) with a set of executable instructions that, when executed cause the processor to: receive a description of a product, service, apparatus, composition, or method from a computing device; extract predetermined details from the inputted description; generate a set of queries based on the extracted details; search at least one database for patents or patent applications associated with the query set; score the search results to provide a relevancy score; and if the relevancy score is over a predetermined value, generate a ranked list of the patents or patent applications to the user.
[0358] In yet another exemplary implementation, the CRM further comprises executable instructions for providing a user interface for inputting the description of the product, service, apparatus, composition or method, as well as extracting information such as product or service features, technical specifications, formulae, or method parameters.
[0359] The present invention has been described here by way of example only. Various modification and variations may be made to these exemplary embodiments without departing from the scope of the invention, which is limited only by the appended claims.
Claims
CLAIMS:
1. A computer-implemented method for fine-tuning a machine learning model, comprising: providing, in a memory, a litigation dataset comprising a plurality of historical patent litigation records, each patent litigation record comprising at least one patent claim, and a litigation outcome corresponding to each of the at least one patent claim; receiving, at a processor in communication with the memory, a product or service content item associated with a candidate patent litigation record in the plurality of historical patent litigation records; updating, at the processor, the candidate patent litigation record based on the product or service content item; and generating, at the processor, a vector database based on the litigation dataset, the vector database used to fine-tune a machine learning model.
2. The computer-implemented method of claim 1 , further comprising: receiving, at a network device in communication with the processor, a user submitted product or service content item associated with a product or service in the candidate patent litigation record.
3. The computer-implemented method of claim 2 wherein the product or service content item comprises a user-submitted URL content item corresponding to the product or service in the candidate patent litigation record, the method further comprising: sending, using a network device in communication with the processor, a content request, the content request based on a URL content item; receiving, using the network device, a content response based on the content request; performing web scraping on the content response to identify at least one image content item and at least one text content item; and updating the candidate patent litigation record by storing the at least one image content item and the at least one text content item in association with the candidate patent litigation record.
4. The computer-implemented method of claim 2 wherein the content item comprises a user-submitted text content item corresponding to the product or service in thecandidate patent litigation record, the method further comprising: storing the user-submitted text content item in association with the candidate patent litigation record.
5. The computer-implemented method of claim 3, further comprising: sending, using the network device, a captioning request to an LLM-system, the captioning request comprising the at least one image content item; receiving, using the network device, a captioning response from the LLM-system based on the captioning request, the captioning response comprising at least one textbased caption based on the at least one image content item; and updating, at the processor, the candidate patent litigation record based on the at least one text-based caption.
6. The computer-implemented method of claim 5, further comprising: determining, at the processor, a hash value corresponding to the content response; updating, at the processor, the candidate patent litigation record with the hash value; and subsequently sending, using a network device in communication with the processor, a second content request based on the URL content item; receiving, using the network device, a second content response based on the second content request; determining, at the processor, that a second hash value corresponding to the second content response is different from the hash value corresponding to the content response; and updating, at the processor, the candidate patent litigation record based on the second content response.
7. The computer implemented method of claim 1 wherein the generating the vector database comprises generating a vector embedding for the product or service content item of the candidate patent litigation record.
8. The computer-implemented method of claim 1 , wherein each patent litigation record in the dataset comprises at least one selected from the group of: a corporate identification and an estimated damages quantum.
9. The computer-implemented method of claim 1 , further comprising: generating, at the processor, a prediction prompt for an LLM-system based on the vector database; sending, using a network device in communication with the processor to the LLM- system, the prediction prompt; and receiving, using the network device, a prompt response comprising at least two search elements.
10. The computer-implemented method of claim 9, further comprising: generating, at the processor, at least two machine learning models based on the vector database, the at least two machine learning models comprising a ranking model and a classification model; and updating, at the processor, each of the at least two search elements based on the at least two machine learning models, each updated search element comprising a rank and a classification.
11. A system for fine-tuning a machine learning model, comprising: a memory, comprising a litigation dataset comprising a plurality of historical patent litigation records, each patent litigation record comprising at least one patent claim, and a litigation outcome corresponding to each of the at least one patent claim; a processor in communication with the memory, configured to: receive a product or service content item corresponding to a candidate patent litigation record in the plurality of historical patent litigation records; update the candidate patent litigation record based on the product or service content item; and generate a vector database based on the litigation dataset, the vector database used to fine-tune a machine learning model.
12. The system of claim 11 , further comprising: a network device in communication with the processor, the network device configured to receive a user submitted product or service content item associated with a product or service in the candidate patent litigation record.
13. The system of claim 12 wherein the product or service content item comprises auser-submitted URL content item corresponding to the product or service in the candidate patent litigation record, the network device further configured to: send a content request, the content request based on a URL content item; receive a content response based on the content request; and the processor further configured to: perform web scraping on the content response to identify at least one image content item and at least one text content item; and update the candidate patent litigation record by storing the at least one image content item and the at least one text content item in association with the candidate patent litigation record.
14. The system of claim 12 wherein the content item comprises a user-submitted text content item corresponding to the product or service in the candidate patent litigation record, the processor further configured to: store the user-submitted text content item in association with the candidate patent litigation record.
15. The system of claim 13, wherein the network device is further configured to: send a captioning request to an LLM-system, the captioning request comprising the at least one image content item; receive a captioning response from the LLM-system based on the captioning request, the captioning response comprising at least one text-based caption based on the at least one image content item; and wherein the processor is further configured to: update the candidate patent litigation record in the memory based on the at least one text-based caption.
16. The system of claim 15, wherein the processor is further configured to: determine a hash value corresponding to the content response; update the candidate patent litigation record in the memory with the hash value; determine that a second hash value corresponding to a second content response is different from the hash value corresponding to the content response; and update the candidate patent litigation record based on the second content response; and wherein the network device is further configured to:subsequently send a second content request based on the URL content item; receive the second content response based on the second content request.
17. The system of claim 11 wherein the generating the vector database comprises generating a vector embedding for each patent litigation record in the plurality of historical patent litigation records.
18. The system of claim 11 , wherein each patent litigation record in the dataset comprises at least one selected from the group of: a corporate identification and an estimated damages quantum.
19. The system of claim 11 , wherein the processor is further configured to: generate a prediction prompt for an LLM-system based on the vector database; and wherein a network device in communication with the processor is configured to: send to the LLM-system, the prediction prompt; and receive a prompt response comprising at least two search elements.
20. The system of claim 19, wherein the processor is configured to: generate at least two machine learning models based on the vector database, the at least two machine learning models comprising a ranking model and a classification model; and update each of the at least two search elements based on the at least two machine learning models, each updated search element comprising a rank and a classification.
21. A computer-implemented method for providing updated search results by fine- tuning of machine learning models using a feedback dataset, comprising: receiving, at a network device from a user device, a search query comprising at least one of a patent document identifier or a product or service identifier; transmitting, from the network device to the user device, a search response user interface comprising a search identifier and at least one search result item; receiving, at the network device from the user device, a feedback item corresponding to a candidate search item in the at least one search result item; storing, in a memory, a feedback item record in a feedback dataset, the feedbackitem record stored in association with the search identifier, the feedback item record comprising the feedback item, at least one of a patent claim associated with the patent document identifier or a product or service description, and the search result item; and transmitting, from the network device to the user device, at least one updated search result item.
22. The computer-implemented method of claim 21 , further comprising: extracting, at a processor in communication with the memory, patent information associated with the patent document identifier; and generating, at the processor, an initial analysis prompt based on the patent information.
23. The computer-implemented method of claim 21 , further comprising: extracting, at a processor in communication with the memory, product or service information associated with the product or service description; and generating, at the processor, an initial analysis prompt based on the product information.
24. The computer-implemented method of any one of claims 22 or 23, wherein the patent information associated with the patent document identifier comprises at least one selected from the group of: claim text; drawing images; caption text for the drawing images; description text; abstract text; forward or backward citations, each citation comprising a patent document identifier; chart images; caption text for the chart images; formulae text; inventor information; assignee information; classification information; prosecution file wrapper text;maintenance payment information; and litigation history text.
25. The computer-implemented method of any one of claims 22 or 23, wherein the product or service information associated with the product or service identifier comprises at least one selected from the group of: a text snippet, a URL, an icon, an image, and a citation.
26. The computer-implemented method of any one of claims 22 or 23, further comprising: transmitting, from the network device to an LLM system, an initial analysis request comprising the initial analysis prompt; receiving, from the network device from the LLM system, an initial analysis response based on the initial analysis request, the initial analysis response comprising at least one remote device search query; transmitting, from the network device to a remote device, the at least one remote device search query; receiving, from the network device to the remote device, at least one remote device search response; and transmitting, from the network device to the user device, the at least one search result item based on the at least one remote device search response.
27. The computer-implemented method of claim 26, further comprising: generating, at the processor, a subsequent analysis prompt based on the initial analysis prompt and the feedback item record; transmitting, from the network device to an LLM system, a subsequent analysis request comprising the subsequent analysis prompt; receiving, from the network device from the LLM system, a subsequent analysis response based on the subsequent analysis request, the subsequent analysis response comprising at least one subsequent remote device search query; transmitting, from the network device to the remote device search provider, the at least one subsequent remote device search query; receiving, from the network device to the remote device search provider, at least one subsequent remote device search response; and transmitting, from the network device to the user device, at least one subsequentsearch result item based on the at least one subsequent remote device search response.
28. The computer-implemented method of claim 21 , wherein the feedback item comprises a rating, the rating corresponding to a feedback control selected by the user.
29. The computer-implemented method of claim 28, wherein the feedback control comprises at least one selected from the group of: two or more button controls, a radio button control, a slider control, a star selector control.
30. The computer-implemented method of claim 21 , further comprising: providing, in a memory, the feedback item dataset comprising a plurality of historical feedback records, each feedback item record stored in association with a corresponding search identifier, the feedback item record comprising a historical feedback item, at least one of a patent claim associated with a corresponding patent document identifier or a product or service description, and a corresponding search result item.31 . The computer-implemented method of claim 30, further comprising: generating, at the processor, at least two machine learning models based on the feedback item dataset, the at least two machine learning models comprising a ranking model and a classification model; and updating, at the processor, each of the at least one search result item based on the at least two machine learning models, each of the at least one search result item comprising a rank and a classification.
32. A system for providing updated search results by fine-tuning machine learning models using a feedback dataset, comprising: a network device configured to: receiving from a user device, a search query comprising at least one of a patent document identifier, a product identifier or service identifier; transmitting to the user device, a search response user interface comprising a search identifier and at least one search result item; and receiving from the user device, a feedback item corresponding to a candidate search item in the at least one search result item; a memory;a processor configured to: store in the memory, a feedback item record in a feedback dataset, the feedback item record stored in association with the search identifier, the feedback item record comprising the feedback item, at least one of a patent claim associated with the patent document identifier or a product or service description, and the search result item; and transmitting, using the network device to the user device, at least one updated search result item based on the feedback item record.
33. The system of claim 32, wherein the processor is configured to: extract patent information associated with the patent document identifier; and generate an initial analysis prompt based on the patent information.
34. The system of claim 32, wherein the processor is configured to: extract product information associated with the product or service description; and generate an initial analysis prompt based on the product or service information.
35. The system of any one of claims 33 or 34, wherein the patent information associated with the patent document identifier comprises at least one selected from the group of: claim text; drawing images; caption text for the drawing images; description text; abstract text; forward or backward citations, each citation comprising a patent document identifier; chart images; caption text for the chart images; formulae text; inventor information; assignee information; classification information; prosecution file wrapper text; maintenance payment information; andlitigation history text.
36. The system of any one of claims 33 or 34, wherein the product or service information associated with the patent document identifier comprises at least one selected from the group of: a text snippet, a URL, an icon, an image, and a citation.
37. The system of any one of claims 33 or 34, wherein the processor is further configured to: transmit using the network device to an LLM system, an initial analysis request comprising the initial analysis prompt; receive from the network device from the LLM system, an initial analysis response based on the initial analysis request, the initial analysis response comprising at least one remote device search query; transmit from the network device to a remote device search provider, the at least one remote device search query; receive from the network device to the remote device search provider, at least one remote device search response; and transmit from the network device to the user device, the at least one search result item based on the at least one remote device search response.
38. The system of claim 35, wherein the processor is further configured to: generate a subsequent analysis prompt based on the initial analysis prompt and the feedback item record; transmit from the network device to an LLM system, a subsequent analysis request comprising the subsequent analysis prompt; receive from the network device from the LLM system, a subsequent analysis response based on the subsequent analysis request, the subsequent analysis response comprising at least one subsequent remote device search query; transmit from the network device to the remote device search provider, the at least one subsequent remote device search query; receive from the network device to the remote device search provider, at least one subsequent remote device search response; and transmit from the network device to the user device, at least one subsequent search result item based on the at least one subsequent remote device search response.
39. The system of claim 32, wherein the feedback item comprises a rating, the rating corresponding to a feedback control selected by the user.
40. The system of claim 38, wherein the feedback control comprises at least one selected from the group of: two or more button controls, a radio button control, a slider control, a star selector control.
41. The system of claim 40, wherein the feedback item dataset further comprises a plurality of historical feedback records, each feedback item record stored in association with a corresponding search identifier, the feedback item record comprising a historical feedback item, at least one of a patent claim associated with a corresponding patent document identifier or a product or service description, and a corresponding search result item.
42. The system of claim 32 wherein the processor is further configured to: generate at least two machine learning models based on the feedback dataset, the at least two machine learning models comprising a ranking model and a classification model; and update each of the at least one search result item based on the at least two machine learning models, each of the at least one search result item comprising a rank and a classification.
43. A method of ranking product-patent or service-patent evaluation search results, the method comprising operating at least one processor to: receive a plurality of search results relevant to a search input, the search input corresponding to a product description or a service description, or a patent claim; for each search result, compare a vector embedding of the search result and a vector embedding of the search input to obtain a relevancy score for the search result; apply one or more ranking refinement models to the search results to obtain refined relevancy scores; and display the search results according to the refined relevancy scores.
44. A method of ranking product-patent or service-patent evaluation search results, the method comprising operating at least one processor to:receive a plurality of search results corresponding to products or services relevant to a patent claim associated with a patent document; calculate a vector embedding of the patent claim; for each search result, calculate a vector embedding of the search result and comparing the vector embedding of the claim and the vector embedding of the search result to obtain a relevancy score for the search result; apply one or more ranking refinement models to the search results to obtain refined relevancy scores; and display the search results according to the refined relevancy scores.
45. The method of claim 44, wherein the one or more ranking refinement models are selected according to one or more of: a subject matter of the patent document and classifications of the patent document.
46. The method of claim 44 or 45, wherein applying the one or more ranking refinement models comprises operating the at least one processor to apply two or more ranking refinement models, and wherein the method further comprises operating the at least one processor to determine a weighting for each of the two or more ranking refinement models according to one or more of: a predetermined weighting, a subject matter of the patent document, classifications of the patent document and a confidence of the ranking refinement model.
47. The method of any one of claims 43 to 46, wherein at least one of the one or more ranking refinement models is a large language model and wherein applying the one or more ranking refinement models comprises generating a guidance prompt for the at least one ranking refinement model based one or more of: a subject matter of the patent document, classifications of the patent document, search parameters, feedback data.
48. The method of any one of claims 43 to 47, wherein applying the one or more ranking refinement models to obtain refined relevancy scores comprises operating the at least one processor to: for each of the one or more ranking refinement models, for each search result, obtain a model relevancy score; and determine the refined relevancy score based on the relevancy score and themodel relevancy score.
49. The method of any one of claims 43 to 48, wherein the plurality of results are URLs of product or service websites and wherein the method further comprises operating the at least one processor to parse each of the URLs to obtain a corresponding product or service description.
50. The method of any one of claims 43 to 49, further comprising operating the at least one processor to: identify the K most relevant search results based on the relevancy scores; and apply the one or more ranking refinement models to the K most relevant search results.
51. The method of any one of claims 43 to 50, wherein applying the one or more ranking refinement models to the search results comprises operating the at least one processor to input a text of the patent claim and a text of a product or service description corresponding to the search result into the one or more ranking refinement models, and wherein the one or more ranking refinement models are configured to determine a measure of similarity between the text of the claim and the text of the product or service description corresponding to the search result.
52. The method of any one of claims 43 to 51 , wherein calculating the vector embedding of the claim comprises operating the at least one processor to: identify one or more portions of the patent corresponding to the claim; and calculate the vector embedding of the patent claim and the one or more portions of the patent document.
53. The method of claim 52, wherein the one or more portions comprise one or more of: a description, a drawing, an equation, a table and a formula.
54. A method of ranking product-patent or service-patent evaluation search results, the method comprising operating at least one processor to: receive a description of a product or a service; convert the description of the product or the description of the service into a vector embedding of the product or service;retrieve from a database, a plurality of vector embeddings of search results corresponding to patent claims; for each search result, compare the vector embedding of the search result and the vector embedding of the product or service description to obtain a relevancy score for the search result; apply one or more ranking refinement models to the search results to obtain refined relevancy scores; and display the search results according to the refined relevancy scores.
55. A system for ranking product-patent evaluation or service-patent evaluation search results, the system comprising at least one processor operable to: receive a plurality of search results corresponding to products or services relevant to a patent claim associated with a patent document; calculate a vector embedding of the patent claim; for each search result, calculate a vector embedding of the search result and comparing the vector embedding of the claim and the vector embedding of the search result to obtain a relevancy score for the search result; apply one or more ranking refinement models to the search results to obtain refined relevancy scores; and display the search results according to the refined relevancy scores.
56. The system of claim 55, wherein the at least one processor is operable to select the one or more ranking refinement models according to one or more of: a subject matter of the patent document and classifications of the patent document.
57. The system of claim 55 or 56, wherein applying the one or more ranking refinement models comprises applying two or more ranking refinement models, and wherein the at least one processor is further operable to determine a weighting for each of the two or more ranking refinement models according to one or more of: a predetermined weighting, a subject matter of the patent document, classifications of the patent document and a confidence of the ranking refinement model.
58. The system of any one of claims 55 to 57, wherein at least one of the one or more ranking refinement models is a large language model and wherein applying the one or more ranking refinement models comprises generating a guidance prompt for the at leastone ranking refinement model based one or more of: a subject matter of the patent document, classifications of the patent document, search parameters, feedback data.
59. The system of any one of claims 55 to 58, wherein applying the one or more ranking refinement models to obtain refined relevancy scores comprises: for each of the one or more ranking refinement models, for each search result, obtaining a model relevancy score; and determining, by the processor, the refined relevancy score based on the relevancy score and the model relevancy score.
60. The system of any one of claims 55 to 59, wherein the plurality of results are URLs of product or service websites and wherein the further comprises parsing each of the URLs to obtain a corresponding product or service description.61 . The system of any one of claims 55 to 60, wherein the at least one processor is further operable to: identify the K most relevant search results based on the relevancy scores; and apply the one or more ranking refinement models to the K most relevant search results.
62. The system of any one of claims 55 to 61 , wherein applying the one or more ranking refinement models to the search results comprises inputting a text of the claim and a text of a product or service description corresponding to the search result into the one or more ranking refinement models, and wherein the one or more ranking refinement models are configured to determine a measure of similarity between the text of the claim and the text of the product or service description corresponding to the search result.
63. The system of any one of claims 55 to 62, wherein calculating the vector embedding of the claim comprises: identifying one or more portions of the patent document corresponding to the patent claim; and calculating the vector embedding of the claim and the one or more portions of the patent document corresponding to the patent claim.
64. The system of claim 62, wherein the one or more portions comprise one or more of: a description, a drawing, an equation, a table and a formula.
65. A computer-readable medium having instructions stored thereon that when executed, cause at least one processor to: receive a plurality of search results corresponding to products or services relevant to a patent claim associated with a patent document; calculate a vector embedding of the claim; for each search result, calculate a vector embedding of the search result and comparing the vector embedding of the claim and the vector embedding of the search result to obtain a relevancy score for the search result; apply one or more ranking refinement models to the search results to obtain refined relevancy scores; and display the search results according to the refined relevancy scores.
66. The computer-readable medium of claim 65, wherein the instructions, when executed, further cause the at least one processor to perform steps of the method of any one of claims 45 to 53.
67. A method for identifying products or services relevant to a target patent claim, the method comprising operating at least one processor to: receive a patent document identifier associated with the target patent claim from a computing device in communication with the at least one processor via a network; obtain the target patent claim; retrieve, from a database in communication with the at least one processor via the network, patent document details associated with the target patent claim, using the patent document identifier; generate a set of queries based on the retrieved patent document details; search remote databases using the set of queries to identify products or services satisfying the query set; score the identified products or services to provide a relevancy score for each identified product or service; and generate a ranked list of the products or services having a relevancy score greater than a predetermined threshold; anddisplay the generated ranked list.
68. The method of claim 67, wherein obtaining the target patent claim comprises operating the at least one processor to: receive a free-form text of the target patent claim from the computing device; or retrieve, from the database, the target patent claim from a patent document associated with the patent document identifier.
69. The method of claim 67, further comprising operating the at least one processor to: when no identified product or service has a relevancy score greater than the predetermined threshold, display an indication indicating a lack of relevant results.
70. The method of claim 67, further comprising operating the at least one processor to: receive a digital image comprising the patent number of the target patent; andExtract the patent number from the digital image using optical character recognition.71 . The method of claim 67, further comprising operating the at least one processor to: access a database comprising patent information to retrieve the patent details associated with the target patent.
72. The method of claim 67, further comprising operating the at least one processor to: analyze the retrieved patent details to identify terms and phrases relevant to the target patent.
73. The method of claim 67, wherein the remote databases comprise one or more of: databases of online marketplaces, databases comprising product catalogs, service offering catalogs, and patent databases.
74. The method of claim 67, further comprising operating the at least one processor to: score the identified products and services using a machine learning algorithm.
75. The method of claim 67, further comprising operating the at least one processor to: display the generated ranked list as a table or a graph; display at least one interactable graphical user interface element, the interactable graphical user interface element causing the at least one processor to filter or sort the ranked list according to a selectable criterion.
76. The method of claim 67, further comprising operating the at least one processor to: determine one or more recommendations.
77. A system for identifying products or services relevant to a target patent claim, the system comprising at least one processor implementing: an input module for receiving a patent document identifier associated with the target patent claim from a computing device in communication with the at least one processor; a patent document detail fetching module for retrieving patent document details associated with the target patent claim using the patent document identifier; a query generation module for generating a set of queries based on the retrieved patent document details; a search module for searching remote databases using the set of queries to identify products or services satisfying the query set; a scoring module for scoring the identified products or services to provide a relevancy score for each identified product or service; a ranking module for generating a ranked list of the products or services having a relevancy score greater than predetermined threshold; and a feedback module for receiving feedback on the ranked list of the products and services from the computing device.
78. The system of claim 77, wherein the input module is further configured to receive a free-form text of the target patent claim from the computing device.
79. The system of claim 77, wherein the patent document detail fetching module is configured to retrieve the target patent claim from a patent document associated with the patent document identifier.
80. The system of claim 77, wherein the patent detail-fetching module comprises a communication interface for connecting to a remote database comprising patent information and wherein the patent detail fetching module is configured to retrieve the patent document details associated with the target patent claim from the remote database.81 . The system of claim 77, wherein the query-generation module comprises a natural language processing component for analyzing the retrieved patent document details to identify terms and phrases relevant to the target patent claim.
82. The system of claim 77, wherein the search module comprises a web crawling component for online searching at least one of: databases comprising online marketplaces, databases comprising product catalogs, service offering catalogs and patent databases.
83. The system of claim 77, wherein the scoring module comprises a machine learning component that uses a pre-trained model to score the identified products and services.
84. The system of claim 77, wherein the ranking module comprises a visualization component configured to display the generated ranked list as a table or a graph and display at least one interactable graphical user interface element, the interactable graphical user interface element causing the at least one processor to filter or sort the ranked list according to a selectable criterion.
85. A non-transitory computer-readable medium having a set of executable instructions stored thereon that, when executed by at least one processor, cause the at least one processor to: receive a patent document identifier associated with a target patent claim from a computing device in communication with the at least one processor via a network; obtain the target patent claim; retrieve, from a database, patent document details associated with the target patent claim, using the patent document identifier; generate a set of queries based on the retrieved patent document details; search remote databases using the set of queries to identify products orservices satisfying the query set; score the identified products or services to provide a relevancy score for each identified product or service; and generate a ranked list of the products or services having a relevancy score greater than a predetermined threshold; and display the generated ranked list.
86. The computer-readable medium of claim 85, wherein the set of executable instructions further cause the at least one processor to: receive a free-form text of the target patent claim from the computing device; or retrieve, from the database, the target patent claim from a patent document associated with the patent document identifier.
87. The non-transitory computer readable medium of claim 85, wherein the set of executable instructions, when executed by the at least one processor, further cause the at least one processor to receive a digital image comprising the patent document identifier associated with the target patent claim and extract the patent document identifier from the digital image using optical character recognition.
88. The non-transitory computer readable medium of claim 85, wherein the set of executable instructions, when executed by the at least one processor, further cause the at least one processor to access a database comprising patent information to retrieve the patent document details associated with the target patent claim.
89. The non-transitory computer readable medium of claim 85, wherein the set of executable instructions, when executed by the at least one processor, further cause the at least one processor to analyze the retrieved patent document details to identify terms and phrases relevant to the target patent claim.
90. The non-transitory computer readable medium of claim 85, wherein the set of executable instructions, when executed by the at least one processor, further cause the at least one processor to score the identified products and services using a machine learning algorithm.91 . The non-transitory computer readable medium of claim 85, wherein the set of executable instructions, when executed by the at least one processor, further cause the at least one processor to: display the generated ranked list as a table or a graph; and display at least one interactable graphical user interface element, the interactable graphical user interface element causing the at least one processor to filter or sort the ranked list according to a selectable criterion.
92. A method for identifying patents relevant to a product or a service, the method comprising operating at least one processor to: receive a description of the product or service from a computing device in communication with the at least one processor; extract description details from the received description; generate a set of search queries based on the extracted description details; search at least one database for patents satisfying the set of search queries; score the patents satisfying the set of search queries to obtain a relevancy score for each patent; generate a ranked list of the patents having a relevancy score greater than a predetermined threshold.
93. The method of claim 88, further comprising operating the at least one processor to: extract information including: product features or service features, technical specifications, or method parameters from the received description.
94. A system for identifying patents relevant to a product or a service, the system comprising operating at least one processor implementing: an input module for receiving a description of the product or the service from a computing device in communication with the at least one processor; a detail extraction module for extracting description details from the received description; a query generation module for generating a set of queries based on the extracted description details; a search module for searching at least one database for patents satisfying the set of search queries; a scoring module for scoring the patents satisfying the set of search queries toobtain a relevancy score for each patent; a list generation module for generating a ranked list of the patents having a relevancy score greater than predetermined threshold.
95. The method of claim 94, wherein the detail extraction module comprises means for extracting information including: product features or service features, technical specifications, or method parameters from the received description.
96. A non-transitory computer-readable medium having a set of executable instructions stored thereon that, when executed by at least one processor, cause the at least one processor to: receive a description of the product or service from a computing device in communication with the at least one processor; extract description details from the received description; generate a set of search queries based on the extracted description details; search at least one database for patents satisfying the set of search queries; score the patents satisfying the set of search queries to obtain a relevancy score for each patent; generate a ranked list of the patents having a relevancy score greater than a predetermined threshold.
97. The non-transitory computer readable medium of claim 96, wherein the set of executable instructions, when executed by the at least one processor, further cause the at least one processor to extract information including: product features or service features, technical specifications, or method parameters from the received description.
Citation Information
Patent Citations
Enterprise internal search engine method based on large language model
CN116775853A
Text recognition method and system for commodity compliance detection
CN118586936A
Patent claim feature mapping
GB2622298A
AI Legal Support System for Public Experience
KR102770195B1
AI Legal Support System for Legal Experts
KR102770197B1