Patent value evaluation method and system
By segmenting, identifying and structuring multi-source unstructured text information, quantified value factor data is generated, which solves the problems of data quality and low efficiency in existing patent value assessment methods and achieves more efficient and accurate patent value assessment.
Patent Information
- Application Number
- CN202510784839.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-12
- Publication Date
- 2025-09-19
AI Technical Summary
Existing patent value assessment methods face challenges such as uneven data quality, reliance on manual experience, insufficient information quantification, and low assessment efficiency when processing multi-source unstructured data.
By acquiring multi-source unstructured text information, using the domain term library and relationship library for word segmentation, identifying key information according to the preset value element extraction rules, and performing structured and quantified processing, quantified value element data is generated.
It achieves the transformation from original unstructured information to structured valuable data that can be used for financial decision-making, significantly reduces dependence on human experience, and improves evaluation efficiency and accuracy.
Smart Images

Figure CN120672203A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of patent value assessment, and in particular to a patent value assessment method and system. Background Art
[0002] When processing patent pledge financing applications submitted by enterprises, financial institutions need to conduct a comprehensive valuation assessment of the patent assets pledged. This assessment aims to accurately determine potential loan amounts, assess associated risk levels, and ultimately formulate reasonable financing terms. To support this core business process, the financial institution has launched and implemented an internal patent valuation process. This process utilizes a data-driven approach. Its core objective is to quantify the value of the proposed patent portfolio by deeply analyzing large amounts of data from multiple sources, including but not limited to patent publications, patent citation data, patent family information, patent legal status change records, and relevant market transaction data. The specific data processing steps typically involve automatically or semi-automatically extracting key technical features from massive amounts of unstructured text, assessing their technological impact by analyzing citation relationships between patents, evaluating the legal validity and stability of the patents, and attempting to correlate and leverage market data on technology licensing or transfer cases to assist in valuation assessments. The ultimate goal of this approach is to generate an objective and quantitative patent valuation report, providing the financial institution with a solid data and information foundation for subsequent financing decisions.
[0003] However, when actually implementing the patent value assessment process described above, even with the adoption of data processing technology, the assessment process still faces significant challenges in many aspects, making it a bottleneck in the entire financing process. First, the sources of patent and market-related data used for assessment are extremely wide-ranging, and these data often exist in a variety of different formats. The data quality varies, often containing noise, missing values, or inconsistent information. For example, in different patent databases or legal announcement platforms, the legal status of the same patent (such as annual fee payment, invalidation declaration, change of right holder, etc.) may be updated with delays or record differences. This data quality issue results in a large amount of manual work required for data cleaning, integration, and verification in the initial stage of the assessment process, namely the data preprocessing stage, which significantly increases the initial time and labor cost of the assessment.
[0004] Secondly, patent value itself is a complex and dynamic concept influenced by multiple factors. These factors include not only technical indicators such as the novelty, advancement, and substitutability of the patented technology itself, but also its potential market application prospects, the overall development trends of the industry to which it belongs, and the potential infringement risks it faces. Many key market and commercial factors are difficult to directly capture or quantify from structured patent data. For example, assessing the market application prospects of a cutting-edge technology or its potential acceptance in a specific industry often requires reading and analyzing a large amount of multi-source, unstructured information, including industry reports, market research data, and expert interview records. The process of acquiring, understanding, and integrating this information is time-consuming and highly dependent on the evaluator's deep professional judgment and industry experience.
[0005] Furthermore, the current pace of technological advancement is extremely rapid, especially in emerging industries. The technical value of a patent can rapidly depreciate or lose market competitiveness with the emergence of new, more advanced technologies. This requires patent valuation methods to be able to promptly and accurately reflect the impact of technological iterations and market changes. However, existing valuation methods often struggle to effectively capture these dynamic changes, resulting in valuation results that lag behind the patent's actual technical and market value.
[0006] Furthermore, a financial institution's own risk appetite, its familiarity and understanding of specific technological fields, and fluctuations and changes in the macroeconomic environment can also indirectly or directly influence its final assessment of patent value and credit decisions. These non-technical considerations often rely more on internal financial institutions' policy guidance, risk management experience, and manual judgment, making them difficult to fully integrate into existing automated, data-driven evaluation processes.
[0007] The combined effects of these factors make the patent valuation process a time-consuming, highly expert, and difficult-to-scale bottleneck. This not only impacts the overall efficiency of financial institutions in processing patent pledge financing applications, but also becomes particularly pronounced when rapidly responding to market opportunities or processing large numbers of applications. This increases operating costs for financial institutions and can cause companies seeking financing to miss out on valuable market or financing opportunities due to lengthy evaluation cycles. Summary of the Invention
[0008] The purpose of the present invention is to provide a patent value assessment method and system, focusing on obtaining specific value elements related to financial pledge decisions from unstructured information, realizing the transformation from original unstructured information to structured value data that can be used for financial decision-making, greatly reducing manual workload and shortening the assessment cycle.
[0009] In a first aspect, the present invention provides a method for evaluating patent value, comprising the following steps: Acquire multi-source unstructured text information; Using a domain term library and a relationship library, the multi-source unstructured text information is segmented to obtain multiple text segmentation results; the domain term library and relationship library contain patent technology field terms, financial industry terms, and legal-related terms and their associated relationships; According to preset value element extraction rules related to patent pledge financing, and in combination with the domain term library and relationship library, identify and extract key information from the multiple text segmentation results; Structuring the extracted key information to generate structured data; Standardizing the market size data in the structured data, and classifying or scoring the market tendency descriptions to obtain quantified value factor data; Based on the value factor data, the patent value is evaluated.
[0010] The patent value assessment method provided by the present invention can automatically and efficiently extract, structure, and quantify financing-related value factors such as patent monetization ability, legal stability, and market recognition from multi-source heterogeneous patent and market-related unstructured information with different data quality, thereby reducing dependence on human experts, shortening the assessment cycle, and improving batch processing capabilities, enabling financial institutions to process patent pledge financing applications more quickly and supporting batch assessment needs.
[0011] In a second aspect, the present invention provides a patent value assessment system, comprising: An acquisition module, configured to acquire multi-source unstructured text information; the multi-source unstructured text information includes patent texts, market report texts, industry news texts, and legal notice texts; A word segmentation module is used to perform word segmentation processing on the multi-source unstructured text information using a domain term library and a relationship library to obtain multiple text segmentation results; the domain term library and the relationship library contain patent technology field terms, financial industry terms and legal terms and their associated relationships; An extraction module, configured to identify and extract key information from the plurality of text segmentation results based on preset value element extraction rules related to patent pledge financing and in combination with the domain term library and relationship library; the key information includes technology application scenarios, market size data, legal status information, and competitor information; The generation module is used to structure the extracted key information and generate structured data containing application scenario fields, market size fields, legal status fields, and competitor information fields; A quantification module is used to standardize the market size data in the structured data and classify or score the market tendency description to obtain quantified value factor data; An evaluation module is used to evaluate the patent value based on the value factor data.
[0012] From the above, it can be seen that the patent value assessment method provided by the present invention, through specialized unstructured information analysis and value element extraction, can automatically and efficiently identify, extract, structure, and quantify key value elements related to patent pledge financing from multi-source heterogeneous patents and market-related unstructured texts. This significantly reduces the evaluation process's dependence on human experts and the time required, and improves the efficiency of data processing and the batch processing capabilities of the evaluation. By timely integrating dynamic information from channels such as market reports and news, the solution helps to more accurately reflect the current market value and monetization potential of patents. Ultimately, this solution provides financial institutions with a fast and reliable data foundation for patent pledge value assessment, improving the efficiency and accuracy of financing decisions.
[0013] Other features and advantages of the present invention will be described in the following description, and in part will become apparent from the description, or will be understood by practicing the embodiments of the present invention. The purposes and other advantages of the present invention can be realized and obtained by the structures particularly pointed out in the written description and the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 A flowchart of a patent value assessment method provided in an embodiment of the present invention.
[0015] Figure 2 A schematic diagram of the structure of a patent value assessment system provided in an embodiment of the present invention.
[0016] Description of labels: 100, acquisition module; 200, word segmentation module; 300, extraction module; 400, generation module; 500, quantification module; 600, evaluation module. DETAILED DESCRIPTION
[0017] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, not all of the embodiments. The components of the embodiments of the present invention generally described and shown in the drawings herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making creative work are within the scope of protection of the present invention.
[0018] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings. At the same time, in the description of the present invention, the terms "first", "second", etc. are used only to distinguish the description and should not be understood as indicating or implying relative importance.
[0019] Reference Attachment Figure 1 The present invention provides a patent value assessment method, comprising the following steps: Acquire multi-source unstructured text information; multi-source unstructured text information includes patent text, market report text, industry news text, and legal notice text; Utilize domain terminology and relationship databases to perform word segmentation on multi-source unstructured text information and obtain multiple text segmentation results; domain terminology and relationship databases contain patent technology field terms, financial industry terms, legal terms, and their relationships; Based on preset value factor extraction rules related to patent pledge financing, and in combination with the domain terminology library and relationship library, key information is identified and extracted from multiple text segmentation results; key information includes technology application scenarios, market size data, legal status information, and competitor information; Structuring the extracted key information to generate structured data containing application scenario fields, market size fields, legal status fields, and competitor information fields; Standardize the market size data in the structured data and classify or score the market tendency description to obtain quantified value factor data; Evaluate patent value based on value factor data.
[0020] Multi-source unstructured text information refers to a collection of text data from various sources and lacks a predefined format. This can include patent disclosures, market research reports, industry analysis news, and legal status announcements. Its primary purpose is to obtain comprehensive, multi-dimensional, original information related to patent value. Domain terminology and relationship databases refer to knowledge collections containing vocabulary from specific professional fields and their interrelationships. For example, domain terminology and relationship databases encompass specialized terms in fields such as patent technology, finance, and law, along with their synonyms, hyponyms, and associations. These databases can be constructed using dictionaries, ontologies, and knowledge graphs, primarily to improve the accuracy and efficiency of processing specialized text information. Value factor extraction rules related to patent pledge financing refer to pre-defined standards or patterns for identifying and extracting specific information from text based on the needs of patent pledge financing assessments. These rules can be implemented using keyword matching, grammatical rules, or machine learning models, primarily to automatically identify key information from text that is valuable for assessment. Key information refers to data or descriptions identified and extracted from the original text that are directly relevant to patent value assessment. These include technical application scenarios, market size data, legal status information, and competitor information, primarily to provide the specific factual basis required for assessment. Structuring refers to the process of converting unstructured or semi-structured information into a data format with clear fields and hierarchies. This can be achieved through methods such as data parsing, field mapping, and database storage. Its primary purpose is to facilitate subsequent calculations and analysis of the extracted information. Standardization refers to the process of converting data from different sources, units, or dimensions to make them comparable. This can be achieved through statistical methods such as Z-Score and Min-Max normalization. It primarily aims to eliminate discrepancies in market size data due to different sources. Classifying or scoring market sentiment descriptions refers to the process of qualitatively or quantitatively analyzing descriptive text expressing market sentiment or trends, classifying them into pre-defined categories (such as positive, neutral, or negative) or assigning numerical scores. This can be achieved through natural language processing techniques such as sentiment analysis and text classification. Its primary purpose is to convert qualitative market descriptions into computable quantitative indicators. Quantified value factor data refers to the numerical representation of various value factor information after standardization, classification, and scoring. It primarily provides input for objective, data-based patent value assessments. Patent value assessment refers to the process of deriving a value judgment or numerical result of patent assets based on the quantified data of various value factors and using a specific assessment model or algorithm. Its main purpose is to complete the value assessment task required for patent pledge financing.
[0021] The core innovation of this application lies in the automatic extraction, structuring and quantification of key value elements that are highly relevant to patent pledge financing by systematically integrating and processing multi-source unstructured text information and combining it with domain knowledge bases and preset rules. This overcomes the challenges of existing methods in processing complex data sources, relying on manual experience, insufficient information quantification, and evaluation efficiency and scalability, and achieves a more objective and efficient patent value assessment.
[0022] Specifically, this solution operates as follows: First, unstructured text information is acquired from various sources, including patent documents, market reports, industry news, and legal notices, to construct the basic data set required for the evaluation. Next, using a knowledge base containing specialized terminology and their relationships in fields such as patent technology, finance, and law, the acquired unstructured text is segmented, breaking the continuous text into semantically meaningful sequences of words or phrases to prepare for subsequent information identification. Then, based on pre-defined value factor extraction rules tailored to the needs of patent pledge financing evaluations and incorporating the expertise provided by the domain terminology and relationship database, key information such as technology application scenarios, market size data, legal status information, and competitor information is identified and extracted from the segmented results. This dispersed key information is then structured into a data structure with clear fields (e.g., application scenario fields and market size fields) to facilitate subsequent automated processing. Furthermore, the market size data within the structured data is standardized to eliminate variations caused by different data sources or units, ensuring comparability. Furthermore, the textual descriptions of market trends are classified or scored, converting them into quantitative indicators. Through these processes, unified, quantified value factor data is generated. Finally, based on this quantified, multi-dimensional value factor data, an evaluation model or algorithm is applied to assess the value of the patent and produce an evaluation result. Through automated processing and quantitative analysis, the entire process reduces reliance on manual experience and improves the efficiency and objectivity of the evaluation.
[0023] As a preferred embodiment, the solution of this application is specifically implemented as follows: A data acquisition module can be deployed to periodically or on-demand acquire text data from sources such as patent databases, market research agency websites, industry news platforms, and legal notice websites through methods such as web crawlers, API interfaces, or database connections. The acquired data is fed into a text processing module, which has a built-in domain knowledge base containing pre-built terminology dictionaries and relationship maps in the fields of patent technology, finance, and law. The text processing module uses this knowledge base to perform pre-processing on the original text, such as word segmentation and part-of-speech tagging. The processed text segmentation results are then fed into an information extraction module, which is equipped with extractors based on rule engines or machine learning models. These extractors, based on preset value factor extraction rules, identify and extract key information from the text, such as descriptions of technical application scenarios, specific market size values or ranges, records of changes in the legal status of patents, and the names of mentioned competitors. The extracted key information is then passed to a data structuring module, which, based on a preset data schema, maps and populates the extracted information into structured data tables or JSON objects containing fields such as application scenarios, market size, legal status, and competitors. The market size field in the structured data is fed into a quantitative processing module, which performs a numerical normalization algorithm (such as the Z-Score) and feeds the market propensity description text into a sentiment analysis or text classification submodule to calculate a sentiment score or perform classification. The quantitative processing module integrates the standardized market size data with the quantified market propensity data, outputting a quantified value factor dataset. Finally, an evaluation module receives this quantified value factor dataset and, based on a pre-defined evaluation model (such as a machine learning-based prediction model or an expert rule-based scoring model), calculates and outputs the patent's assessed value.
[0024] Through the above solution, this application can effectively solve the problem of difficulty in processing multi-source unstructured data in patent value assessment, realize the automation and accurate extraction of key information, convert difficult-to-quantify market and business information into calculable values, significantly reduce dependence on manual experience, improve assessment efficiency and objectivity, and thus realize the large-scale processing of the patent value assessment process.
[0025] In certain embodiments, the steps of identifying and extracting key information from multiple text segmentation results based on preset value element extraction rules related to patent pledge financing and in combination with a domain term library and a relationship library include: Based on the preset value element extraction rules related to patent pledge financing, and in combination with the domain terminology library and relationship library, the initial key information is identified and extracted from multiple text segmentation results; Extract the publishing agency information of the information source from the industry news text and market report text, and obtain the reputation score of the publishing agency information by querying the preset reputation evaluation database; Based on the reputation score, the market size data, technology application scenarios, legal status information, and competitor information in the initial key information are weighted for credibility to obtain weighted key information; According to the reputation score, the weighted key information with a credibility lower than a preset credibility threshold is filtered to obtain the final key information; the credibility threshold is determined based on the reputation score of the issuing organization information.
[0026] Pre-defined value factor extraction rules related to patent pledge financing refer to a series of patterns, grammatical structures, or semantic rules defined manually or through machine learning methods based on the needs of patent pledge financing assessments, used to identify and extract specific types of information from text. These rules can be implemented using regular expression-based rule sets, extraction templates based on dependency parsing, or named entity recognition and relationship extraction models based on deep learning models. Domain terminology and relationship libraries refer to knowledge resources containing specialized vocabulary and interrelationships within specific fields (such as patent technology, finance, and law). These libraries can be implemented using ontologies, concept graphs, or relational databases. Multiple text segmentation results refer to word sequences and related linguistic information obtained by processing raw unstructured text with a segmentation tool. These results can be implemented using dictionary-based, statistical model-based, or deep learning model-based segmentation methods. Initial key information refers to raw information fragments that are initially identified and extracted from the segmentation results based on the preset extraction rules and the domain knowledge base, but have not yet been reliability-assessed or processed. These information can be implemented using structured data tables, sets of key-value pairs, or lists of objects. The publishing organization information of an information source refers to the name or logo of the author, publisher, or publishing unit of a text (especially industry news and market reports). This information can be implemented as a string, organization code, or unified identifier. A pre-established reputation evaluation database refers to a pre-established database or knowledge base used to store and query the reputation ratings or scores of various information publishing organizations. This database can be implemented as a relational database, graph database, or key-value store. A reputation score is a numerical value or rating derived from a quantitative assessment of the reliability, professionalism, or authority of an information publishing organization. This score can be calculated using methods such as expert scoring, historical data analysis, or user feedback aggregation. Credibility weighting is the process of assigning different weights to extracted key information based on the reputation score of the information source to reflect its reliability. This weighting can be implemented using linear weighting, nonlinear weighting, or weighting based on confidence intervals. A credibility threshold is a criterion used to filter out low-credibility information. Only information with a credibility above this threshold is retained. This threshold can be implemented using a fixed value, a dynamically adjusted value, or a value calculated based on a statistical distribution.
[0027] This solution addresses the diverse and uneven quality of multi-source unstructured text information by integrating source reliability assessment and processing into the basic key information extraction process. By extracting publisher information from industry news and market reports and obtaining reputation scores by querying a pre-defined reputation evaluation database, the initially extracted key information is subsequently weighted and filtered based on this reputation score. By weighting the credibility of key information susceptible to source influence, such as market size data, technology application scenarios, legal status information, and competitor information, information from highly reputable organizations is given higher weight, making it more influential in information aggregation. Furthermore, by determining a credibility threshold based on the reputation score and filtering out information with a weighted credibility below this threshold, the overall reliability of the key information ultimately used for the evaluation is ensured. This weighting and filtering mechanism based on source reputation effectively mitigates the negative impact of low-quality or unreliable information on the evaluation results, improving the accuracy and reliability of the extracted results. This process, combined with the basic key information extraction steps, forms a more robust information acquisition process, providing a more reliable data foundation for subsequent structured processing, quantification and final patent value assessment.
[0028] In one embodiment, the step of identifying and extracting key information from multiple text segmentation results based on preset value element extraction rules related to patent pledge financing, combined with a domain terminology library and a relationship library, can be implemented as follows: First, a rule-based and dictionary-based extractor, combined with preset value element extraction rules, a domain terminology library and a relationship library, is used to scan text data such as patent texts, market reports, industry news, and legal announcements that have undergone word segmentation processing, and preliminarily identifies and extracts all key information fragments that meet the rules, such as market size values, technical terms, and legal status descriptions, to obtain an initial set of key information. Next, for the key information extracted from the industry news text and market report text, named entity recognition technology or pattern matching rules are used to identify and extract the name of the organization that published the news or report from the text, such as "CCID Consultants," "Bloomberg," "Xinhua News Agency," etc., as the publishing organization information of the information source. Then, by querying a pre-built reputation assessment database that stores the reputation scores of known organizations, for example, the reputation score of "CCID Consultants" is set to 0.9, and the reputation score of "an unknown tabloid" is set to 0.3. Based on the credibility score of the issuing agency, the corresponding market size data, technology application scenario description, legal status information, and competitor information in the initial key information are weighted for credibility. For example, the original data value can be multiplied by a weight factor calculated based on the credibility score, or the priority of the information in subsequent processing can be adjusted. Finally, a credibility threshold is set based on the credibility score. For example, if the credibility score is lower than 0.5, the credibility threshold of the key information from that source is set to a lower value. If the credibility of the weighted key information is lower than the threshold, the key information is removed from the set, resulting in the final set of key information used for subsequent evaluation.
[0029] By introducing a reputation assessment mechanism for information sources and weighting and filtering the extracted key information based on the reputation score, this solution can effectively identify and reduce the impact of noise from unreliable sources on the accuracy of key information extraction. This makes the key information ultimately used in patent value assessment more reliable, thereby improving the accuracy of patent value assessment results.
[0030] In some embodiments, the step of structuring the extracted key information to generate structured data including an application scenario field, a market size field, a legal status field, and a competitor information field includes: A1. Extract technical keywords from the application scenario field and determine the technical field labels corresponding to the technical keywords based on the preset patent classification system; A2. Based on the technical field label, query the preset industry standard database to obtain industry standard information related to the technical field label; A3. Determine whether industry standard information already exists in the application scenario field. If not, add the industry standard information to the application scenario field to supplement the application scenario description information and update the structured data including the application scenario field, market size field, legal status field, and competitor information field.
[0031] An application scenario field refers to textual descriptions of potential application areas, industries, or scenarios within structured data. This field can be implemented as a string, a text block, or a field within a specific data structure. Technical keywords refer to words or phrases identified from the application scenario description that represent the core content or application direction of the technology. These keywords can be extracted using keyword extraction algorithms within natural language processing (e.g., TF-IDF, TextRank, or deep learning-based methods). A patent classification system refers to a classification framework that systematically divides and organizes patent technologies, such as the International Patent Classification (IPC), the Concordat Classification (CPC), or national patent classification systems. These can be implemented using tree-like, hierarchical, or graph-like data models. A technical field label refers to an identifier or code within the patent classification system that corresponds to an extracted technical keyword and represents a specific technical field. These labels can be implemented using classification numbers, category names, or a combination of these. An industry standards database refers to a collection of industry standards, specifications, and technical requirements. These standards are typically issued by industry associations, standardization organizations, or government agencies. These databases can be implemented using relational databases, document databases, or knowledge graphs. Industry standard information refers to the standard text, standard number, issuing agency, scope of application and other information related to specific technical fields obtained from the industry standard database. It can be implemented using structured records, unstructured text or linked data.
[0032] Based on the structured processing of the extracted key information, this solution further refines and improves the application scenario description information in the structured data. By linking technical fields and industry standards, it makes it more closely aligned with actual industry applications, thereby improving the accuracy and comprehensiveness of the judgment of technology application prospects in patent value assessment. Specifically, step A1 extracts technical keywords from the application scenario field in the structured data and maps these keywords to standardized technical field labels based on a preset patent classification system. This transforms the unstructured application scenario description into standardized technical field information that can be used for subsequent related queries, laying the foundation for subsequent steps. Step A2, based on the technical field labels determined in step A1, queries a preset industry standard database to obtain industry standard information related to the technical field. This industry standard information is an important basis for measuring technology maturity, market acceptance, and potential application specifications. Step A3 determines whether the acquired industry standard information is already reflected in the original application scenario field. If not, this important industry standard information is added to the application scenario field to supplement and enrich the application scenario description. In this way, even if the original text does not explicitly mention the relevant industry standards, standard information closely related to actual industry applications can be incorporated into the structured data through association with technical fields, making the application scenario description more comprehensive and specific, and providing more accurate and valuable reference information for subsequent patent value assessment based on structured data. This supplementation and improvement of application scenario information enables subsequent quantitative assessments based on structured data to more fully consider the fit between the technology and actual industry needs and the maturity of the technology, thereby improving the reliability of patent value assessments.
[0033] In a specific embodiment, the step of structuring the extracted key information to generate structured data containing application scenario fields, market size fields, legal status fields, and competitor information fields can be implemented as follows: First, in step A1, a statistical method, such as the TF-IDF algorithm, can be used to extract highly weighted terms from the text description of the application scenario field as technical keywords. Then, using a pre-built mapping table or model that associates technical keywords with IPC classification numbers, the extracted technical keywords are mapped to one or more IPC classification numbers as technical field labels. Next, in step A2, based on the obtained IPC classification number, a database storing industry standard information corresponding to each IPC classification number is searched. This database may include information such as the standard name, standard number, release date, and applicable industries. For example, if the technical field label is "G06F 17 / 30" (database structure), the database standards related to this classification number, such as SQL standards and NoSQL standards, are searched. Subsequently, in step A3, the retrieved industry standard information is compared with the original text of the application scenario field to determine whether the standard information has been mentioned. If an important related standard (such as a key interface standard or data format standard) is found to be missing in the original description, the standard information (such as the standard name and standard number) is added to the end of the application scenario field or a dedicated subfield to update the entire structured data record.
[0034] Through the above-mentioned technical means, this solution can effectively supplement and improve the information of the application scenario field in the structured data, and take into account the industry standards closely related to the technical field. This makes the description of the application scenario of the patent technology more specific and comprehensive, and makes up for the problem of missing industry standard information that may be caused by relying solely on the original text. Therefore, when conducting subsequent patent value assessments, the market potential, maturity and feasibility of the patent technology can be more accurately judged based on more complete application scenario information that is closer to actual industrial applications, thereby improving the accuracy and reliability of the patent value assessment results.
[0035] In some embodiments, the specific steps in step A1 include: A11. Extract technical keywords from the application scenario field and determine the corresponding technical field labels based on the preset patent classification system. A12. Determine whether the technical field label corresponding to the technical keyword has been successfully determined. If so, proceed to step A2. If not, use a fuzzy matching algorithm based on the technical keyword to search for similar technical field labels within the pre-set patent classification system. The label with the highest similarity is selected as a candidate technical field label. A13. Calculate the semantic similarity between the technical keyword and the candidate technical field label, and determine whether the semantic similarity is greater than a preset similarity threshold. If so, use the candidate technical field label as the technical field label and execute step A2. Otherwise, terminate the process and stop executing subsequent steps A2-A3 to supplement the industry standard information.
[0036] A fuzzy matching algorithm is a technique for finding patterns in a dataset that don't exactly match a given pattern but are sufficiently similar. This can be achieved using string edit distance algorithms (such as the Levenshtein distance), N-gram similarity algorithms, or similarity calculation methods based on vector space. Semantic similarity measures the degree of closeness in meaning between two words, phrases, or text fragments. This can be achieved using vector cosine similarity calculated using word embedding models (such as Word2Vec and GloVe), path distance calculations based on knowledge graphs, or semantic encoding methods based on deep learning models. A preset similarity threshold is a numerical limit set in advance when making similarity judgments, used to distinguish between "sufficiently similar" and "not similar enough." This threshold can be set and adjusted based on the needs and experience of the actual application scenario.
[0037] The solution of this application further incorporates a judgment mechanism, a fuzzy matching mechanism, and a semantic similarity verification mechanism, building upon the basic steps of extracting technical keywords from the application scenario field and determining technical field labels based on a pre-defined patent classification system. Specifically, a direct match is first attempted (step A11), which is the most direct and effective approach. If the direct match fails to successfully determine the technical field label, the process does not abort, but instead enters an alternative search path (step A12). In this alternative path, a fuzzy matching algorithm is employed to conduct a broader search across the entire pre-defined patent classification system, searching for labels similar to the technical keywords. The label with the highest similarity is selected as a potential candidate. This step expands the search scope and increases the likelihood of finding relevant labels. However, the fuzzy matching results may not always accurately reflect the true meaning of the technical keywords. Therefore, after obtaining candidate labels, a semantic similarity calculation is further introduced (step A13). By calculating the semantic similarity between the technical keywords and the candidate technical field labels, the validity of the candidate labels can be verified at the semantic level. Only when the semantic similarity exceeds a pre-defined threshold is the candidate label considered reliable and adopted as the final technical field label, proceeding to the subsequent steps of obtaining industry standard information and supplementing the application scenario description. If the semantic similarity is insufficient, the candidate label is discarded and the subsequent supplementation process is stopped to avoid the introduction of irrelevant industry standard information. Through this combined process of judgment-fuzzy matching-semantic verification, the solution of this application can attempt to determine the technical field label in a more flexible and semantically verified manner when direct and accurate matching is not possible, thereby improving the success rate and accuracy of label determination, making the subsequent supplementation of industry standard information more effective, and further improving the application scenario description in the structured data, providing a more comprehensive information basis for patent value assessment.
[0038] For example, in one specific embodiment, the technical keyword extracted from the application scenario field is "smart home energy management." First, an attempt is made to directly search for a classification code that exactly matches "smart home energy management" within a pre-defined patent classification system (e.g., the IPC classification), such as "G05B 15 / 02" (program-based control systems, particularly programmable logic controllers) or "H02J 13 / 00" (control or regulation of power networks). If direct matching fails to find a suitable classification code, a fuzzy matching process is initiated. An N-gram-based similarity calculation method can be used to compare the technical keyword "smart home energy management" with all tags in the patent classification system to find the tag with the highest similarity. For example, candidate tags such as "smart grid energy management" and "home automation system" may be found. Assume that the candidate tag with the highest similarity is "smart grid energy management." Next, the semantic similarity between the technical keyword "smart home energy management" and the candidate tag "smart grid energy management" is calculated. A pre-trained word embedding model (e.g., the BERT model) can be used to encode the two phrases into vectors, and then the cosine similarity between them is calculated. If the calculated semantic similarity is greater than a preset similarity threshold (e.g., 0.7), "Smart Grid Energy Management" is identified as a technical field tag, and subsequent steps are performed to query industry standard information related to this tag. If the semantic similarity is less than the threshold, the candidate tag is considered inaccurate and subsequent industry standard information supplementation is stopped.
[0039] By introducing fuzzy matching and semantic similarity judgment mechanisms, the solution of this application can attempt to find potentially relevant tags in a more flexible way when technical keywords cannot be directly and accurately matched with the preset classification system labels, and ensure the accuracy of the tags through semantic verification. This improves the success rate and accuracy of determining technical field tags from the application scenario field, especially when dealing with emerging technologies or non-standard expressions, thereby more effectively utilizing industry standard information to supplement the application scenario description, making the structured data more complete and providing more comprehensive information for subsequent patent value assessment.
[0040] In certain embodiments, the steps of normalizing the market size data in the structured data and classifying or scoring the market tendency description to obtain quantified value factor data include: Obtaining data type information of the market size data and determining whether the data type information is numerical data; if so, performing Z-Score normalization processing on the market size data according to a preset normalization rule to obtain normalized market size data; if not, converting the market size data into numerical data and performing Z-Score normalization processing on the converted market size data according to a preset normalization rule to obtain normalized market size data; Obtain text information describing market trends, calculate the sentiment score of the text information using a preset sentiment dictionary, and classify the text information into three categories: positive, neutral, and negative according to preset classification rules; Perform weighted calculation based on the sentiment score, the preset sentiment score weight, and the preset classification weight of the corresponding type to obtain quantified market tendency data; Integrate the standardized market size data and the quantified market tendency data to obtain the quantified value factor data.
[0041] Pre-set normalization rules refer to mathematical formulas or algorithms used to convert data to a uniform dimension. This can be achieved through maximum and minimum value normalization, Z-score normalization, or decimal scaling. Z-score normalization is a commonly used data normalization method that converts the data into a distribution with a mean of 0 and a standard deviation of 1 by calculating the difference between the original data and the mean and dividing it by the standard deviation. A pre-set sentiment lexicon refers to a database containing words and their corresponding sentiment polarity or intensity scores. This can be implemented using a rule-based sentiment lexicon or a machine learning-based sentiment lexicon. A sentiment score is a numerical value calculated by analyzing sentiment words and their combinations within a text, reflecting the overall sentiment orientation and intensity of the text. Pre-set classification rules refer to the logic or algorithm used to classify text information into pre-set categories (such as positive, neutral, and negative) based on its content or sentiment score. This can be implemented using threshold-based rules or machine learning-based classification models. A pre-set sentiment score weight is a coefficient assigned to the sentiment score during the weighting calculation, reflecting its importance in quantifying market sentiment. The preset classification weights for each category (e.g., positive, neutral, negative) are coefficients assigned to different categories in a weighted calculation, reflecting their contribution or influence on market trends. Weighted calculations involve multiplying multiple values (e.g., sentiment scores and category information) by their corresponding weights and then performing summation or averaging to produce a composite value. Integration refers to the merging or linking of two or more independent data sets to form a unified dataset.
[0042] This solution specifies how to quantify market size data and market propensity descriptions in structured data. This addresses the issues of non-numeric data that may prevent direct standardization, and the inability to effectively utilize sentiment information in the classification results of market propensity descriptions, resulting in inaccurate quantification. Specifically, for market size data, the data type information is first obtained and a determination is made as to whether it is numeric. Based on this data type, if it is numeric, Z-score normalization is performed on it according to pre-set normalization rules to obtain standardized market size data. This allows numerical market size data of different magnitudes and units to be converted to a unified scale, making them comparable. If the data is not numeric, an attempt is made to convert it to numeric data first, and then Z-score normalization is performed on the converted data according to pre-set normalization rules to obtain standardized market size data. This addresses the issue of non-numeric market size data being unable to be directly calculated and compared numerically, and through standardization, it is made usable for subsequent quantitative evaluation. For market propensity descriptions, the textual information is first obtained. Using a preset sentiment dictionary, sentiment analysis is performed on text information, calculating a sentiment score. The text information is then categorized into three categories: positive, neutral, and negative, based on preset classification rules. By calculating the sentiment score using the preset sentiment dictionary, qualitative market propensity descriptions are converted into quantitative scores reflecting the intensity and direction of sentiment. Furthermore, the classification is categorized into preset categories, providing a basis for subsequent weighted calculations. After obtaining the sentiment score and categorization, a weighted calculation is performed based on the sentiment score, the preset sentiment score weight, and the corresponding category weight to obtain quantified market propensity data. This combines the intensity of the sentiment score with the category information and introduces weights for weighted calculation to produce a comprehensive quantitative indicator that more comprehensively and objectively reflects market propensity, resolving the issue of qualitative textual information being difficult to directly use for quantitative assessment. Finally, the standardized market size data and the quantified market propensity data are integrated to obtain quantified value factor data. This step combines the differently processed but quantified market size and market propensity data into a unified, quantitative dataset containing multiple aspects of market information, providing direct input for the subsequent patent value assessment model.Based on acquiring multi-source unstructured text information, performing word segmentation, extracting key information, and then performing structural processing, this solution further refines and accurately quantifies the market size and market tendency information in structured data. By processing non-numerical data and combining sentiment scores with classification for weighted calculations, the market information extracted from the text can be more effectively converted into numerical values that can be used for evaluation. Furthermore, through sentiment scores and weight coefficients, the quantitative results are closer to the actual situation and more comprehensively reflect market tendencies, thereby supporting more accurate and objective evaluations.
[0043] For example, when evaluating the value of a patent, the system extracts market size data from a market report, stating, "The market size is expected to reach trillions of yuan over the next five years," and also describes the market sentiment as, "This technology is expected to revolutionize the industry landscape, and the market outlook is extremely optimistic." First, the system obtains the data type information for the market size data, determining that it is non-numeric. The system converts "trillions of yuan" to numeric data, for example, to 1 trillion yuan, and then performs Z-score normalization based on pre-set normalization rules to obtain the normalized market size data. Simultaneously, the system obtains textual information describing the market sentiment as, "This technology is expected to revolutionize the industry landscape, and the market outlook is extremely optimistic." Using a pre-set sentiment dictionary, the system identifies sentiment terms such as "promising," "radical change," and "extremely optimistic," and calculates a sentiment score for the text, for example, 0.9. Based on pre-set classification rules, the text is classified as positive. Assuming the preset sentiment score weight is 0.6 and the preset positive category weight is 0.4, the system performs a weighted calculation based on the sentiment score of 0.9, the sentiment score weight of 0.6, and the positive category weight of 0.4 to obtain quantified market propensity data. For example, the calculation result is 0.9*0.6+1*0.4=0.94 (assuming the corresponding value of the positive category is 1). Finally, the system integrates the standardized market size data and the quantified market propensity data of 0.94 to obtain quantified value factor data for subsequent patent value assessment.
[0044] Through the above technical solution, this application solves the problem that market size data may contain non-numerical data, which makes it impossible to directly perform standardization processing. This allows different forms of market size data to be effectively standardized and comparable. At the same time, this application solves the problem that the classification results of market tendency descriptions cannot effectively utilize emotional information, resulting in inaccurate quantitative results. By combining emotional scores and classifications for weighted calculations, market tendencies are quantified more comprehensively and accurately, making the quantitative results closer to actual market sentiment and providing more reliable and accurate input for subsequent patent value assessments.
[0045] In some embodiments, the steps of obtaining text information describing market trends and calculating a sentiment score of the text information using a preset sentiment dictionary include: Preprocess the text information describing market trends; preprocessing includes removing stop words, word segmentation, and part-of-speech tagging; For the pre-processed text information, match it in the sentiment dictionary to identify the sentiment words contained in the text information; the sentiment dictionary contains multiple sentiment words and their corresponding sentiment polarity and intensity scores; Based on the identified sentiment words, find their corresponding sentiment polarity and intensity scores in the sentiment dictionary; Based on the sentiment polarity and intensity scores, a weighted summation algorithm is used to calculate the sentiment score of the text information.
[0046] Preprocessing refers to the process of normalizing and cleaning raw text. This can be achieved through techniques such as text cleaning, format conversion, and encoding unification. Stop word removal involves removing frequently occurring words from a text that contribute little to the text's meaning. This can be achieved by matching and deleting words against a preset stop word list. Word segmentation involves dividing a continuous text sequence into semantically meaningful word units. This can be achieved through methods based on dictionaries, statistical models, or deep learning. Part-of-speech tagging involves identifying the grammatical category of each word in a text. This can be achieved through methods such as hidden Markov models, conditional random fields, or neural network models. A sentiment lexicon is a database containing words and their corresponding sentiment attributes. It can include information such as words, sentiment polarity, and sentiment intensity. Sentiment polarity refers to the direction of emotion expressed by a word or text, which can be expressed as positive, negative, or neutral. Intensity scoring quantifies the strength of the emotion expressed by a word or text and can be a numerical value. A weighted summation algorithm is a calculation method that multiplies multiple numerical values by their respective weights and then adds them together. Weights can be determined based on a word's importance, position, or relationship to other words.
[0047] This application aims to more accurately quantify the sentiment expressed in the text by performing sophisticated processing and calculations on the textual information describing market trends, thereby improving the accuracy of market trend data quantification and, in turn, enhancing the reliability of patent value assessments. Specifically, the original market trend description text is preprocessed by removing stop words, performing word segmentation, and performing part-of-speech tagging. This converts the raw text into a structured word sequence, removes noise, and prepares clean data for subsequent analysis. Next, the processed word sequence is searched and matched against a pre-set sentiment lexicon to identify key words that actually express sentiment in the text. Based on the identified sentiment words, the pre-set sentiment polarity (positive, negative, etc.) and intensity score for each sentiment word are retrieved from the sentiment lexicon. Finally, based on these sentiment polarity and intensity scores, a weighted summation algorithm is used to comprehensively consider the influence of all sentiment words in the text to calculate an overall sentiment score for the entire text. This approach, through meticulous linguistic processing and lexicon-based quantification of the text, combined with a more sophisticated computational approach, can more accurately capture the complex sentiment expressed in the text, overcoming the shortcomings of simply counting sentiment words. This more accurate sentiment score calculation provides a more reliable basis for the subsequent quantification of market tendency data, thereby making the integrated quantified value factor data more accurate, and ultimately improving the reliability of patent value assessment.
[0048] For example, when processing a piece of text describing market sentiment, one can first utilize open-source natural language processing toolkits, such as the Jieba library, for word segmentation and part-of-speech tagging. This can then be filtered using a commonly used Chinese stop word list to obtain a preprocessed word list. This word list is then compared against a pre-built sentiment lexicon containing a large number of Chinese sentiment terms and their corresponding polarity and intensity values to identify sentiment-expressing words in the text, such as "optimistic," "worried," "growth," and "decline." Based on the matched sentiment terms, the sentiment polarity (e.g., "optimistic" is positive, "worried" is negative) and corresponding intensity values (e.g., "optimistic" has an intensity of 0.8, "worried" has an intensity of 0.9) are extracted from the sentiment lexicon. Finally, a weighted summation algorithm is used to calculate the overall sentiment score of the text using these polarity and intensity values. For example, positive terms can be assigned positive weights and negative terms negative weights. Weights can be adjusted based on the term's position in the text and whether it is modified by degree adverbs or negation words. Finally, the weighted scores of all sentiment terms are summed to obtain a single value reflecting the overall sentiment of the text.
[0049] By preprocessing the textual information describing market trends and removing stop words, word segmentation, and part-of-speech tagging, the accuracy of subsequent sentiment word recognition can be improved. The preprocessed textual information is matched and identified within a sentiment dictionary, accurately locating the emotional expressions within the text. The identified sentiment words are searched for their corresponding sentiment polarity and intensity scores within the sentiment dictionary, providing basic data for sentiment quantification. Based on the sentiment polarity and intensity scores, a weighted summation algorithm is used to calculate the sentiment score of the textual information. This method comprehensively considers the influence of multiple sentiment words to produce a more refined and accurate text sentiment score. This more accurate sentiment score calculation can more realistically reflect the market's inclination toward patented technologies, providing more reliable data for subsequent market trend quantification, thereby improving the accuracy of patent value assessment.
[0050] Reference Attachment Figure 2 The present invention provides a patent value assessment system, comprising: The acquisition module 100 is used to acquire multi-source unstructured text information; the multi-source unstructured text information includes patent texts, market report texts, industry news texts, and legal notice texts; The word segmentation module 200 is used to perform word segmentation processing on multi-source unstructured text information using a domain terminology library and a relationship library, and obtain multiple text segmentation results; the domain terminology library and the relationship library contain patent technology field terms, financial industry terms, and legal terms and their associated relationships; Extraction module 300 is used to identify and extract key information from multiple text segmentation results based on preset value element extraction rules related to patent pledge financing and in combination with the domain term library and relationship library; key information includes technology application scenarios, market size data, legal status information, and competitor information; The generation module 400 is used to structure the extracted key information to generate structured data including application scenario fields, market size fields, legal status fields, and competitor information fields; The quantification module 500 is used to standardize the market size data in the structured data and classify or score the market tendency description to obtain quantified value factor data; The evaluation module 600 is used to evaluate the patent value based on the value factor data.
[0051] In certain embodiments, the extraction module 300 is configured to identify and extract key information from multiple text segmentation results based on preset value element extraction rules related to patent pledge financing and in combination with a domain term library and a relationship library: Based on the preset value element extraction rules related to patent pledge financing, and in combination with the domain terminology library and relationship library, the initial key information is identified and extracted from multiple text segmentation results; Extract the publishing agency information of the information source from the industry news text and market report text, and obtain the reputation score of the publishing agency information by querying the preset reputation evaluation database; Based on the reputation score, the market size data, technology application scenarios, legal status information, and competitor information in the initial key information are weighted for credibility to obtain weighted key information; According to the reputation score, the weighted key information with a credibility lower than a preset credibility threshold is filtered to obtain the final key information; the credibility threshold is determined based on the reputation score of the issuing organization information.
[0052] In some embodiments, when the generation module 400 is used to structure the extracted key information and generate structured data including an application scenario field, a market size field, a legal status field, and a competitor information field, the following is executed: A1. Extract technical keywords from the application scenario field and determine the technical field labels corresponding to the technical keywords based on the preset patent classification system; A2. Based on the technical field label, query the preset industry standard database to obtain industry standard information related to the technical field label; A3. Determine whether industry standard information already exists in the application scenario field. If not, add the industry standard information to the application scenario field to supplement the application scenario description information and update the structured data including the application scenario field, market size field, legal status field, and competitor information field.
[0053] In some embodiments, the generation module 400 is configured to extract technical keywords from the application scenario field and determine technical field labels corresponding to the technical keywords based on a preset patent classification system. A11. Extract technical keywords from the application scenario field and determine the corresponding technical field labels based on the preset patent classification system. A12. Determine whether the technical field label corresponding to the technical keyword has been successfully determined. If so, proceed to step A2. If not, use a fuzzy matching algorithm based on the technical keyword to search for similar technical field labels within the pre-set patent classification system. The label with the highest similarity is selected as a candidate technical field label. A13. Calculate the semantic similarity between the technical keyword and the candidate technical field label, and determine whether the semantic similarity is greater than a preset similarity threshold. If so, use the candidate technical field label as the technical field label and execute step A2. Otherwise, terminate the process and stop executing subsequent steps A2-A3 to supplement the industry standard information.
[0054] In certain embodiments, the quantification module 500 performs the following steps when standardizing the market size data in the structured data and classifying or scoring the market tendency description to obtain the quantified value factor data: Obtaining data type information of the market size data and determining whether the data type information is numerical data; if so, performing Z-Score normalization processing on the market size data according to a preset normalization rule to obtain normalized market size data; if not, converting the market size data into numerical data and performing Z-Score normalization processing on the converted market size data according to a preset normalization rule to obtain normalized market size data; Obtain text information describing market trends, calculate the sentiment score of the text information using a preset sentiment dictionary, and classify the text information into three categories: positive, neutral, and negative according to preset classification rules; Perform weighted calculation based on the sentiment score, the preset sentiment score weight, and the preset classification weight of the corresponding type to obtain quantified market tendency data; Integrate the standardized market size data and the quantified market tendency data to obtain the quantified value factor data.
[0055] In some embodiments, the quantification module 500 performs the following when obtaining text information describing market trends and calculating the sentiment score of the text information using a preset sentiment dictionary: Preprocess the text information describing market trends; preprocessing includes removing stop words, word segmentation, and part-of-speech tagging; For the pre-processed text information, match it in the sentiment dictionary to identify the sentiment words contained in the text information; the sentiment dictionary contains multiple sentiment words and their corresponding sentiment polarity and intensity scores; Based on the identified sentiment words, find their corresponding sentiment polarity and intensity scores in the sentiment dictionary; Based on the sentiment polarity and intensity scores, a weighted summation algorithm is used to calculate the sentiment score of the text information.
[0056] In this document, relational terms such as first and second, etc. are used merely to distinguish one entity or operation from another entity or operation, but do not necessarily require or imply any actual relationship or order between these entities or operations.
[0057] The foregoing description is merely an embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A patent value assessment method, characterized in that: The following steps are involved: Acquire multi-source unstructured text information; Using a domain term library and a relationship library, the multi-source unstructured text information is segmented to obtain multiple text segmentation results; the domain term library and relationship library contain patent technology field terms, financial industry terms, and legal-related terms and their associated relationships; According to preset value element extraction rules related to patent pledge financing, and in combination with the domain term library and relationship library, identify and extract key information from the multiple text segmentation results; Structuring the extracted key information to generate structured data; Standardizing the market size data in the structured data, and classifying or scoring the market tendency descriptions to obtain quantified value factor data; Based on the value factor data, the patent value is evaluated.
2. The patent value assessment method according to claim 1, characterized in that: The multi-source unstructured text information includes patent texts, market report texts, industry news texts and legal announcement texts.
3. The patent value assessment method according to claim 2, characterized in that: The key information includes technology application scenarios, market size data, legal status information and competitor information.
4. The patent value assessment method according to claim 3, characterized in that: According to preset value element extraction rules related to patent pledge financing, and in combination with the domain term library and relationship library, the steps of identifying and extracting key information from the multiple text segmentation results include: According to preset value element extraction rules related to patent pledge financing, and in combination with the domain term library and relationship library, initial key information is identified and extracted from the multiple text segmentation results; Extracting information about the publishing organization of the information source from the industry news text and the market report text, and obtaining a reputation score of the publishing organization information by querying a preset reputation evaluation database; Based on the reputation score, the market size data, technology application scenarios, legal status information, and competitor information in the initial key information are weighted for credibility to obtain weighted key information; According to the reputation score, the key information with a credibility lower than a preset credibility threshold in the weighted key information is filtered to obtain the final key information; the credibility threshold is determined based on the reputation score of the issuing organization information.
5. The patent value assessment method according to claim 3, characterized in that: The steps of structuring the extracted key information and generating structured data include: The extracted key information is structured to generate structured data containing application scenario fields, market size fields, legal status fields, and competitor information fields.
6. The patent value assessment method according to claim 5, characterized in that: The steps for structuring the extracted key information and generating structured data containing application scenario fields, market size fields, legal status fields, and competitor information fields include: A1. Extract technical keywords from the application scenario field and determine the technical field labels corresponding to the technical keywords based on a preset patent classification system; A2. According to the technical field label, query the preset industry standard database to obtain industry standard information related to the technical field label; A3. Determine whether the industry standard information already exists in the application scenario field; if not, add the industry standard information to the application scenario field to supplement the application scenario description information, and update the structured data including the application scenario field, market size field, legal status field, and competitor information field.
7. The patent value assessment method according to claim 6, characterized in that: The specific steps in step A1 include: A11. Extract technical keywords from the application scenario field and, based on a preset patent classification system, determine the technical field labels corresponding to the technical keywords; A12. Determine whether the technical field label corresponding to the technical keyword has been successfully determined; if so, proceed to step A2; if not, use a fuzzy matching algorithm based on the technical keyword to search for similar technical field labels within the preset patent classification system, and select the label with the highest similarity as a candidate technical field label; A13. Calculate the semantic similarity between the technical keyword and the candidate technical field label, and determine whether the semantic similarity is greater than a preset similarity threshold; if so, use the candidate technical field label as the technical field label and execute step A2; otherwise, terminate the process and stop executing subsequent steps A2-A3 to supplement the industry standard information.
8. The patent value assessment method according to claim 1, characterized in that: The steps of standardizing the market size data in the structured data and classifying or scoring the market tendency description to obtain quantified value factor data include: Obtaining data type information of the market size data, and determining whether the data type information is numerical data; if so, performing Z-Score normalization processing on the market size data according to a preset normalization rule to obtain normalized market size data; if not, converting the market size data into numerical data, and performing Z-Score normalization processing on the converted market size data according to a preset normalization rule to obtain normalized market size data; Obtaining text information describing the market tendency, calculating a sentiment score of the text information using a preset sentiment dictionary, and classifying the text information into three categories: positive, neutral, and negative according to preset classification rules; Performing weighted calculation based on the sentiment score, the preset sentiment score weight, and the preset classification weight of the corresponding type to obtain quantified market tendency data; The standardized market size data and the quantified market tendency data are integrated to obtain quantified value factor data.
9. The patent value assessment method according to claim 8, characterized in that: The steps of obtaining the text information describing the market tendency and calculating the sentiment score of the text information using a preset sentiment dictionary include: Preprocessing the text information describing the market tendency; the preprocessing includes removing stop words, segmenting words, and tagging parts of speech; For the pre-processed text information, matching is performed in the sentiment dictionary to identify the sentiment words contained in the text information; the sentiment dictionary contains a plurality of sentiment words and their corresponding sentiment polarity and intensity scores; According to the identified sentiment words, searching for the corresponding sentiment polarity and intensity score in the sentiment dictionary; Based on the sentiment polarity and intensity scores, a weighted summation algorithm is used to calculate the sentiment score of the text information.
10. A patent value assessment system, characterized in that: include: An acquisition module is used to acquire multi-source unstructured text information; the multi-source unstructured text information includes patent texts, market report texts, industry news texts and legal announcement texts; A word segmentation module is used to perform word segmentation processing on the multi-source unstructured text information using a domain term library and a relationship library to obtain multiple text segmentation results; the domain term library and the relationship library contain patent technology field terms, financial industry terms and legal terms and their associated relationships; An extraction module, configured to identify and extract key information from the plurality of text segmentation results based on preset value element extraction rules related to patent pledge financing and in combination with the domain term library and relationship library; the key information includes technology application scenarios, market size data, legal status information, and competitor information; The generation module is used to structure the extracted key information and generate structured data containing application scenario fields, market size fields, legal status fields, and competitor information fields; A quantification module is used to standardize the market size data in the structured data and classify or score the market tendency description to obtain quantified value factor data; An evaluation module is used to evaluate the patent value based on the value factor data.