Enterprise innovation ability portrait description method and device, electronic equipment and medium
By preprocessing relevant enterprise information and evaluating patent evolution technology indicators, combined with peer cosine similarity calculation and Delphi principal component analysis, the problem of low accuracy of evaluation results in existing methods has been solved, achieving accurate evaluation of enterprise innovation capabilities and improving the accuracy and credibility of evaluation results.
Patent Information
- Application Number
- CN202511447693.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-11
- Publication Date
- 2026-02-17
AI Technical Summary
Existing methods for profiling enterprise innovation capabilities are insufficient to meet the needs of precise assessment, resulting in low accuracy. Traditional patent valuation methods do not adequately consider dynamic dimensions, and the assessment results fail to reflect the cutting-edge nature and evolution potential of the technology.
By preprocessing relevant enterprise information, screening core data, and integrating patent text information, R&D investment data, etc., and combining patent evolution technology indicator evaluation and peer cosine similarity calculation, an explanatory framework is constructed using the Delphi method and principal component analysis to evaluate the enterprise's innovation capability.
It enables precise assessment of enterprises' innovation capabilities, accurately reflects the cutting-edge nature and evolutionary potential of technologies, improves the accuracy and credibility of assessment results, and provides more accurate basis for government policy-making and financial decision-making.
Smart Images

Figure CN121544092A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of data processing technology, and more specifically, to a method, apparatus, electronic device, and medium for depicting a profile of an enterprise's innovation capabilities. Background Technology
[0002] Against the backdrop of rapid development in the technology industry, the innovation capability of technology companies has become a key basis for measuring the core competitiveness of enterprises, supporting industrial upgrading and the rational allocation of financial resources. As a core tool for quantitatively assessing the innovation level of enterprises, the accuracy and comprehensiveness of enterprise innovation capability profiles directly affect the formulation of government industrial policies, credit decisions of financial institutions and the enterprise's own R&D strategic planning. It plays an irreplaceable role in promoting the efficient flow of technology resources and optimizing the innovation ecosystem.
[0003] Currently, various practical methods have been developed within the industry for profiling corporate innovation capabilities. Regarding data collection and indicator selection, existing methods primarily construct evaluation systems based on fundamental data such as corporate R&D investment, financial revenue, and the number of intellectual property rights. Some methods introduce indicators such as patent citation counts, number of patent families, and legal status to supplement the dimensions of patent value assessment. At the data processing and modeling level, some studies use text mining techniques to extract keywords from patent abstracts or employ topic models to mine technical information from patent texts. Meanwhile, some methods leverage big data technology to integrate multi-source corporate data to construct a profiling framework. Other solutions attempt to integrate scientific and technological data from multiple perspectives or achieve specialized patent value assessments through structured fields.
[0004] However, existing methods for profiling corporate innovation capabilities still have significant limitations and cannot meet the needs of precise assessment. First, existing methods have a limited scope for capturing corporate innovation characteristics, resulting in low accuracy in assessment results that reflect the true level and development trend of corporate innovation capabilities. Second, traditional patent valuation methods rely solely on static indicators such as citations and patent families, failing to adequately consider dynamic dimensions, making it difficult for patent valuation results to reflect the cutting-edge nature and evolutionary potential of the technology. Summary of the Invention
[0005] The problem addressed by this invention is how to improve the accuracy of profiling a company's innovation capabilities.
[0006] To address the aforementioned issues, this invention provides a method, apparatus, electronic device, and medium for depicting a profile of an enterprise's innovation capabilities.
[0007] In a first aspect, the present invention provides a method for depicting a corporate innovation capability profile, including: Preprocessing the previously acquired enterprise-related information yields evaluation parameters; Based on the patent vectors of the past years in the evaluation parameters, the patent evolution technology indicators are evaluated, and the evaluation results are mapped to obtain the periodic score. Determine the word vector similarity of at least one peer company in the first year, determine the word vector similarity in the second year according to the evaluation parameters, obtain the peer cosine similarity between companies based on the word vector similarity of the first year and the word vector similarity of the second year, and map the peer cosine similarity to obtain a similarity score; Based on the explanatory framework constructed using the Delphi method and principal component analysis, the enterprise's innovation capability is obtained according to the evaluation parameters, the period score, and the similarity score.
[0008] Optionally, the evaluation parameters include the average sentence vector; the preprocessing of the pre-acquired enterprise-related information to obtain the evaluation parameters includes: A hybrid analyzer is used to extract keywords from the patent information in the enterprise-related information to obtain word vectors; Based on the patent information, the inverse document frequency of each word vector and its word frequency in the corresponding sentence are determined, and the corresponding weight is obtained by multiplying the word frequency and the inverse document frequency. The average sentence vector is obtained by weighting and summing the word vectors according to their corresponding weights.
[0009] Optionally, the evaluation of patent evolution technology indicators based on the patent vectors from previous years in the evaluation parameters includes: The types and number of keywords in the patent information of the relevant enterprise information for each year are counted, and a patent vector is generated by combining the corresponding patent text length. A word standardization vector is obtained based on the patent vector and the number of patents in the corresponding year. The vocabulary change is obtained based on the standardized vocabulary vectors from two consecutive years. The keywords from previous years are deduplicated to obtain word vectors, and the patent evolution technology index is obtained based on the word vectors and the word change.
[0010] Optionally, determining the word vector similarity for at least one peer company in the first year, and determining the word vector similarity for the second year based on the evaluation parameters, includes: Based on the patent information of the peer companies, the comparison vector for each year is extracted and organized, and the word vector similarity for the first year is obtained based on the comparison vectors of the two adjacent years. The word vector similarity for the second year is obtained based on the patent vectors from the two adjacent years.
[0011] Optionally, obtaining the peer cosine similarity between enterprises based on the word vector similarity of the first year and the word vector similarity of the second year includes: The first cosine similarity between each enterprise and itself is obtained based on the word vector similarity of the first year and the word vector similarity of the second year. The first cosine similarity of the corresponding enterprises is weighted based on their industry level. The peer cosine similarity is obtained by weighting the weights after division and the corresponding first cosine similarity.
[0012] Optionally, the peer companies include leading companies in the same industry.
[0013] Optionally, the method for depicting a corporate innovation capability profile further includes: The operating data in the relevant information of the enterprise is retrieved, and multiple credit assessment values are obtained by inputting them into multiple assessment models. The assessment models are obtained by training a preset model with historical operating data and historical credit assessment values. The preset model is constructed using different network models. The multiple credit assessment values are combined to obtain the final assessment value.
[0014] Secondly, the present invention provides a device for depicting a corporate innovation capability profile, comprising: The preprocessing module is used to preprocess the pre-acquired enterprise-related information to obtain evaluation parameters; The evaluation module is used to evaluate the patent evolution technology indicators based on the patent vectors of previous years in the evaluation parameters, and to map the evaluation results to obtain a periodic score. The processing module is used to determine the word vector similarity of at least one peer company in the first year, determine the word vector similarity of the second year according to the evaluation parameters, obtain the peer cosine similarity between companies based on the word vector similarity of the first year and the word vector similarity of the second year, and map the peer cosine similarity to obtain a similarity score. The description module is used to determine the enterprise's innovation capability based on the evaluation parameters, the period score, and the similarity score, using an interpretive framework constructed according to the Delphi method and principal component analysis.
[0015] Thirdly, the present invention provides an electronic device, including a memory and a processor; The memory is used to store computer programs; The processor is configured to implement the enterprise innovation capability profiling method as described in the first aspect when executing the computer program.
[0016] Fourthly, the present invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the enterprise innovation capability profiling method as described in the first aspect.
[0017] The beneficial effects of the enterprise innovation capability profiling method of this invention are as follows: By preprocessing relevant enterprise information, it can screen and integrate core data strongly related to enterprise innovation capability, such as patent text information, R&D investment data, and science and technology award information, eliminating redundant information and abnormal data interference. This provides a standardized and high-quality input foundation for subsequent patent evolution technology indicator evaluation, peer cosine similarity calculation, and enterprise innovation capability evaluation, effectively solving the problem of low accuracy of evaluation results caused by inconsistent data quality in existing methods, and ensuring the reliability and consistency of subsequent evaluation stages. Relying on the patent vectors of previous years in the evaluation parameters, the evaluation focuses on the dynamic evolution characteristics of patent technology. By quantifying the degree of evolution of patent technology over the years and mapping it to periodic scores, it can accurately capture the position of patent technology in its life cycle, such as its cutting-edge nature and maturity. Compared with existing traditional patent value evaluation methods that only rely on static indicators such as citations and family, this method effectively considers the dynamic dimensions of patent technology, enabling the evaluation results to truly reflect the cutting-edge nature and evolution potential of the technology, and making up for the shortcomings of existing methods in considering dynamic dimensions. By calculating and integrating the year-specific word vector similarity among peer companies to obtain the peer cosine similarity, the consistency and differentiation between the target company and its peers in terms of technological direction can be quantified. Introducing peer companies as a reference further enriches the evaluation dimensions of corporate innovation capabilities, avoiding the problem of a single evaluation perspective due to the lack of industry references in existing methods. This makes the subsequent corporate innovation capability evaluation results more reflective of the company's relative innovation level within the industry, improving the industry adaptability and accuracy of the evaluation results. Based on the explanatory framework constructed using the Delphi method and principal component analysis, expert experience and data patterns can be fully integrated to scientifically weight and integrate evaluation parameters, period scores, and similarity scores. On the one hand, the expert experience of the Delphi method ensures that the evaluation logic aligns with the actual innovation characteristics of the industry; on the other hand, the data patterns of principal component analysis eliminate redundant indicators and optimize weight allocation, avoiding the problem of biased evaluation results caused by unreasonable indicator weight settings in existing methods. The final output of corporate innovation capability results possesses both data-supported objectivity and conformity to industry practice, significantly improving the accuracy and credibility of the evaluation results.This invention establishes a complete enterprise innovation capability assessment system. On the one hand, it ensures data quality through preprocessing and supplements dynamic dimensions by combining patent evolution technology indicators, thus solving the core problems of low accuracy and difficulty in reflecting the cutting-edge nature and evolution potential of existing methods. On the other hand, it introduces industry references through peer cosine similarity calculation and achieves the scientific integration of multi-dimensional indicators based on the interpretation framework constructed by the Delphi method and principal component analysis. This ensures that the final enterprise innovation capability assessment results can accurately reflect the enterprise's own innovation level and technological evolution trend, as well as the enterprise's relative competitive position in the industry. This provides a more accurate and reliable basis for government industrial policy formulation, financial institution credit decisions, and enterprise's own R&D strategic planning, effectively promoting the efficient flow of scientific and technological resources and the optimization of the innovation ecosystem. Attached Figure Description
[0018] Figure 1 This is a flowchart illustrating the method for depicting a profile of an enterprise's innovation capabilities according to an embodiment of the present invention. Figure 2 This is a schematic diagram of the enterprise innovation capability profiling device according to an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed Implementation
[0019] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the accompanying drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0020] It should be understood that the various steps described in the method embodiments of the present invention may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present invention is not limited in this respect.
[0021] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to"; the term "based on" means "at least partially based on"; the term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments"; and the term "optionally" means "optional embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first," "second," etc., mentioned in this invention are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0022] It should be noted that the terms "one" and "more" used in this invention are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0023] The names of the messages or information exchanged between the multiple devices in the embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of these messages or information.
[0024] like Figure 1 As shown in the figure, an embodiment of the present invention provides a method for depicting a corporate innovation capability profile, comprising: Step S1: Preprocess the previously acquired enterprise-related information to obtain evaluation parameters.
[0025] Specifically, relevant enterprise information is collected in advance from three core sources: enterprise patent databases, annual business reports, and science and technology policy documents. This information includes patent text information such as patent titles, abstracts, and first claims, as well as enterprise operating data such as previous year's R&D expenses, previous year's operating revenue, R&D personnel ratio, and enterprise qualifications and honors, such as government science and technology certifications, the number of science and technology projects undertaken, and intellectual property infringement records. The data is then processed through deduplication, outlier correction, and standardization, ultimately integrating these into evaluation parameters that include patent characteristics, operating indicators, and qualification indicators.
[0026] Step S2: Evaluate the patent evolution technology index based on the patent vectors of previous years in the evaluation parameters, and map the evaluation results to obtain the period score.
[0027] Specifically, from the evaluation parameters, structured patent text data of the target company for each year are extracted. A word vector model (such as a pre-trained word vector model) is used to process the patent texts for each year, converting each patent text into a corresponding patent vector. Then, the arithmetic mean of all patent vectors for each year is taken to obtain the target company's patent vectors for each year, such as the 2020 patent vector, the 2021 patent vector, and so on. Based on the patent vectors for each year, patent evolution technology indicators are determined. Then, the original values of the patent evolution technology indicators for all companies in the industry are collected, and the global maximum and minimum values of the original values are determined. A linear normalization method is used to map the original values of the target company's patent evolution technology level for each year to, for example, the interval [0, 100], to obtain a periodic score. It should be noted that the closer the original value is to the global maximum value, and the closer the periodic score is to 100, the more advanced the patent technology and the faster its evolution. The closer the original value is to the global minimum value, and the closer the periodic score is to 0, the more mature or declining the patent technology.
[0028] Step S3: Determine the word vector similarity of at least one peer company in the first year, determine the word vector similarity of the second year according to the evaluation parameters, obtain the peer cosine similarity between companies based on the word vector similarity of the first year and the word vector similarity of the second year, and map the peer cosine similarity to obtain a similarity score.
[0029] Specifically, the business scope of the target company is extracted from the annual business report information of the evaluation parameters. Based on the business scope of the target company, at least one peer company of the same industry and scale is identified. Patent text data of the peer company over the years is obtained, and word vector models consistent with those of the target company are used to generate year-specific word vectors for the peer company. The cosine similarity between the word vector of a peer company in a given year and the word vectors of adjacent years is calculated to obtain the first year's word vector similarity, such as the cosine similarity between the word vectors of the peer company in 2020 and 2021, the cosine similarity between the word vectors of the peer company in 2021 and 2022, and so on. Structured patent text data of the target company over the years is extracted from the evaluation parameters, and the same word vector model is used to generate year-specific word vectors for the target company. The cosine similarity between the word vector of the target company in a given year and the word vectors of adjacent years is calculated to obtain the second year's word vector similarity, such as the cosine similarity between the word vectors of the target company in 2020 and 2021, the cosine similarity between the word vectors of the target company in 2021 and 2022, and so on. Based on the cosine similarity expression formula, the cosine similarity between companies is obtained according to the word vector similarity of the first year and the word vector similarity of the second year. The cosine similarity between companies in the industry is collected, and its global maximum and minimum values are determined. The linear normalization method is used to map the cosine similarity between companies to the interval such as [0, 100] to obtain the similarity score. The higher the score, the more consistent the technical direction of the target company and the peer companies are, and the stronger the adaptability of the technology in the industry.
[0030] Step S4: Based on the explanatory framework constructed using the Delphi method and principal component analysis, the enterprise's innovation capability is obtained according to the evaluation parameters, the period score, and the similarity score.
[0031] Specifically, the Delphi method was employed, and an expert team composed of scholars in the technology field, corporate technology leaders, and risk control experts from financial institutions was invited to conduct multiple rounds of anonymous scoring on the importance of evaluation parameters (including patent features, operational indicators, and qualification indicators), period scores, and similarity scores to determine the initial weights of each indicator. Principal component analysis was used to reduce the dimensionality of the multi-dimensional indicators in the evaluation parameters, calculate the variance contribution rate of each indicator, adjust the initial weights based on the variance contribution rate, and eliminate redundant indicators such as those with a variance contribution rate below 5%. Finally, an interpretive framework containing indicator weight allocation rules was constructed. Core indicators, such as R&D expenses, the proportion of R&D personnel, and the number of science and technology awards received, were extracted from the evaluation parameters. Combined with period scores and similarity scores, each indicator was assigned a corresponding weight according to the weight allocation rules in the interpretive framework. The standardized value of each indicator was multiplied by its corresponding weight to obtain the weighted score of each indicator. The weighted scores of all indicators were summed to obtain the original value of the enterprise's innovation capability, as shown in the following formula: , Among them, Y i X represents the raw value of a company's innovation capability. i The evaluation indicators related to the innovation capabilities of technology companies are shown in Table 1. i This represents the industry-lifecycle adaptive weighting, where n represents the number of evaluation metrics.
[0032] Table 1 - Assessment Indicators Related to Enterprise Innovation Capability
[0033] After obtaining the initial values, a grading standard is determined based on the distribution of the original values of enterprise innovation capabilities across the entire industry. For example, the top 20% of the original values are considered excellent, 20%-60% are considered good, and 60%-100% are considered qualified. The original values of enterprise innovation capabilities are then converted into corresponding enterprise innovation capability levels, such as excellent, good, and qualified, according to the grading standard. Alternatively, the original values can be directly retained as the quantitative result of enterprise innovation capabilities.
[0034] In this embodiment, by preprocessing relevant enterprise information, core data strongly related to enterprise innovation capabilities, such as patent text information, R&D investment data, and science and technology award information, can be screened and integrated. Redundant information and abnormal data interference are eliminated, providing a standardized and high-quality input foundation for subsequent patent evolution technology indicator evaluation, peer cosine similarity calculation, and enterprise innovation capability evaluation. This effectively solves the problem of low accuracy of evaluation results caused by inconsistent data quality in existing methods, ensuring the reliability and consistency of subsequent evaluation stages. Based on the patent vectors of previous years in the evaluation parameters, the evaluation focuses on the dynamic evolution characteristics of patent technology. By quantifying the degree of evolution of patent technology over the years and mapping it to periodic scores, the position of patent technology in its life cycle, such as its cutting-edge nature and maturity, can be accurately captured. Compared with existing traditional patent value evaluation methods that only rely on static indicators such as citations and patent families, this method effectively considers the dynamic dimensions of patent technology, enabling the evaluation results to truly reflect the cutting-edge nature and evolution potential of the technology, and making up for the shortcomings of existing methods in considering dynamic dimensions. By calculating and integrating the year-specific word vector similarity among peer companies to obtain the peer cosine similarity, the consistency and differentiation between the target company and its peers in terms of technological direction can be quantified. Introducing peer companies as a reference further enriches the evaluation dimensions of corporate innovation capabilities, avoiding the problem of a single evaluation perspective due to the lack of industry references in existing methods. This makes the subsequent corporate innovation capability evaluation results more reflective of the company's relative innovation level within the industry, improving the industry adaptability and accuracy of the evaluation results. Based on the explanatory framework constructed using the Delphi method and principal component analysis, expert experience and data patterns can be fully integrated to scientifically weight and integrate evaluation parameters, period scores, and similarity scores. On the one hand, the expert experience of the Delphi method ensures that the evaluation logic aligns with the actual innovation characteristics of the industry; on the other hand, the data patterns of principal component analysis eliminate redundant indicators and optimize weight allocation, avoiding the problem of biased evaluation results caused by unreasonable indicator weight settings in existing methods. The final output of corporate innovation capability results possesses both data-supported objectivity and conformity to industry practice, significantly improving the accuracy and credibility of the evaluation results.
[0035] The embodiments of the present invention form a complete enterprise innovation ability evaluation system. On the one hand, by ensuring data quality in the preprocessing link and combining patent evolution technical indicators to evaluate and supplement dynamic dimensions, the core problems of low accuracy of evaluation results and difficulty in reflecting technological frontiers and evolution potential in existing methods are solved; on the other hand, by introducing industry references through the calculation of peer cosine similarity and relying on the interpretation framework constructed by the Delphi method and the principal component analysis method to achieve the scientific integration of multi-dimensional indicators, the final evaluation results of enterprise innovation ability can not only accurately reflect the enterprise's own innovation level and technological evolution trend, but also reflect the enterprise's relative competitive position in the industry, providing a more accurate and reliable basis for the formulation of government industrial policies, the credit decisions of financial institutions and the enterprise's own R & D strategic planning, and effectively promoting the efficient transfer of scientific and technological resources and the optimization of the innovation ecosystem.
[0036] Optionally, the evaluation parameter includes the average sentence vector; the preprocessing of the pre-acquired enterprise-related information to obtain evaluation parameters includes: Use a hybrid analyzer to extract keywords from the patent information in the enterprise-related information to obtain word vectors.
[0037] Specifically, the hybrid analyzer can be a BERT-Jieba hybrid tokenizer enhanced by a science and technology finance thesaurus, that is, based on the Jieba tokenizer framework, integrating the semantic understanding ability of the BERT pre-trained model, and at the same time importing the preset science and technology finance thesaurus, including exclusive vocabulary related to the innovation of science and technology enterprises such as "invention patent", "R & D investment", "technology cycle", etc., into the tokenizer dictionary to ensure the accurate recognition of technical terms and industry-specific expressions in the patent text. Screen out the patent information from the enterprise-related information, and the patent information specifically includes the patent title, patent abstract, and text of the first claim; conduct preliminary cleaning on the screened patent text, removing stop words without actual semantics such as "of", "in", "for", etc., punctuation marks and special characters such as "①", " / ", etc., to obtain a pure patent text segment. Input the pure patent text segment into the configured hybrid analyzer, complete text tokenization through the Jieba tokenizer, and at the same time perform semantic encoding on each keyword through the BERT pre-trained model to generate real number vectors with unified dimensions, where each keyword corresponds to a unique word vector, and finally obtain a set of word vectors of the patent text.
[0038] Determine the inverse document frequency of each word vector based on the patent information, as well as the word frequency in the corresponding sentence, and obtain the corresponding weight according to the product of the word frequency and the inverse document frequency.
[0039] Specifically, taking a single sentence in the patent text as a unit, the frequency of each keyword corresponding to a word vector in the current sentence is counted, and then divided by the total number of keywords in that sentence to obtain the word frequency. Based on the patent text corpus constructed in the preprocessing stage, the total number of patent documents in the corpus and the number of patent documents containing the keyword are counted, and the inverse document frequency is obtained through logarithmic operation. The word frequency corresponding to each word vector is multiplied by the inverse document frequency to obtain the weight of that word vector.
[0040] The average sentence vector is obtained by weighting and summing the word vectors according to their corresponding weights.
[0041] Specifically, for a single sentence in the patent text, the set of word vectors and their corresponding weights are obtained. For each dimension of each word vector, its corresponding weight is multiplied, and then the weighted values of all word vectors for each dimension are summed to obtain a summation vector. To eliminate the influence of different sentence lengths (differences in the number of keywords) on the summation vector, L2 normalization is performed on the summation vector to obtain the average sentence vector. For all sentences in the target company's patent text (including the title, abstract, and the first claim), the above weighted summation and normalization operations are performed to obtain the average sentence vector of all sentences. Then, the arithmetic mean of the average sentence vectors of all sentences is taken to obtain the final average sentence vector of the patent, which is included in the evaluation parameters for subsequent patent evolution technical indicator evaluation and peer cosine similarity calculation.
[0042] Optionally, the evaluation of patent evolution technology indicators based on the patent vectors from previous years in the evaluation parameters includes: The types and number of keywords in the patent information of the relevant enterprise information for each year are counted, and a patent vector is generated by combining the corresponding patent text length. A vocabulary standardization vector is obtained based on the patent vector and the number of patents in the corresponding year.
[0043] Specifically, patent information (including patent titles, abstracts, and first claims) categorized by application year is selected from the aforementioned enterprise information. For each year t, a BERT-Jieba hybrid word segmenter enhanced with a science and technology finance thesaurus (it should be noted that this segmentation tool is consistent with the one used in the evaluation parameter preprocessing stage to ensure uniform keyword extraction rules) is used to extract keywords from all patent texts for that year, resulting in the keyword set {w} for that year. 1,t w 2,t , ..., w n,t}, where n is the total number of keyword types appearing in patents that year; simultaneously, the number of times each keyword wn, t appears in each patent k that year is counted, forming a keyword quantity statistics matrix. For each patent k that year, a single patent's "keyword-quantity" vector V is constructed with the keyword set of that year as the dimension. k,t=[v 1,t v 2,t , ..., v n,t ], where v n,t Indicates keyword w n,t The number of occurrences in patent k; to eliminate the interference of patent text length differences on the vector, V k,t Divide by the text length of the patent, i.e., the total number of keywords in patent k, and then use V. k,t The standardized patent vector is obtained by multiplying the vector with the all-1 vector. The total number of patents in year t is calculated as |P t (This can be obtained from the patent metadata of the enterprise's relevant information), sum all the standardized single patent vectors for that year, and then divide by the total number of patents |Pt| to obtain the vocabulary standardized vector Z for year t. t This reduces the impact of varying annual patent numbers. The formula can be expressed as: .
[0044] The vocabulary change is obtained based on the standardized vocabulary vectors from two consecutive years.
[0045] Specifically, for the patent application year sequence of the target company, two adjacent years t and t-1 are selected, and the corresponding vocabulary normalization vector Z is extracted. t (Year t) and Z t-1 (Year t-1). Using the relative rate of change formula, calculate the change of each keyword between two consecutive years, forming a vocabulary change vector Δt=[Δt]. 1,t , △ 2,t , ..., △ n,t] , is a vector of word change percentages, representing the changes in word n in year t-1 and year t-2. The percentage change in quantity between years can be expressed using the formula: .
[0046] The keywords from previous years are deduplicated to obtain word vectors, and the patent evolution technology index is obtained based on the word vectors and the word change.
[0047] Specifically, for a specific patent k of the target company to be evaluated, with an application year t, the keywords of the patent text are extracted, and the keywords are deduplicated to remove duplicate keywords, resulting in a deduplicated keyword set {w} for the patent. k1 w k2 , ..., w km Using the keyword dimension of the vocabulary change vector Δt obtained in step two for year t as a benchmark, construct the vocabulary vector B of this patent. k,t =[b 1,t b 2,t, ..., b m,t] To eliminate the impact of differences in the number of keywords within the patent itself on subsequent calculations, the word vector B... k,t Dividing by the total number of unique keywords in the patent, i.e., by the dot product of Bk,t and the all-1 vector, yields the standardized vocabulary vector. The standardized vocabulary vector is then multiplied by the vocabulary change vector Δt in year t, and multiplied by a scaling factor of 10000 (to amplify the numerical difference for easier subsequent score mapping) to obtain the original value of the evolutionary technology level of patent k, PRETk,t, which can be expressed by the formula: .
[0048] Further for enterprises In the The average value of the rapid evolution technology level index of relevant patents in the year was used to obtain the enterprise's [information / results] in the [year]. The rapid evolution of technology indicators RET in 2010 i,t : , Among them, K i This represents the total number of invention patents granted to company i in year t.
[0049] Optionally, determining the word vector similarity of at least one peer company in the first year, and determining the word vector similarity of the second year based on the evaluation parameters, includes: Based on the patent information of the peer companies, the comparison vectors for each year are extracted and organized, and the word vector similarity for the first year is obtained based on the comparison vectors of the two adjacent years.
[0050] Specifically, firstly, the business scope of the target company is extracted from the annual business report data of the enterprise. Combined with the industry classification standards in the science and technology policy text, the industry track to which the target company belongs is determined. Then, companies in the same track and of similar size as the target company are selected as peer companies, with a selection of at least five to ensure statistical validity. Secondly, patent information of each peer company is collected from public patent databases. This information includes the patent titles, abstracts, first claims, and application years of each peer company's patents over the years. Patents with legal statuses of "invalid" or "rejected" are removed, retaining only valid patent texts. A BERT-Jieba hybrid word segmenter enhanced with a science and technology finance lexicon is used to extract keywords from all valid patent texts of the peer company for a single year. After removing stop words, a set of patent keywords for that year is obtained. The word vector of each keyword is obtained using a Tencent AI-Lab pre-trained model. Based on the TF-IDF weighted average method, a sentence vector for a single patent is synthesized. The arithmetic mean of the sentence vectors of all valid patents of the peer company for a given year is taken to obtain the comparison vector for that year, such as the comparison vector V of peer company A in 2021. 对,A,2021 2022 Comparison Vector V对,A,2022 The vector dimensions are consistent with the target company's patent vectors to ensure compatibility in subsequent similarity calculations. For each peer company, the comparison vectors of its two adjacent years are selected, and the word vector similarity of the first year is calculated based on the cosine similarity algorithm.
[0051] The word vector similarity for the second year is obtained based on the patent vectors from the two adjacent years.
[0052] Specifically, the method for calculating the word vector similarity in the second year is the same as that in the first year, and will not be repeated here.
[0053] Optionally, obtaining the peer cosine similarity between enterprises based on the word vector similarity of the first year and the word vector similarity of the second year includes: The first cosine similarity between each enterprise and itself is obtained based on the word vector similarity of the first year and the word vector similarity of the second year.
[0054] The first cosine similarity of the corresponding enterprises is weighted based on their industry level. The peer cosine similarity is obtained by weighting the weights after division and the corresponding first cosine similarity.
[0055] Specifically, based on the cosine similarity algorithm, the first cosine similarity between each enterprise and itself is obtained according to the word vector similarity of the first year and the word vector similarity of the second year. The core indicators of an enterprise's industry level are clearly defined. Based on dimensions such as industry potential and technological output, industry technology concentration and enterprise industry position are selected as the basis for weighting, and corresponding comprehensive weighting is performed. Based on the comprehensive industry level weights, the first cosine similarity of each comparison group is weighted and summed to obtain the final peer cosine similarity. Peer cosine similarity integrates the consistency of technological change pace and industry level differences, reflecting both the dynamic technological connections between enterprises and highlighting the reference value of important enterprises within the industry, thus overcoming the limitations of static and indiscriminate weighting in traditional similarity calculations.
[0056] Optionally, the peer companies include leading companies in the same industry.
[0057] Specifically, the leading companies in the industry can be determined based on the company's overall strength or by established criteria, such as having the largest market share or the best profitability.
[0058] Optionally, the method for depicting a corporate innovation capability profile further includes: The operating data in the relevant information of the enterprise is retrieved, and multiple credit assessment values are obtained by inputting them into multiple assessment models. The assessment models are obtained by training a preset model with historical operating data and historical credit assessment values. The preset model is constructed using different network models.
[0059] Specifically, firstly, the scope of operational data in the enterprise's relevant information should be clearly defined. Operational data strongly correlated with credit assessment should be retrieved from the enterprise's annual business report, financial statements, and tax declaration records. This includes: liquidity indicators (current ratio, quick ratio), asset status (debt-to-equity ratio, fixed asset ratio), cash flow indicators (net cash flow from operating activities, cash flow growth rate), and profitability indicators (net profit, sales profit margin). Secondly, the retrieved operational data should be preprocessed. Outlier correction and standardization can be used. For outlier data exceeding a reasonable range (such as a debt-to-equity ratio > 200%), correction should be made by referring to the average data of enterprises of similar size in the same industry. All operational data should be standardized and mapped to the [0, 1] interval using Min-Max to eliminate differences in magnitude (such as "net profit (in ten thousand yuan)" and "current ratio (multiples)"), resulting in a standardized operational dataset to ensure a consistent data format for input into the model. Pre-set models for different network models should be constructed, specifically up to 12 models, which can be built using logistic regression models, random forest models, deep neural networks (DNN), etc. Historical operating data (consistent with current operating data dimensions) and corresponding historical credit assessment values (the credit limit coefficient ultimately approved by financial institutions, ranging from [0, 1]) of technology companies in the same industry were collected from the historical credit database of financial institutions in recent years to construct a training dataset (sample size ≥ 5000 records to ensure the effectiveness of model training). The training dataset was divided into a training set and a validation set in a 7:3 ratio, and 12 preset models were trained separately—using "historical operating data" as the input feature and "historical credit assessment value" as the target variable, and the Adam optimizer was used to minimize the mean squared error (MSE) loss function; during the training process, hyperparameters (such as learning rate and number of iterations) were adjusted in real time through the validation set. Training was stopped when the MSE of the validation set did not decrease for 5 consecutive rounds, and finally 12 trained assessment models were obtained, each of which could output a credit assessment value (i.e., credit limit coefficient) in the range of [0, 1]. The preprocessed standardized operating data of the target companies were input one by one into the 12 trained assessment models, and finally a set of credit assessment values containing 12 values was obtained, providing basic data for subsequent processing.
[0060] The multiple credit assessment values are combined to obtain the final assessment value.
[0061] Specifically, multiple credit assessment values can be combined using a weighted average method to obtain a final assessment value. This final assessment value can serve as an interpretable and implementable comprehensive patent value assessment tool for financial institutions when conducting intellectual property pledge financing, providing a unified and stable credit benchmark for technology valuation.
[0062] like Figure 2 As shown, an embodiment of the present invention provides a device 200 for depicting a corporate innovation capability profile, comprising: Preprocessing module 210 is used to preprocess the pre-acquired enterprise-related information to obtain evaluation parameters; Evaluation module 220 is used to evaluate the patent evolution technology index based on the patent vectors of previous years in the evaluation parameters, and map the evaluation results to obtain a periodic score; Processing module 230 is used to determine the word vector similarity of at least one peer company in the first year, determine the word vector similarity of the second year according to the evaluation parameters, obtain the peer cosine similarity between companies based on the word vector similarity of the first year and the word vector similarity of the second year, and map the peer cosine similarity to obtain a similarity score. The description module 240 is used to determine the enterprise's innovation capability based on the evaluation parameters, the period score, and the similarity score, according to the explanatory framework constructed based on the Delphi method and principal component analysis.
[0063] like Figure 3 As shown, an electronic device 300 provided in this embodiment of the invention includes a memory 310 and a processor 320; the memory 310 is used to store a computer program; the processor 320 is used to implement the enterprise innovation capability profiling method as described above when the computer program is executed.
[0064] Alternatively, an electronic device 300 includes a memory 310 and a processor 320 coupled to the memory 310; the memory 310 is configured to store a computer program; and the processor 320 is configured to perform the following operations when the computer program is executed: Preprocessing the previously acquired enterprise-related information yields evaluation parameters; Based on the patent vectors of the past years in the evaluation parameters, the patent evolution technology indicators are evaluated, and the evaluation results are mapped to obtain the periodic score. Determine the word vector similarity of at least one peer company in the first year, determine the word vector similarity in the second year according to the evaluation parameters, obtain the peer cosine similarity between companies based on the word vector similarity of the first year and the word vector similarity of the second year, and map the peer cosine similarity to obtain a similarity score; Based on the explanatory framework constructed using the Delphi method and principal component analysis, the enterprise's innovation capability is obtained according to the evaluation parameters, the period score, and the similarity score.
[0065] This invention provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the enterprise innovation capability profiling method described above.
[0066] Alternatively, a non-volatile computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the following operations: Preprocessing the previously acquired enterprise-related information yields evaluation parameters; Based on the patent vectors of the past years in the evaluation parameters, the patent evolution technology indicators are evaluated, and the evaluation results are mapped to obtain the periodic score. Determine the word vector similarity of at least one peer company in the first year, determine the word vector similarity in the second year according to the evaluation parameters, obtain the peer cosine similarity between companies based on the word vector similarity of the first year and the word vector similarity of the second year, and map the peer cosine similarity to obtain a similarity score; Based on the explanatory framework constructed using the Delphi method and principal component analysis, the enterprise's innovation capability is obtained according to the evaluation parameters, the period score, and the similarity score.
[0067] The present invention will now be described an electronic device 300 that can serve as a server or client of the present invention, which is an example of a hardware device that can be applied to various aspects of the present invention. Electronic device 300 is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic device 300 can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0068] Electronic device 300 includes a computing unit that can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) or a computer program loaded from a storage unit into random access memory (RAM). The RAM may also store various programs and data required for device operation. The computing unit, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.
[0069] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc. In this application, the units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of the present invention according to actual needs. Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated units can be implemented in hardware or as software functional units.
[0070] While the present invention has been disclosed above, its scope of protection is not limited thereto. Those skilled in the art can make various changes and modifications without departing from the spirit and scope of the present invention, and all such changes and modifications will fall within the scope of protection of the present invention.
Claims
1. An enterprise innovation capability profiling method, characterized in that, The method comprises the following steps: pre-processing the pre-acquired enterprise-related information to obtain evaluation parameters; evaluating patent evolution technical indicators based on patent vectors of the evaluation parameters of each year, and mapping the evaluation results to obtain period scores; determining a first-year word vector similarity of at least one peer enterprise, determining a second-year word vector similarity based on the evaluation parameters, obtaining a peer cosine similarity between enterprises based on the first-year word vector similarity and the second-year word vector similarity, and mapping the peer cosine similarity to obtain a similarity score; obtaining the enterprise innovation capability based on an interpretation framework constructed according to the Delphi method and the principal component analysis method, the evaluation parameters, the period scores, and the similarity score. 2.The enterprise innovation capability profiling method of claim 1, wherein, The evaluation parameters include average sentence vectors; the pre-processing of the pre-acquired enterprise-related information to obtain evaluation parameters comprises the following steps: extracting keywords from patent information in the enterprise-related information using a hybrid analyzer to obtain word vectors; determining the inverse document frequency of each word vector based on the patent information, and the word frequency in the corresponding sentence, obtaining the corresponding weight based on the product of the word frequency and the inverse document frequency; weighting and summing the word vectors according to the corresponding weight to obtain the average sentence vector. 3.The enterprise innovation capability profiling method of claim 1, wherein, The evaluation of patent evolution technical indicators based on the evaluation parameters of each year comprises the following steps: counting the types and quantities of keywords of patent information in the enterprise-related information of each year, generating a patent vector in combination with the corresponding patent text length, obtaining a lexical standardization vector based on the patent vector and the number of patents of the corresponding year; obtaining the lexical change amount based on the lexical standardization vectors of adjacent two years; de-duplicating the keywords of each year to obtain a lexical vector, and obtaining the patent evolution technical indicators based on the lexical vector and the lexical change amount. 4.The enterprise innovation capability profiling method of claim 3, wherein, The determination of the first-year word vector similarity of at least one peer enterprise and the determination of the second-year word vector similarity based on the evaluation parameters comprise the following steps: extracting and organizing the comparison vectors of each year based on the patent information of the peer enterprises, and obtaining the first-year word vector similarity based on the comparison vectors of adjacent two years; obtaining the second-year word vector similarity based on the patent vectors of adjacent two years.
5. The enterprise innovation capability profiling method of claim 4, wherein, The obtaining of the peer cosine similarity between enterprises based on the first-year word vector similarity and the second-year word vector similarity comprises the following steps: obtaining the first cosine similarity between each peer enterprise and the enterprise based on the corresponding first-year word vector similarity and the second-year word vector similarity; dividing the weight of the corresponding first cosine similarity based on the industry level of the enterprise; performing weighted calculation on the divided weight and the corresponding first cosine similarity to obtain the peer cosine similarity. 6.The enterprise innovation capability profiling method of claim 1, wherein, The peer enterprises include peer leading enterprises. 7.The enterprise innovation capability profiling method of claim 1, wherein, Further comprising: retrieve the operating data in the enterprise-related information, input a plurality of evaluation models to obtain a plurality of credit evaluation values, wherein the evaluation models are trained by historical operating data and historical credit evaluation values to obtain a preset model, and the preset model is constructed by using different network models; The plurality of credit evaluation values are arranged to obtain a final evaluation value.
8. An enterprise innovation capability profiling apparatus, characterized by comprising: The method comprises the following steps: A preprocessing module is configured to preprocess the pre-acquired enterprise-related information to obtain evaluation parameters; An evaluation module is configured to perform patent evolution technology index evaluation based on the patent vectors of the evaluation parameters of different years, and map the evaluation results to obtain period scores; A processing module is configured to determine the first-year word vector similarity of at least one peer enterprise, determine the second-year word vector similarity based on the evaluation parameters, obtain the peer cosine similarity between enterprises based on the first-year word vector similarity and the second-year word vector similarity, and map the peer cosine similarity to obtain a similarity score; A delineation module is configured to obtain the enterprise innovation capability based on the evaluation parameters, the period scores, and the similarity scores according to an interpretation framework constructed based on the Delphi method and the principal component analysis method.
9. An electronic device, comprising: The method comprises a memory and a processor; The memory is configured to store a computer program; The processor is configured to implement the enterprise innovation capability portrait delineation method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and when the computer program is executed by the processor, the enterprise innovation capability portrait delineation method according to any one of claims 1 to 7 is implemented.