Enterprise scientific and technological innovation data comprehensive optimization management method and system

By identifying the technology segment and analyzing its innovativeness, the system accurately divides science and technology innovation materials into segments, uses TF-IDF and SVM models to identify the types of science and technology innovation, and calculates the breakthrough index. This solves the problems of misclassification of types and waste of resources in enterprise science and technology innovation data management, and achieves efficient resource allocation and innovation evaluation.

CN122133945APending Publication Date: 2026-06-02西藏自治区科技创业服务中心

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
西藏自治区科技创业服务中心
Filing Date
2026-01-09
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing technologies in enterprise science and technology innovation data management suffer from problems such as misjudgment of innovation types, waste of resources, and low innovation efficiency. They fail to effectively distinguish between material and process innovations and do not optimize resource allocation according to feasibility and innovativeness.

Method used

By identifying the science and technology section and analyzing its innovativeness, we can accurately divide the effective and practical sections of science and technology innovation materials. We can also use TF-IDF and SVM classification models to identify the types of science and technology innovation, calculate the breakthrough index, and generate innovative optimization schemes ranked by priority.

Benefits of technology

It has improved the efficiency of science and technology innovation resource allocation and the accuracy of innovation decision-making, avoided resource waste, enhanced the accuracy of data foundation and the precision of innovation evaluation, and reduced the difficulty of cross-type comparison.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122133945A_ABST
    Figure CN122133945A_ABST
Patent Text Reader

Abstract

This invention discloses a comprehensive optimization management method and system for enterprise science and technology innovation data, relating to the field of data optimization management technology. The method includes technology segment identification and innovation analysis. First, in the technology segment identification stage, target science and technology innovation materials are collected and converted into standardized text. Through core technology identification, technology application identification, and similarity analysis, effective segments, practical segments, and unknown segments are accurately divided. In the innovation analysis stage, based on effective and practical segments, the SVM classification model is used to identify science and technology innovation types and summarize them into target science and technology innovation material groups. Combining historical science and technology innovation data, the comprehensive difference of numerical parameters and the difference in non-numerical features of each material within the group are calculated to obtain a breakthrough index. Then, combined with the feasibility of existing technologies, a priority index is calculated to generate innovation optimization schemes ranked by priority. This invention improves the efficiency of science and technology innovation resource allocation and the accuracy of innovation decision-making.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data optimization and management technology, specifically to a comprehensive optimization and management method and system for enterprise scientific and technological innovation data. Background Technology

[0002] In corporate technological innovation activities, scientific and technological innovation data is a core asset supporting technology research and development, innovation decision-making, and resource allocation. Therefore, a comprehensive optimization management method and system for corporate scientific and technological innovation data is needed.

[0003] Currently, most technological innovations are judged by researchers based on their personal experience, and are classified as material innovations simply because they contain material names, without being combined with the application scenarios of the technology for collaborative verification. If a material contains novel components, but its core innovation is the optimization of the process for medical minimally invasive equipment, it is easy to misjudge it as a material innovation rather than a process innovation. This classification bias directly affects the subsequent summarization of similar material groups, resulting in a lack of comparability among materials within the group.

[0004] The existing technology often prioritizes the most innovative solutions without considering their feasibility or analyzing their compatibility with existing equipment and processes. When highly innovative solutions require the import of scarce equipment, the existing technology fails to assess the feasibility of acquiring such equipment and directly recommends solutions that cannot be implemented due to unavailable resources, resulting in a waste of R&D resources.

[0005] Existing technologies do not rank multiple candidate optimization schemes by calculating priority indices, but instead use methods such as average resource allocation or random selection to simultaneously promote multiple innovative schemes. This fails to distinguish between high-breakthrough-and-high-feasibility and high-breakthrough-and-low-feasibility priorities, resulting in the dispersed investment of limited R&D budgets and an inability to concentrate resources to promote the rapid implementation of high-value schemes, leading to low efficiency in scientific and technological innovation. Summary of the Invention

[0006] To address the aforementioned technical shortcomings, the present invention aims to provide a comprehensive optimization management method and system for enterprise science and technology innovation data.

[0007] To solve the above-mentioned technical problems, the present invention adopts the following technical solution: The present invention provides a comprehensive optimization management method for enterprise science and technology innovation data, including the following steps: Step 1, Science and Technology Segment Identification: Collect target science and technology innovation materials, analyze the target science and technology innovation materials, and obtain each effective segment, each practical segment and each unknown segment of the target science and technology innovation materials.

[0008] Step 2, Innovation Analysis: Based on the practical and effective paragraphs of the target science and technology innovation materials, identify the science and technology innovation types of the target science and technology innovation materials, and then summarize them to obtain the target science and technology innovation material group. Conduct innovation analysis on the effective, practical and unknown paragraphs of each science and technology innovation material in the target science and technology innovation material group to obtain the breakthrough index of each science and technology innovation material in the target science and technology innovation material group, and then set the innovation optimization scheme of the target science and technology innovation material group.

[0009] Preferably, the similarity analysis of each available effective paragraph and each available practical paragraph of the target scientific and technological innovation material is carried out as follows: feature vectors are constructed for each available effective paragraph, each available practical paragraph and each paragraph of the standardized text of the target scientific and technological innovation material. The feature vectors of each original paragraph, each available effective paragraph and each available practical paragraph of the target scientific and technological innovation material are obtained by using the TF-IDF algorithm.

[0010] The cosine similarity between the feature vectors of each original paragraph and the feature vectors of each available valid paragraph is calculated to obtain the cosine similarity between each original paragraph and each available valid paragraph. The mean is calculated to obtain the average cosine similarity between each available valid paragraph. Each available valid paragraph with an average cosine similarity greater than a preset average cosine similarity threshold is recorded as a valid paragraph, and each available valid paragraph with an average cosine similarity less than or equal to the preset average cosine similarity threshold is recorded as an unknown paragraph.

[0011] The cosine similarity between the feature vectors of each original paragraph and the feature vectors of each usable paragraph is calculated to obtain the cosine similarity between each usable paragraph and the original paragraph. The mean is calculated to obtain the average cosine similarity between each usable paragraph. Each usable paragraph with an average cosine similarity greater than a preset average cosine similarity threshold is recorded as a usable paragraph, and each usable paragraph with an average cosine similarity less than or equal to the preset average cosine similarity threshold is recorded as an unknown paragraph.

[0012] Preferably, the specific process for obtaining the breakthrough index of each scientific and technological innovation material in the target scientific and technological innovation material group is as follows: The historical mean parameters corresponding to various numerical parameters of the target scientific and technological innovation material group, the effective paragraph parameters of each scientific and technological innovation material, the practical paragraph parameters of each scientific and technological innovation material, and the unknown paragraph parameters of each scientific and technological innovation material are calculated using single-parameter relative differences. This yields the relative differences of effective paragraphs, practical paragraphs, and unknown paragraphs corresponding to various numerical parameters of each scientific and technological innovation material in the target scientific and technological innovation material group. These are then weighted and calculated to obtain the comprehensive difference of numerical parameters for each scientific and technological innovation material in the target scientific and technological innovation material group.

[0013] By using feature structured representation, non-numerical features are transformed into keyword sequences, thereby obtaining the historical feature sequences corresponding to various non-numerical features of the target science and technology innovation material group and the feature sequences of each science and technology innovation material. The historical feature sequences are recorded as the baseline sequences. Using the Levenshtein edit distance algorithm, the difference steps of each science and technology innovation material corresponding to various non-numerical features of the target science and technology innovation material group are obtained. Each difference step is assigned a value to obtain the baseline sequence length and the edit distance of each science and technology innovation material corresponding to various non-numerical features of the target science and technology innovation material group.

[0014] Divide the edit distance of each science and technology innovation material corresponding to each type of non-numerical feature of the target science and technology innovation material group by the length of the baseline sequence to obtain the difference degree of each type of non-numerical feature of each science and technology innovation material in the target science and technology innovation material group. After weighted calculation, the difference degree of non-numerical feature of each science and technology innovation material in the target science and technology innovation material group is obtained.

[0015] The breakthrough index of each scientific and technological innovation material in the target group is obtained by normalizing the comprehensive difference of numerical parameters and the difference of non-numerical characteristics after normalization and weighting.

[0016] On the other hand, the present invention provides a comprehensive optimization and management system for enterprise science and technology innovation data, including the following modules: a technology segment identification module, used to collect target science and technology innovation materials, analyze the target science and technology innovation materials, and obtain each effective segment, each practical segment and each unknown segment of the target science and technology innovation materials.

[0017] The innovation analysis module identifies the innovation type of the target science and technology innovation materials based on the practical and effective paragraphs of each material, and then summarizes them to obtain the target science and technology innovation material group. The module performs innovation analysis on the effective, practical and unknown paragraphs of each material in the target science and technology innovation material group, obtains the breakthrough index of each material in the target science and technology innovation material group, and then sets the innovation optimization scheme for the target science and technology innovation material group.

[0018] The beneficial effects of this invention are as follows: 1. This invention first collects target scientific and technological innovation materials and converts them into standardized text through the technology segment identification stage. Through core technology identification, technology application identification, and similarity analysis, it accurately divides effective segments, practical segments, and unknown segments. In the innovation analysis stage, based on effective and practical segments, it identifies the types of scientific and technological innovations through the SVM classification model and summarizes them into target scientific and technological innovation material groups. Combining historical scientific and technological innovation data, it calculates the comprehensive difference of numerical parameters and the difference of non-numerical features of each scientific and technological innovation material in the group to obtain a breakthrough index. Then, it calculates the priority index by combining the feasibility of existing technologies and generates an innovation optimization scheme ranked by priority. This invention improves the efficiency of scientific and technological innovation resource allocation and the accuracy of innovation decision-making.

[0019] 2. This invention employs a dual-path design—using semantic keywords identified through core technologies, multi-index weighting, and keyword replacement migration based on technology applications—to separate core technology information and technology application information in scientific and technological innovation materials into independent effective and practical paragraphs. This avoids the omission of core data caused by information overlap in traditional manual classification. Simultaneously, by identifying unknown paragraphs through similarity analysis, it prevents potential innovative information from being misjudged as invalid information and eliminates redundant content, ensuring that subsequent innovation evaluation focuses only on high-value data and improving the accuracy of the data foundation.

[0020] 3. This invention trains an SVM classification model based on historical scientific and technological innovation data. By inputting the feature vectors of effective and practical paragraphs, it can output a quantitative type probability, replacing the traditional experience-based judgment mode. For example, for scientific and technological innovation materials with improved material composition and fine-tuned processes, the model can accurately identify them as material innovations by the material composition ratio of technical feature weights and the matching degree of application scenarios, providing an objective basis for the subsequent summary and comparison of similar material groups.

[0021] 4. This invention summarizes the target scientific and technological innovation materials into groups according to the type of scientific and technological innovation, so that the materials in the group have comparability in terms of technical direction and application scenarios. This reduces the difficulty of comparing process innovation and product innovation caused by cross-type mixed arrangement in traditional management, improves the ability of subsequent innovation analysis to focus on the differences in core technologies within the group, makes the comparison dimensions more accurate, and improves the quantitative effect of evaluating breakthroughs.

[0022] 5. This invention combines the breakthrough index with the feasibility similarity of existing technologies to calculate the priority index, generating optimized solutions ranked by priority. Solutions with high breakthrough but low feasibility and solutions with lower priority than medium breakthrough but high feasibility can be prioritized for enterprises to invest resources to promote the implementation of highly feasible solutions, reducing the waste of resources caused by innovation-only thinking that solutions cannot be implemented. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a schematic diagram of the implementation steps of the method of the present invention.

[0025] Figure 2 This is a schematic diagram of the system structure connection of the present invention. Detailed Implementation

[0026] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0027] according to Figure 1As shown, the present invention provides a comprehensive optimization management method for enterprise science and technology innovation data, including the following steps: Step 1, Science and Technology Segment Identification: Collect target science and technology innovation materials, analyze the target science and technology innovation materials, and obtain each effective segment, each practical segment and each unknown segment of the target science and technology innovation materials.

[0028] In one specific embodiment, the collection of target scientific and technological innovation materials is carried out as follows: collect the target scientific and technological innovation materials, and convert them into editable text using OCR technology and text extraction tools.

[0029] It should be noted that scientific and technological innovation materials include, but are not limited to, research and development documents for new composite materials and patent applications for process improvement.

[0030] In one specific embodiment, the analysis of the target scientific and technological innovation material is carried out as follows: the target scientific and technological innovation material is a standardized text of the target scientific and technological innovation material; the core technology of the standardized text of the target scientific and technological innovation material is identified to obtain each usable and effective paragraph of the target scientific and technological innovation material; the technology application of the standardized text of the target scientific and technological innovation material is identified to obtain each usable and practical paragraph; and similarity analysis is performed on each usable and effective paragraph and each usable and practical paragraph of the target scientific and technological innovation material to obtain each effective paragraph, each practical paragraph and each unknown paragraph of the target scientific and technological innovation material.

[0031] In one specific embodiment, the core technology identification process for the standardized text of the target scientific and technological innovation material is as follows: Core technology keywords are identified using semantic recognition technology to obtain the core technology keywords for each sentence in each paragraph of the target scientific and technological innovation material. The percentage of core technology keywords per sentence in each paragraph of the target scientific and technological innovation material is statistically analyzed. Sentences with a percentage greater than a preset threshold are recorded as core technology key sentences in each paragraph of the target scientific and technological innovation material. Two adjacent core technology key sentences are recorded as two core technology key sentences in the same core technology group. Core technology groups are then classified in this way to obtain the core technology groups for each paragraph of the target scientific and technological innovation material.

[0032] It should be noted that semantic recognition technology includes, but is not limited to, the BERT algorithm based on pre-trained language models and the Naive Bayes algorithm based on supervised learning.

[0033] The core technical keywords include, but are not limited to, material composition, experimental parameters, process steps, and performance indicators.

[0034] The number of core technology keywords in each core technology group of each section of the target science and technology innovation material was statistically analyzed, and the core technology group with the largest number of core technology keywords was recorded as the largest core technology group in each section of the target science and technology innovation material.

[0035] The percentage of the largest core technology group keywords, the percentage of core technology keywords per paragraph, the percentage of the average percentage of core technology per sentence, and the percentage of the total text of core technology in each paragraph of the target science and technology innovation material were statistically obtained. After normalization and weighted calculation, the core technology usage index of each paragraph of the target science and technology innovation material was obtained. Each paragraph with a core technology usage index greater than the preset core technology usage index threshold was recorded as a usable and effective paragraph of the target science and technology innovation material.

[0036] It should be noted that the weighted calculation process after normalization is as follows: Normalization is performed using the minimum-maximum normalization method to obtain the normalized values ​​of the maximum core technology group keyword ratio, the single-segment core technology keyword ratio, the average single-sentence core technology ratio, and the total core technology text ratio for each segment of the target scientific and technological innovation material. The weight factors of the maximum core technology group keyword ratio, the single-segment core technology keyword ratio, the average single-sentence core technology ratio, and the total core technology text ratio are then obtained from the database. These weight factors are multiplied by the corresponding weight factors, and the results are then summed to obtain the core technology usage index for each segment of the target scientific and technological innovation material.

[0037] The preset core technology index threshold refers to a quantitative judgment standard value that is pre-set when identifying core technologies in standardized texts of target science and technology innovation materials. This threshold is based on the characteristics of the enterprise's science and technology innovation field, the quality benchmark of historical science and technology innovation data, and the needs for extracting core technology information. This threshold is a key critical value for measuring whether a paragraph contains sufficient core technology information. The specific value is set by the staff.

[0038] The weight factors for the largest core technology group keywords, the weight factors for the proportion of core technology keywords in a single paragraph, the weight factors for the proportion of average core technology sentences, and the weight factors for the proportion of total core technology text are set by staff based on experience.

[0039] In one specific embodiment, the technical application identification of the standardized text of the target scientific and technological innovation material is specifically analyzed as follows: Technical application keywords are identified using semantic recognition technology to obtain the technical application keywords for each sentence in each paragraph of the target scientific and technological innovation material. The percentage of each sentence containing technical application keywords in each paragraph of the target scientific and technological innovation material is statistically analyzed. Sentences with a percentage greater than a preset threshold are recorded as key technical application sentences in each paragraph of the target scientific and technological innovation material. Two adjacent key technical application sentences are recorded as two key technical application sentences in the same technical application group. Technical application groups are then classified to obtain the technical application groups for each paragraph of the target scientific and technological innovation material.

[0040] It should be noted that the key terms for the application of technology include, but are not limited to, application scenarios, practical value, and implementation conditions.

[0041] The text volume of technical application keywords in each technical application group of each section of the target scientific and technological innovation material was statistically obtained, and the technical application group with the largest text volume of technical application keywords was recorded as the largest technical application group in each section of the target scientific and technological innovation material.

[0042] The percentage of the largest group of technical application keywords, the percentage of technical application keywords per paragraph, the percentage of average technical application per sentence, and the percentage of total technical application text in each paragraph of the target scientific and technological innovation material were statistically obtained. After normalization and weighted calculation, the technical application usage index of each paragraph of the target scientific and technological innovation material was obtained. Each paragraph with a technical application usage index greater than the preset technical application usage index threshold was recorded as a usable and practical paragraph of the target scientific and technological innovation material.

[0043] It should be noted that the weighted calculation process after normalization is as follows: Normalization is performed using the minimum-maximum normalization method to obtain the normalized values ​​of the maximum proportion of keywords in the technology application group, the single-segment proportion of keywords in technology application, the average single-sentence proportion of keywords in technology application, and the total proportion of text in technology application for each segment of the target scientific and technological innovation material. The weighting factors for the maximum proportion of keywords in the technology application group, the single-segment proportion of keywords in technology application, the average single-sentence proportion of keywords in technology application, and the total proportion of text in technology application are obtained from the database. These weighting factors are then multiplied by the corresponding weighting factors, and the results are summed to obtain the technology application usage index for each segment of the target scientific and technological innovation material.

[0044] The preset technology application threshold refers to the quantitative judgment threshold set in advance when identifying the technology application of standardized texts of target science and technology innovation materials. This threshold is based on the enterprise's science and technology innovation implementation needs, technology application information evaluation standards, and historical technology application data characteristics. It is the core standard for measuring whether a paragraph contains sufficient value of technology application information. The specific value is set by the staff.

[0045] The weighting factors for the percentage of keywords in the largest technology application group, the percentage of keywords in a single paragraph in the technology application group, the percentage of keywords in a single sentence in the average technology application group, and the percentage of total text in the technology application group are set by staff based on experience.

[0046] In one specific embodiment, the similarity analysis of each available effective paragraph and each available practical paragraph of the target scientific and technological innovation material is performed as follows: feature vectors are constructed for each available effective paragraph, each available practical paragraph, and each paragraph of the standardized text of the target scientific and technological innovation material. The feature vectors of each original paragraph, each available effective paragraph, and each available practical paragraph of the target scientific and technological innovation material are obtained by using the TF-IDF algorithm.

[0047] The cosine similarity between the feature vectors of each original paragraph and the feature vectors of each available valid paragraph is calculated to obtain the cosine similarity between each original paragraph and each available valid paragraph. The mean is calculated to obtain the average cosine similarity between each available valid paragraph. Each available valid paragraph with an average cosine similarity greater than a preset average cosine similarity threshold is recorded as a valid paragraph, and each available valid paragraph with an average cosine similarity less than or equal to the preset average cosine similarity threshold is recorded as an unknown paragraph.

[0048] It should be noted that the preset mean cosine similarity threshold refers to a quantitative matching threshold set in advance based on the text feature matching accuracy requirements, the consistency standard of scientific and technological innovation data information, and the historical similarity analysis results when performing similarity analysis on the available effective paragraphs, available practical paragraphs, and the original paragraphs of the target scientific and technological innovation materials. This threshold is used to measure the degree of feature matching between the available effective paragraphs, available practical paragraphs, and the original paragraphs. It is the core judgment standard for distinguishing between effective paragraphs and unknown paragraphs, and can also be used to distinguish between practical paragraphs and unknown paragraphs. The specific value is set by the staff.

[0049] The cosine similarity between the feature vectors of each original paragraph and the feature vectors of each usable paragraph is calculated to obtain the cosine similarity between each usable paragraph and the original paragraph. The mean is calculated to obtain the average cosine similarity between each usable paragraph. Each usable paragraph with an average cosine similarity greater than a preset average cosine similarity threshold is recorded as a usable paragraph, and each usable paragraph with an average cosine similarity less than or equal to the preset average cosine similarity threshold is recorded as an unknown paragraph.

[0050] Step 2, Innovation Analysis: Based on the practical and effective paragraphs of the target science and technology innovation materials, identify the science and technology innovation types of the target science and technology innovation materials, and then summarize them to obtain the target science and technology innovation material group. Conduct innovation analysis on the effective, practical and unknown paragraphs of each science and technology innovation material in the target science and technology innovation material group to obtain the breakthrough index of each science and technology innovation material in the target science and technology innovation material group, and then set the innovation optimization scheme of the target science and technology innovation material group.

[0051] In one specific embodiment, the identification process for the scientific and technological innovation type of the target scientific and technological innovation material is as follows: the feature vectors of each original paragraph of the historical scientific and technological innovation material and the innovation tags corresponding to the material are used to train an SVM classification model to obtain the SVM classification model of the scientific and technological innovation material: input the feature vectors of each paragraph of the scientific and technological innovation material, and output the probability of each preset type of the scientific and technological innovation material.

[0052] Input each effective paragraph and each practical paragraph of the target science and technology innovation material into the SVM classification model of science and technology innovation materials to obtain the probability of each preset type of the target science and technology innovation material. Record the preset type with the highest probability as the preset type of the target science and technology innovation material, thereby identifying the science and technology innovation type of the target science and technology innovation material.

[0053] In one specific embodiment, the innovative analysis is performed as follows: historical scientific and technological innovation data of the target scientific and technological innovation material group is obtained from the experimental database. The historical scientific and technological innovation data of the target scientific and technological innovation material group includes historical average parameters of various numerical technologies and various non-numerical technical features. The corresponding numerical parameters and non-numerical features are extracted from the effective paragraphs of the target scientific and technological innovation materials. The dimensions are aligned with the parameters and features of the historical scientific and technological innovation data to obtain the historical average parameters of various numerical parameters of the target scientific and technological innovation material group, the effective paragraph parameters of each scientific and technological innovation material, the practical paragraph parameters of each scientific and technological innovation material, and the unknown paragraph parameters of each scientific and technological innovation material. At the same time, the historical features and features of each scientific and technological innovation material corresponding to various non-numerical features of the target scientific and technological innovation material group are obtained.

[0054] It should be noted that the numerical technical mean parameters include, but are not limited to, fracture toughness, high temperature resistance, and reaction efficiency, while the non-numerical technical characteristics include, but are not limited to, cleaning, coating, and sintering.

[0055] In one specific embodiment, the process of obtaining the breakthrough index of each scientific and technological innovation material in the target scientific and technological innovation material group is as follows: The historical mean parameters, effective paragraph parameters, practical paragraph parameters, and unknown paragraph parameters of each scientific and technological innovation material in the target scientific and technological innovation material group are calculated using single-parameter relative differences. This yields the relative differences of effective paragraphs, practical paragraphs, and unknown paragraphs corresponding to various numerical parameters of each scientific and technological innovation material in the target scientific and technological innovation material group. These differences are then weighted and calculated to obtain the comprehensive difference of numerical parameters for each scientific and technological innovation material in the target scientific and technological innovation material group.

[0056] It should be noted that the calculation process for the single-parameter relative difference is as follows: subtract the corresponding historical mean parameter from the effective paragraph parameter, practical paragraph parameter, and unknown paragraph parameter of each scientific and technological innovation material corresponding to the various numerical parameters of the target scientific and technological innovation material group, and then divide them by the corresponding historical mean parameter after absolute value conversion to obtain the relative difference of the effective paragraph, the relative difference of the practical paragraph, and the relative difference of the unknown paragraph corresponding to the various numerical parameters of each scientific and technological innovation material in the target scientific and technological innovation material group.

[0057] The weighted calculation process is as follows: Obtain the weighting factors for the relative differences of valid paragraphs, the relative differences of practical paragraphs, and the relative differences of unknown paragraphs from the database. Multiply the relative differences of valid paragraphs, practical paragraphs, and unknown paragraphs corresponding to various numerical parameters of each material in the target science and technology innovation material group by their respective weighting factors, and then sum them to obtain the relative differences of paragraphs for various numerical parameters of each material in the target science and technology innovation material group. Obtain the weighting factors corresponding to various numerical parameters from the database, multiply the relative differences of paragraphs for various numerical parameters of each material in the target science and technology innovation material group by their respective weighting factors, and then sum them to obtain the comprehensive difference of numerical parameters of each material in the target science and technology innovation material group.

[0058] The weighting factors for the relative differences of effective paragraphs, the relative differences of practical paragraphs, the relative differences of unknown paragraphs, and various numerical parameters are set by staff based on experience.

[0059] By using feature structured representation, non-numerical features are transformed into keyword sequences, thereby obtaining the historical feature sequences corresponding to various non-numerical features of the target science and technology innovation material group and the feature sequences of each science and technology innovation material. The historical feature sequences are recorded as the baseline sequences. Using the Levenshtein edit distance algorithm, the difference steps of each science and technology innovation material corresponding to various non-numerical features of the target science and technology innovation material group are obtained. Each difference step is assigned a value to obtain the baseline sequence length and the edit distance of each science and technology innovation material corresponding to various non-numerical features of the target science and technology innovation material group.

[0060] It should be noted that the process of assigning values ​​to each difference step is as follows: Based on the Levenshtein edit distance algorithm, staff first set correction factors for effective paragraphs, useful paragraphs, and unknown paragraphs to obtain the correction factors for effective paragraphs, useful paragraphs, and unknown paragraphs. Next, staff set costs for each type of difference to obtain the edit distance for each type of difference. Finally, the baseline sequence length of the target non-numerical feature is preset, and each difference step of the target non-numerical feature of the target scientific and technological innovation material is classified to obtain the edit distance for each difference step of the target non-numerical feature of the target scientific and technological innovation material. Simultaneously, a text extraction tool is used to obtain the paragraph in the target scientific and technological innovation material where each difference step text first appears, which is recorded as the marker paragraph of each difference step text. This yields the type of the marker paragraph for each difference step text, including effective paragraphs, useful paragraphs, unknown paragraphs, and other paragraphs. No correction factor is assigned to other paragraphs. The edit distance of each difference step of the target non-numerical feature of the target scientific and technological innovation material is multiplied by the corresponding correction factor and then summed to obtain the edit distance of the target non-numerical feature of the target scientific and technological innovation material. This process yields the edit distance of each scientific and technological innovation material corresponding to each type of non-numerical feature of the target scientific and technological innovation material group.

[0061] Divide the edit distance of each science and technology innovation material corresponding to each type of non-numerical feature of the target science and technology innovation material group by the length of the baseline sequence to obtain the difference degree of each type of non-numerical feature of each science and technology innovation material in the target science and technology innovation material group. After weighted calculation, the difference degree of non-numerical feature of each science and technology innovation material in the target science and technology innovation material group is obtained.

[0062] It should be noted that the weighted calculation process is as follows: obtain the weight factors of various non-numerical features from the database, multiply the differences of various non-numerical features of each scientific and technological innovation material in the target scientific and technological innovation material group by the corresponding weight factors, and then add them together to obtain the differences of non-numerical features of each scientific and technological innovation material in the target scientific and technological innovation material group.

[0063] The weighting factors for various non-numerical features are set by staff based on experience.

[0064] The breakthrough index of each scientific and technological innovation material in the target group is obtained by normalizing the comprehensive difference of numerical parameters and the difference of non-numerical characteristics after normalization and weighting.

[0065] It should be noted that the weighted calculation process after normalization is as follows: normalization is performed using the minimum-maximum normalization method to obtain the normalized value of the comprehensive difference of numerical parameters and the normalized value of the difference of non-numerical features for each scientific and technological innovation material in the target scientific and technological innovation material group. The weight factors of the comprehensive difference of numerical parameters and the weight factors of the difference of non-numerical features are obtained from the database. The normalized values ​​of the comprehensive difference of numerical parameters and the normalized values ​​of the difference of non-numerical features for each scientific and technological innovation material in the target scientific and technological innovation material group are multiplied by the corresponding weight factors, and then added together to obtain the breakthrough index of each scientific and technological innovation material in the target scientific and technological innovation material group.

[0066] The weighting factor for the comprehensive difference of numerical parameters and the weighting factor for the difference of non-numerical features are set by staff based on experience.

[0067] In one specific embodiment, the innovative optimization scheme for setting the target scientific and technological innovation material group is specifically set as follows: each scientific and technological innovation material with a breakthrough index greater than a preset breakthrough index is recorded as a usable scientific and technological innovation material in the target scientific and technological innovation material group. Then, the non-numerical features of each difference step of each usable scientific and technological innovation material in the target scientific and technological innovation material group are obtained. The non-numerical features of each step of the prior art are obtained from the experimental database. Cosine similarity is calculated with the non-numerical features of each difference step of each usable scientific and technological innovation material in the target scientific and technological innovation material group to obtain the feasible similarity of each usable scientific and technological innovation material in the target scientific and technological innovation material group. The feasible similarity and breakthrough index of each usable scientific and technological innovation material in the target scientific and technological innovation material group are normalized and then weighted to obtain the priority index of each usable scientific and technological innovation material in the target scientific and technological innovation material group.

[0068] It should be noted that the preset breakthrough index refers to a quantitative screening threshold set in advance during the process of setting up the innovation optimization scheme of the target science and technology innovation material group, in combination with the enterprise's science and technology innovation resource investment standards, innovation value assessment benchmarks and historical science and technology innovation project breakthrough data. This threshold is used to measure the degree of technological breakthrough of science and technology innovation materials. It is the core judgment standard for selecting usable science and technology innovation materials with actual innovation value from the target science and technology innovation material group that are worthy of further resource allocation. The specific value is set by the staff.

[0069] The weighted calculation process after normalization is as follows: Normalization is performed using the minimum-maximum normalization method to obtain the feasible similarity normalization value and breakthrough index normalization value of each available scientific and technological innovation material in the target scientific and technological innovation material group. The feasible similarity weight factor and breakthrough index weight factor are obtained from the database. The feasible similarity normalization value and breakthrough index normalization value of each available scientific and technological innovation material in the target scientific and technological innovation material group are multiplied by the corresponding weight factor and then added together to obtain the priority index of each available scientific and technological innovation material in the target scientific and technological innovation material group.

[0070] The feasible similarity weighting factor and the breakthrough index weighting factor were set by staff based on experience.

[0071] The innovative optimization scheme for the target scientific and technological innovation materials group is as follows: sort the available scientific and technological innovation materials in the target scientific and technological innovation materials group in descending order of priority index to obtain the sequence of available scientific and technological innovation materials in the target scientific and technological innovation materials group. When the target scientific and technological innovation materials type needs to be optimized, provide scientific and technological innovation materials according to the sequence of available scientific and technological innovation materials.

[0072] according to Figure 2 As shown, the present invention provides an enterprise science and technology innovation data comprehensive optimization management system, including the following modules: a technology segment identification module, an innovation analysis module, and a database.

[0073] The technology segment identification module and the innovation analysis module are connected, and both the technology segment identification module and the innovation analysis module are connected to the database.

[0074] The science and technology segment identification module is used to collect target science and technology innovation materials, analyze the target science and technology innovation materials, and obtain each effective segment, each practical segment, and each unknown segment of the target science and technology innovation materials.

[0075] The innovation analysis module identifies the innovation type of the target science and technology innovation materials based on the practical and effective paragraphs of each material, and then summarizes them to obtain the target science and technology innovation material group. The module performs innovation analysis on the effective, practical and unknown paragraphs of each material in the target science and technology innovation material group, obtains the breakthrough index of each material in the target science and technology innovation material group, and then sets the innovation optimization scheme for the target science and technology innovation material group.

[0076] The database stores the following weight factors: the weight factor for the largest core technology group keywords, the weight factor for the percentage of core technology keywords in a single paragraph, the weight factor for the percentage of core technology keywords in a single sentence, the weight factor for the percentage of core technology total text, the weight factor for the percentage of the largest technology application group keywords, the weight factor for the percentage of technology application keywords in a single paragraph, the weight factor for the percentage of technology application in a single sentence, the weight factor for the percentage of technology application total text, the weight factor for the relative difference of effective paragraphs, the weight factor for the relative difference of practical paragraphs, the weight factor for the relative difference of unknown paragraphs, the weight factors corresponding to various numerical parameters, the weight factors for various non-numerical features, the weight factor for the comprehensive difference of numerical parameters, the weight factor for the difference of non-numerical features, the weight factor for feasible similarity, the weight factor for breakthrough index, and the experimental database.

[0077] The experimental database is used to store historical scientific and technological innovation data of the target scientific and technological innovation material group and non-numerical features of each step of the existing technology.

[0078] The OCR technology, text extraction tools, BERT algorithm based on pre-trained language models, Naive Bayes algorithm based on supervised learning, min-max normalization method, TF-IDF algorithm, cosine similarity calculation, SVM classification model and Levenshtein edit distance algorithm described in this invention are all existing technologies and can be found on the Internet, so they will not be described in detail here.

[0079] The examples described in this invention are not limited to the specific embodiments listed above. The examples are merely illustrative to facilitate understanding of the invention and do not constitute a limitation on the scope of protection of this invention. Any modifications, equivalent substitutions, etc., made within the spirit and principles of this invention should be included within the scope of protection.

[0080] The above description is merely an example and illustration of the concept of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described or use similar methods to replace them, as long as they do not deviate from the concept of the invention or exceed the scope defined in this specification, they should all fall within the protection scope of the present invention.

Claims

1. A comprehensive optimization management method for enterprise scientific and technological innovation data, characterized in that, Includes the following steps: Step 1: Identification of Science and Technology Sections: Collect target science and technology innovation materials, analyze the target science and technology innovation materials, and obtain each effective section, each practical section, and each unknown section of the target science and technology innovation materials; Step 2, Innovation Analysis: Based on the practical and effective paragraphs of the target science and technology innovation materials, identify the science and technology innovation types of the target science and technology innovation materials, and then summarize them to obtain the target science and technology innovation material group. Conduct innovation analysis on the effective, practical and unknown paragraphs of each science and technology innovation material in the target science and technology innovation material group to obtain the breakthrough index of each science and technology innovation material in the target science and technology innovation material group, and then set the innovation optimization scheme of the target science and technology innovation material group.

2. The method for comprehensive optimization and management of enterprise scientific and technological innovation data according to claim 1, characterized in that, The analysis of the target scientific and technological innovation materials is conducted in the following specific process: The target science and technology innovation material is a standardized text of the target science and technology innovation material. The core technology is identified in the standardized text of the target science and technology innovation material to obtain each usable and effective paragraph. The technology application is identified in the standardized text of the target science and technology innovation material to obtain each usable and practical paragraph. The similarity analysis is performed on each usable and effective paragraph and each usable and practical paragraph to obtain each effective paragraph, each practical paragraph and each unknown paragraph of the target science and technology innovation material.

3. The method for comprehensive optimization and management of enterprise scientific and technological innovation data according to claim 2, characterized in that, The core technology identification process for the standardized text of the target scientific and technological innovation materials is as follows: By using semantic recognition technology, core technology keywords are identified to obtain the core technology keywords of each sentence in each paragraph of the target science and technology innovation material. The percentage of core technology keywords per sentence in each paragraph of the target science and technology innovation material is calculated. Sentences with a percentage greater than a preset threshold are recorded as core technology key sentences in each paragraph of the target science and technology innovation material. Two adjacent core technology key sentences are recorded as two core technology key sentences in the same core technology group. Core technology groups are classified in this way to obtain the core technology groups in each paragraph of the target science and technology innovation material. The text volume of core technology keywords in each core technology group of each section of the target science and technology innovation material was statistically obtained, and the core technology group with the largest text volume of core technology keywords was recorded as the largest core technology group in each section of the target science and technology innovation material. The percentage of the largest core technology group keywords, the percentage of core technology keywords per paragraph, the percentage of the average percentage of core technology per sentence, and the percentage of the total text of core technology in each paragraph of the target science and technology innovation material were statistically obtained. After normalization and weighted calculation, the core technology usage index of each paragraph of the target science and technology innovation material was obtained. Each paragraph with a core technology usage index greater than the preset core technology usage index threshold was recorded as a usable and effective paragraph of the target science and technology innovation material.

4. The method for comprehensive optimization and management of enterprise scientific and technological innovation data according to claim 2, characterized in that, The specific analysis process for applying technical identification to the standardized text of the target scientific and technological innovation materials is as follows: Based on the method of identifying core technologies in the standardized text of the target scientific and technological innovation materials, the technical application identification is performed on the standardized text of the target scientific and technological innovation materials. The keyword identification of core technologies is replaced with the keyword identification of technical applications, thereby obtaining each usable and practical paragraph of the target scientific and technological innovation materials.

5. The method for comprehensive optimization and management of enterprise scientific and technological innovation data according to claim 4, characterized in that, The similarity analysis of each available effective paragraph and each available practical paragraph of the target scientific and technological innovation material is performed as follows: Feature vectors are constructed for each available effective paragraph, each available practical paragraph, and each paragraph of the standardized text in the target scientific and technological innovation material. The feature vectors of each original paragraph, each available effective paragraph, and each available practical paragraph in the target scientific and technological innovation material are obtained by using the TF-IDF algorithm. The cosine similarity between the feature vector of each original paragraph and the feature vector of each available effective paragraph is calculated to obtain the cosine similarity between the original paragraphs of each available effective paragraph. The mean is calculated to obtain the average cosine similarity of each available effective paragraph. Each available effective paragraph with an average cosine similarity greater than the preset average cosine similarity threshold is recorded as an effective paragraph, and each available effective paragraph with an average cosine similarity less than or equal to the preset average cosine similarity threshold is recorded as an unknown paragraph. The cosine similarity between the feature vectors of each original paragraph and the feature vectors of each usable paragraph is calculated to obtain the cosine similarity between each usable paragraph and the original paragraph. The mean is calculated to obtain the average cosine similarity between each usable paragraph. Each usable paragraph with an average cosine similarity greater than a preset average cosine similarity threshold is recorded as a usable paragraph, and each usable paragraph with an average cosine similarity less than or equal to the preset average cosine similarity threshold is recorded as an unknown paragraph.

6. The method for comprehensive optimization and management of enterprise scientific and technological innovation data according to claim 1, characterized in that, The specific identification process for determining the science and technology innovation type of the target science and technology innovation material is as follows: The feature vectors of each original paragraph of historical science and technology innovation materials and the corresponding innovation tags of the materials are used to train an SVM classification model to obtain the SVM classification model of science and technology innovation materials: input the feature vectors of each paragraph of the science and technology innovation materials, and output the probability of each preset type of the science and technology innovation materials. Input each effective paragraph and each practical paragraph of the target science and technology innovation material into the SVM classification model of science and technology innovation materials to obtain the probability of each preset type of the target science and technology innovation material. Record the preset type with the highest probability as the preset type of the target science and technology innovation material, thereby identifying the science and technology innovation type of the target science and technology innovation material.

7. The method for comprehensive optimization and management of enterprise scientific and technological innovation data according to claim 1, characterized in that, The innovative analysis is conducted, and the specific analysis process is as follows: Historical scientific and technological innovation data of the target scientific and technological innovation material group were obtained from the experimental database. The historical scientific and technological innovation data of the target scientific and technological innovation material group includes historical average parameters of various numerical technologies and various non-numerical technologies. The corresponding numerical parameters and non-numerical features were extracted from the effective paragraphs of the target scientific and technological innovation materials. The parameters and features of the target scientific and technological innovation materials were aligned with the parameters and features of the historical scientific and technological innovation data. In this way, the historical average parameters of various numerical parameters of the target scientific and technological innovation material group, the effective paragraph parameters of each scientific and technological innovation material, the practical paragraph parameters of each scientific and technological innovation material, and the unknown paragraph parameters of each scientific and technological innovation material were obtained. At the same time, the historical features and features of each scientific and technological innovation material corresponding to various non-numerical features of the target scientific and technological innovation material group were obtained.

8. The method for comprehensive optimization and management of enterprise scientific and technological innovation data according to claim 7, characterized in that, The specific process for obtaining the breakthrough index of each scientific and technological innovation material in the target scientific and technological innovation material group is as follows: The historical mean parameters, effective paragraph parameters, practical paragraph parameters, and unknown paragraph parameters of various numerical parameters of the target scientific and technological innovation materials group are calculated by performing single-parameter relative differences. This yields the relative differences of effective paragraphs, practical paragraphs, and unknown paragraphs of various numerical parameters of each scientific and technological innovation material in the target scientific and technological innovation materials group. After weighted calculation, the comprehensive difference of numerical parameters of each scientific and technological innovation material in the target scientific and technological innovation materials group is obtained. By using feature structured representation, non-numerical features are transformed into keyword sequences, thereby obtaining the historical feature sequences and feature sequences of each science and technology innovation material group corresponding to various non-numerical features. The historical feature sequences are recorded as the baseline sequences. The Levenshtein edit distance algorithm is used to obtain the difference steps of each science and technology innovation material corresponding to various non-numerical features of the target science and technology innovation material group. Values ​​are assigned to each difference step to obtain the baseline sequence length and edit distance of each science and technology innovation material corresponding to various non-numerical features of the target science and technology innovation material group. Divide the edit distance of each science and technology innovation material corresponding to each non-numerical feature of the target science and technology innovation material group by the length of the baseline sequence to obtain the difference degree of each non-numerical feature of each science and technology innovation material in the target science and technology innovation material group. After weighted calculation, the difference degree of non-numerical feature of each science and technology innovation material in the target science and technology innovation material group is obtained. The breakthrough index of each scientific and technological innovation material in the target group is obtained by normalizing the comprehensive difference of numerical parameters and the difference of non-numerical characteristics after normalization and weighting.

9. The method for comprehensive optimization and management of enterprise scientific and technological innovation data according to claim 8, characterized in that, The innovative optimization scheme for setting the target scientific and technological materials group is specifically set up as follows: Each scientific and technological innovation material with a breakthrough index greater than the preset breakthrough index is denoted as a usable scientific and technological innovation material in the target scientific and technological innovation material group. Then, the non-numerical features of each difference step of each usable scientific and technological innovation material in the target scientific and technological innovation material group are obtained. The non-numerical features of each step of the prior art are obtained from the experimental database and cosine similarity is calculated with the non-numerical features of each difference step of each usable scientific and technological innovation material in the target scientific and technological innovation material group to obtain the feasible similarity of each usable scientific and technological innovation material in the target scientific and technological innovation material group. The feasible similarity and breakthrough index of each usable scientific and technological innovation material in the target scientific and technological innovation material group are normalized and then weighted to obtain the priority index of each usable scientific and technological innovation material in the target scientific and technological innovation material group. The innovative optimization scheme for the target scientific and technological innovation materials group is as follows: sort the available scientific and technological innovation materials in the target scientific and technological innovation materials group in descending order of priority index to obtain the sequence of available scientific and technological innovation materials in the target scientific and technological innovation materials group. When the target scientific and technological innovation materials type needs to be optimized, provide scientific and technological innovation materials according to the sequence of available scientific and technological innovation materials.

10. A method and management system for comprehensive optimization management of enterprise science and technology innovation data according to any one of claims 1-9, characterized in that, Includes the following modules: The science and technology segment identification module is used to collect target science and technology innovation materials, analyze the target science and technology innovation materials, and obtain each effective segment, each practical segment, and each unknown segment of the target science and technology innovation materials; The innovation analysis module identifies the innovation type of the target science and technology innovation materials based on the practical and effective paragraphs of each material, and then summarizes them to obtain the target science and technology innovation material group. The module performs innovation analysis on the effective, practical and unknown paragraphs of each material in the target science and technology innovation material group, obtains the breakthrough index of each material in the target science and technology innovation material group, and then sets the innovation optimization scheme for the target science and technology innovation material group.