Construction method and system for patent industrialization potential prediction model of medical institution

By constructing a predictive model for the commercialization potential of medical institution patents based on the random forest algorithm, the problem of the inability to assess the commercialization potential of patents in existing technologies is solved, and an accurate prediction of the commercialization potential of patents is achieved, providing a fast and objective means of commercialization evaluation.

CN120875110APending Publication Date: 2025-10-31SHANGHAI PUDONG HOSPITAL +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510738235.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing technologies cannot effectively predict the industrialization potential of medical institution patents, resulting in a large number of patents failing to be transformed into practical applications, and a lack of objective evaluation methods.

Method used

A machine learning-based model for predicting the commercialization potential of medical institution patents was constructed. The random forest algorithm was adopted, and the accuracy of the model was constructed and verified through descriptive analysis of 12 indicators and comparison of various machine learning classification algorithms. The random forest model was then used to predict the probability of patent commercialization.

Benefits of technology

It enables accurate prediction of the patent commercialization potential of medical institutions, and can quickly and objectively identify patents with high and low commercialization potential, providing an effective basis for commercialization evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120875110A_ABST
    Figure CN120875110A_ABST
Patent Text Reader

Abstract

The invention provides a construction method of a medical institution patent industrialization potential prediction model and a medical institution patent industrialization potential prediction system. The construction method of the patent industrialization potential prediction model of the medical institution comprises the steps of 1, performing applicability analysis on machine classification learning algorithms; 2, performing comparative analysis on various machine learning classification models, and selecting the machine learning classification algorithm with the optimal model performance; 3, constructing a medical institution patent industrialization potential prediction model by adopting a machine learning classification algorithm with optimal model performance; and 4, performing multi-aspect verification on the medical institution patent industrialization potential prediction model. By using the medical institution patent industrialization potential prediction model, the medical institution patent convertible probability can be simply, objectively, quickly and accurately predicted. If the predicted industrialization probability is greater than or equal to 70%, the authorized patent of the medical institution has the industrialization potential, that is, the patent transformation possibility is relatively high.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of machine learning technology, and in particular to a method and system for constructing a model for predicting the commercialization potential of patents in medical institutions. Background Technology

[0002] Using valid invention patents held by hospitals nationwide as the denominator and valid invention patents that have not yet been commercialized as the numerator, as of December 2022, patents held by Chinese medical institutions that have not been commercialized, also known as "dormant patents," accounted for 95.53% of the total. Driving the commercialization of these dormant patents has become a major need for improving my country's healthcare system. Faced with a massive number of dormant patents in medical institutions and the continuous emergence of new patent resources, there is an urgent need to provide a solution that can objectively and accurately predict the commercialization potential of medical institution patents, thereby providing an objective basis for the commercialization of scientific and technological achievements in medical institutions. Summary of the Invention

[0003] To meet the above-mentioned technical requirements, the first aspect of the present invention provides a method for constructing a prediction model for the industrialization potential of medical institution patents, comprising:

[0004] Step 1: Applicability Analysis of Machine Classification Learning Algorithm: All sample data from the global medical institution authorized invention patent sample database constructed in this application are grouped into transformed and non-transformed groups, i.e., transformed group and non-transformed group. Descriptive analysis is performed on the sample data of 12 indicators of the industrialization potential of medical institution patents using the mean, variance, maximum, minimum, median, and quartile statistics. The rank-sum test is used to analyze the differences in the overall distribution of the 12 indicators between the two groups of data to determine the applicability of the machine classification learning algorithm. If the P-value of the indicators of the sample data is <0.010 after the rank-sum test, it indicates that there are differences in the overall distribution of the transformed and non-transformed sample data, and it is suitable to use the machine learning classification algorithm to establish a predictive model of the industrialization potential of medical institution patents.

[0005] Step 2: Comparative analysis of multiple machine learning classification models: When the applicability of machine classification learning algorithms meets the conditions, the Scikit-learn machine learning library is used, with random_state set to 1. Six machine learning classification algorithms, namely AdaBoost, K-Nearest Neighbors, Naive Bayes, Random Forest, Decision Tree and Support Vector Machine, are selected for classification performance comparison experiments. The six machine learning classification algorithms are internally validated based on precision, recall, F1 score and AUC area, and the machine learning classification algorithm with the best model performance is selected.

[0006] Step 3: Construct a predictive model for the commercialization potential of medical institution patents using the machine learning classification algorithm with optimal model performance;

[0007] Step 4: Verify the prediction model for the industrialization potential of medical institution patents from multiple perspectives.

[0008] Furthermore, in step 3, the machine learning classification algorithm with the best model performance is random forest.

[0009] Furthermore, in step 3, the method for constructing the random forest model includes:

[0010] Step 3.1: Data Sampling: Randomly select 20% of the data from the original patent dataset as the training set for each tree through bootstrapping, and use the remaining patent data as the internal validation set;

[0011] Step 3.2: Feature selection: During the training process of each tree, a subset of features are randomly selected for the selection of split nodes;

[0012] Step 3.3: Constructing Decision Trees: Construct decision trees based on the selected features, allowing each tree to grow as much as possible without pruning;

[0013] Step 3.4: Prediction: For a new input sample, each tree will give a prediction result; use random grid search to search for hyperparameters of the random forest model to obtain the optimal parameter model. The optimal model hyperparameters are: the maximum number of iterations for the weak learner is 50, the maximum depth of the decision tree is 7, the minimum number of samples required for internal node repartition is 130, and the minimum number of samples for leaf nodes is 30.

[0014] Step 3.5: Results Summary: For classification problems, a simple voting method, i.e., majority voting, is usually used. The final prediction result is the majority vote of all decision tree prediction results.

[0015] Furthermore, in step 4, the verification method for the random forest model includes:

[0016] Step 4.1: Internal Validation: Based on the internal validation set, the random forest model is internally validated in terms of precision, recall, F1 score, and AUC area.

[0017] Step 4.2: External validation based on unconverted patent sample data: The random forest model is used to evaluate the industrialization potential of unconverted authorized invention patents of domestic and foreign medical institutions in the sample database, and the evaluation result of each patent, i.e., the conversion probability, is output. A bar chart of the conversion probability of authorized invention patents of domestic and foreign medical institutions is drawn, and the data is fitted. If the conversion probability of authorized invention patents of domestic and foreign medical institutions both show a logarithmic curve distribution trend, it indicates that the random forest model has predictive accuracy.

[0018] Step 4.3: External validation based on converted patent sample data: Collect patent data from several medical institutions that have been actually converted, including the conversion contract amount and the most relevant indicators of the industrialization potential of 12 medical institution patents corresponding to the patent technology, to form an external data validation set for validating the effectiveness of the random forest model. Calculate the average conversion probability of all patent data. If the average conversion probability is greater than 0.5, it indicates that the random forest model of this application has predictive accuracy for the overall external data.

[0019] The second aspect of this application provides a system for predicting the industrialization potential of medical institution patents, which includes:

[0020] The data input module is used to manually input or automatically retrieve 12 industrialization potential indicators of the medical institution's patents to be transformed from a local database and / or a third-party database. These 12 indicators include the number of standardized patent citations, the number of patent document citations, inventor weight, number of patent lawsuits, patents in three locations, patent family size, number of IPC classification numbers, number of claims, number of hospital beds, number of priorities, patent maintenance period, and technology lifecycle. Inventor weight is based on the inventor weight of the first inventor of each patent among the authorized invention patents of medical institutions worldwide. The calculation formula is: Inventor Weight = Total number of invention patents applied for by the first inventor in a certain field ÷ Total number of invention patent applications in that field, where "field" is represented by the main classification number of each patent sample, accurate to the subclass.

[0021] The patent convertibility probability prediction module integrates the medical institution patent industrialization potential prediction model obtained by the construction method of the above-mentioned medical institution patent industrialization potential prediction model, and uses the medical institution patent industrialization potential prediction model to predict the industrialization probability of the current patent to be converted.

[0022] The evaluation result output module is used to output the industrialization probability of the current patent to be transformed, as well as the evaluation result of whether it has industrialization potential.

[0023] Furthermore, if the current patent's probability of industrialization is greater than or equal to a specific threshold, the evaluation result output module assesses that the current patent has industrialization potential; otherwise, it assesses that the current patent does not have industrialization potential.

[0024] Furthermore, the specific threshold is 70%.

[0025] Compared with existing technologies, the above technical solution has the following advantages:

[0026] This application uses 12 indicators of the commercialization potential of medical institution patents and selects the best-performing random forest algorithm from six machine learning classification algorithms to construct a predictive model for the commercialization potential of medical institution patents. This random forest model has shown good predictive performance through both internal and external validation. Using this predictive model, the probability of commercialization of medical institution patents can be predicted simply, objectively, quickly, and accurately. If the predicted commercialization probability is ≥70%, it means that the authorized patent of the medical institution has commercialization potential, i.e., a high probability of patent commercialization. Attached Figure Description

[0027] Figure 1 This is a schematic diagram illustrating the technical principle of this application.

[0028] Figure 2 This is a trend chart showing the probability distribution of patent commercialization in domestic medical institutions.

[0029] Figure 3 This is a trend chart showing the probability distribution of patent commercialization for foreign medical institutions.

[0030] Figure 4 External data validation results of the prediction model for the industrialization potential of medical institution patents. Detailed Implementation

[0031] The advantages of the present invention are further illustrated below with reference to the accompanying drawings and specific embodiments. Those skilled in the art should understand that the following detailed description is illustrative rather than restrictive and should not be construed as limiting the scope of protection of the present invention.

[0032] Example 1: Method for constructing a predictive model for the industrialization potential of medical institution patents

[0033] like Figure 1 As shown in the figure, this embodiment provides a method for constructing a predictive model for the industrialization potential of medical institution patents, including:

[0034] Step 1: Applicability analysis of machine classification learning algorithms.

[0035] This application utilizes various machine learning classification algorithms to construct a predictive model for the commercialization potential of medical institution patents. All 57,508 global hospital authorized invention patent sample data from the global medical institution authorized invention patent sample database (including the incoPat and Derwent patent databases) constructed in this embodiment are grouped into those that have been converted and those that have not. Descriptive analysis is performed on the sample data for 12 indicators using statistical measures such as mean, variance, maximum, minimum, median, and quartiles. The rank-sum test is then used to analyze the differences in the overall distribution of the 12 indicators between the two groups of data to determine the applicability of the machine classification learning algorithm.

[0036] A descriptive analysis of the sample data for 12 indicators from a transformation perspective revealed that the Z-scores of the rank-sum test were all less than 0, indicating that the overall distribution of the data in the transformed group was greater than that in the non-transformed group. Furthermore, the rank-mean value of the technology lifecycle of the transformed patent sample group was greater than that of the non-transformed group. This differs from the results of rank-sum tests on data groups from both domestic and international sources. Although the number of patent litigation cases contains significant information, the rank-mean values ​​of the transformed and non-transformed patent samples did not differ considerably. Since the P-values ​​of the sample data indicators after the rank-sum test were all <0.010, it indicates a difference in the overall distribution of the transformed and non-transformed sample data. Therefore, a machine learning classification algorithm is suitable for establishing a predictive model for the patent commercialization potential of medical institutions.

[0037] Step 2: Comparative analysis of multiple machine learning classification models.

[0038] The ultimate task of this machine learning model is to predict the commercialization potential of authorized invention patents of medical institutions using a binary classification model. A high probability of commercialization indicates high potential for industrialization, while a low probability indicates low potential. The model's performance is evaluated using the F1 score based on the confusion matrix and the ROC curve. Using the Scikit-learn (sklearn) machine learning library with random_state set to 1, six machine learning classification algorithms—AdaBoost, K-Nearest Neighbors, Naive Bayes, Random Forest, Decision Tree, and Support Vector Machine—were compared. The constructed prediction model showed good identification results for patent commercialization potential under each machine learning classification algorithm, as shown in Table 1. The AUC of Support Vector Machine and Naive Bayes algorithms ranged from 0.600 to 0.650, outperforming the random model and demonstrating a certain level of accuracy. The AUC of Decision Tree, K-Nearest Neighbors, AdaBoost, and Random Forest algorithms ranged from 0.700 to 0.900, showing good identification accuracy. It can be observed that among the six methods, the prediction model based on the random forest algorithm exhibits the best performance, demonstrating the application value of the random forest algorithm in identifying the potential for patent commercialization. Therefore, this application adopts the random forest algorithm to construct a prediction model for the potential for patent commercialization in medical institutions.

[0039] Table 1. Comparison of performance scores of six machine learning classification algorithms for assessing patent commercialization potential.

[0040]

[0041] Accuracy represents the proportion of correctly classified samples out of the total number of samples. It is the ratio of the number of correctly classified medical institution patents to the total number of samples.

[0042] ,

[0043] Precision (P-value) represents the ratio of the number of patents correctly classified into a particular category to the number of patents actually classified into that category.

[0044] ,

[0045] Recall (R value) represents the ratio of the number of samples correctly classified by the model to the total number of samples in a specific patent conversion category.

[0046] ,

[0047] The F1 score is a weighted average of precision and recall. In practical applications, evaluating the performance of classification models requires considering both precision and recall; therefore, a weighted harmonic average of the two is used as the evaluation metric. Its maximum value is 1, and its minimum value is 0; a higher value indicates better predictive performance.

[0048] ,

[0049] The ROC curve, also known as the "Receiver Operating Characteristic" curve, is primarily used to evaluate the prediction accuracy of X for Y. Based on the classifier's prediction results, a threshold is adjusted from 0 to a maximum. Initially, samples from the global database of authorized medical institution patents constructed in this application are used as positive examples for prediction. As the threshold increases, the number of positive examples predicted by the classifier decreases until finally no samples are positive. During this process, the values ​​of two key quantities are calculated at each stage, and these are plotted on the x and y axes to form the "ROC curve." This method is simple, intuitive, and allows for the visual observation and analysis of the classifier's accuracy.

[0050] When comparing classifiers, if the ROC curve of one classifier is completely "enclosed" by the curve of another, it can be asserted that the latter performs better than the former. However, if the ROC curves of the two classifiers intersect, it is difficult to generally conclude which is superior. In this case, if a comparison must be made, a more reasonable criterion is to compare the area under the ROC curves, i.e., the AUC area. The AUC area is a performance indicator for measuring the quality of a classifier; the larger the AUC area value, the better the model's predictive performance on the prospects of patent commercialization for medical institutions.

[0051] The four main elements of the confusion matrix in Table 2 are used to characterize the classification of the predicted commercialization prospects of medical institution authorized invention patents in the test set. The true positive represents the number of medical institution authorized invention patents that have been commercialized and are also judged as commercialized by the classifier; the true negative represents the number of medical institution authorized invention patents that have not been commercialized and are also judged as uncommercialized by the classifier; the pseudo positive represents the number of medical institution authorized invention patents that have not been commercialized but are judged as commercialized by the classifier; and the pseudo negative represents the number of medical institution authorized invention patents that have been commercialized but are judged as uncommercialized by the classifier.

[0052] Table 2. Confusion Matrix of the Classification of Prospects for Commercialization of Authorized Invention Patents in Medical Institutions

[0053]

[0054] Step 3: Construct a model to predict the potential for the commercialization of medical institution patents using the machine learning classification algorithm with the best model performance.

[0055] A classification performance comparison experiment was conducted using the random forest machine learning classification algorithm. The standardized data of the influencing factors of the commercialization potential of medical institution patents were assigned 0 / 1 variables for classification prediction tasks to avoid overfitting problems that may occur in different types of machine algorithms. The training set and prediction set were randomly selected from the patent data at a ratio of 8:2 to conduct training experiments on the prediction prospects of medical institution patent commercialization.

[0056] The method for constructing the random forest model includes:

[0057] Step 3.1: Data Sampling: Randomly select 20% of the data from the original patent dataset as the training set for each tree through bootstrapping, and use the remaining patent data as the internal validation set;

[0058] Step 3.2: Feature selection: During the training process of each tree, a subset of features are randomly selected for the selection of split nodes;

[0059] Step 3.3: Constructing Decision Trees: Construct decision trees based on the selected features, allowing each tree to grow as much as possible without pruning;

[0060] Step 3.4: Prediction: For a new input sample, each tree will give a prediction result; the parameters of the random forest model are: the maximum number of iterations for the weak learner is 50, the maximum depth of the decision tree is 7, the minimum number of samples required for the internal node to be split is 130, and the minimum number of samples for the leaf node is 30.

[0061] Step 3.5: Results Summary: For classification problems, a simple voting method, i.e., majority voting, is usually used.

[0062] Step 4: Verify the prediction model for the industrialization potential of medical institution patents from multiple perspectives.

[0063] Validation methods for random forest models include:

[0064] Step 4.1: The internal validation model showed good prediction results.

[0065] Based on the internal validation set, the random forest model was internally validated in terms of precision, recall, F1 score, and AUC area.

[0066] As shown in Table 1, the random forest model has a precision of 0.871, a recall of 0.874, an F1 score of 0.869, and an AUC area of ​​0.900.

[0067] Step 4.2: External Validation Based on Unconverted Patent Sample Data

[0068] The results obtained from the predictive model of the commercialization potential of medical institution patents are consistent with existing literature and general experience.

[0069] The random forest model is used to evaluate the industrialization potential of domestic and foreign untransformed authorized invention patents in the sample database. The evaluation result of each patent, i.e., the transformation probability, is output. A bar chart of the transformability of authorized invention patents of domestic and foreign medical institutions is drawn, and the data is fitted. If the transformability probabilities of authorized invention patents of domestic and foreign medical institutions all show a logarithmic curve distribution trend, it indicates that the random forest model has predictive accuracy.

[0070] Based on the selection of the optimal machine learning classification model, a random forest model was used to predict the commercialization prospects of uncommercialized authorized invention patents from domestic and international medical institutions in the sample database. The prediction result for each patent, i.e., the commercialization probability, was output. Based on the output commercialization probabilities, a 10-level patent commercialization probability threshold was set using a 10-point standard evaluation method. The results are shown in Table 3. Table 3 shows that 1% (516 patents) of the uncommercialized authorized invention patents from medical institutions have a commercialization prospect level of 10, with a commercialization probability exceeding 90%. The content of these patents can be considered the medical technologies with the greatest development prospects and industrialization potential. Overall, the proportion of patents at level 7 and above, i.e., a commercialization probability of over 70%, is only 4.788% (2072 / 43275). This indicates that from a commercialization perspective, the number of patents with good commercialization prospects in medical institutions is relatively small. Meanwhile, patents at levels 1 to 3 accounted for 70.084% (30,329 / 43,275), reflecting the low probability and difficulty of patent commercialization for most medical institutions. Nearly 70% of authorized invention patents held by medical institutions risk becoming "permanently dormant patents." Further stratifying the data, if patents at level 7 (with a commercialization probability of 70% or higher) are defined as high-value patents, the proportion of high-value patents held by foreign medical institutions (7.603%) is higher than that of domestic medical institutions (4.788%).

[0071] Table 3. Predicted conversion probability of unconverted patents in medical institutions

[0072]

[0073] Based on the prediction of the commercialization prospects of patents in medical institutions, a bar chart of the commercialization potential of authorized invention patents in domestic and foreign medical institutions was created, and the data was fitted. Figure 2 and Figure 3 It is evident that the probability of commercialization of authorized invention patents held by domestic and international medical institutions exhibits a logarithmic curve distribution. Fitting a logarithmic function equation to the probability curve of commercialization of authorized invention patents held by domestic medical institutions yields the following result: , The logarithmic function equation was fitted to the probability curve of the convertibility of authorized invention patents of foreign medical institutions, and the result is as follows: , , Both values ​​are greater than 0.9, indicating a good fit between the two curves. The logarithmic curve of the probability of patent commercialization for foreign medical institutions is steeper than that for domestic medical institutions. However, both domestic and foreign medical institutions face the same problem in patent commercialization: most patents cannot be commercialized. Studies have shown that patent value usually follows a log-normal curve distribution trend (see: NAIR SS, MATHEW M, NAG D. Dynamics between patent latent variables and patent price[J]. Technovation, 2011,31(12):648-654.). That is, in a certain field, only a small portion of patents have commercialization potential and most patents have low value. The results obtained by the patent commercialization potential prediction model in this application are consistent with existing literature and general experience.

[0074] Step 4.3: External Validation Based on Transformed Patent Sample Data

[0075] Based on the established prediction model, the conversion probability of 368 medical institution patent conversion data containing transfer prices was predicted. It was found that as the conversion amount increases, the conversion probability also increases, which is consistent with the actual situation.

[0076] To further verify the effectiveness of the model, data on 368 patents that had been converted from patents in 18 medical institutions in Shanghai were collected. Each patent included the amount of the conversion contract. Data on 12 key influencing factors of the commercialization potential of medical institution patents corresponding to these 368 patents were also collected, forming an external data validation set to verify the effectiveness of the predictive model for the commercialization potential of medical institution patents established in this application. The validation results showed that the average conversion probability of the 368 patent data was 0.608, greater than 0.5, and the F1 score of the prediction model was 0.959, indicating that the model has a good evaluation effect for the overall external data. The data was stratified according to the conversion probabilities in Table 3, forming an abscissa. The primary axis (blue bar chart) represents the average conversion contract amount (ten thousand yuan) of the patents contained in each stratum, and the secondary axis (orange line chart) represents the average conversion probability of each stratum, forming... Figure 4 ,from Figure 4 It can be observed that as the conversion amount increases, the conversion probability also increases.

[0077] Example 2: Medical Institution Patent Industrialization Potential Prediction System

[0078] This embodiment provides a patent commercialization potential prediction system for medical institutions, which includes: a data input module, a patent commercialization probability prediction module, an evaluation result output module, a user interface, and a patent prediction database.

[0079] The data input module is used to obtain 12 industrialization potential indicators of current medical institution patents through manual input or automatic acquisition from local databases (such as patent prediction databases) and / or third-party databases. The 12 industrialization potential indicators include the number of standardized patent citations, the number of patent document citations, inventor weight, number of patent litigations, patents in three locations, patent family size, number of IPC classification numbers, number of claims, number of hospital beds, number of priorities, patent maintenance period, and technology life cycle. The inventor weight is based on the inventor weight of the first inventor of each patent in the global medical institution's authorized invention patents. The calculation formula is: Inventor weight = Total number of invention patents applied for by the first inventor in a certain field ÷ Total number of invention patent applications in that field, where "field" is represented by the main classification number of each patent sample, accurate to the subclass.

[0080] The patent conversion probability prediction module integrates the medical institution patent industrialization potential prediction model obtained in Example 1, and uses the medical institution patent industrialization potential prediction model to predict the conversion probability of the current patent.

[0081] The evaluation results output module outputs the conversion probability of the current patent to be converted, displaying the evaluation results. For example, if the predicted patent conversion probability is ≥70%, it means that the patent granted by the medical institution has high patent industrialization potential and a high probability of patent conversion.

[0082] The user interface is used for login by individuals and administrators. Individuals can register and log in using their mobile phone number. The administrator interface can review individual registration information and grant individual accounts prediction and assessment permissions; the personal information page allows users to view submitted prediction records and download conversion probability prediction results.

[0083] The patent prediction database is used to retain already evaluated patent data and has functions such as querying and adding patent prediction data.

[0084] It should be noted that the embodiments of the present invention have better implementability and are not intended to limit the present invention in any way. Any person skilled in the art may use the above-disclosed technical content to change or modify it into equivalent effective embodiments. However, any modifications or equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solution of the present invention shall still fall within the scope of the technical solution of the present invention.

Claims

1. A method for constructing a predictive model for the industrialization potential of medical institution patents, characterized in that, include: Step 1: Applicability Analysis of Machine Classification Learning Algorithm: All sample data from the global medical institution authorized invention patent sample database constructed in this application are grouped into transformed and non-transformed groups, i.e., transformed group and non-transformed group. Descriptive analysis is performed on the sample data of 12 indicators of the industrialization potential of medical institution patents using the mean, variance, maximum, minimum, median, and quartile statistics. The rank-sum test is used to analyze the differences in the overall distribution of the 12 indicators between the two groups of data to determine the applicability of the machine classification learning algorithm. If the P-value of the indicators of the sample data is <0.010 after the rank-sum test, it indicates that there are differences in the overall distribution of the transformed and non-transformed sample data, and it is suitable to use the machine learning classification algorithm to establish a predictive model of the industrialization potential of medical institution patents. Step 2: Comparative analysis of multiple machine learning classification models: When the applicability of machine classification learning algorithms meets the conditions, the Scikit-learn machine learning library is used, with random_state set to 1. Six machine learning classification algorithms, namely AdaBoost, K-Nearest Neighbors, Naive Bayes, Random Forest, Decision Tree and Support Vector Machine, are selected for classification performance comparison experiments. The six machine learning classification algorithms are internally validated based on precision, recall, F1 score and AUC area, and the machine learning classification algorithm with the best model performance is selected. Step 3: Construct a predictive model for the commercialization potential of medical institution patents using the machine learning classification algorithm with optimal model performance; Step 4: Verify the prediction model for the industrialization potential of medical institution patents from multiple perspectives.

2. The method for constructing the medical institution patent industrialization potential prediction model as described in claim 1, characterized in that, In step 3, the machine learning classification algorithm with the best model performance is random forest.

3. The method for constructing the medical institution patent industrialization potential prediction model as described in claim 2, characterized in that, In step 3, the method for constructing the random forest model includes: Step 3.1: Data Sampling: Randomly select 20% of the data from the original patent dataset as the training set for each tree through bootstrapping, and use the remaining patent data as the internal validation set; Step 3.2: Feature Selection: During the training process of each tree, a subset of features are randomly selected for the selection of split nodes; Step 3.3: Constructing Decision Trees: Construct decision trees based on the selected features, allowing each tree to grow as much as possible without pruning; Step 3.4: Prediction: For a new input sample, each tree will give a prediction result; use random grid search to search for hyperparameters of the random forest model to obtain the optimal parameter model. The optimal model hyperparameters are: the maximum number of iterations for the weak learner is 50, the maximum depth of the decision tree is 7, the minimum number of samples required for internal node re-split is 130, and the minimum number of samples for leaf nodes is 30. Step 3.5: Results Summary: For classification problems, a simple voting method, i.e., majority voting, is usually used. The final prediction result is the majority vote of all decision tree prediction results.

4. The method for constructing the medical institution patent industrialization potential prediction model as described in claim 2, characterized in that, In step 4, the verification method for the random forest model includes: Step 4.1: Internal Validation: Based on the internal validation set, the random forest model is internally validated in terms of precision, recall, F1 score, and AUC area. Step 4.2: External validation based on unconverted patent sample data: The random forest model is used to evaluate the industrialization potential of unconverted authorized invention patents of domestic and foreign medical institutions in the sample database, and the evaluation result of each patent, i.e., the conversion probability, is output. A bar chart of the conversion probability of authorized invention patents of domestic and foreign medical institutions is drawn, and the data is fitted. If the conversion probability of authorized invention patents of domestic and foreign medical institutions both show a logarithmic curve distribution trend, it indicates that the random forest model has predictive accuracy. Step 4.3: External validation based on converted patent sample data: Collect patent data from several medical institutions that have been actually converted, including the conversion contract amount and the most relevant indicators of the industrialization potential of 12 medical institution patents corresponding to the patent technology, to form an external data validation set for validating the effectiveness of the random forest model. Calculate the average conversion probability of all patent data. If the average conversion probability is greater than 0.5, it indicates that the random forest model of this application has predictive accuracy for the overall external data.

5. A system for predicting the industrialization potential of medical institution patents, characterized in that, include: The data input module is used to obtain 12 industrialization potential indicators of the medical institution's patents to be transformed by manual input or automatic acquisition from local databases and / or third-party databases. The 12 industrialization potential indicators include the number of standardized patent citations, the number of patent literature citations, the inventor's weight, the number of patent lawsuits, patents in three locations, the size of the patent family, the number of IPC classification numbers, the number of claims, the number of hospital beds, the number of priorities, the patent maintenance period, and the technology life cycle. The inventor's weight is based on the inventor's weight of the first inventor of each patent in the global medical institution's authorized invention patents. The calculation formula is: Inventor's weight = Total number of invention patents applied for by the first inventor in a certain field ÷ Total number of invention patent applications in that field, where "field" is represented by the main classification number of each patent sample, accurate to the subclass. The patent convertibility probability prediction module integrates a medical institution patent industrialization potential prediction model obtained by the method described in any one of claims 1 to 4, and uses the medical institution patent industrialization potential prediction model to predict the industrialization probability of the current patent to be converted. The evaluation result output module is used to output the industrialization probability of the current patent to be transformed, as well as the evaluation result of whether it has industrialization potential.

6. The medical institution patent industrialization potential prediction system as described in claim 5, characterized in that, If the probability of commercialization of the current patent is greater than or equal to a certain threshold, the evaluation result output module assesses that the current patent has commercialization potential; otherwise, it assesses that the current patent does not have commercialization potential.

7. The medical institution patent industrialization potential prediction system as described in claim 6, characterized in that, The specific threshold is 70%.