Enterprise type determination method, apparatus, device, medium, and product

By acquiring enterprise data, calculating and filtering feature variable weights, and using the Cartesian product algorithm to determine enterprise types, the problem of insufficient feature variable fields when commercial banks classify cooperative enterprises is solved, thus improving the accuracy and interpretability of classification.

CN115809425BActive Publication Date: 2026-07-31SHANGHAI PUDONG DEVELOPMENT BANK
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI PUDONG DEVELOPMENT BANK
Filing Date
2022-12-23
Publication Date
2026-07-31

AI Technical Summary

Technical Problem

When commercial banks classify the technological level of their partner companies, existing methods may result in an insufficient number of characteristic variable fields, affecting the accuracy of classification.

Method used

By acquiring enterprise data, identifying multiple characteristic variables, calculating the weight of each characteristic variable, selecting the second characteristic variable, and using the Cartesian product algorithm to calculate the weight of the target combination, the enterprise type is finally determined based on the preset weight range.

Benefits of technology

This effectively solves the problem of insufficient number of feature variable fields, and improves the accuracy and interpretability of enterprise type classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115809425B_ABST
    Figure CN115809425B_ABST
Patent Text Reader

Abstract

This application relates to a method, apparatus, computer equipment, storage medium, and computer program product for determining enterprise type. The method includes: firstly, acquiring enterprise data and determining multiple first characteristic variables corresponding to the enterprise data. The enterprise data includes basic enterprise information, enterprise technical capability data, and enterprise scientific and technological achievements data. Then, based on a preset grouping and characteristic variable judgment conditions for each first characteristic variable, calculating a first weight for each first characteristic variable, and filtering the first characteristic variables according to the first weights to obtain second characteristic variables. Next, based on the Cartesian product algorithm, calculating a second weight for a target combination of preset groups of different second characteristic variables according to the preset groupings of each second characteristic variable. Finally, determining the enterprise type based on a preset weight range and the second weight. The method provided by this application can effectively solve the problem of insufficient number of characteristic variable fields.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information science and technology, and in particular to a method, apparatus, computer equipment, storage medium, and computer program product for determining enterprise type. Background Technology

[0002] Commercial banks provide their partner companies with application programming interfaces (APIs) for interconnection, thereby exporting their financial service and information technology capabilities to better serve these companies. As the number of partner companies increases and the technological levels of these companies vary significantly, it is necessary for commercial banks to categorize their partner companies based on their technological capabilities.

[0003] Currently, commercial banks often use methods such as outlier handling and missing value imputation to process feature variables and classify the technological level of partner companies based on the processed feature variables. This method may result in problems such as missing fields or insufficient number of fields in the processed feature variables. Summary of the Invention

[0004] Therefore, it is necessary to provide a method, apparatus, computer equipment, computer-readable storage medium, and computer program product for determining enterprise type that can solve the problem of insufficient number of feature variable fields, in order to address the above-mentioned technical issues.

[0005] Firstly, this application provides a method for determining enterprise type, the method comprising:

[0006] Acquire enterprise data and determine multiple first feature variables corresponding to the enterprise data, wherein the enterprise data includes basic enterprise information, enterprise technical capability data, and enterprise scientific and technological achievement data;

[0007] Calculate the first weight of each first feature variable based on the preset grouping and feature variable judgment conditions for each first feature variable;

[0008] The second feature variable is obtained by filtering the first feature variable according to the first weight;

[0009] Based on the Cartesian product algorithm, the second weight of the target combination of the preset groups of different second feature variables is calculated according to the preset grouping of each second feature variable;

[0010] The enterprise type is determined based on the preset weight range and the second weight.

[0011] In one embodiment, calculating the first weight of each first feature variable based on the preset grouping and feature variable judgment conditions for each first feature variable includes:

[0012] The third weight of each preset group is determined based on the judgment conditions of the aforementioned feature variables;

[0013] Calculate the first weight of each first feature variable based on all the third weights corresponding to each first feature variable.

[0014] In one embodiment, the step of filtering the first feature variable according to the first weight to obtain the second feature variable includes:

[0015] The first weights corresponding to different first feature variables are sorted in ascending order of weight;

[0016] Delete the first feature variables corresponding to the first weight of the first number of sorting results;

[0017] The remaining first characteristic variable is determined as the second characteristic variable.

[0018] In one embodiment, the step of calculating the second weight of the target combination of different second feature variables based on the Cartesian product algorithm, according to the preset grouping of each second feature variable, includes:

[0019] The target combination is obtained by combining the preset groups of different second feature variables according to the Cartesian product algorithm.

[0020] The second weight is calculated based on the number of enterprise data corresponding to the target combination.

[0021] In one embodiment, determining the enterprise type based on a preset weight range and the second weight includes:

[0022] Calculate the fourth weight corresponding to the enterprise based on the second weight of all target combinations corresponding to the enterprise;

[0023] The enterprise type is determined based on the fourth weight and the preset weight range.

[0024] In one embodiment, before calculating the first weight of each first feature variable based on the preset grouping and feature variable judgment conditions for each first feature variable, the method further includes:

[0025] Place all enterprise data corresponding to each first feature variable in a one-dimensional coordinate system;

[0026] Based on the distribution of all enterprise data corresponding to each first feature variable in the one-dimensional coordinate system, calculate the local reachability density of each enterprise data corresponding to each first feature variable in the one-dimensional coordinate system.

[0027] The outlier factor for each enterprise data corresponding to each first feature variable is determined based on the local reachability density.

[0028] Determine the magnitude of the outlier factor and the preset value. If the outlier factor is greater than the preset value, delete the enterprise data corresponding to the outlier factor.

[0029] After deleting the enterprise data corresponding to the outlier factor, the step of calculating the first weight of each first feature variable based on the preset grouping and feature variable judgment conditions of each first feature variable is performed.

[0030] Secondly, this application also provides an apparatus for determining enterprise type, the apparatus comprising:

[0031] The acquisition module is used to acquire enterprise data and determine multiple first feature variables corresponding to the enterprise data. The enterprise data includes basic enterprise information, enterprise technical capability data, and enterprise scientific and technological achievement data.

[0032] The first calculation module is used to calculate the first weight of each first feature variable based on the preset grouping and feature variable judgment conditions for each first feature variable;

[0033] The filtering module is used to filter the first feature variable according to the first weight to obtain the second feature variable;

[0034] The second calculation module is used to calculate the second weight of the target combination of the preset groups of different second feature variables based on the Cartesian product algorithm and according to the preset grouping of each second feature variable;

[0035] The determination module is used to determine the enterprise type based on a preset weight range and the second weight.

[0036] Thirdly, this application also provides a computer device. The computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the methods in any of the above embodiments.

[0037] Fourthly, this application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program thereon, which, when executed by a processor, implements the steps of the methods in any of the above embodiments.

[0038] Fifthly, this application also provides a computer program product. The computer program product includes a computer program that, when executed by a processor, implements the steps of the methods in any of the above embodiments.

[0039] The aforementioned method, apparatus, computer equipment, storage medium, and computer program product for determining enterprise type first acquires enterprise data and determines multiple first characteristic variables corresponding to the enterprise data. Enterprise data includes basic enterprise information, enterprise technical capability data, and enterprise scientific and technological achievements data. Then, based on preset groupings and characteristic variable judgment conditions for each first characteristic variable, a first weight is calculated for each first characteristic variable. Second characteristic variables are obtained by filtering the first characteristic variables based on the first weights. Next, based on the Cartesian product algorithm, a second weight is calculated for the target combination of preset groups of different second characteristic variables according to the preset groupings of each second characteristic variable. Finally, the enterprise type is determined based on the preset weight range and the second weight. The method provided in this application, which calculates the second weight of the target combination based on the Cartesian product and the preset groupings of the second characteristic variables, can effectively solve the problem of insufficient number of characteristic variable fields. Attached Figure Description

[0040] Figure 1 This is a flowchart illustrating a method for determining enterprise type in one embodiment;

[0041] Figure 2 This is a flowchart illustrating the first weight calculation method in one embodiment;

[0042] Figure 3 This is a structural block diagram of an enterprise type determination device in one embodiment;

[0043] Figure 4 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation

[0044] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0045] In one embodiment, such as Figure 1 As shown, a method for determining enterprise type is provided. This embodiment illustrates the application of this method to a terminal. It is understood that this method can also be applied to a server, and further to a system including both a terminal and a server, and is implemented through interaction between the terminal and the server. In this embodiment, the method includes the following steps:

[0046] S102. Obtain enterprise data and determine the various primary characteristic variables corresponding to the enterprise data. Enterprise data includes basic enterprise information, enterprise technical capability data, and enterprise scientific and technological achievement data.

[0047] Enterprise data is used to characterize the technological capabilities of enterprises cooperating with commercial banks. The first characteristic variable corresponding to the basic information of an enterprise may include the industry in which the enterprise operates, the nature of the enterprise, the size of the enterprise, the efficiency of the enterprise, the products and services, the sales of products, the market share, and the information of the entrepreneur. The first characteristic variable corresponding to the enterprise's technological capability data may include the scale of research and development, the research and development model, the information of technical personnel, the technological investment, and the cooperation between industry and academia. Among them, the information of technical personnel may include the total number of technical personnel, their age, education, professional title, and the mobility information of technical personnel. The technological investment may include the financial investment and the source of funding. The enterprise's scientific and technological achievements data may include projects, patents, papers, scientific and technological achievements awards, honors, technological income, and academic influence ranking information.

[0048] S104. Calculate the first weight of each first feature variable based on the preset grouping and feature variable judgment conditions for each first feature variable.

[0049] Preset grouping refers to the pre-defined grouping of the first characteristic variable based on the actual situation and project experience of the cooperating enterprises. For example, the nature of the enterprise can be divided into state-owned enterprises, foreign-funded enterprises, and private enterprises, and the R&D model can be divided into independent R&D, commissioned R&D, and collaborative R&D. The characteristic variable judgment condition is used to determine the third weight of each preset group. For example, the characteristic variable judgment condition may include that the enterprise is a non-technology industry enterprise, the R&D expenditure is less than 3% of the sales revenue, and the enterprise has no technological achievements. When a commercial bank's cooperating enterprise meets any of the above judgment conditions, the cooperating enterprise can be defined as a bad sample.

[0050] Specifically, the WOE (Weight of Evidence) value of each preset group for each first characteristic variable is calculated based on the judgment conditions of the characteristic variable. Then, the IV (Information Value) value of each preset group is calculated based on the WOE value of each preset group. The IV value of each preset group is the third weight of each preset group. The sum of the third weights of all preset groups for each first characteristic variable is determined as the first weight of that first characteristic variable.

[0051] S106. The second feature variable is obtained by filtering the first feature variable according to the first weight.

[0052] After determining the first weights of all first feature variables, sort all the first weights in ascending order, delete the first feature variables corresponding to the first weights of the first weights in the sorting results (a preset number of times), and determine the first feature variables corresponding to the remaining first weights as second feature variables.

[0053] S108. Based on the Cartesian product algorithm, calculate the second weight of the target combination of different second feature variables according to the preset grouping of each second feature variable.

[0054] The target combination refers to the combination of different pre-defined groups of different feature variables. For example, based on the Cartesian product algorithm, the pre-defined groups of enterprise nature and R&D mode can be combined to obtain the nine combined feature variables shown in Table 1:

[0055] Table 1:

[0056] State-owned enterprises State-owned enterprises independently develop State-owned enterprises commissioned research and development State-owned enterprise collaborative research and development Foreign-invested enterprises Foreign-funded independent research and development Foreign-funded commissioned R&D Foreign-funded collaborative research and development private enterprises Private independent research and development Private commissioned research and development Private sector collaborative research and development

[0057] By calculating the IV value of each combination of characteristic variables, where the IV value of each combination of characteristic variables is the second weight of each target combination and the combination of characteristic variables is the target combination, the proportion of enterprises of different types adopting each R&D model can be determined. Based on the determined proportion, the R&D capabilities of enterprises of different types can be determined.

[0058] For discrete and continuous feature variables, such as the professional titles of researchers and the average number of patents held by researchers, the Cartesian product algorithm can be used to combine the preset groups of professional titles and the number of patents to obtain the nine combined feature variables shown in Table 2.

[0059] Table 2:

[0060] Intermediate professional title Intermediate &> 5 pieces Intermediate &>10 pieces Intermediate &>15 pieces Associate senior professional title Subtropical high &>5 pieces Subtropical high &>10 pieces Subtropical high-pressure system (15 items) Senior professional title Advanced > 5 items Advanced &>10 items Advanced &>15 items

[0061] By calculating the IV value of each combination of characteristic variables, the average research capability of researchers within the enterprise can be determined.

[0062] S110. Determine the enterprise type based on the preset weight range and the second weight.

[0063] The preset weight range is used to determine the enterprise type based on the second weight corresponding to the preset group. For example, the second weights for the three target groups corresponding to state-owned enterprises in the enterprise type are: 0.1 for independent R&D of state-owned enterprises, 0.15 for commissioned R&D of state-owned enterprises, and 0.15 for cooperative R&D of state-owned enterprises. Therefore, the second weight corresponding to state-owned enterprises is 0.4. The second weights for researchers with intermediate professional titles corresponding to different numbers of patents are: 0.15 for more than 5 patents, 0.1 for more than 10 patents, and 0.1 for more than 15 patents. Therefore, the second weight corresponding to researchers with intermediate professional titles is 0.35. If the preset weight range is that the second weight of state-owned enterprises is greater than 0.5 and the second weight corresponding to researchers with intermediate professional titles is greater than 0.3, the enterprise is identified as a technology enterprise. However, the second weights of the above enterprises do not meet the preset weight range for technology enterprises. Therefore, the above enterprises are non-technology enterprises.

[0064] The aforementioned method for determining enterprise type first acquires enterprise data and determines various primary characteristic variables corresponding to the enterprise data. Enterprise data includes basic enterprise information, enterprise technical capabilities, and enterprise scientific and technological achievements. Then, based on the preset grouping of each primary characteristic variable and the characteristic variable judgment conditions, a primary weight is calculated for each primary characteristic variable. Secondary characteristic variables are obtained by filtering the primary characteristic variables according to the primary weights. Next, based on the Cartesian product algorithm, a secondary weight is calculated for the target combination of different secondary characteristic variables according to the preset grouping of each secondary characteristic variable. Finally, the enterprise type is determined based on the preset weight range and the secondary weight. The method provided in this application, which calculates the secondary weight of the target combination based on the Cartesian product and the preset grouping of the secondary characteristic variables, effectively solves the problem of insufficient number of characteristic variable fields.

[0065] In some embodiments, such as Figure 2 As shown, Figure 2 This is a flowchart illustrating a first weight calculation method in one embodiment. The method calculates the first weight of each first feature variable based on a preset grouping and feature variable judgment conditions, including: determining the third weight of each preset group based on the feature variable judgment conditions; and calculating the first weight of each first feature variable based on all the third weights corresponding to each first feature variable.

[0066] In this step, after determining the third weight of each preset group of each first feature variable according to the feature variable judgment conditions, the sum of all the third weights corresponding to each first feature variable is determined as the first weight of that first feature variable.

[0067] The method provided in this step determines the weight value of the first feature variable based on the weight values ​​of the preset grouping, laying the foundation for the subsequent process of determining the enterprise type.

[0068] In some embodiments, filtering the first feature variables according to the first weight to obtain the second feature variables includes: sorting the first weights corresponding to different first feature variables in ascending order of weight; deleting the first feature variables corresponding to the first preset number of first weights in the sorting result; and determining the remaining first feature variables as the second feature variables.

[0069] In this step, for example, the first characteristic variables of a certain enterprise include enterprise efficiency, R&D scale and number of patents. The first weight of enterprise efficiency is 0.1, the first weight of R&D scale is 0.2, and the first weight of number of patents is 0.3. If the preset number is 1, then enterprise efficiency with the smallest first weight will be removed from the first characteristic variable, and R&D scale and number of patents will be determined as the second characteristic variables.

[0070] The method provided in this step removes the first feature variable with a smaller weight, which can improve the accuracy of subsequent determination of enterprise type.

[0071] In some embodiments, based on the Cartesian product algorithm, the second weight of the target combination of the preset groups of different second feature variables is calculated according to the preset grouping of each second feature variable, including: combining the preset groups of different second feature variables according to the Cartesian product algorithm to obtain the target combination; and calculating the second weight according to the number of enterprise data corresponding to the target combination.

[0072] In this step, the IV value of each target combination is calculated based on the number of enterprise data corresponding to the target combination. The IV value of each target combination is the second weight of each target combination.

[0073] The method provided in this step determines the target combination based on the Cartesian product algorithm, making the obtained target combination more interpretable.

[0074] In some embodiments, determining the enterprise type based on a preset weight range and the second weight includes: calculating a fourth weight corresponding to the enterprise based on the second weight of all target combinations corresponding to the enterprise; and determining the enterprise type based on the fourth weight and the preset weight range.

[0075] In this step, the second weight of each preset group of each second characteristic variable of the enterprise can be determined according to the second weight of the target combination. The second weight of each preset group is the fourth weight of each preset group. Then, by judging the preset weight range in which the fourth weight is located, it is determined whether the enterprise is a technology enterprise.

[0076] The method provided in this step can accurately and efficiently determine whether a company is a technology company.

[0077] Before calculating the first weight of each first characteristic variable based on the preset grouping and characteristic variable judgment conditions for each first characteristic variable, the method further includes: placing all enterprise data corresponding to each first characteristic variable in a one-dimensional coordinate system; calculating the local reachability density of each enterprise data corresponding to each first characteristic variable in the one-dimensional coordinate system based on the distribution of all enterprise data corresponding to each first characteristic variable in the one-dimensional coordinate system; determining the outlier factor of each enterprise data corresponding to each first characteristic variable based on the local reachability density; judging the magnitude of the outlier factor and the preset value, and deleting the enterprise data corresponding to the outlier factor if the outlier factor is greater than the preset value; after deleting the enterprise data corresponding to the outlier factor, performing the step of calculating the first weight of each first characteristic variable based on the preset grouping and characteristic variable judgment conditions for each first characteristic variable.

[0078] In this step, the LOF (Local Outlier Factor) algorithm is used to delete abnormal data from the enterprise data.

[0079] The method provided in this step can remove abnormal data from enterprise data, reducing errors in the subsequent process of determining the enterprise type.

[0080] In one embodiment, this application provides another method for determining enterprise type, comprising the following steps:

[0081] (1) Select data source

[0082] For open banking technology companies, data sources are collected in three aspects: basic enterprise information, enterprise technology information, and enterprise technology achievements. The more data collected, the more accurate the feedback on the enterprise's technological level.

[0083] Basic enterprise information includes information on industry, size, profitability, and products and services; as well as market research information such as product sales, market share, and entrepreneur information.

[0084] A company's technological capabilities include its R&D scale (R&D investment, funding sources, government subsidies), R&D model (e.g., independent R&D, commissioned R&D, collaborative R&D), total number of technical personnel, their age, education, professional titles, and technical personnel turnover information. It also includes information on technological investment (funding input, funding sources) and cooperation with industry-academia institutions.

[0085] Enterprise scientific and technological achievements include information on collected projects, patents, papers, scientific and technological awards, honors, technology income, and academic influence rankings.

[0086] (2) Sample definition

[0087] The definition of a bad sample differs from traditional overdue indicators for credit companies. The state of a bad sample is as follows:

[0088] Company basic information: Non-technology industry company.

[0089] Company technical information: There is no R&D investment, and R&D expenditure is less than 3% of sales revenue.

[0090] Enterprise's scientific and technological achievements: No achievements recorded, R&D revenue.

[0091] If any of the above conditions are met, the sample is considered a bad sample from a technology company that is a partner of the open banking program, and such samples are marked as "rejected".

[0092] (3) Data cleaning

[0093] For data integrity, it is necessary to find the intersection between data sources to ensure the integrity of the data from each data source. We use the credit information of the enterprise's basic information as the main table, and perform a left join on the data of the enterprise's technical capabilities and scientific and technological achievements. This can avoid additional missing values ​​after the data is merged.

[0094] For outliers, the LOF method is used to find and remove them.

[0095] (4) Feature Engineering Construction

[0096] Enterprise status indicators include: enterprise nature, size, enterprise efficiency, products and services, market research information, product sales, market share, years of establishment, and entrepreneur information.

[0097] R&D capability indicators include: R&D model (e.g., independent R&D, commissioned R&D, collaborative R&D), R&D funding, funding sources, government subsidies, total number of technical personnel, age, education, professional titles, and technical personnel mobility information. Also includes information on technology input (funding input and funding sources) and industry-academia collaboration.

[0098] Scientific research achievement indicators: collect information on projects, patents, papers, scientific and technological achievements awards, honors, technology income, and academic influence rankings.

[0099] (5) Data dimensionality reduction

[0100] 1) The variables are encoded using WOE encoding technology in order to increase the interpretability and non-linearity of the variables.

[0101] 2) Using the optimal IV binning method, the number of bins, the upper and lower limit intervals, and the number of samples in each interval are obtained.

[0102] 3) Variable selection: Here we use the filtering method to reduce the dimensionality of variables and select variables suitable for inclusion in the model. First, we select from the perspective of missing data and delete variables with many missing values. Then we select variance variables and remove variables with small variance to improve the predictive ability of variables.

[0103] (6) Data Upgrading

[0104] Upscaling discrete variables, for example, enterprise status indicators—enterprise type (state-owned enterprise, foreign-invested enterprise, private enterprise); and R&D capability indicators—R&D mode (independent R&D, commissioned R&D, collaborative R&D), can be combined using Cartesian products to obtain nine combinations as shown in Table 1. After combination, the dimension is upgraded to 3*3=9 dimensional Cartesian product feature variables, which have better interpretability. For example, business personnel can judge that the technological level of independent R&D in state-owned enterprises is significantly higher than that of collaborative R&D in private enterprises. After combination, WOE encoding is used, followed by training with a Logistic regression model (a generalized linear model) to obtain the score of this combination.

[0105] Upgrading the dimension between discrete and continuous variables, for example, R&D capability indicators - professional titles, with values ​​(intermediate, associate senior, senior); R&D capability indicators - patents. By combining them through Cartesian product, we can obtain 9 combinations as shown in Table 2. Compared with the previous methods, the feature processing method of Cartesian product can obtain more obvious nonlinear features, rather than through the form of variable "addition".

[0106] (7) Algorithm Dimensionality Upgrading

[0107] This embodiment uses a decision tree algorithm as an example, selecting three variables: year of establishment, whether industry-academia collaboration has been established, and academic influence; generating four leaf nodes, i.e., producing four rules, with a decision tree depth of two levels. The algorithm is illustrated as follows: if year of establishment < a certain year & industry-academia collaboration established = yes, then pass; if year of establishment < a certain year & academic influence = high, then pass. This decision tree-based rule achieves an effect similar to increasing the dimensionality of data through Cartesian product.

[0108] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0109] Based on the same inventive concept, this application also provides an apparatus for determining enterprise type to implement the above-described enterprise type determination method. The solution provided by this apparatus is similar to the implementation described in the above-described method; therefore, the specific limitations in one or more embodiments of the enterprise type determination apparatus provided below can be found in the limitations of the enterprise type determination method described above, and will not be repeated here.

[0110] In one embodiment, such as Figure 3 As shown, a business type determination device 300 is provided, including: an acquisition module 301, a first calculation module 302, a filtering module 303, a second calculation module 304, and a determination module 305, wherein:

[0111] The acquisition module 301 is used to acquire enterprise data and determine multiple first feature variables corresponding to the enterprise data. The enterprise data includes basic enterprise information, enterprise technical capability data, and enterprise scientific and technological achievement data.

[0112] The first calculation module 302 is used to calculate the first weight of each first feature variable according to the preset grouping and feature variable judgment conditions of each first feature variable;

[0113] The filtering module 303 is used to filter the first feature variable according to the first weight to obtain the second feature variable;

[0114] The second calculation module 304 is used to calculate the second weight of the target combination of the preset groups of different second feature variables based on the Cartesian product algorithm and according to the preset grouping of each second feature variable;

[0115] The determination module 305 is used to determine the enterprise type based on the preset weight range and the second weight.

[0116] In some embodiments, the first calculation module 302 is further configured to: determine the third weight of each preset group according to the feature variable judgment condition; and calculate the first weight of each first feature variable according to all the third weights corresponding to each first feature variable.

[0117] In some embodiments, the filtering module 303 is further configured to: sort the first weights corresponding to different first feature variables in ascending order of weight; delete the first feature variables corresponding to the first preset number of first weights in the sorting result; and determine the remaining first feature variables as second feature variables.

[0118] In some embodiments, the second calculation module 304 is further configured to: combine preset groups of different second feature variables according to the Cartesian product algorithm to obtain the target combination; and calculate the second weight according to the number of enterprise data corresponding to the target combination.

[0119] In some embodiments, the determining module 305 is further configured to: calculate a fourth weight corresponding to the enterprise based on a second weight of all target combinations corresponding to the enterprise; and determine the enterprise type based on the fourth weight and the preset weight range.

[0120] In some embodiments, the enterprise type determination device 300 is specifically used for: placing all enterprise data corresponding to each first feature variable in a one-dimensional coordinate system; calculating the local reachability density of each enterprise data corresponding to each first feature variable in the one-dimensional coordinate system based on the distribution of all enterprise data corresponding to each first feature variable in the one-dimensional coordinate system; determining the outlier factor of each enterprise data corresponding to each first feature variable based on the local reachability density; judging the magnitude of the outlier factor and a preset value; if the outlier factor is greater than the preset value, deleting the enterprise data corresponding to the outlier factor; after deleting the enterprise data corresponding to the outlier factor, performing the step of calculating the first weight of each first feature variable based on the preset grouping and feature variable judgment conditions of each first feature variable.

[0121] Each module in the aforementioned enterprise type determination device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0122] In one embodiment, a computer device is provided, which may be a terminal, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, input / output interfaces, a communication interface, a display unit, and an input device. The processor, memory, and input / output interfaces are connected via a system bus, and the communication interface, display unit, and input device are also connected to the system bus via the input / output interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The input / output interfaces are used for exchanging information between the processor and external devices. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, mobile cellular networks, NFC (Near Field Communication), or other technologies. When executed by the processor, the computer program implements a method for determining an enterprise type. The display unit is used to form a visually visible image and can be a display screen, a projection device, or a virtual reality imaging device. The display screen can be an LCD screen or an e-ink screen. The input device of the computer device can be a touch layer covering the display screen, or buttons, trackballs, or touchpads set on the casing of the computer device, or external keyboards, touchpads, or mice, etc.

[0123] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0124] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to perform the following steps: acquiring enterprise data and determining multiple first feature variables corresponding to the enterprise data, the enterprise data including basic enterprise information, enterprise technical capability data, and enterprise scientific and technological achievement data; calculating a first weight for each first feature variable based on a preset grouping and feature variable judgment conditions; filtering the first feature variables according to the first weights to obtain second feature variables; calculating a second weight for a target combination of preset groups of different second feature variables based on the Cartesian product algorithm and the preset grouping of each second feature variable; and determining the enterprise type according to a preset weight range and the second weights.

[0125] In one embodiment, the processor, when executing a computer program, calculates a first weight for each first feature variable based on a preset grouping and a feature variable judgment condition, including: determining a third weight for each preset group based on the feature variable judgment condition; and calculating the first weight for each first feature variable based on all the third weights corresponding to each first feature variable.

[0126] In one embodiment, the process of the processor executing a computer program to filter the first feature variables according to the first weight to obtain the second feature variables includes: sorting the first weights corresponding to different first feature variables in ascending order of weight; deleting the first feature variables corresponding to the first preset number of first weights in the sorting result; and determining the remaining first feature variables as the second feature variables.

[0127] In one embodiment, the Cartesian product-based algorithm implemented by the processor when executing a computer program calculates a second weight for a target combination of preset groups of different second feature variables according to the preset grouping of each second feature variable, including: combining the preset groups of different second feature variables according to the Cartesian product algorithm to obtain the target combination; and calculating the second weight according to the number of enterprise data corresponding to the target combination.

[0128] In one embodiment, the process of determining the enterprise type based on a preset weight range and the second weight when the processor executes a computer program includes: calculating a fourth weight corresponding to the enterprise based on the second weight of all target combinations corresponding to the enterprise; and determining the enterprise type based on the fourth weight and the preset weight range.

[0129] In one embodiment, before the processor executes the computer program to calculate the first weight of each first feature variable based on a preset grouping and feature variable judgment conditions, the method further includes: placing all enterprise data corresponding to each first feature variable in a one-dimensional coordinate system; calculating the local reachability density of each enterprise data corresponding to each first feature variable in the one-dimensional coordinate system based on the distribution of all enterprise data corresponding to each first feature variable in the one-dimensional coordinate system; determining the outlier factor of each enterprise data corresponding to each first feature variable based on the local reachability density; determining the magnitude of the outlier factor and a preset value; if the outlier factor is greater than the preset value, deleting the enterprise data corresponding to the outlier factor; and after deleting the enterprise data corresponding to the outlier factor, performing the step of calculating the first weight of each first feature variable based on a preset grouping and feature variable judgment conditions.

[0130] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a processor, the computer program performs the following steps: acquiring enterprise data and determining multiple first feature variables corresponding to the enterprise data, the enterprise data including basic enterprise information, enterprise technical capability data, and enterprise scientific and technological achievement data; calculating a first weight for each first feature variable based on a preset grouping and feature variable judgment conditions; filtering the first feature variables according to the first weights to obtain second feature variables; calculating a second weight for a target combination of preset groups of different second feature variables based on the Cartesian product algorithm and the preset grouping of each second feature variable; and determining the enterprise type according to a preset weight range and the second weights.

[0131] In one embodiment, when a computer program is executed by a processor, it calculates a first weight for each first feature variable based on a preset grouping and a feature variable judgment condition, including: determining a third weight for each preset grouping based on the feature variable judgment condition; and calculating the first weight for each first feature variable based on all the third weights corresponding to each first feature variable.

[0132] In one embodiment, when a computer program is executed by a processor, the process of filtering the first feature variables according to the first weight to obtain the second feature variables includes: sorting the first weights corresponding to different first feature variables in ascending order of weight; deleting the first feature variables corresponding to the first preset number of first weights in the sorting result; and determining the remaining first feature variables as the second feature variables.

[0133] In one embodiment, when a computer program is executed by a processor, it implements a Cartesian product-based algorithm to calculate a second weight of a target combination of preset groups of different second feature variables according to the preset grouping of each second feature variable, including: combining the preset groups of different second feature variables according to the Cartesian product algorithm to obtain the target combination; and calculating the second weight according to the number of enterprise data corresponding to the target combination.

[0134] In one embodiment, the process of determining the enterprise type based on a preset weight range and the second weight when the computer program is executed by a processor includes: calculating a fourth weight corresponding to the enterprise based on the second weight of all target combinations corresponding to the enterprise; and determining the enterprise type based on the fourth weight and the preset weight range.

[0135] In one embodiment, before the computer program, when executed by a processor, calculates the first weight of each first feature variable based on a preset grouping and feature variable judgment conditions, the method further includes: placing all enterprise data corresponding to each first feature variable in a one-dimensional coordinate system; calculating the local reachability density of each enterprise data corresponding to each first feature variable in the one-dimensional coordinate system based on the distribution of all enterprise data corresponding to each first feature variable in the one-dimensional coordinate system; determining the outlier factor of each enterprise data corresponding to each first feature variable based on the local reachability density; determining the magnitude of the outlier factor and a preset value; if the outlier factor is greater than the preset value, deleting the enterprise data corresponding to the outlier factor; and after deleting the enterprise data corresponding to the outlier factor, performing the step of calculating the first weight of each first feature variable based on a preset grouping and feature variable judgment conditions.

[0136] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, performs the following steps: acquiring enterprise data and determining multiple first feature variables corresponding to the enterprise data, the enterprise data including basic enterprise information, enterprise technical capability data, and enterprise scientific and technological achievements data; calculating a first weight for each first feature variable based on a preset grouping and feature variable judgment conditions; filtering the first feature variables according to the first weights to obtain second feature variables; calculating a second weight for a target combination of preset groups of different second feature variables based on the Cartesian product algorithm and the preset grouping of each second feature variable; and determining the enterprise type according to a preset weight range and the second weights.

[0137] In one embodiment, when a computer program is executed by a processor, it calculates a first weight for each first feature variable based on a preset grouping and a feature variable judgment condition, including: determining a third weight for each preset grouping based on the feature variable judgment condition; and calculating the first weight for each first feature variable based on all the third weights corresponding to each first feature variable.

[0138] In one embodiment, when a computer program is executed by a processor, the process of filtering the first feature variables according to the first weight to obtain the second feature variables includes: sorting the first weights corresponding to different first feature variables in ascending order of weight; deleting the first feature variables corresponding to the first preset number of first weights in the sorting result; and determining the remaining first feature variables as the second feature variables.

[0139] In one embodiment, when a computer program is executed by a processor, it implements a Cartesian product-based algorithm to calculate a second weight of a target combination of preset groups of different second feature variables according to the preset grouping of each second feature variable, including: combining the preset groups of different second feature variables according to the Cartesian product algorithm to obtain the target combination; and calculating the second weight according to the number of enterprise data corresponding to the target combination.

[0140] In one embodiment, the process of determining the enterprise type based on a preset weight range and the second weight when the computer program is executed by a processor includes: calculating a fourth weight corresponding to the enterprise based on the second weight of all target combinations corresponding to the enterprise; and determining the enterprise type based on the fourth weight and the preset weight range.

[0141] In one embodiment, before the computer program, when executed by a processor, calculates the first weight of each first feature variable based on a preset grouping and feature variable judgment conditions, the method further includes: placing all enterprise data corresponding to each first feature variable in a one-dimensional coordinate system; calculating the local reachability density of each enterprise data corresponding to each first feature variable in the one-dimensional coordinate system based on the distribution of all enterprise data corresponding to each first feature variable in the one-dimensional coordinate system; determining the outlier factor of each enterprise data corresponding to each first feature variable based on the local reachability density; determining the magnitude of the outlier factor and a preset value; if the outlier factor is greater than the preset value, deleting the enterprise data corresponding to the outlier factor; and after deleting the enterprise data corresponding to the outlier factor, performing the step of calculating the first weight of each first feature variable based on a preset grouping and feature variable judgment conditions.

[0142] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data shall comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0143] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0144] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0145] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for determining enterprise type, characterized in that, The method includes: Acquire enterprise data and determine multiple first feature variables corresponding to the enterprise data, wherein the enterprise data includes basic enterprise information, enterprise technical capability data, and enterprise scientific and technological achievement data; Calculate the first weight of each first feature variable based on the preset grouping and feature variable judgment conditions for each first feature variable; The step of calculating the first weight of each first feature variable based on the preset grouping and feature variable judgment conditions includes: The evidence weight (WOE) value of each preset group for each first feature variable is calculated based on the judgment conditions of the feature variable, and the information value (IV) value of each preset group is calculated based on the WOE value of each preset group. The IV value of each preset group is the third weight of each preset group. The sum of the third weights of all preset groups for each first feature variable is determined as the first weight of the first feature variable. The second feature variable is obtained by filtering the first feature variable according to the first weight; Based on the Cartesian product algorithm, according to the preset grouping of each second feature variable, the second weight of the target combination of the preset grouping of different second feature variables is calculated; wherein, the target combination refers to the combination of different preset groups of different feature variables. The proportion of enterprises adopting each R&D model among enterprises of different types is determined based on the second weight of the target combination; the R&D capabilities of enterprises of different types are determined based on the determined proportion. The enterprise type is determined based on the preset weight range and the second weight.

2. The method according to claim 1, characterized in that, The step of calculating the first weight of each first feature variable based on the preset grouping and feature variable judgment conditions includes: The third weight of each preset group is determined based on the judgment conditions of the aforementioned feature variables; Calculate the first weight of each first feature variable based on all the third weights corresponding to each first feature variable.

3. The method according to claim 1, characterized in that, The step of filtering the first feature variable according to the first weight to obtain the second feature variable includes: The first weights corresponding to different first feature variables are sorted in ascending order of weight; Delete the first feature variables corresponding to the first weight of the first number of sorting results; The remaining first characteristic variable is determined as the second characteristic variable.

4. The method according to claim 1, characterized in that, The Cartesian product algorithm, based on the preset grouping of each second feature variable, calculates the second weight of the target combination of preset groups of different second feature variables, including: The target combination is obtained by combining the preset groups of different second feature variables according to the Cartesian product algorithm. The second weight is calculated based on the number of enterprise data corresponding to the target combination.

5. The method according to claim 1, characterized in that, The step of determining the enterprise type based on the preset weight range and the second weight includes: Calculate the fourth weight corresponding to the enterprise based on the second weight of all target combinations corresponding to the enterprise; The enterprise type is determined based on the fourth weight and the preset weight range.

6. The method according to claim 1, characterized in that, Before calculating the first weight of each first feature variable based on the preset grouping and feature variable judgment conditions for each first feature variable, the method further includes: Place all enterprise data corresponding to each first characteristic variable in a one-dimensional coordinate system; Based on the distribution of all enterprise data corresponding to each first feature variable in the one-dimensional coordinate system, calculate the local reachability density of each enterprise data corresponding to each first feature variable in the one-dimensional coordinate system. The outlier factor for each enterprise data corresponding to each first feature variable is determined based on the local reachability density. Determine the magnitude of the outlier factor and the preset value. If the outlier factor is greater than the preset value, delete the enterprise data corresponding to the outlier factor. After deleting the enterprise data corresponding to the outlier factor, the step of calculating the first weight of each first feature variable based on the preset grouping and feature variable judgment conditions of each first feature variable is performed.

7. An apparatus for determining enterprise type, characterized in that, The device includes: The acquisition module is used to acquire enterprise data and determine multiple first feature variables corresponding to the enterprise data. The enterprise data includes basic enterprise information, enterprise technical capability data, and enterprise scientific and technological achievement data. The first calculation module is used to calculate the first weight of each first feature variable according to the preset grouping and feature variable judgment conditions of each first feature variable; The first calculation module is further configured to calculate the evidence weight (WOE) value of each preset group for each first feature variable according to the feature variable judgment condition, and calculate the information value (IV) value of each preset group according to the WOE value of each preset group, wherein the IV value of each preset group is the third weight of each preset group; and determine the first weight of the first feature variable by summing the third weights of all preset groups for each first feature variable. The filtering module is used to filter the first feature variable according to the first weight to obtain the second feature variable; The second calculation module is used to calculate the second weight of the target combination of the preset groups of different second feature variables based on the Cartesian product algorithm and according to the preset grouping of each second feature variable; wherein, the target combination refers to the combination of different preset groups of different feature variables. The second calculation module is further configured to determine the proportion of enterprises adopting each R&D model among enterprises of different enterprise types based on the second weight of the target combination; and to determine the R&D capabilities of enterprises of different enterprise types based on the determined proportion. The determination module is used to determine the enterprise type based on a preset weight range and the second weight.

8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.