A method and system for enterprise evaluation based on big data analysis
Through the enterprise evaluation method based on big data analysis, the division of the enterprise data is divided into components and the processing strategy is adjusted, and the problem of inefficient data analysis in the existing technology is solved, achieving more efficient and accurate enterprise evaluation.
Patent Information
- Application Number
- CN202510342281.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-21
- Publication Date
- 2025-08-26
- Estimated Expiration
- 2045-03-21
AI Technical Summary
In the prior art, when obtaining enterprise-related information through network crawlers, the targeted subset analysis process cannot be determined based on the information processing situation at different stages in the actual information acquisition process, resulting in inefficient data analysis in the enterprise evaluation process.
The enterprise evaluation method based on big data analysis is adopted, and the data of the target evaluation enterprise is obtained, divided into different target evaluation subsets, and the subset processing strategy is determined based on the sharded execution status of each subset. The first confidence analysis method or the second confidence analysis method is used for analysis, and the subset task scheduling and concurrent requests are adjusted to improve the data analysis efficiency.
It improves the data analysis efficiency and the effectiveness of information acquisition in the enterprise evaluation process, ensures that the analysis process complies with actual information, and improves the accuracy and efficiency of information acquisition.
Smart Images

Figure CN119849993B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of enterprise evaluation, and in particular to an enterprise evaluation method and system based on big data analysis. Background Art
[0002] The enterprise evaluation process requires obtaining a large amount of enterprise-related information. Acquiring enterprise-related information through crawler technology can effectively reduce human resource consumption and time consumption in the information acquisition stage. Based on the determined key content and crawler technology, information mining is continuously performed. However, the mined information often contains invalid data for enterprise evaluation, which can easily lead to poor data analysis efficiency in the enterprise evaluation process. Therefore, how to conduct targeted analysis based on the task execution status of different subsets in the actual data acquisition process to ensure the data analysis efficiency of the subsequent enterprise evaluation process is an urgent problem to be solved by technical personnel in this field.
[0003] Chinese patent publication number CN113239253A discloses a web crawler implementation method suitable for execution in a web crawler implementation system, the method comprising the following steps: a message queue server receives a web crawler task from a user; the message queue server establishes multiple message queues based on multiple query keywords of the web crawler task; a web page collection server crawls multiple web pages related to the query keyword from the Internet based on a query keyword, and sends the URLs of the multiple web pages to the message queue server; the message queue server receives the URLs of the multiple web pages and stores the URLs of the multiple web pages in a message queue corresponding to the target query keyword; a content crawler obtains the URLs of the multiple web pages related to the query keyword from a message queue, downloads the multiple web pages based on the URL of each web page to obtain multiple web pages, and crawls information from the multiple web pages to generate web crawler results. However, the above scheme has the following problems: it fails to determine a targeted subset analysis process based on the information processing conditions at different stages of the actual information acquisition process, resulting in low effectiveness of the acquired enterprise-related information, which in turn leads to low data analysis efficiency in the subsequent enterprise evaluation process. Summary of the Invention
[0004] To this end, the present invention provides an enterprise evaluation method and system based on big data analysis, which is used to overcome the problem in the prior art that a targeted subset analysis process is determined according to the information processing conditions at different stages in the actual information acquisition process, resulting in low effectiveness of the acquired enterprise-related information, and further leading to low data analysis efficiency in the subsequent enterprise evaluation process.
[0005] To achieve the above objectives, the present invention provides an enterprise evaluation method based on big data analysis, comprising:
[0006] Obtain target enterprise data of target assessment enterprises;
[0007] Dividing the target enterprise data of the target assessment enterprise to obtain different target assessment subsets;
[0008] Determine a subset processing strategy based on the shard execution status of each target evaluation subset, wherein the subset processing strategy is to analyze the target evaluation subset using the first confidence analysis method or the second confidence analysis method;
[0009] In the first confidence analysis method, the shard quality parameters of a type of evaluation subset are determined based on the subset execution efficiency and the reference correlation index. The efficiency adjustment method is determined based on the shard quality parameters, and the confidence reference coefficient of each relevant analysis text is determined. The efficiency adjustment method is to perform subset task scheduling for the type of evaluation subset, or to adjust the number of concurrent requests based on the subset execution efficiency.
[0010] In the second confidence analysis method, the subset execution coefficient of the second type of evaluation subset is determined according to the reference correlation index and the topic correlation ratio, and the set setting method and the effective evaluation strategy are determined according to the subset execution coefficient. The effective evaluation strategy is to determine the confidence reference coefficient based on the reference correlation matching coefficient and the key evaluation parameters, or to determine the confidence reference coefficient based on the access matching coefficient and the reference influence coefficient.
[0011] Under the conditions of subset analysis, enterprise evaluation is carried out for the target evaluation enterprises. In the enterprise evaluation, the final evaluation score of each target evaluation data is determined based on the data value of the target evaluation data and the evaluation weight coefficient corresponding to the relevant analysis text.
[0012] Furthermore, for a single target evaluation subset, if the target evaluation subset is in a shard execution state where the extended effective coefficient is less than or equal to a preset extended effective coefficient or the execution complete coefficient is less than or equal to a preset execution complete coefficient, the target evaluation subset is analyzed using a first confidence analysis method, including:
[0013] Determine a sharding quality parameter of a class of evaluation subsets according to subset execution efficiency and reference correlation index;
[0014] Determine whether to perform subset task scheduling for a type of evaluation subset according to the shard quality parameter;
[0015] Under the subset confirmation condition, the confidence reference coefficient of each relevant analysis text is determined based on the topic enrichment parameter;
[0016] The subset confirmation condition is that a type of evaluation subset does not need to be scheduled for a subset task or has completed the subset task;
[0017] The first type of evaluation subset is a target evaluation subset in a first type of shard execution state.
[0018] Furthermore, for a single type of evaluation subset, if the slice quality parameter of the type of evaluation subset is less than the preset slice quality parameter, subset task scheduling is performed for the type of evaluation subset;
[0019] Determine the URLs to be segmented for this evaluation subset based on the extended correlation coefficient and the topic-independent parameter;
[0020] The target allocation subset is determined based on the key words corresponding to the URL to be split and its extended URL and the subset processing efficiency.
[0021] Furthermore, for a single first-class evaluation subset, if the shard quality parameter of the first-class evaluation subset is greater than the preset shard quality parameter, the number of concurrent requests is increased and adjusted according to the subset execution efficiency;
[0022] Determine whether to perform secondary adjustment on the number of concurrent requests based on the efficiency optimization parameters;
[0023] The increase in the number of concurrent requests is negatively correlated with the subset execution efficiency.
[0024] Furthermore, for a single target evaluation subset, if the target evaluation subset is in a second-class shard execution state where the extension validity coefficient is greater than the preset extension validity coefficient and the execution completeness coefficient is greater than the preset execution completeness coefficient, the target evaluation subset is analyzed using a second confidence analysis method, including:
[0025] Determine the subset execution coefficient of the second-category evaluation subset based on the reference relevance index and the proportion of subject relevance;
[0026] Determine the set setting method and effective evaluation strategy based on the subset execution coefficient;
[0027] The set setting method is to determine the relevant matching set according to the subject association parameter and the set association index or to determine the access matching set according to the execution correlation coefficient and the access correlation coefficient;
[0028] The second-category evaluation subset is a target evaluation subset in a second-category shard execution state.
[0029] Furthermore, for a single second-category evaluation subset, if the subset execution coefficient of the second-category evaluation subset is greater than the preset subset execution coefficient, a relevant matching set is determined according to the topic association parameter and the set association index, and a confidence reference coefficient of the relevant analysis text obtained by the second-category evaluation subset in the current subset analysis cycle is determined based on the reference relevant matching coefficient of each evaluation matching set and the key evaluation parameter;
[0030] The correlation matching coefficient of each key word in any correlation matching set is greater than the preset correlation matching coefficient.
[0031] Furthermore, for a single second-category evaluation subset, if the subset execution coefficient of the second-category evaluation subset is less than or equal to the preset subset execution coefficient, an access matching set is determined based on the execution correlation coefficient and the access correlation coefficient, and a confidence reference coefficient of the relevant analysis text obtained by the second-category evaluation subset in the current subset analysis cycle is determined based on the access matching coefficient and the corresponding reference influence coefficient of each access matching set;
[0032] The access matching coefficient in any access matching set is greater than the preset access matching coefficient.
[0033] Furthermore, under the subset analysis condition, the information acquisition quality of each target evaluation set is determined based on the effective reference ratio and information overlap parameters;
[0034] The information acquisition quality is positively correlated with the effective reference ratio and the information overlap parameter respectively;
[0035] The subset analysis condition is that each target evaluation subset completes the task effective analysis or performs effective analysis.
[0036] Further, the final evaluation score of each target evaluation data of the target evaluation enterprise is determined based on each data value corresponding to the target evaluation data in the obtained relevant analysis text and the evaluation weight coefficient corresponding to the relevant analysis text;
[0037] For a single relevant analysis text, its corresponding evaluation weight coefficient is positively correlated with the information acquisition quality of the target evaluation subset to which the relevant analysis text belongs.
[0038] The present invention also provides a system for applying the enterprise evaluation based on big data analysis, comprising:
[0039] A subset evaluation module for determining a subset processing strategy by expanding effective coefficients and executing complete coefficients in response to subset evaluation conditions;
[0040] a first confidence analysis module connected to the subset evaluation module, configured to respond to the shard quality parameter of the adjustment judgment condition to determine an efficiency adjustment method and complete the determination of the confidence reference coefficient of each relevant analysis text, wherein the efficiency adjustment method is to perform subset task scheduling for a class of evaluation subsets, or to adjust the number of concurrent requests based on the subset execution efficiency;
[0041] a second confidence analysis module connected to the subset evaluation module, for responding to the subset execution coefficient of the set judgment condition to determine a set setting mode and an effective evaluation strategy, wherein the effective evaluation strategy is to determine a confidence reference coefficient based on a reference related matching coefficient and a key evaluation parameter, or to determine a confidence reference coefficient based on an access matching coefficient and a reference influence coefficient;
[0042] An evaluation execution module, which is connected to the first confidence analysis module and the second confidence analysis module, is used to determine the final evaluation score of each target evaluation data based on the data value of the target evaluation data and the evaluation weight coefficient corresponding to the relevant analysis text.
[0043] Compared with the prior art, the beneficial effect of the present invention lies in that the technical solution of the present invention periodically judges the shard execution status of each target evaluation subset based on the extended effective coefficient and the execution complete coefficient, and determines the subset processing strategy based on the shard execution status, so that the subsequent analysis process of the subset retrieval results is more in line with the actual information situation. While ensuring the accuracy of the judgment on the effectiveness of the obtained information, the processing efficiency of the judgment process is guaranteed. The present invention improves the data analysis efficiency of the subsequent enterprise evaluation process.
[0044] Furthermore, the present invention determines the fragmented execution status of each target evaluation subset based on the extended effectiveness coefficient and the execution completeness coefficient, which is used to characterize the execution status of each target evaluation subset for the assigned information capture task, and also characterizes the data richness of the acquired web page information, so that the subsequent targeted subset processing strategy has a stronger match. The present invention improves the accuracy of the effectiveness evaluation of the acquired web page information.
[0045] Furthermore, the present invention determines an optimization method for targeted subset allocation based on fragmentation quality parameters for target evaluation subsets with poor execution of the assigned information capture tasks. If the fragmentation quality parameters of a type of evaluation subset are small, it means that the correlation and mutual compensation of the URLs constituting the current subset are poor, and targeted URL scheduling is performed for this. For a type of evaluation subset with a larger fragmentation quality parameter, the subset execution efficiency is monitored in real time to optimize the number of concurrent requests and improve data acquisition efficiency. The present invention takes into account the efficiency of the information acquisition process while ensuring the accuracy of the effectiveness evaluation of the acquired web page information.
[0046] Furthermore, the present invention aims at the target evaluation subsets that have good performance in executing the assigned information capture tasks and have obtained good quality web page information, and determines targeted set setting methods and effective evaluation strategies for the subject keywords based on the subset execution coefficients of such target evaluation subsets. For key subject words with poor relevance, the sets are divided according to the access situation, so that the setting methods of different matching sets are more in line with the actual information situation. The present invention improves the accuracy of the effectiveness evaluation of the obtained web page information. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 A schematic diagram of the enterprise evaluation method based on big data analysis of the present invention;
[0048] Figure 2 A flowchart of a subset processing strategy for determining a target evaluation subset based on a slice execution state according to the present invention;
[0049] Figure 3 A flow chart of a method for determining a set setting according to a subset execution coefficient according to the present invention;
[0050] Figure 4 This is a module connection diagram of the enterprise evaluation system based on big data analysis of the present invention. DETAILED DESCRIPTION
[0051] In order to make the objects and advantages of the present invention more clearly understood, the present invention is further described below in conjunction with embodiments; it should be understood that the specific embodiments described herein are merely used to explain the present invention and are not intended to limit the present invention.
[0052] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the scope of protection of the present invention.
[0053] It should be noted that, in the description of the present invention, terms such as "up", "down", "left", "right", "inside", and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or element must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present invention.
[0054] Furthermore, it should be noted that, in the description of the present invention, unless otherwise expressly specified or limited, the terms "mounted," "connected," and "connected" should be understood in a broad sense. For example, they may refer to fixed connections, detachable connections, or integral connections; mechanical connections or electrical connections; direct connections or indirect connections through an intermediate medium; and internal communication between two components. Those skilled in the art will understand the specific meanings of the above terms in the present invention based on specific circumstances.
[0055] See also Figures 1 to 3 As shown, the present invention provides an enterprise evaluation method based on big data analysis, comprising:
[0056] Obtain target enterprise data of target assessment enterprises;
[0057] Dividing the target enterprise data of the target assessment enterprise to obtain different target assessment subsets;
[0058] Determine a subset processing strategy based on the shard execution status of each target evaluation subset, wherein the subset processing strategy is to analyze the target evaluation subset using the first confidence analysis method or the second confidence analysis method;
[0059] In the first confidence analysis method, the shard quality parameters of a type of evaluation subset are determined based on the subset execution efficiency and the reference correlation index. The efficiency adjustment method is determined based on the shard quality parameters, and the confidence reference coefficient of each relevant analysis text is determined. The efficiency adjustment method is to perform subset task scheduling for the type of evaluation subset, or to adjust the number of concurrent requests based on the subset execution efficiency.
[0060] In the second confidence analysis method, the subset execution coefficient of the second type of evaluation subset is determined according to the reference correlation index and the topic correlation ratio, and the set setting method and the effective evaluation strategy are determined according to the subset execution coefficient. The effective evaluation strategy is to determine the confidence reference coefficient based on the reference correlation matching coefficient and the key evaluation parameters, or to determine the confidence reference coefficient based on the access matching coefficient and the reference influence coefficient.
[0061] Under the conditions of subset analysis, enterprise evaluation is carried out for the target evaluation enterprises. In the enterprise evaluation, the final evaluation score of each target evaluation data is determined based on the data value of the target evaluation data and the evaluation weight coefficient corresponding to the relevant analysis text.
[0062] Among them, the present invention is applied to the process of enterprise evaluation for target evaluation enterprises. The target evaluation enterprise is an enterprise that needs to be evaluated. The target enterprise data is a number of initial URLs. For a single target evaluation enterprise, the present invention determines a number of initial URLs based on the evaluation direction of the target evaluation enterprise and the business areas involved, and divides the initial URLs of the target evaluation enterprise to obtain a number of target evaluation subsets. In each subset analysis cycle, the processing node corresponding to the target evaluation subset can determine a new URL based on the initial URL in the target evaluation subset and the obtained text content, and record the newly determined URL as an extended URL in the corresponding target evaluation subset. Any target evaluation subset is a collection of a number of initial URLs and extended URLs. Before the new URL is determined, there are only initial URLs in the target evaluation subset. The initial URL and the extended URL contained in each target evaluation subset are recorded as the relevant URLs of the corresponding target evaluation subset. How to determine the initial URL based on the evaluation direction of the target evaluation enterprise and the business areas involved, and how to determine the new URL based on the initial URL and the text content obtained each time is easy for technical personnel in this field to understand and will not be elaborated here. The subject keywords contained in the relevant analysis text in the present invention can be set by the user according to the actual evaluation enterprise. For example, the main business of the target evaluation enterprise is green petrochemicals and new materials. The subject keywords include but are not limited to: anti-corrosion coating preparation, ink production, natural gas, coating stirring, printing ink, petroleum products, mixing devices, water-based coatings, ink preparation, coating processing, coating production, preparation process, coating combination, waterproof coating preparation and positive electrode materials.
[0063] The present invention applies a cyclic subset analysis cycle. The duration of the subset analysis cycle can be determined by the user. The higher the user's requirements for data analysis efficiency in the enterprise evaluation process, the shorter the subset analysis cycle. A subset analysis cycle duration is provided, which is 5 minutes. At the end of each subset analysis cycle, the shard execution status of each target evaluation subset is determined based on the expansion effectiveness coefficient and the execution completeness coefficient, and the subset processing strategy is determined based on the shard execution status.
[0064] The present invention uses several historical evaluation records, and any historical evaluation record records at least one extended effectiveness coefficient, execution completeness coefficient, sharding quality parameter, segmentation priority coefficient, correlation index, correlation matching coefficient, access matching coefficient and confidence reference coefficient in the enterprise evaluation process for the target evaluation enterprise, and each historical evaluation record corresponds to a qualified mark, which records whether the data analysis efficiency in the enterprise evaluation process meets user requirements. It can be understood that users can determine whether the data analysis efficiency in the enterprise evaluation process meets the requirements based on self-set indicators. For example, self-set indicators can be but are not limited to evaluation analysis efficiency. The evaluation analysis efficiency is the time spent determining the enterprise evaluation results after completing the acquisition of enterprise information.
[0065] Specifically, for a single target evaluation subset, if the target evaluation subset is in a shard execution state where the extension validity coefficient is less than or equal to the preset extension validity coefficient or the execution completion coefficient is less than or equal to the preset execution completion coefficient, the first confidence analysis method is used to analyze the target evaluation subset, including:
[0066] Determine a sharding quality parameter of a class of evaluation subsets according to subset execution efficiency and reference correlation index;
[0067] Determine whether to perform subset task scheduling for a type of evaluation subset according to the shard quality parameter;
[0068] Under the subset confirmation condition, the confidence reference coefficient of each relevant analysis text is determined based on the topic enrichment parameter;
[0069] The subset confirmation condition is that a type of evaluation subset does not need to be scheduled for a subset task or has completed the subset task;
[0070] The first type of evaluation subset is a target evaluation subset in a first type of shard execution state.
[0071] Among them, for a single target evaluation subset, the extension effectiveness coefficient = ln (subset extension parameter × extension effectiveness parameter), the subset extension parameter is the number of extended URLs obtained in the current subset analysis cycle, the extension effectiveness parameter is the average number of key subject words of each extended URL obtained in the current subset analysis cycle, the execution completion coefficient = (the number of extended URLs obtained in the previous subset analysis cycle - the number of extended URLs obtained in the current subset analysis cycle) / the number of extended URLs obtained in the previous subset analysis cycle, the text content corresponding to the web page information determined based on each relevant URL contained in the target evaluation subset is recorded as the relevant analysis text, key subject words are extracted for each relevant information text, and the correlation index between the key subject words of the two types of evaluation subsets is detected. For any two key subject words, the correlation index is the number of times the above two key subject words exist in the same trusted reference text at the same time in each historical evaluation record, and the reference correlation index , n is the number of key words in the second category evaluation subset, is the correlation index between the key subject words in the i-th group. For a single related analysis text, the confidence reference coefficient is positively correlated with the topic enrichment parameter, which is the number of different key subject words in the related analysis text.
[0072] The values of the preset extension effectiveness coefficient and the preset execution complete coefficient can be determined by the user according to the actual work scenario. For example, the user can set them according to historical evaluation records. The higher the user's requirements for the data analysis efficiency in the enterprise evaluation process, the smaller the value of the preset extension effectiveness coefficient and the smaller the value of the preset execution complete coefficient. A method for determining the value of the preset extension effectiveness coefficient is provided, and the average value of the extension effectiveness coefficients of the target evaluation subset in a type of shard execution state in the historical evaluation records that meet the user's requirements for the data analysis efficiency in the enterprise evaluation process is recorded as the preset extension effectiveness coefficient. A method for determining the value of the preset execution complete coefficient is provided, and the average value of the execution complete coefficients of the target evaluation subset in a type of shard execution state in the historical evaluation records that meet the user's requirements for the data analysis efficiency in the enterprise evaluation process is recorded as the preset execution complete coefficient.
[0073] Specifically, for a single type of evaluation subset, if the slice quality parameter of the type of evaluation subset is less than or equal to the preset slice quality parameter, subset task scheduling is performed for the type of evaluation subset;
[0074] Determine the URLs to be segmented for the evaluation subset based on the extended correlation coefficient and the topic-independent parameter;
[0075] The target allocation subset is determined based on the key words corresponding to the URL to be split and its extended URL and the subset processing efficiency.
[0076] Among them, for a single first-class evaluation subset, the shard quality parameter is the product of the subset execution efficiency and the reference correlation index, and the subset execution efficiency is the product of the extension effectiveness coefficient and the execution completeness coefficient. The value of the preset shard quality parameter can be determined by the user according to the actual work scenario. For example, the user can set it according to the historical evaluation records. The higher the user's requirements for the data analysis efficiency in the enterprise evaluation process, the smaller the value of the preset shard quality parameter. A method for setting the value of the preset shard quality parameter is provided, and the historical evaluation records for subset task scheduling for a class of evaluation subsets are recorded as quality reference records, and the average value of the shard quality parameters of each class of evaluation subsets in the quality reference records that meet the user's requirements for the data analysis efficiency in the enterprise evaluation process is recorded as the preset shard quality parameter;
[0077] If any evaluation subset needs to be scheduled for subset tasks, for a single related URL in the evaluation subset, the segmentation priority coefficient = ln (topic independent parameter / extended correlation coefficient), the extended correlation coefficient is the number of extended URLs determined by the related URL, the topic independent parameter = 1 / the number of composite topics of the related URL, for the key topic word corresponding to a single related URL, if the key topic word exists in other related URLs in the evaluation subset, then the key topic word is determined to be the composite topic of the related URL, and if the segmentation priority coefficient of the related URL is greater than the preset segmentation priority coefficient, then the related URL is scheduled to other target evaluation subsets, and The relevant URL is recorded as the URL to be split. For a single URL to be split, the allocation priority coefficient of each target evaluation subset is determined according to the key subject words corresponding to the URL to be split and its extended URL and the subset processing efficiency. For a single target evaluation subset, the allocation priority coefficient = ln (topic relevance index × subset processing efficiency of the target evaluation subset), the topic relevance index = the number of key subject words that coexist in the relevant analysis text corresponding to the URL to be split and its extended URL and the target evaluation subset / the number of key subject words that exist in the relevant analysis text corresponding to the URL to be split and its extended URL. The target evaluation subset with the largest allocation priority coefficient is recorded as the target allocation subset;
[0078] The value of the preset segmentation priority coefficient can be determined by the user according to the actual work scenario. For example, the user can set it according to historical evaluation records. The higher the user's requirements for the data analysis efficiency in the enterprise evaluation process, the smaller the value of the preset segmentation priority coefficient. A method for determining the value of the preset segmentation priority coefficient is provided, and the historical evaluation records for subset task scheduling for a type of evaluation subset are recorded as quality reference records. The average value of the segmentation priority coefficients of each to-be-segmented URL in the quality reference records that meet the user's requirements for data analysis efficiency in the enterprise evaluation process is recorded as the preset segmentation priority coefficient.
[0079] Specifically, for a single first-class evaluation subset, if the shard quality parameter of the first-class evaluation subset is greater than the preset shard quality parameter, the number of concurrent requests is increased and adjusted based on the subset execution efficiency;
[0080] Determine whether to perform secondary adjustment on the number of concurrent requests based on the efficiency optimization parameters;
[0081] The increase in the number of concurrent requests is negatively correlated with the subset execution efficiency.
[0082] Among them, for a type of evaluation subset for which the number of concurrent requests is adjusted, the efficiency optimization parameter is the difference between the subset execution efficiency of the next subset evaluation cycle and the subset execution efficiency of the current subset analysis cycle. If the efficiency optimization parameter is less than the preset efficiency optimization parameter, the number of concurrent requests is reduced and adjusted. The reduction value of the concurrent request number is negatively correlated with the efficiency optimization parameter. How to adjust the number of concurrent requests is easy to understand for those skilled in the art and will not be elaborated here.
[0083] Specifically, for a single target evaluation subset, if the target evaluation subset is in a second-class shard execution state where the expansion validity coefficient is greater than the preset expansion validity coefficient and the execution completeness coefficient is greater than the preset execution completeness coefficient, the second confidence analysis method is used to analyze the target evaluation subset, including:
[0084] Determine the subset execution coefficient of the second-category evaluation subset based on the reference relevance index and the proportion of subject relevance;
[0085] Determine the set setting method and effective evaluation strategy based on the subset execution coefficient;
[0086] The set setting method is to determine the relevant matching set according to the subject association parameter and the set association index or to determine the access matching set according to the execution correlation coefficient and the access correlation coefficient;
[0087] The second-category evaluation subset is a target evaluation subset in a second-category shard execution state.
[0088] Among them, for a single second-category evaluation subset, the subset execution coefficient = ln (reference correlation index × topic correlation ratio), topic correlation ratio = the number of key topic words with associated topic words / the number of key topic words in the second-category evaluation subset. For any two key topic words, if the correlation index between the two key topic words is greater than the preset correlation index, it is determined that there is a correlation relationship between the above set of key topic words, and each is recorded as the other's associated topic words;
[0089] The value of the preset association index can be determined by the user according to the actual work scenario. For example, the user can set it according to historical evaluation records. The higher the user's requirements for the data analysis efficiency in the enterprise evaluation process, the larger the value of the preset association index. A method for determining the value of the preset association index is provided, and the average value of the association index of the associated keywords in the historical evaluation records that meet the user's requirements for the data analysis efficiency in the enterprise evaluation process is recorded as the preset association index.
[0090] Specifically, for a single second-category evaluation subset, if the subset execution coefficient of the second-category evaluation subset is greater than the preset subset execution coefficient, the relevant matching set is determined according to the topic association parameter and the set association index, and the confidence reference coefficient of the relevant analysis text obtained by the second-category evaluation subset in the current subset analysis cycle is determined based on the reference relevant matching coefficient of each evaluation matching set and the key evaluation parameter;
[0091] The correlation matching coefficient of each key word in any correlation matching set is greater than the preset correlation matching coefficient.
[0092] Among them, for a single second-category evaluation subset, if a relevant matching set is determined based on a topic association parameter and a set association index, for any key subject word in the second-category evaluation subset, the relevant matching coefficient is the product of the topic association parameter and the set association index, wherein the topic association parameter is the number of associated subject words of the key subject word in the second-category evaluation subset, and the set association index is the average value of the association index between the key subject word and other key subject words in the second-category evaluation subset;
[0093] The values of the preset correlation matching coefficient and the preset subset execution coefficient can be determined by the user according to the actual work scenario. For example, the user can set them according to historical evaluation records. The higher the user's requirements for data analysis efficiency in the enterprise evaluation process, the larger the value of the preset correlation matching coefficient. A method for determining the value of the preset correlation matching coefficient is provided. The historical evaluation records of the relevant matching set determined according to the subject association parameter and the set association index are recorded as relevant matching reference records. The minimum value of the correlation matching coefficient of each key subject word in the relevant matching set in the relevant matching reference records that meet the user's requirements for data analysis efficiency in the enterprise evaluation process is recorded as the preset correlation matching coefficient. A method for determining the value of the preset subset execution coefficient is provided. The minimum value of the subset execution coefficient of the second category of evaluation subsets in the relevant matching reference records that meet the user's requirements for data analysis efficiency in the enterprise evaluation process is recorded as the preset subset execution coefficient.
[0094] The relevant analysis text is the relevant analysis text of each target evaluation subset within the current subset analysis cycle. For a single relevant analysis text, the confidence reference coefficient is the sum of the products of the reference relevant matching coefficients of each relevant matching set and the corresponding set weight coefficient. The reference relevant matching coefficient is the average value of the relevant matching coefficients of each key subject word in the relevant matching set. For a single relevant matching set, its corresponding set weight coefficient is positively correlated with the overlap evaluation parameter of the relevant matching set. The key evaluation parameter is the product of the overlap topic ratio and the overlap correlation coefficient. The overlap topic ratio = the number of key subject words in the relevant matching set that exist in the relevant analysis text / the number of key subject words in the relevant matching set. The overlap correlation coefficient is the number of combinations of key subject words that have an association relationship among the key subject words in the relevant matching set that exist in the relevant analysis text.
[0095] Specifically, for a single second-category evaluation subset, if the subset execution coefficient of the second-category evaluation subset is less than or equal to the preset subset execution coefficient, an access matching set is determined based on the execution correlation coefficient and the access correlation coefficient, and a confidence reference coefficient of the relevant analysis text obtained by the second-category evaluation subset in the current subset analysis cycle is determined based on the access matching coefficient and the corresponding reference influence coefficient of each access matching set;
[0096] The access matching coefficient in any access matching set is greater than the preset access matching coefficient.
[0097] Among them, for a single second-category evaluation subset, if the access matching set is determined based on the execution correlation coefficient and the access correlation coefficient, for any access matching set, the access matching coefficient = execution correlation coefficient / access correlation coefficient, the execution correlation coefficient is the average value of the execution-related parameters of each key word in the access matching set, for a single key word, the execution-related parameter = the number of execution-related URLs / the number of related URLs of the second-category evaluation subset, the execution-related URL is the related URL that has an execution relationship with the related URL corresponding to the key word in the second-category evaluation subset, if a related URL is determined based on the initial URL or other extended URL in the corresponding target evaluation subset, then it is judged that the related URL has an execution relationship with the initial URL or other extended URL in the corresponding target evaluation subset; the access correlation coefficient , m is the number of related URLs that have execution relationships with the related URLs corresponding to the key word, The reference access interval between the jth related URL that has an execution relationship with the related URL corresponding to the keyword and the related URL corresponding to the keyword. For any two related URLs, the reference access interval is the average duration of each user's access to the web pages corresponding to the two related URLs during a single visit.
[0098] The value of the preset access matching coefficient can be determined by the user according to the actual work scenario. For example, the user can set it according to historical evaluation records. The higher the user's requirements for the data analysis efficiency in the enterprise evaluation process, the larger the value of the preset access matching coefficient. A method for determining the value of the preset access matching coefficient is provided. The historical evaluation records of the access matching set determined according to the execution correlation coefficient and the access correlation coefficient are recorded as access matching reference records. The minimum value of the access matching coefficient of the access matching set in the access matching reference records that meets the user's requirements for the data analysis efficiency in the enterprise evaluation process is recorded as the preset access matching coefficient.
[0099] For a single relevant analysis text, the confidence reference coefficient is the sum of the products of the access matching coefficients of each access matching set in the corresponding second-category evaluation subset and the corresponding reference influence coefficients. For a single access matching set, its corresponding reference influence coefficient is positively correlated with the behavior evaluation parameter of the access matching set. The behavior evaluation parameter is the product of topic coverage and key access parameters. The topic coverage = the number of key subject words existing in the relevant analysis text / the number of key subject words in the access matching set. The key access parameter is the average value of the regional retention time of each key subject word existing in the relevant analysis text. For a single key subject word, the regional retention time is the average value of the length of time the mouse marker stays in the area where the key subject word exists when the user visits the relevant URL corresponding to the relevant analysis text.
[0100] Specifically, under the subset analysis condition, the information acquisition quality of each target evaluation set is determined based on the effective reference ratio and information overlap parameters;
[0101] The information acquisition quality is positively correlated with the effective reference ratio and the information overlap parameter respectively;
[0102] The subset analysis condition is that each target evaluation subset completes the task effective analysis or performs effective analysis.
[0103] Specifically, the final evaluation score of each target evaluation data of the target evaluation enterprise is determined based on each data value corresponding to the target evaluation data in the obtained relevant analysis text and the evaluation weight coefficient corresponding to the relevant analysis text;
[0104] For a single relevant analysis text, its corresponding evaluation weight coefficient is positively correlated with the information acquisition quality of the target evaluation subset to which the relevant analysis text belongs.
[0105] Among them, for a single target evaluation set, the information acquisition quality is the sum of the effective reference ratio and the information overlap parameter, the effective reference ratio = the number of confident reference texts / the number of related analysis texts obtained by the target evaluation set, the confident reference text is the related analysis text whose confidence reference coefficient is greater than the preset confidence reference coefficient, and the information overlap parameter is the number of key keywords whose number of related analysis texts is greater than the preset number of texts;
[0106] The user can determine the final evaluation score of each target evaluation data based on the data information in the relevant analysis text and the evaluation weight coefficient corresponding to each relevant analysis text. In addition, the user can determine the method of determining the final evaluation score according to actual needs. For example, for a single target evaluation data, the target evaluation data must remain stable, and the data value corresponding to the target evaluation data in the relevant analysis text is obtained, and the final evaluation score is obtained. , t is the number of evaluation data of this project in each relevant analysis text obtained, The data difference value corresponding to the s-th target evaluation data in the project target evaluation data existing in each relevant analysis text obtained is the data value corresponding to the s-th target evaluation data and the data reference value of the project target evaluation data. The data reference value is the average value of the data values corresponding to each project target evaluation data in each relevant analysis text obtained. is the evaluation weight coefficient corresponding to the relevant analysis text corresponding to the sth target evaluation data in the target evaluation data in each relevant analysis text;
[0107] For a single target evaluation data, if the relevant analysis texts for the target evaluation data only belong to a single target evaluation subset, then the evaluation weight coefficient corresponding to each relevant analysis text is positively correlated with the information acquisition quality of the target evaluation subset. If the relevant analysis texts for the target evaluation data belong to multiple target evaluation subsets, for the relevant analysis texts for the target evaluation data in a single target evaluation subset, the corresponding evaluation weight coefficient is positively correlated with the distribution information acquisition quality of the target evaluation subset. The distribution information acquisition quality is the product of the information acquisition quality of the target evaluation subset and the proportion of relevant texts in the target evaluation subset. The proportion of relevant texts = the number of relevant analysis texts for the target evaluation data in the target evaluation subset / the number of relevant analysis texts for the target evaluation data. How to determine the abnormal coefficients of target evaluation data of different categories based on the acquired data information and the evaluation weight coefficient is easy for those skilled in the art to understand and will not be elaborated on here.
[0108] The values of the preset confidence reference coefficient and the preset number of texts can be determined by the user according to the actual work scenario. For example, the user can set them according to historical evaluation records. The higher the user's requirements for the data analysis efficiency in the enterprise evaluation process, the larger the value of the preset confidence reference coefficient and the smaller the value of the preset number of texts. A method for determining the value of the preset confidence reference coefficient is provided, and the minimum value of the confidence reference coefficient of each confidence reference text in the historical evaluation records that meets the user's requirements for the data analysis efficiency in the enterprise evaluation process is recorded as the preset confidence reference coefficient. A value of the preset number of texts is provided, and the value of the preset number of texts is 30% of the number of relevant analysis texts in the target evaluation subset.
[0109] See also Figure 4 As shown in FIG, which is a module connection diagram of the enterprise evaluation system based on big data analysis of the present invention, the present invention provides an enterprise evaluation system based on big data analysis, including:
[0110] A subset evaluation module for determining a subset processing strategy by expanding effective coefficients and executing complete coefficients in response to subset evaluation conditions;
[0111] a first confidence analysis module connected to the subset evaluation module, configured to respond to the shard quality parameter of the adjustment judgment condition to determine an efficiency adjustment method and complete the determination of the confidence reference coefficient of each relevant analysis text, wherein the efficiency adjustment method is to perform subset task scheduling for a class of evaluation subsets, or to adjust the number of concurrent requests based on the subset execution efficiency;
[0112] a second confidence analysis module connected to the subset evaluation module, for responding to the subset execution coefficient of the set judgment condition to determine a set setting mode and an effective evaluation strategy, wherein the effective evaluation strategy is to determine a confidence reference coefficient based on a reference related matching coefficient and a key evaluation parameter, or to determine a confidence reference coefficient based on an access matching coefficient and a reference influence coefficient;
[0113] An evaluation execution module, which is connected to the first confidence analysis module and the second confidence analysis module, is used to determine the final evaluation score of each target evaluation data based on the data value of the target evaluation data and the evaluation weight coefficient corresponding to the relevant analysis text.
[0114] Thus far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art may make equivalent changes or substitutions to the relevant technical features, and the technical solutions after such changes or substitutions will fall within the scope of protection of the present invention.
[0115] The foregoing description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Those skilled in the art will readily appreciate that the present invention is susceptible to various modifications and variations. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present invention are intended to be within the scope of protection of the present invention.
Claims
1. A method for enterprise evaluation based on big data analysis, characterized in that: include: Obtain target enterprise data of target assessment enterprises; Dividing the target enterprise data of the target assessment enterprise to obtain different target assessment subsets; Determine a subset processing strategy based on the shard execution status of each target evaluation subset, wherein the subset processing strategy is to analyze the target evaluation subset using the first confidence analysis method or the second confidence analysis method; In the first confidence analysis method, the shard quality parameters of a type of evaluation subset are determined based on the subset execution efficiency and the reference correlation index. The efficiency adjustment method is determined based on the shard quality parameters, and the confidence reference coefficient of each relevant analysis text is determined. The efficiency adjustment method is to perform subset task scheduling for the type of evaluation subset, or to adjust the number of concurrent requests based on the subset execution efficiency. In the second confidence analysis method, the subset execution coefficient of the second type of evaluation subset is determined according to the reference correlation index and the topic correlation ratio, and the set setting method and the effective evaluation strategy are determined according to the subset execution coefficient. The effective evaluation strategy is to determine the confidence reference coefficient based on the reference correlation matching coefficient and the key evaluation parameters, or to determine the confidence reference coefficient based on the access matching coefficient and the reference influence coefficient. Under the subset analysis condition, enterprise evaluation is conducted on the target evaluation enterprise. In the enterprise evaluation, the final evaluation score of each target evaluation data is determined based on the data value of the target evaluation data and the evaluation weight coefficient corresponding to the relevant analysis text; The subset execution efficiency is the product of the extension effectiveness coefficient and the execution completeness coefficient. For a single target evaluation subset, the extension effectiveness coefficient = ln (subset extension parameter × extension effectiveness parameter), where the subset extension parameter is the number of extended URLs obtained in the current subset analysis cycle, the extension effectiveness parameter is the average number of key subject words of each extended URL obtained in the current subset analysis cycle, and the execution completeness coefficient = (the number of extended URLs obtained in the previous subset analysis cycle - the number of extended URLs obtained in the current subset analysis cycle) / the number of extended URLs obtained in the previous subset analysis cycle; For a single second-category evaluation subset, the subset execution coefficient = ln (reference relevance index × topic relevance ratio), topic relevance ratio = the number of key topic words with associated topic words / the number of key topic words in the second-category evaluation subset, reference relevance index , n is the number of key words in the second category evaluation subset, n is greater than or equal to 2, is the correlation index between the key words in group i; For any key subject word in a single second-category evaluation subset, the correlation matching coefficient is the product of the subject association parameter and the set association index, wherein the subject association parameter is the number of associated subject words of the key subject word in the second-category evaluation subset, the set association index is the average of the association indexes between the key subject word and other key subject words in the second-category evaluation subset, and the reference correlation matching coefficient is the average of the correlation matching coefficients of the key subject words in the correlation matching set; For a single relevant matching set, the key evaluation parameter is the product of the overlapping topic ratio and the overlapping correlation coefficient, where the overlapping topic ratio = the number of key subject terms in the relevant matching set that exist in the relevant analysis text / the number of key subject terms in the relevant matching set. The overlapping correlation coefficient is the number of key subject terms in the relevant matching set that have an associated relationship among the key subject terms in the relevant analysis text. For any access matching set, the access matching coefficient = execution correlation coefficient / access correlation coefficient, the execution correlation coefficient is the average value of the execution correlation parameters of each key word in the access matching set, the access correlation coefficient , m is the number of related URLs that have execution relationships with the related URLs corresponding to the key word, It is the reference access interval between the jth related URL that has an execution relationship with the related URL corresponding to the keyword and the related URL corresponding to the keyword. For any two related URLs, the reference access interval is the average of the interval durations between the visits of each user to the web pages corresponding to the above two related URLs during a single visit.
2. The enterprise evaluation method based on big data analysis according to claim 1, characterized in that: For a single target evaluation subset, if the target evaluation subset is in a shard execution state where the extension validity coefficient is less than or equal to the preset extension validity coefficient or the execution completion coefficient is less than or equal to the preset execution completion coefficient, the first confidence analysis method is used to analyze the target evaluation subset, including: Determine a sharding quality parameter of a class of evaluation subsets according to subset execution efficiency and reference correlation index; Determine whether to perform subset task scheduling for a type of evaluation subset according to the shard quality parameter; Under the subset confirmation condition, the confidence reference coefficient of each relevant analysis text is determined based on the topic enrichment parameter; The subset confirmation condition is that a type of evaluation subset does not need to be scheduled for a subset task or has completed the subset task; The first type of evaluation subset is a target evaluation subset in a first type of shard execution state.
3. The enterprise evaluation method based on big data analysis according to claim 2, characterized in that: For a single type of evaluation subset, if the shard quality parameter of the type of evaluation subset is less than the preset shard quality parameter, subset task scheduling is performed for the type of evaluation subset; Determine the URLs to be segmented for this evaluation subset based on the extended correlation coefficient and the topic-independent parameter; The target allocation subset is determined based on the key words corresponding to the URL to be split and its extended URL and the subset processing efficiency.
4. The enterprise evaluation method based on big data analysis according to claim 3 is characterized in that: For a single first-class evaluation subset, if the shard quality parameter of the first-class evaluation subset is greater than the preset shard quality parameter, the number of concurrent requests is increased based on the subset execution efficiency; Determine whether to perform secondary adjustment on the number of concurrent requests based on the efficiency optimization parameters; The increase in the number of concurrent requests is negatively correlated with the subset execution efficiency.
5. The enterprise evaluation method based on big data analysis according to claim 4 is characterized in that: For a single target evaluation subset, if the target evaluation subset is in a second-class shard execution state where the expansion validity coefficient is greater than the preset expansion validity coefficient and the execution completeness coefficient is greater than the preset execution completeness coefficient, the second confidence analysis method is used to analyze the target evaluation subset, including: Determine the subset execution coefficient of the second-category evaluation subset based on the reference relevance index and the proportion of subject relevance; Determine the set setting method and effective evaluation strategy based on the subset execution coefficient; The set setting method is to determine the relevant matching set according to the subject association parameter and the set association index or to determine the access matching set according to the execution correlation coefficient and the access correlation coefficient; The second-category evaluation subset is a target evaluation subset in a second-category shard execution state.
6. The enterprise evaluation method based on big data analysis according to claim 5, characterized in that: For a single second-category evaluation subset, if the subset execution coefficient of the second-category evaluation subset is greater than the preset subset execution coefficient, the relevant matching set is determined based on the topic association parameter and the set association index, and the confidence reference coefficient of the relevant analysis text obtained by the second-category evaluation subset in the current subset analysis cycle is determined based on the reference relevant matching coefficient of each evaluation matching set and the key evaluation parameter; The correlation matching coefficient of each key word in any correlation matching set is greater than the preset correlation matching coefficient.
7. The enterprise evaluation method based on big data analysis according to claim 6, characterized in that: For a single second-category evaluation subset, if the subset execution coefficient of the second-category evaluation subset is less than or equal to the preset subset execution coefficient, an access matching set is determined based on the execution correlation coefficient and the access correlation coefficient, and a confidence reference coefficient of the relevant analysis text obtained by the second-category evaluation subset in the current subset analysis cycle is determined based on the access matching coefficient and the corresponding reference influence coefficient of each access matching set; The access matching coefficient in any access matching set is greater than the preset access matching coefficient.
8. The enterprise evaluation method based on big data analysis according to claim 7 is characterized in that: Under the subset analysis condition, the information acquisition quality of each target evaluation set is determined based on the effective reference ratio and information overlap parameters; The information acquisition quality is positively correlated with the effective reference ratio and the information overlap parameter respectively; The subset analysis condition is that each target evaluation subset completes the task effective analysis or performs effective analysis.
9. The enterprise evaluation method based on big data analysis according to claim 8, characterized in that: Determine the final evaluation score of each target evaluation data of the target evaluation enterprise based on the data values corresponding to the target evaluation data in the obtained relevant analysis text and the evaluation weight coefficient corresponding to the relevant analysis text; For a single relevant analysis text, its corresponding evaluation weight coefficient is positively correlated with the information acquisition quality of the target evaluation subset to which the relevant analysis text belongs.
10. An enterprise evaluation system using the enterprise evaluation method based on big data analysis according to any one of claims 1 to 9, characterized in that: include: A subset evaluation module for determining a subset processing strategy by expanding effective coefficients and executing complete coefficients in response to subset evaluation conditions; a first confidence analysis module connected to the subset evaluation module, configured to respond to the shard quality parameter of the adjustment judgment condition to determine an efficiency adjustment method and complete the determination of the confidence reference coefficient of each relevant analysis text, wherein the efficiency adjustment method is to perform subset task scheduling for a class of evaluation subsets, or to adjust the number of concurrent requests based on the subset execution efficiency; a second confidence analysis module connected to the subset evaluation module, for responding to the subset execution coefficient of the set judgment condition to determine a set setting mode and an effective evaluation strategy, wherein the effective evaluation strategy is to determine a confidence reference coefficient based on a reference related matching coefficient and a key evaluation parameter, or to determine a confidence reference coefficient based on an access matching coefficient and a reference influence coefficient; An evaluation execution module, which is connected to the first confidence analysis module and the second confidence analysis module, is used to determine the final evaluation score of each target evaluation data based on the data value of the target evaluation data and the evaluation weight coefficient corresponding to the relevant analysis text.
Citation Information
Patent Citations
Web crawler implementation method and system, computing equipment and storage medium
CN113239253A
Enterprise operation risk assessment system and method based on data mining algorithm
CN117172530A
Risk assessment for industrial systems using big data
US20140337086A1