Optimization Method and System for the Architecture of a Health Big Data Platform Based on Cloud Computing
By optimizing the data transmission and storage processes of the health big data platform, segmenting data into fragments and performing architectural modulation, the problem of unreasonable data storage is solved, the platform's data transmission and storage efficiency is improved, the cost is reduced, and resource utilization and platform stability are improved.
Patent Information
- Application Number
- CN202510396653.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-04-01
AI Technical Summary
The existing health big data platform has problems such as excessive data volume and unreasonable storage structure in the data storage process, resulting in low data transmission and storage efficiency, high risk of resource waste, and affecting the platform's operating efficiency and data service quality.
By analyzing the data transmission process of the health big data platform, the data transmission architecture is optimized and configured, the data is divided into several storage fragments and the storage architecture is modulated, and the resource allocation is dynamically adjusted in combination with elastic computing resource optimization management.
It improves data transmission efficiency and stability, reduces storage costs, enhances data storage and retrieval efficiency, improves resource utilization, and ensures the continuity and stability of the platform's task processing during peak periods.
Smart Images

Figure CN119917549B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of electronic digital data processing, and specifically to an optimization method and system for the architecture of a health big data platform based on cloud computing. Background Art
[0002] In the field of health big data platforms, with the rapid development of information technology and the growth of medical and health data, efficiently and securely storing, processing, and analyzing massive amounts of health data has become one of the key technologies for improving the quality of medical services and the level of health management. Health big data covers various types of data from individual physiology, diseases to environmental factors, and has the characteristics of diversity, high dimensionality, and complexity. The health big data platform based on cloud computing realizes the distributed storage and real-time processing of data through the powerful computing and storage capabilities of cloud computing, and can deeply mine and intelligently predict health data, thus promoting the intelligent and modernization process of medical and health services.
[0003] The prior art, such as the automatic optimization method for configuring a big data processing platform announced in the invention patent with the publication number CN113032033B, collects a configuration parameter training set according to the following steps: randomly generate a set of configuration parameters for the big data processing platform, run this set of configuration parameters, and monitor the execution time t. When the execution time t exceeds the dynamically set maximum allowable time Tmax, terminate the operation; during the operation of the configuration parameters, determine whether to exit the process of collecting the configuration parameter training set according to the fluctuation of the execution time t; after deciding to exit the process of collecting the configuration parameter training set, select the set of configuration parameters with the shortest execution time among all successfully run configuration parameters as the optimal configuration parameters.
[0004] The prior art, such as the overall architecture optimization system and optimization method for an integrated data platform announced in the invention patent with the publication number CN110046143B, includes a database. The database is electrically connected to a database operation status monitoring module in an output manner. The database operation status monitoring module is electrically connected to a management end and an optimization scheme selection module in an output manner. Both the hard disk and the server are bidirectionally electrically connected to a distribution adjustment module. The distribution adjustment module is bidirectionally electrically connected to an optimization model generation module. The adjustment command confirmation module is electrically connected to the management end in an input manner. The distribution adjustment module is electrically connected to a process optimization module in an output manner. The process optimization module is electrically connected to the server in an output manner.
[0005] Combined with the above solutions, it is found that in the current field of optimizing the health big data architecture, usually only the data collection process and the module optimization process are analyzed. However, in the data storage process, there will be problems such as being unable to manage finely due to excessive data volume and unreasonable storage structure, which not only affects the efficiency and stability of data transmission and storage, but also may pose a risk of resource waste, affecting the operation efficiency of the health big data platform and the quality of data services, and further affecting the utilization rate of resources. Summary of the Invention
[0006] In view of the deficiencies of the prior art, the present invention provides a method and system for optimizing the architecture of a health big data platform based on cloud computing, which can effectively solve the problems involved in the above-mentioned background technology.
[0007] To achieve the above object, the present invention is realized through the following technical solutions: A method for optimizing the architecture of a health big data platform based on cloud computing includes analyzing the data transmission process of the health big data platform to obtain data transmission optimization information of the health big data platform, and performing data transmission architecture optimization configuration on the health big data platform based on the data transmission optimization information of the health big data platform.
[0008] Statistical health data after data transmission architecture optimization configuration, and divide it into several segments according to time, denoted as each health data storage segment.
[0009] Analyze the query data of each health data storage segment, and then perform storage architecture modulation on the data of each health data storage segment of the health big data platform.
[0010] Statistical each task node of the health big data platform after storage architecture modulation, denoted as each task processing node, monitor and analyze the task processing process of each task processing node, and then perform elastic computing resource optimization management on the health big data platform.
[0011] In a second aspect of the present invention, there is provided a system for optimizing the architecture of a health big data platform based on cloud computing, including: a data transmission architecture optimization module, configured to analyze the data transmission process of the health big data platform to obtain data transmission optimization information of the health big data platform, and perform data transmission architecture optimization configuration on the health big data platform based on the data transmission optimization information of the health big data platform.
[0012] A data storage division module, configured to statistically analyze the health data after data transmission architecture optimization configuration, and divide it into several segments according to time, denoted as each health data storage segment.
[0013] A storage architecture modulation module, configured to analyze the query data of each health data storage segment, and then perform storage architecture modulation on the data of each health data storage segment of the health big data platform.
[0014] A computing resource optimization management module, configured to statistically analyze each task node of the health big data platform after storage architecture modulation, denoted as each task processing node, monitor and analyze the task processing process of each task processing node, and then perform elastic computing resource optimization management on the health big data platform.
[0015] The present invention has the following beneficial effects:
[0016] (1) The present invention provides an optimization method for the architecture of a health big data platform based on cloud computing. First, the data transmission process is analyzed, and then the data transmission architecture is optimized and configured, which can reduce service interruptions caused by transmission problems. Then, the data of each health data storage segment is modulated in the storage architecture, which helps to improve the flexibility of data storage and reduce the storage cost. Finally, elastic computing resource optimization management is carried out, which can dynamically adjust resources according to actual needs and improve resource utilization.
[0017] (2) By optimizing and configuring the data transmission architecture of the health big data platform according to the data transmission optimization information of the health big data platform, the present invention helps to improve the efficiency of data transmission, reduce latency and packet loss phenomena during data transmission, enhance the stability and reliability of data transmission, thereby improving the accuracy of health data. At the same time, the optimized data transmission architecture can improve the platform's ability to process high-concurrency data and reduce the risk of system failures caused by data transmission problems.
[0018] (3) By analyzing the query data of each health data storage segment and then modulating the data of each health data storage segment of the health big data platform in the storage architecture, the present invention improves the efficiency of data storage and retrieval. Through the refined management and optimization of storage segments, the response time of data queries is reduced, the overall performance of the platform is improved, the utilization rate of storage resources is optimized, and the cost of data storage is reduced.
[0019] (4) By monitoring and analyzing the task processing process of each task processing node and then carrying out elastic computing resource optimization management on the health big data platform, the present invention effectively improves the platform's resource scheduling ability, ensuring that the platform can dynamically adjust computing resources during data peak periods and task load changes to adapt to changing needs. This not only improves resource utilization but also enhances the stability and scalability of the platform, ensuring the continuity and efficiency of task processing.
[0020] Of course, it is not necessary for any product implementing the present invention to achieve all the above-mentioned advantages simultaneously. Brief Description of the Drawings
[0021] Figure 1 It is a schematic flowchart of the method of the present invention.
[0022] Figure 2 It is a schematic diagram of the connection of system modules of the present invention. Detailed Embodiments
[0023] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0024] Please refer to Figure 1 As shown, a technical solution is provided in the first aspect of the embodiment of the present invention: an optimization method for the architecture of a health big data platform based on cloud computing, including analyzing the data transmission process of the health big data platform to obtain data transmission optimization information of the health big data platform, and performing data transmission architecture optimization configuration on the health big data platform based on the data transmission optimization information of the health big data platform.
[0025] It should be noted that the health big data platform is a comprehensive system integrating the collection, storage, processing, analysis, and sharing of health-related information.
[0026] Statistical health data after data transmission optimization configuration, and divide it into several segments according to time, denoted as each health data storage segment.
[0027] Analyze the query data of each health data storage segment, and then modulate the storage architecture of the data of each health data storage segment of the health big data platform.
[0028] Statistical each task node of the health big data platform after storage architecture modulation, denoted as each task processing node, monitor and analyze the task processing process of each task processing node, and then perform elastic computing resource optimization management on the health big data platform.
[0029] Specifically, analyzing the data transmission process of the health big data platform, the specific analysis process is: analyzing the data transmission process of the health big data platform to obtain the data transmission parameters of the health big data platform, and the data transmission parameters of the health big data platform include the average data transmission delay duration of the health big data platform within a preset period, the maximum data transmission coefficient of the network link, the average utilization rate of the data packet payload, the data transmission anomaly coefficient, and the average data transmission rate.
[0030] Based on the data transmission parameters of the health big data platform, a data transmission efficiency evaluation value of the health big data platform is processed, and the data transmission efficiency evaluation value of the health big data platform is used to comprehensively quantify the overall performance of data transmission.
[0031] It should be noted that the average data transmission delay duration refers to the average time it takes for data to travel from the sender to the receiver. By using network analysis tools (such as Wireshark, TCPdump) to record the delay duration of each data packet, summing up the delay durations of all data packets, and then dividing by the total number of data packets, the average data transmission delay duration is obtained. The maximum data transmission coefficient of the network link refers to the ratio of the maximum data transmission rate that the network link can achieve within a preset period to the theoretical maximum transmission rate. Network bandwidth testing tools (such as iperf, Speedtest.net) can be used to measure the actual maximum data transmission rate of the network link within a preset period, and then the maximum data transmission coefficient of the network link is obtained. The average utilization rate of the data packet payload refers to the average value of the ratio of the payload (actual data) in each data packet to the total size of the data packet within a preset period. By using network analysis tools to record the payload and the total size of each data packet, the utilization rate of the payload of each data packet can be obtained, and then the average utilization rate of the data packet payload is obtained through mean processing. The data transmission anomaly coefficient reflects the abnormal conditions (such as data packet loss, errors) that occur during data transmission within a preset period. By using network monitoring tools (such as Nagios, Zabbix) to count the number of data transmission anomaly events that occur within a preset period, and removing the unit from the ratio of the number of data transmission anomaly events to the total number of transmissions, the data transmission anomaly coefficient is obtained. The average data transmission rate of the healthy big data platform within a preset period can be measured using network bandwidth testing tools.
[0032] Specifically, for the data transmission efficiency evaluation value of the healthy big data platform, the specific analysis conditions are as follows:
[0033] ;
[0034] In the formula, represents the data transmission efficiency evaluation value of the healthy big data platform, represents the average data transmission delay duration of the healthy big data platform, represents the data transmission efficiency correction factor corresponding to the set average data transmission delay duration per unit data, represents the maximum data transmission coefficient of the network link of the healthy big data platform, represents the data transmission efficiency correction factor corresponding to the set maximum data transmission coefficient of the network link, represents the average utilization rate of the data packet payload of the healthy big data platform, represents the data transmission efficiency correction factor corresponding to the set average utilization rate of the data packet payload per unit, represents the data transmission anomaly coefficient of the healthy big data platform, represents the data transmission efficiency correction factor corresponding to the set data transmission anomaly coefficient, represents the average data transmission rate of the health big data platform represents the data transmission efficiency correction factor corresponding to the set unit data transmission rate, and e represents the natural constant
[0035] It should be added that in this embodiment, the data transmission efficiency correction factors corresponding to the preset average unit data transmission delay duration, the maximum data transmission coefficient of the network link, the average utilization rate of the unit data packet payload, the data transmission anomaly coefficient, and the unit data transmission rate are obtained from the platform architecture optimization database
[0036] It should be explained that the data transmission efficiency correction factors corresponding to the average unit data transmission delay duration, the maximum data transmission coefficient of the network link, the average utilization rate of the unit data packet payload, the data transmission anomaly coefficient, and the unit data transmission rate are respectively used to adjust the importance of the average data transmission delay duration, the maximum data transmission coefficient of the network link, the average utilization rate of the data packet payload, the data transmission anomaly coefficient, and the average data transmission rate of the health big data platform in the process of analyzing the data transmission efficiency evaluation value of the health big data platform. For example, there is a preset mapping relationship between the data transmission parameters of the health big data platform and the corresponding data transmission efficiency correction factors in the platform architecture optimization database. Through the preset mapping relationship, the data transmission efficiency correction factors corresponding to the real-time data transmission parameters of the health big data platform can be matched. The average data transmission delay duration, the maximum data transmission coefficient of the network link, the average utilization rate of the data packet payload, the data transmission anomaly coefficient, and the average data transmission rate of the health big data platform are respectively matched with the preset mapping relationship to obtain the data transmission efficiency correction factors corresponding to the average unit data transmission delay duration, the maximum data transmission coefficient of the network link, the average utilization rate of the unit data packet payload, the data transmission anomaly coefficient, and the unit data transmission rate
[0037] In this implementation plan, there is a correlation among the average data transmission delay duration, the maximum data transmission coefficient of the network link, the average utilization rate of the data packet payload, the data transmission anomaly coefficient, and the average data transmission rate of the health big data platform, and they do not exist independently. For example, if the maximum data transmission coefficient of the network link is low, even if the average data transmission delay duration is very low, the average data transmission rate will still decrease. If the data transmission anomaly coefficient is high, it indicates that the network link is unstable, with many errors, packet losses, or retries. Each retransmission will cause an increase in the average data transmission delay duration and may reduce the actual maximum transmission rate, affecting the maximum data transmission coefficient of the network link. A low average utilization rate of the data packet payload may lead to an increased burden on the network link, increasing the incidence of anomalies and resulting in an increase in the data transmission anomaly coefficient. By comprehensively analyzing the data transmission efficiency evaluation value of the health big data platform, it is possible to help identify the data transmission performance bottleneck of the health big data platform, understand the maximum transmission capacity and actual utilization rate of the network link, and thus allocate network resources more effectively.
[0038] Specifically, the process of obtaining the data transmission optimization information of the health big data platform is as follows: The data transmission optimization information of the health big data platform includes performing data transmission optimization and not performing data transmission optimization.
[0039] Compare the data transmission efficiency evaluation value of the health big data platform with the set data transmission efficiency evaluation threshold. If the data transmission efficiency evaluation value of the health big data platform is higher than or equal to the set data transmission efficiency evaluation threshold, mark the data transmission optimization information of the health big data platform as not performing data transmission optimization; otherwise, mark the data transmission optimization information of the health big data platform as performing data transmission optimization.
[0040] Specifically, based on the data transmission optimization information of the health big data platform, perform an optimized configuration of the data transmission architecture of the health big data platform. The specific process is as follows: Extract the data transmission optimization information of the health big data platform. If the data transmission optimization information of the health big data platform is not performing data transmission optimization, continue to perform the optimized configuration of the data transmission architecture with the current data transmission compression ratio.
[0041] If the data transmission optimization information of the health big data platform is performing data transmission optimization, add the current data transmission compression ratio to the set data transmission compression ratio correction value to obtain the target data transmission compression ratio, and adjust the data transmission compression ratio of the health big data platform to the target data transmission compression ratio for the optimized configuration of the data transmission architecture.
[0042] It should be noted that analyzing the data transmission process of the health big data platform can reflect the efficiency, stability, and reliability of data transmission. If subsequent data transmission is carried out at the original default data transmission compression ratio when the data transmission performance is unqualified, it may lead to an increase in data transmission latency, a decrease in data transmission quality, and even the risk of data transmission interruption or data loss. Therefore, it is necessary to optimize the configuration of the data transmission architecture in combination with the data transmission compression ratio correction value to ensure the smoothness of data transmission and the accuracy of data processing, providing data support for optimizing the overall data transmission performance of the health big data platform.
[0043] Specifically, the query data of each health data storage segment is analyzed. The specific analysis process is as follows: Based on the data transmission efficiency evaluation value of the health big data platform, the average data transmission efficiency evaluation value of the health big data platform is statistically obtained, and the data storage evaluation impact factor is matched according to the average data transmission efficiency evaluation value of the health big data platform.
[0044] It should be explained that the average data transmission efficiency evaluation value of the health big data platform refers to the average of the data transmission efficiency evaluation values of the health big data platform within each preset period. The data storage evaluation impact factor is matched according to the average data transmission efficiency evaluation value of the health big data platform. The specific process is to match the average data transmission efficiency evaluation value of the health big data platform with the data storage evaluation impact factors corresponding to each data transmission efficiency average evaluation value interval stored in the platform architecture optimization database, and statistically obtain the data storage evaluation impact factor corresponding to the interval where the average data transmission efficiency evaluation value of the health big data platform is located, denoted as the data storage evaluation impact factor. The larger the average data transmission efficiency evaluation value of the health big data platform, the better the quality of the data transmission process, and the smaller the impact on the subsequent storage evaluation. Therefore, the smaller the matched data storage evaluation impact factor.
[0045] The query data of each health data storage segment includes the query frequency, query response time fluctuation coefficient, query failure rate, and query concurrency coefficient within the preset period.
[0046] It should be noted that the query frequency of each health data storage segment refers to the number of times of querying each health data storage segment within a preset period. The number of queries for each health data storage segment within the preset period can be statistically obtained from the monitoring log, and the ratio of the number of queries to the duration of the preset period is counted as the query frequency of each health data storage segment. The query response time fluctuation coefficient is used to measure the stability of the query response time. A database performance monitoring tool is used to measure the response duration of each query within the preset period. The average response duration is obtained by averaging the response durations. The sum of the squares of the differences between the response durations of each query and the average response duration is calculated, and then divided by the total number of queries, and then the square root is taken to obtain the standard deviation. The ratio of the obtained standard deviation to the average response duration is processed to remove the unit to obtain the query response time fluctuation coefficient. The query failure rate refers to the ratio of the number of failed query attempts to the total number of queries within the preset period. The total number of queries and the number of failed queries within the preset period can be statistically obtained from the monitoring log, and the ratio of the number of failed queries to the total number of queries is recorded as the query failure rate. The query concurrency coefficient is an index that measures the ability of a data storage segment to process multiple query requests within a preset period. By monitoring the log within the preset period, the time points of all query requests are recorded, and the time point with the largest number of concurrent query requests occurring simultaneously within the preset period is found. The query concurrency coefficient is the value after removing the unit of the number of concurrent requests at this time point.
[0047] Comprehensive analysis is performed on the query data of each health data storage segment to obtain the data storage characterization value of each health data storage segment, and the data storage characterization value of each health data storage segment is used to comprehensively quantify the query status of the data storage segment.
[0048] In this embodiment, the data storage characterization value of each health data storage segment can be obtained through the following analysis method, and the specific analysis conditions are as follows:
[0049] ;
[0050] In the formula, represents the data storage characterization value of the i-th health data storage segment, represents the query frequency of the i-th health data storage segment, represents the defined query frequency of the i-th health data storage segment, represents the query response time fluctuation coefficient of the i-th health data storage segment, represents the data storage evaluation factor corresponding to the set query response time fluctuation coefficient, represents the query failure rate of the i-th health data storage segment, represents the data storage evaluation factor corresponding to the set unit query failure rate, Represents the query concurrency coefficient of the i-th health data storage segment, Represents the defined query concurrency coefficient of the i-th health data storage segment, where i represents the number of each health data storage segment, and n represents the total number of health data storage segments.
[0051] It should be added that in this embodiment, the data storage evaluation factors corresponding to the preset query response time fluctuation coefficient and the data storage evaluation factors corresponding to the unit query failure rate are obtained from the platform architecture optimization database.
[0052] It should be explained that the data storage evaluation factors corresponding to the query response time fluctuation coefficient and the unit query failure rate are respectively used to adjust the importance of the query response time fluctuation coefficient and the query failure rate of each health data storage segment in the process of analyzing the data storage characterization value of each health data storage segment. For example, there is a preset mapping relationship between the query data of each health data storage segment and the corresponding data storage evaluation factor in the platform architecture optimization database. Through the preset mapping relationship, the data storage evaluation factor corresponding to the query data of each health data storage segment in real time can be matched. The query response time fluctuation coefficient and the query failure rate of each health data storage segment are respectively matched with the preset mapping relationship to obtain the data storage evaluation factors corresponding to the query response time fluctuation coefficient and the unit query failure rate.
[0053] In this implementation plan, there is a correlation among the query frequency, query response time fluctuation coefficient, query failure rate, and query concurrency coefficient of each health data storage segment, and they do not exist independently. For example, a high query frequency usually increases the load on the platform. Especially when query requests are concentrated, the platform may have unstable response times because it cannot process all query requests in time, resulting in an increase in the query response time fluctuation coefficient. At the same time, it may cause query timeouts, database connection failures, or query requests to be discarded, thus increasing the query failure rate. A large query response time fluctuation coefficient is often an indication of uneven platform load, insufficient resources, or bottlenecks, which may lead to query failures. When the response time is too long, the query may time out or be interrupted, and the query failure rate will increase. When the query concurrency is too high, resource competition may occur, resulting in unstable response times, thus increasing the query response time fluctuation coefficient. By comprehensively analyzing the data storage characterization values of each health data storage segment, the utilization status of the health data storage segments can be evaluated more accurately, and the storage resources can be allocated more reasonably, thereby improving the response efficiency of the overall storage system.
[0054] Based on the data storage characterization values of each health data storage segment, combined with the data storage evaluation impact factors, the data storage comprehensive index values of each health data storage segment are obtained through comprehensive processing. The data storage comprehensive index values of each health data storage segment are used to comprehensively quantify the query resource utilization status of the data storage segment.
[0055] It should be added that the product of the data storage characterization value of each health data storage segment and the data storage evaluation impact factor is denoted as the data storage comprehensive index value of each health data storage segment.
[0056] Specifically, the data of each health data storage segment of the health big data platform is subjected to storage architecture modulation. The specific process is as follows: The data storage comprehensive index value of each health data storage segment is compared with the set data storage comprehensive index threshold. If the data storage comprehensive index value of a certain health data storage segment is higher than or equal to the set data storage comprehensive index threshold, the health data storage segment is marked as a health data high-frequency query storage segment. Thus, all health data high-frequency query storage segments are obtained through traversal, and the data in each health data high-frequency query storage segment is transmitted to the high-performance cloud computing service platform.
[0057] If the data storage comprehensive index value of a certain health data storage segment is lower than the set data storage comprehensive index threshold, the health data storage segment is marked as a health data low-frequency query storage segment. Thus, all health data low-frequency query storage segments are obtained through traversal, and the data in each health data low-frequency query storage segment is transmitted to the normal performance storage service platform.
[0058] It should be added that the high-performance cloud computing service platform refers to a cloud storage service that provides fast data access, high throughput, and low latency. The normal performance storage service platform refers to a cloud storage service that meets general business requirements and has a relatively low cost.
[0059] It should be explained that placing the health data storage segments with frequent queries on the high-performance cloud computing service platform can ensure that the data can be retrieved and accessed faster, thereby improving the overall data query efficiency. By storing the data separately according to the query frequency, the storage resources can be more reasonably allocated, avoiding wasting resources on storing infrequently accessed data on the high-performance storage platform, and at the same time, the resources can be effectively utilized on the normal performance storage service platform.
[0060] Specifically, the task processing process of each task processing node is monitored and analyzed. The specific analysis process is as follows: The task processing process of each task processing node is monitored and analyzed to obtain the task processing data of each task processing node. The task processing data of each task processing node includes the task execution time fluctuation coefficient, task failure rate, resource saturation, and task average execution duration of each task processing node within a preset period.
[0061] It should be noted that the task execution time fluctuation coefficient reflects the stability of task processing time. System monitoring tools (such as Prometheus, Grafana) can be used to record the execution time of each task within a preset period, perform mean processing to obtain the average execution time, sum the squares of the differences between the execution time of each task and the average execution time to obtain a value, divide this value by the number of tasks and then take the square root to obtain the standard deviation of the task execution time. The value obtained by dividing the standard deviation of the task execution time by the average execution time is recorded as the task execution time fluctuation coefficient. The task failure rate refers to the ratio of the number of task execution failures to the total number of executions within a preset period. The total number of task executions and the number of failures within a preset period can be monitored and recorded through a log analysis tool (such as ELK Stack, Splunk), and the ratio of the number of failures to the total number of executions is recorded as the task failure rate. The resource saturation refers to the degree of utilization of task processing node resources (such as CPU, memory, disk I / O). The resource usage amount and the total resource amount of the task processing node within a preset period are monitored using the built-in monitoring tool of the platform, and the ratio of the resource usage amount to the total resource amount is recorded as the resource saturation.
[0062] Based on the comprehensive analysis of the task processing data of each task processing node, the task processing basic evaluation value of each task processing node is obtained, and the task processing basic evaluation value of each task processing node is used to comprehensively quantify the task processing load of each task processing node.
[0063] In this embodiment, the task processing basic evaluation value of each task processing node can be obtained through the following analysis method, and the specific analysis conditions are as follows:
[0064] ;
[0065] In the formula, represents the task processing basic evaluation value of the j-th task processing node, represents the task execution time fluctuation coefficient of the j-th task processing node, represents the task processing efficiency evaluation factor corresponding to the set task execution time fluctuation coefficient, represents the task failure rate of the j-th task processing node, represents the task processing efficiency evaluation factor corresponding to the set unit task failure rate, represents the resource saturation of the j-th task processing node, represents the defined resource saturation of the j-th task processing node, represents the average task execution duration of the j-th task processing node, represents the defined task execution duration of the j-th task processing node, and j represents the number of each task processing node. , m represents the total number of task processing nodes.
[0066] It should be added that in this embodiment, the task processing efficiency evaluation factor corresponding to the preset task execution time fluctuation coefficient and the task processing efficiency evaluation factor corresponding to the unit task failure rate are obtained from the platform architecture optimization database.
[0067] It should be explained that the task processing efficiency evaluation factors corresponding to the task execution time fluctuation coefficient and the unit task failure rate are respectively used to adjust the importance degrees of the task execution time fluctuation coefficient and the task failure rate of each task processing node in the process of analyzing the task processing basic evaluation value of each task processing node. For example, there is a preset mapping relationship between the task processing data of each task processing node and the corresponding task processing efficiency evaluation factor in the platform architecture optimization database. Through the preset mapping relationship, the task processing efficiency evaluation factor corresponding to the task processing data of each task processing node in real time can be matched. The task execution time fluctuation coefficient and the task failure rate of each task processing node are respectively matched with the preset mapping relationship to obtain the task processing efficiency evaluation factors corresponding to the task execution time fluctuation coefficient and the unit task failure rate.
[0068] In this implementation solution, there is a correlation among the task execution time fluctuation coefficient, the task failure rate, the resource saturation degree, and the average task execution duration of each task processing node, and they do not exist independently. For example, when the task execution time fluctuation coefficient is large, the execution time of the task is unstable, which may cause the execution time of the task to exceed the average value, thereby increasing the average task execution duration. When the platform resources are close to saturation, the task may not be able to obtain sufficient resources in time to complete the execution, resulting in task failure. Therefore, a higher resource saturation degree is usually accompanied by a higher task failure rate. A large task execution time fluctuation coefficient usually exacerbates the risk of task failure. Long-term fluctuations in the task execution time may cause the task to time out, thereby resulting in task failure. A higher task failure rate may mean that the resource configuration of the platform is unreasonable, resulting in an increase in the time consumption during task execution, and further lengthening the average task execution duration. Comprehensively analyzing the task processing basic evaluation value of each task processing node can help identify the key factors affecting task processing performance, and thus can improve the overall efficiency of task processing.
[0069] Specifically, for the elastic computing resource optimization management of the health big data platform, the specific process is as follows: Extract the task processing basic evaluation value of each task processing node, compare the task processing basic evaluation value of each task processing node with the set first task processing basic evaluation threshold. If the task processing basic evaluation value of a certain task processing node is higher than or equal to the set first task processing basic evaluation threshold, then mark this task processing node as a high task volume processing node.
[0070] If the task processing basic evaluation value of a certain task processing node is lower than the set first task processing basic evaluation threshold, then compare the task processing basic evaluation value of this task processing node with the set second task processing basic evaluation threshold. If the task processing basic evaluation value of this task processing node is higher than or equal to the set second task processing basic evaluation threshold, then mark this task processing node as a normal task volume processing node.
[0071] If the task processing basic evaluation value of this task processing node is lower than the set second task processing basic evaluation threshold, then mark this task processing node as a low task volume processing node.
[0072] Perform elastic computing resource optimization management on the task processing nodes marked as high task volume processing nodes and low task volume processing nodes.
[0073] It should be noted that the specific process of performing elastic computing resource optimization management on the task processing nodes marked as high task volume processing nodes and low task volume processing nodes is as follows: by sending prompt information to the management terminal, prompting to add data processing nodes to the high task volume processing nodes, increasing computing resources (such as increasing resources such as CPU and memory), increasing network bandwidth, prompting to reduce data processing nodes for the low task volume processing nodes, reducing computing resources (such as reducing resources such as CPU and memory), reducing network bandwidth, and at the same time using a load balancer to automatically allocate new tasks to the low task volume processing nodes, which helps to improve resource utilization efficiency, reduce platform operation costs, and ensure that the platform maintains high performance and stability in the face of volatile loads.
[0074] Such as Figure 2 , the second aspect of the present invention provides a health big data platform architecture optimization system based on cloud computing, including: a data transmission architecture optimization module, which is used to analyze the data transmission process of the health big data platform to obtain data transmission optimization information of the health big data platform, and perform data transmission architecture optimization configuration on the health big data platform based on the data transmission optimization information of the health big data platform.
[0075] A data storage division module, which is used to count the health data after the data transmission optimization configuration and divide it into several segments according to time, denoted as each health data storage segment.
[0076] A storage architecture modulation module, which is used to analyze the query data of each health data storage segment, and then perform storage architecture modulation on the data of each health data storage segment of the health big data platform.
[0077] A computing resource optimization management module, which is used to count each task node of the health big data platform after the storage architecture modulation, denoted as each task processing node, monitor and analyze the task processing process of each task processing node, and then perform elastic computing resource optimization management on the health big data platform.
[0078] It should be noted that the health big data platform architecture optimization system based on cloud computing further includes a platform architecture optimization database for storing a first parameter set, a second parameter set, and a third parameter set obtained by analyzing historical data.
[0079] The first parameter set includes a data transmission efficiency correction factor corresponding to the average delay duration of unit data transmission, a data transmission efficiency correction factor corresponding to the maximum data transmission coefficient of the network link, a data transmission efficiency correction factor corresponding to the average utilization rate of the effective payload of the unit data packet, a data transmission efficiency correction factor corresponding to the data transmission anomaly coefficient, a data transmission efficiency correction factor corresponding to the unit data transmission rate, a data transmission efficiency evaluation threshold, and a data transmission compression ratio correction value.
[0080] The second parameter set includes a data storage evaluation impact factor corresponding to each data transmission efficiency average evaluation value interval, the defined query frequency of each health data storage segment, a data storage evaluation factor corresponding to the query response time fluctuation coefficient, a data storage evaluation factor corresponding to the unit query failure rate, the defined query concurrency coefficient of each health data storage segment, and a data storage comprehensive index threshold.
[0081] The third parameter set includes a task processing efficiency evaluation factor corresponding to the task execution time fluctuation coefficient, a task processing efficiency evaluation factor corresponding to the unit task failure rate, the defined resource saturation of each task processing node, the defined task execution duration of each task processing node, a task processing basic first evaluation threshold, and a task processing basic second evaluation threshold.
[0082] It should be noted that in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.
[0083] The preferred embodiments of the present invention disclosed above are only used to help explain the present invention. The preferred embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the present invention, so that those skilled in the art can well understand and utilize the present invention.
Claims
1. An optimization method for the architecture of a health big data platform based on cloud computing, characterized in that Including: Analyze the data transmission process of the health big data platform to obtain the data transmission optimization information of the health big data platform, and optimize the configuration of the data transmission architecture of the health big data platform based on the data transmission optimization information of the health big data platform; Statistically analyze the health data after the data transmission optimization configuration, and divide it into several segments according to time, denoted as each health data storage segment; Analyze the query data of each health data storage segment, and then modulate the data storage architecture of each health data storage segment of the health big data platform; Statistically analyze each task node of the health big data platform after the data storage architecture modulation, denoted as each task processing node, monitor and analyze the task processing process of each task processing node, and then optimize the management of the elastic computing resources of the health big data platform; The analysis of the data transmission process of the health big data platform, the specific analysis process is: Analyze the data transmission process of the health big data platform to obtain the data transmission parameters of the health big data platform. The data transmission parameters of the health big data platform include the average data transmission delay duration of the health big data platform within a preset period, the maximum data transmission coefficient of the network link, the average utilization rate of the data packet payload, the data transmission anomaly coefficient, and the average data transmission rate; Process the data transmission parameters of the health big data platform to obtain the data transmission efficiency evaluation value of the health big data platform. The data transmission efficiency evaluation value of the health big data platform is used to comprehensively quantify the overall performance of data transmission; The analysis of the query data of each health data storage segment, the specific analysis process is: Based on the data transmission efficiency evaluation value of the health big data platform, statistically obtain the average data transmission efficiency evaluation value of the health big data platform, and match the data storage evaluation influence factor according to the average data transmission efficiency evaluation value of the health big data platform; The query data of each health data storage segment includes the query frequency, query response time fluctuation coefficient, query failure rate, and query concurrency coefficient of each health data storage segment within a preset period; Comprehensively analyze the query data of each health data storage segment to obtain the data storage characterization value of each health data storage segment. The data storage characterization value of each health data storage segment is used to comprehensively quantify the query status of the data storage segment; Based on the data storage characterization value of each health data storage segment, combined with the data storage evaluation influence factor, comprehensively process to obtain the data storage comprehensive index value of each health data storage segment. The data storage comprehensive index value of each health data storage segment is used to comprehensively quantify the query resource utilization status of the data storage segment.
2. The method for optimizing the architecture of a health big data platform based on cloud computing according to claim 1, wherein: The process of obtaining the data transmission optimization information of the health big data platform is specifically: The data transmission optimization information of the health big data platform includes executing data transmission optimization and not executing data transmission optimization; Compare the data transmission efficiency evaluation value of the health big data platform with the set data transmission efficiency evaluation threshold. If the data transmission efficiency evaluation value of the health big data platform is higher than or equal to the set data transmission efficiency evaluation threshold, mark the data transmission optimization information of the health big data platform as not performing data transmission optimization; otherwise, mark the data transmission optimization information of the health big data platform as performing data transmission optimization.
3. The method for optimizing the architecture of a health big data platform based on cloud computing according to claim 1, characterized in that: Perform data transmission architecture optimization configuration on the health big data platform based on the data transmission optimization information of the health big data platform. The specific process is as follows: Extract the data transmission optimization information of the health big data platform. If the data transmission optimization information of the health big data platform is not performing data transmission optimization, continue to perform data transmission architecture optimization configuration with the current data transmission compression ratio. If the data transmission optimization information of the health big data platform is performing data transmission optimization, add the current data transmission compression ratio to the set data transmission compression ratio correction value to obtain the target data transmission compression ratio, and adjust the data transmission compression ratio of the health big data platform to the target data transmission compression ratio for data transmission architecture optimization configuration.
4. The method for optimizing the architecture of a health big data platform based on cloud computing according to claim 1, characterized in that: Perform storage architecture modulation on the data of each health data storage segment of the health big data platform. The specific process is as follows: Compare the data storage comprehensive index value of each health data storage segment with the set data storage comprehensive index threshold. If the data storage comprehensive index value of a certain health data storage segment is higher than or equal to the set data storage comprehensive index threshold, mark the health data storage segment as a health data high-frequency query storage segment, traverse to obtain all health data high-frequency query storage segments, and transmit the data in each health data high-frequency query storage segment to the high-performance cloud computing service platform. If the data storage comprehensive index value of a certain health data storage segment is lower than the set data storage comprehensive index threshold, mark the health data storage segment as a health data low-frequency query storage segment, traverse to obtain all health data low-frequency query storage segments, and transmit the data in each health data low-frequency query storage segment to the normal performance storage service platform.
5. The method for optimizing the architecture of a health big data platform based on cloud computing according to claim 1, wherein: Monitor and analyze the task processing process of each task processing node. The specific analysis process is as follows: Monitor and analyze the task processing process of each task processing node to obtain the task processing data of each task processing node. Comprehensively analyze the task processing data of each task processing node to obtain the task processing basic evaluation value of each task processing node. The task processing basic evaluation value of each task processing node is used to comprehensively quantify the task processing load of each task processing node.
6. The method for optimizing the architecture of a health big data platform based on cloud computing according to claim 1, wherein: Perform elastic computing resource optimization management on the health big data platform. The specific process is as follows: Extract the task processing basic evaluation value of each task processing node, compare the task processing basic evaluation value of each task processing node with the set task processing basic first evaluation threshold. If the task processing basic evaluation value of a certain task processing node is higher than or equal to the set task processing basic first evaluation threshold, mark the task processing node as a high task volume processing node. If the task processing basic evaluation value of a certain task processing node is lower than the set first task processing basic evaluation threshold, then compare the task processing basic evaluation value of this task processing node with the set second task processing basic evaluation threshold. If the task processing basic evaluation value of this task processing node is higher than or equal to the set second task processing basic evaluation threshold, then mark this task processing node as a normal task volume processing node; If the task processing basic evaluation value of this task processing node is lower than the set second task processing basic evaluation threshold, then mark this task processing node as a low task volume processing node; Perform elastic computing resource optimization management on the task processing nodes marked as high task volume processing nodes and low task volume processing nodes.
7. The method for optimizing the architecture of a health big data platform based on cloud computing according to claim 5, characterized in that: The task processing data of each task processing node includes the task execution time fluctuation coefficient, task failure rate, resource saturation, and average task execution duration of each task processing node within a preset period.
8. A health big data platform architecture optimization system based on cloud computing, which applies the health big data platform architecture optimization method according to any one of claims 1-7, characterized in that Including: A data transmission architecture optimization module, which is used to analyze the data transmission process of the health big data platform to obtain the data transmission optimization information of the health big data platform, and perform data transmission architecture optimization configuration on the health big data platform based on the data transmission optimization information of the health big data platform; A data storage division module, which is used to count the health data after the data transmission optimization configuration and divide it into several segments according to time, denoted as each health data storage segment; A storage architecture modulation module, which is used to analyze the query data of each health data storage segment, and then perform storage architecture modulation on the data of each health data storage segment of the health big data platform; A computing resource optimization management module, which is used to count each task node of the health big data platform after the storage architecture modulation, denoted as each task processing node, monitor and analyze the task processing process of each task processing node, and then perform elastic computing resource optimization management on the health big data platform.
Citation Information
Patent Citations
An integrated data platform architecture optimization system and optimization method
CN110046143B
An automatic optimization method for big data processing platform configuration
CN113032033B
Method and system for improving data interaction efficiency of hyper-converged architecture
CN119576527A