Data weight clustering and marking method, device and storage medium based on large data sets

By establishing a sign data record database on a large data set, eliminating free points using clustering algorithms, calculating and sorting the weight clustering points of sign data, the problem of large deviations in cluster analysis results in the existing technology is solved, and accurate clustering of medical data and reasonable allocation of resources is achieved.

CN119207823BActive Publication Date: 2025-07-11SHANGHAI MEISI PHARM TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411489442.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-24
Publication Date
2025-07-11
Estimated Expiration
2044-10-24

AI Technical Summary

Technical Problem

The existing weight clustering marking technology has the problem that the cluster point analysis results are large in the medical field and it is difficult to uniformly normalize different medical data. In particular, the existence of free points affects the accuracy of clustering analysis.

Method used

By establishing a sign data record database based on the big data set, counting the weight set of sign data, using the clustering algorithm to exclude free coordinate points, calculate the weight clustering points, and sorting and numbering cluster values of different sign data, and finally clustering marking is performed.

Benefits of technology

It improves the accuracy and rationality of clustering analysis, can unify the standards for different sign data, screen out important and secondary sign data, reduce unnecessary monitoring, and improve the work efficiency of medical staff.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119207823B_ABST
    Figure CN119207823B_ABST
Patent Text Reader

Abstract

The present invention discloses a data weight clustering and marking method, device and storage medium based on a large data set, relating to the technical field of weight clustering and marking, and comprising the following steps: based on the large data set, statistically obtaining a weight set of physical sign data; analyzing the weight set through a clustering algorithm to obtain weight clustering points of the physical sign data; calculating clustering values of the physical sign data based on the weight clustering points, and then clustering and marking the physical sign data according to the clustering values of different physical sign data; the present invention is used to solve the problems that the existing weight clustering and marking technology still has a large deviation in the analysis result of clustering points and it is difficult to perform clustering analysis on different medical data, resulting in an inaccurate final clustering and marking result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of weighted clustering and marking, and specifically to a method, device and storage medium for weighted clustering and marking of data based on a large dataset. Background Art

[0002] The weighted clustering and marking technology refers to a method that combines data weight assignment, clustering analysis and marking process, and is used to effectively classify and mark the data in a dataset. In this technology, each data point is grouped and aggregated according to its importance.

[0003] The existing weighted clustering and marking technologies usually directly perform clustering analysis after constructing a distribution map. There are usually several free points in the distribution map. The proportion of such points is very small, but it will affect the result of the clustering points. The clustering algorithm itself is an abstract algorithm and there is no absolutely correct result. However, the existence of free points will increase the error of the analysis result of the clustering points. Only by performing clustering analysis on the dense area can the accuracy of the clustering points be improved. Moreover, when the existing weighted clustering and marking technologies are applied in the medical field, it is difficult to normalize different medical data. Different data have different units and different standards, and it is difficult to unify them during clustering analysis. For example, in the patent application with the publication number CN109978007A, a method for extracting disease risk factors based on attribute weight clustering is disclosed. This solution directly performs clustering analysis without excluding the influence of free points, resulting in a large deviation in the result of the clustering points obtained by the analysis. The existing weighted clustering and marking technologies also have problems of large deviation in the analysis result of the clustering points and difficulty in performing clustering analysis on different medical data, resulting in inaccurate final clustering and marking results. Summary of the Invention

[0004] The present invention aims to solve at least one of the technical problems in the prior art to some extent. By establishing a physical sign data record database based on a large dataset to store each data record, then analyzing and statistically obtaining the weight set of the physical sign data based on the physical sign data record database, then analyzing the weight set, excluding free coordinate points to obtain a clustering group, then analyzing the clustering group through a clustering algorithm to obtain the weighted clustering points of the physical sign data, then calculating the clustering value of the physical sign data based on the weighted clustering points, then sorting and numbering the clustering values of different physical sign data to obtain a clustering queue, and finally performing clustering and marking on the physical sign data based on the clustering queue, so as to solve the problems that the existing weighted clustering and marking technologies have large deviation in the analysis result of the clustering points and difficulty in performing clustering analysis on different medical data, resulting in inaccurate final clustering and marking results.

[0005] To achieve the above object, in a first aspect, the present application provides a data weight clustering and marking method based on a large data set, including the following steps:

[0006] Based on the large data set, statistically analyze the weight set of the physical sign data;

[0007] Analyze the weight set through a clustering algorithm to obtain the weight clustering points of the physical sign data;

[0008] Calculate the clustering value of the physical sign data based on the weight clustering points, and then cluster and mark the physical sign data according to the clustering values of different physical sign data.

[0009] Further, based on the large data set, statistically analyzing the weight set of the physical sign data includes the following sub-steps:

[0010] Based on the large data set, establish a physical sign data record database to store each data record;

[0011] Based on the physical sign data record database, analyze and statistically analyze the weight set of the physical sign data.

[0012] Further, based on the large data set, establishing a physical sign data record database to store each data record includes the following sub-steps:

[0013] For any physical sign data, establish a physical sign data record database, and a corresponding physical sign data record database is established for each physical sign data;

[0014] The data records stored in the physical sign data record database include a data number, a data weight, a data value, and an evaluation value;

[0015] The data number is the sorting number of the data record in the physical sign data record database; the data weight is the weight corresponding to the physical sign data in the field where this data record acts; the data value is the value of the physical sign data of the patient monitored by this data record; the evaluation value is the result finally calculated by the health evaluation calculation participated by the data record.

[0016] Further, based on the physical sign data record database, analyzing and statistically analyzing the weight set of the physical sign data includes the following sub-steps:

[0017] Calculate the data value divided by the evaluation value, and mark the calculation result as the data proportion;

[0018] Mark the data weight and the data proportion in a data record as a set of physical sign weight groups;

[0019] Statistically analyze all the physical sign weight groups in the same physical sign data record database, and integrate them into a weight set;

[0020] There is a corresponding weight set for each physical sign data.

[0021] Further, analyzing the weight set through a clustering algorithm to obtain the weight clustering points of the physical sign data includes the following sub-steps:

[0022] Taking the physical sign weight as the X-axis and the data proportion as the Y-axis, establish a plane rectangular coordinate system that only includes the first quadrant, named the data weight distribution diagram. Each weight set constructs an independent data weight distribution diagram;

[0023] Enter all the physical sign weight groups within the same weight set into the data weight distribution diagram, name the coordinate points among them as data weight points. Taking any data weight point as the center and the first correlation distance as the radius, construct a circle, named the correlation circle. Check whether there are data weight points within the correlation circle. If there are, output a data correlation signal; if not, output a data non-correlation signal;

[0024] If a data correlation signal is output, mark the data weight points within the correlation circle as correlation points, then take the correlation points as the center and the first correlation distance as the radius to construct a correlation circle, and judge again;

[0025] When no data correlation signal is output anymore, mark both the data weight point serving as the center and the correlation points as clustering coordinate points. At the same time, summarize the clustering coordinate points obtained from this analysis into a clustering group, and they will no longer be used as data weight points for subsequent analysis of clustering coordinate points;

[0026] Reconstruct the correlation circle to analyze the clustering group until all data weight points are transformed into clustering coordinate points and then stop the analysis. At this time, several clustering groups are obtained. Mark the number of clustering coordinate points in the clustering group as the clustering quantity. Judge whether the clustering quantity is less than or equal to the first error quantity. If so, output a clustering error signal; if not, output a clustering correct signal;

[0027] If a clustering error signal is output, dissolve this clustering group, mark the clustering coordinate points in it as free coordinate points, and analyze each clustering group to obtain several free coordinate points;

[0028] Analyze the remaining clustering groups through the mean shift clustering algorithm, and mark the analyzed center point as the weight clustering point.

[0029] Further, calculating the clustering value of the physical sign data based on the weight clustering point, and then clustering and marking the physical sign data according to the clustering values of different physical sign data includes the following sub-steps:

[0030] Calculate the clustering value of the physical sign data based on the weight clustering point, then sort and number the clustering values of different physical sign data to obtain a clustering queue;

[0031] Perform clustering marking on the physical sign data based on the clustering queue.

[0032] Further, calculate the clustering value of the physical sign data based on the weighted clustering points, and then sort and number the clustering values of different physical sign data to obtain the clustering queue, including the following sub-steps:

[0033] For any physical sign data, calculate the average value of the weighted clustering points of the physical sign data, and mark the calculation result as the clustering value;

[0034] Sort and number the clustering values of all physical sign data in ascending order, and represent them by the symbol P n where n is a non-zero natural number and n is the serial number of P. Mark the maximum value of n as max(n). After sorting and numbering, the clustering queue is obtained.

[0035] Further, clustering and marking the physical sign data based on the clustering queue includes the following sub-steps:

[0036] Take n of P in the clustering queue as the horizontal axis and the clustering value as the vertical axis to establish a plane rectangular coordinate system that only includes the first quadrant, named the classification marking graph; n Take n of P in the clustering queue as the horizontal axis and the clustering value as the vertical axis to establish a plane rectangular coordinate system that only includes the first quadrant, named the classification marking graph;

[0037] Analyze the classification marking graph through the mean shift clustering algorithm, mark the obtained center points as classification marking points, and obtain several classification marking points through analysis;

[0038] Sort and number the classification marking points in ascending order, and represent them by the symbol T i where i is a non-zero natural number and i is the serial number of T. Mark the maximum value of i as max(i), and establish max(i) clustering groups, which are respectively named the i-th group and represented by the symbol R i ;

[0039] Starting from n = 1 and i = 1, calculate |P n -T i |, and mark the calculation result as S(n, i), where (n, i) is the serial number of S. Judge whether i is equal to max(i). If not, then set i + 1 and calculate S(n, i) again. If not, then set n + 1 and reset i to 1, and recalculate S(n, i);

[0040] Obtain several S(n, i) through calculation. Starting from n = 1, find the minimum value in S(n, i), mark the i among them as h, and at the same time include P n into R h , where R h represents R when i = h i ; Judge whether n is equal to max(n). If so, stop searching. If not, then set n + 1 and re-search and group;

[0041] Calculate max(i) / 2, mark the calculation result as f, and mark the R where i≥f i as important physical sign data, and mark the R where i<f i as minor physical sign data.

[0042] In a second aspect, the present application provides an electronic device, including a processor and a memory. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the steps in the above method are run.

[0043] In a third aspect, the present application provides a storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method are run.

[0044] Advantages of the present invention: By establishing a physical sign data record database based on a large dataset to store each data record, and then analyzing and statistically calculating the weight set of physical sign data based on the physical sign data record database, and then analyzing the weight set, clustering groups are obtained after excluding free coordinate points, and then the clustering algorithm is used to analyze the clustering groups to obtain the weight clustering points of physical sign data. The advantage is that the free coordinate points in the data weight distribution diagram are removed, and then the clustering algorithm is used to analyze the clustering groups, which can effectively improve the accuracy of clustering analysis, exclude the influence of small-proportion points on the clustering analysis result, and improve the accuracy and rationality of weight clustering;

[0045] The present invention calculates the clustering value of physical sign data based on the weight clustering points, then sorts and numbers the clustering values of different physical sign data to obtain a clustering queue, and finally performs clustering marking on the physical sign data based on the clustering queue. The advantage is that through the previous clustering analysis, the weight clustering points of different physical sign data are obtained. The weight clustering points unify different physical sign data, and are no longer limited by different units and standards between different physical sign data. And because the physical sign data in the human body is complex and diverse, there are the same physical sign data in different organs, but the influence of some physical sign data in most organs or tissues is small. If not marked, there may be a situation where medical staff spend time and effort monitoring it, but with little effect. By performing clustering marking on all physical sign data, screening out important physical sign data and minor physical sign data, medical staff can reduce the monitoring of minor physical sign data and spend more energy and resources monitoring important physical sign data, improving the accuracy and effectiveness of weight clustering marking. Description of the Drawings

[0046] Figure 1 is a flowchart of the steps of the method of the present invention;

[0047] Figure 2Data weight distribution diagram of the present invention;

[0048] Figure 3 Schematic diagram of the associated circle and associated points of the present invention;

[0049] Figure 4 Schematic diagram of the distribution of clustering groups of the present invention;

[0050] Figure 5 Data weight distribution diagram of the present invention after removing free coordinate points;

[0051] Figure 6 Schematic diagram of the weight clustering points of the present invention;

[0052] Figure 7 Classification marker diagram of the present invention;

[0053] Figure 8 Schematic diagram of the classification marker points of the present invention. Detailed implementation manners

[0054] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0055] Example 1, please refer to Figure 1 As shown, the present application provides a data weight clustering and marking method based on a large data set, including the following steps:

[0056] Step S1, based on the large data set, statistically analyze the weight set of the physical sign data; Step S1 includes the following sub-steps:

[0057] Step S101, based on the large data set, establish a physical sign data record database to store each data record;

[0058] Step S101 includes the following sub-steps:

[0059] Step S101.1, for any physical sign data, establish a physical sign data record database. The physical sign data includes various data for evaluating the health status of different organs in the human body, such as respiratory rate, blood pressure, alanine aminotransferase, alkaline phosphatase, etc. A corresponding physical sign data record database is established for each physical sign data;

[0060] In specific implementation, the physical sign data in this embodiment are monitoring data for evaluating human health in various tissues and organs of the human body. For example, the respiratory rate can be used to evaluate the patient's lung health status and the patient's overall health status. For example, blood pressure can evaluate the health status of the patient's heart, and alanine aminotransferase and alkaline phosphatase can be used to evaluate the health status of the patient's liver. Some physical sign data can evaluate the health status of multiple tissues or organs in the patient's body at the same time. When the physical sign data is applied in different aspects, the data weight will also change accordingly. Many physical sign data have little impact in other parts except in special tissues or organs and are not sufficient to reflect the health status of the human body. However, medical staff may spend time and energy monitoring them, wasting human and material resources. This embodiment aims to find out the physical sign data that play an important role in the human body through clustering and marking, so that medical staff can spend more human and material resources monitoring important physical sign data and improve the effect of patient physical sign monitoring;

[0061] Step S101.2, the data records stored in the physical sign data record database include a data number, a data weight, a data value, and an evaluation value;

[0062] The data number is the sorting number of the data record in the physical sign data record database; the data weight is the weight corresponding to the physical sign data in the field where this data record acts; the data value is the value of the physical sign data of the patient monitored by this data record; the evaluation value is the result finally calculated by the health assessment calculation participated by the data record;

[0063] In specific implementation, taking the physical sign data glucose as an example, part of the data of the physical sign data record database of glucose constructed based on the large dataset is shown in Table 1 below:

[0064] Table 1 Part of the data of the physical sign data record database of glucose

[0065]

[0066] When physical sign data is used to evaluate the user's health status, in addition to the basic threshold comparison, an evaluation value for evaluating the patient's health status will also be calculated through different health index calculation formulas. When evaluating different tissues, organs or physiological characteristics of the patient, the health index calculation formulas are all different. Therefore, there are also different data weights for physical sign data. The data in the physical sign data record database are directly obtained by connecting with the large dataset of the hospital, and the large dataset stores a huge amount of medical staff data;

[0067] Step S102, analyze and statistically calculate the weight set of physical sign data based on the physical sign data record database;

[0068] Step S102 includes the following sub-steps:

[0069] Step S102.1: Calculate the data value divided by the evaluation value, and mark the calculation result as the data ratio.

[0070] Step S102.2: Mark the data weight and the data ratio in a data record as a set of physical sign weight groups.

[0071] Step S102.3: Statistically analyze all the physical sign weight groups in the physical sign data record database of the same physical sign, and integrate them into a weight set.

[0072] Step S102.4: Each physical sign data has a corresponding weight set.

[0073] In a specific implementation, taking Table 1 as an example, in the data No. 100001, the data value is 3.9 and the evaluation value is 36.79. By calculation, the data ratio is 0.106, that is, 10.6%. The calculation result is reserved to three decimal places. In the data No. 100001, the data weight is 0.33, that is, (0.33, 10.6%) is a set of physical sign weight groups. Statistically analyze the physical sign weight groups of each data in the physical sign data record database of glucose, and integrate to obtain a weight set. Since the weight set obtained in this embodiment is the weight set of glucose, it is named the glucose weight set. The weight sets of different physical sign data are independent of each other. The glucose weight set includes several physical sign weight groups such as (0.33, 10.6%), (0.34, 14.2%), and (0.58, 12.1%).

[0074] Step S2: Analyze the weight set through a clustering algorithm to obtain the weight clustering points of the physical sign data. Step S2 includes the following sub-steps:

[0075] Please refer to Figure 2 As shown, in Step S201, establish a plane rectangular coordinate system that only includes the first quadrant with the physical sign weight as the X-axis and the data ratio as the Y-axis, and name it the data weight distribution diagram. Each weight set constructs an independent data weight distribution diagram.

[0076] Please refer to Figure 3 As shown, in Step S202, enter all the physical sign weight groups in the same weight set into the data weight distribution diagram, name the coordinate points therein as data weight points, take any data weight point as the center of the circle, and the first correlation distance as the radius to construct a circle, named the correlation circle. Check whether there are data weight points within the correlation circle. If there are, output a data correlation signal. If not, output a data non-correlation signal.

[0077] Step S203: If a data correlation signal is output, mark the data weight points within the correlation circle as correlation points, then take the correlation points as the center of the circle and the first correlation distance as the radius to construct a correlation circle, and judge again.

[0078] Step S204, when the data association signal is no longer output, mark all the data weight points and associated points that are used as the center of the circle as clustering coordinate points. At the same time, classify the clustering coordinate points obtained from this analysis into a clustering group, and they will no longer be used as data weight points for subsequent analysis of clustering coordinate points;

[0079] In specific implementation, the data weight distribution diagram is as Figure 2 shown. Figure 2 Each coordinate point in it is a physical sign weight recombination group, Figure 2 which is only the data weight distribution diagram of the physical sign data "glucose"; the first association distance is obtained from historical analysis. The function of the first association distance is to perform preprocessing before clustering analysis on the data weight distribution diagram to find the free coordinate points in the data weight distribution diagram, so as to achieve the purpose of improving the accuracy of clustering analysis; the setting range of the first association distance is between 0.02 and 0.1, and preferably the first association distance is set to 0.05. The relationship between the association circle and the associated points is as Figure 3 shown. Figure 3 In it, the black circles are the data weight points within the association circle, that is, the associated points, and the data weight points outside the association circle are not associated points;

[0080] Please refer to Figure 4 shown. Step S205, reconstruct the association circle to analyze the clustering group until all the data weight points are transformed into clustering coordinate points and then stop the analysis. At this time, several clustering groups are obtained. Mark the number of clustering coordinate points in the clustering group as the clustering number, and judge whether the clustering number is less than or equal to the first error number. If so, output a clustering error signal; if not, output a clustering correct signal;

[0081] Step S206, if a clustering error signal is output, dissolve this clustering group, mark the clustering coordinate points in it as free coordinate points, and analyze each clustering group to obtain several free coordinate points;

[0082] Please refer to Figures 5 to 6 shown. Step S207, analyze the remaining clustering groups through the mean shift clustering algorithm, and mark the obtained center point as the weight clustering point;

[0083] In specific implementation, usually, at least 5 clustering coordinate points are required in a clustering group to form a group. Therefore, the first error number is set to 5; the distribution of the clustering groups is as Figure 4 shown. Figure 4 In it, the three largest clustering groups are marked with different colors respectively, and the remaining uncolored coordinate points are all free coordinate points. By removing the free coordinate points, directly use the existing mean shift clustering algorithm for Figure 5Performing clustering analysis can significantly improve the processing speed and accuracy of clustering analysis; by analyzing, weighted clustering points are obtained as Figure 6 shown, where the weighted clustering points are 0.2895, 0.6263, and 0.7206 respectively. Retaining four decimal places, the weighted clustering points only consider the values on the X-axis and do not consider the values on the Y-axis.

[0084] Step S3: Calculate the clustering value of the physical sign data based on the weighted clustering points, and then perform clustering marking on the physical sign data according to the clustering values of different physical sign data; Step S3 includes the following sub-steps:

[0085] Step S301: Calculate the clustering value of the physical sign data based on the weighted clustering points, then sort and number the clustering values of different physical sign data to obtain a clustering queue;

[0086] Step S301 includes the following sub-steps:

[0087] Step S301.1: For any physical sign data, calculate the average value of the weighted clustering points of the physical sign data, and mark the calculation result as the clustering value;

[0088] Step S301.2: Sort and number the clustering values of all physical sign data in ascending order, represented by the symbol P n where n is a non-zero natural number and n is the serial number of P. Mark the maximum value of n as max(n). After sorting and numbering, a clustering queue is obtained;

[0089] In a specific implementation, the calculated clustering value is 0.5455, and the calculation result is retained to four decimal places. The clustering value calculated in this embodiment is the glucose clustering value. When sorting, it is necessary to integrate the clustering values of all physical sign data for sorting and numbering. Through sorting and numbering, P1 to P 88 are obtained, where 1 ≤ n ≤ 88 and max(n) = 88;

[0090] Step S302: Perform clustering marking on the physical sign data based on the clustering queue;

[0091] Step S302 includes the following sub-steps:

[0092] Please refer to Figures 7 to 8 shown. Step S302.1: Establish a plane rectangular coordinate system that only includes the first quadrant with n in P in the clustering queue as the horizontal axis and the clustering value as the vertical axis, and name it the classification marking graph; n Step S302.2: Analyze the classification marking graph through the mean shift clustering algorithm, mark the obtained center point as the classification marking point, and obtain several classification marking points through analysis;

[0093] Step S302.2: Analyze the classification marking graph through the mean shift clustering algorithm, mark the obtained center point as the classification marking point, and obtain several classification marking points through analysis;

[0094] Step S302.3, sort the classification marker points in ascending order and number them, denoted by symbol T i where i is a non-zero natural number and i is the serial number of T. Mark the maximum value of i as max(i), and establish max(i) clustering groups, named the i-th group respectively, denoted by symbol R i ;

[0095] In specific implementation, the constructed classification marker graph is as shown in Figure 7 . Analyze the classification marker graph through the mean shift clustering algorithm to obtain the classification marker points as shown in Figure 8 . Obtain T1 to T6 through sorting and numbering, 1 ≤ i ≤ 6, max(i) = 6, and establish 6 clustering groups, namely the first clustering group, the second clustering group, the third clustering group, the fourth clustering group, the fifth clustering group, and the sixth clustering group, that is, R1 to R6;

[0096] Step S302.4, start from n = 1 and i = 1, calculate |P n -T i |, mark the calculation result as S(n,i), where (n,i) is the serial number of S. Judge whether i is equal to max(i). If not, then set i + 1 and calculate S(n,i) again. If not, then set n + 1 and reset i to 1, and recalculate S(n,i);

[0097] Step S302.5, obtain several S(n,i) through calculation. Start from n = 1, find the minimum value in S(n,i), mark the i among them as h, and at the same time incorporate P n into R h , where R h represents R i with i = h. Judge whether n is equal to max(n). If so, stop searching. If not, then set n + 1 and re-perform searching and grouping;

[0098] Step S302.6, calculate max(i) / 2, mark the calculation result as f, mark the R i with i ≥ f as the important physical sign data, and mark the R i with i < f as the secondary physical sign data;

[0099] In a specific implementation, in this embodiment, P1 is 0.0738 and T1 is 0.1024. By calculation, S(1,1) = 0.0286 is obtained. For P1, S(1,1), S(1,2), S(1,3), S(1,4), S(1,5), and S(1,6) are calculated. By comparison, S(1,1) is the smallest. At this time, i = 1, that is, h = 1. P1 is incorporated into R1, that is, P1 is incorporated into the first clustering group. The corresponding physical sign data of P1 is growth hormone. That is, the physical sign data "growth hormone" is classified into the first clustering group. And so on, analyze the clustering groups of P2 to P 88 ; Calculate max(i) / 2 = f = 3. Mark the physical sign data in R3, R4, R5, and R6 as important physical sign data, and mark the physical sign data in R1 and R2 as secondary physical sign data; When medical staff monitor the health status of patients, they focus on monitoring important physical sign data.

[0100] Embodiment 2. This embodiment provides an electronic device, which may include: a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus. The memory stores computer-readable instructions, and the processor can call the instructions in the memory. When the computer-readable instructions are executed by the processor, the steps in the data weight clustering and marking method based on a large dataset are run to achieve the following functions: Based on a large dataset, count the weight set of physical sign data; Analyze the weight set through a clustering algorithm to obtain the weight clustering points of physical sign data; Calculate the clustering values of physical sign data based on the weight clustering points, and then cluster and mark the physical sign data according to the clustering values of different physical sign data.

[0101] In addition, when the logical instructions in the above-mentioned memory can be implemented in the form of software function units and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. And the aforementioned storage medium includes: USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs, etc., which can store program codes.

[0102] Embodiment 3. The present application further provides a computer program product. The computer program product includes a computer program stored on a computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the data weight clustering and marking method based on a large data set provided by each of the above methods. The method includes: based on the large data set, statistically obtaining a weight set of physical sign data; analyzing the weight set through a clustering algorithm to obtain weight clustering points of the physical sign data; calculating a clustering value of the physical sign data based on the weight clustering points, and then clustering and marking the physical sign data according to the clustering values of different physical sign data.

[0103] Embodiment 4. The present application further provides a computer-readable storage medium. The present application provides a storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps in the above data weight clustering and marking method based on a large data set are run to achieve the following functions: based on the large data set, statistically obtaining a weight set of physical sign data; analyzing the weight set through a clustering algorithm to obtain weight clustering points of the physical sign data; calculating a clustering value of the physical sign data based on the weight clustering points, and then clustering and marking the physical sign data according to the clustering values of different physical sign data.

[0104] Through the description of the above embodiments, the embodiments of the present invention can be provided as a method, a system or a computer program product. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0105] In the embodiments provided by the present application, it should be understood that the disclosed system or method can be implemented in other ways. The above-described embodiments are merely illustrative. For example, the division of modules or units is only a logical function division, and there may be other division methods in actual implementation. For another example, multiple modules or units can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of systems, modules, and units can be in an electrical, mechanical, or other form.

[0106] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.

Claims

1. A method for clustering and labeling data weights based on a large dataset, characterized in that, It includes the following steps: Based on a large dataset, statistically calculate the weight set of physical sign data; Analyze the weight set through a clustering algorithm to obtain the weight clustering points of physical sign data; Calculate the clustering value of physical sign data based on the weight clustering points, and then perform clustering marking on the physical sign data according to the clustering values of different physical sign data; Calculate the clustering value of physical sign data based on the weight clustering points, then sort and number the clustering values of different physical sign data to obtain a clustering queue, which includes the following sub-steps: For any physical sign data, calculate the average value of the weight clustering points of the physical sign data, and mark the calculation result as the clustering value; Sort and number the clustering values of all physical sign data in ascending order, denoted by the symbol P n where n is a non-zero natural number and n is the serial number of P. Mark the maximum value of n as max(n). After sorting and numbering, a clustering queue is obtained; Performing clustering marking on the physical sign data based on the clustering queue includes the following sub-steps: With the n of P in the clustering queue n as the horizontal axis and the clustering value as the vertical axis, a rectangular coordinate system including only the first quadrant is established and named the classification marker graph; Analyze the classification marking graph through a mean shift clustering algorithm, mark the obtained center points as classification marking points, and obtain several classification marking points through analysis; Sort and number the classification marker points in ascending order, and represent them by the symbol T i where i is a non-zero natural number and i is the serial number of T. Mark the maximum value of i as max(i), and establish max(i) clustering groups, which are respectively named the i-th group and represented by the symbol R i Indicate; Start from n = 1 and i = 1, calculate |P n -T i |, mark the calculation result as S(n, i), where (n, i) is the serial number of S. Determine whether i is equal to max(i). If not, then increment i by 1 and calculate S(n, i) again. If not, then increment n by 1 and reset i to 1, and recalculate S(n, i); A number of S(n, i) are obtained through calculation. Starting from n = 1, the minimum value in S(n, i) is found, and the corresponding i is marked as h. At the same time, P n is included in R h , where R h represents the R when i = h i ; Determine whether n is equal to max(n). If so, stop the search. If not, increment n by 1 and restart the search and grouping; Calculate max(i) / 2, mark the calculation result as f, and mark R where i≥f i as important physical sign data, and mark R where i<f i as secondary physical sign data.

2. The method for clustering and marking data weights based on a large dataset according to claim 1, wherein, Based on a large dataset, statistically calculating the weight set of physical sign data includes the following sub-steps: Based on a large dataset, establish a physical sign data record database to store each data record; Analyze and statistically calculate the weight set of physical sign data based on the physical sign data record database.

3. The data weight clustering and marking method based on a large data set according to claim 2, wherein Based on a large dataset, establishing a physical sign data record database to store each data record includes the following sub-steps: For any physical sign data, establish a physical sign data record database, and each physical sign data has a corresponding physical sign data record database; The data records stored in the physical sign data record database include data numbers, data weights, data values, and evaluation values; The data number is the sorting number of the data record in the physical sign data record database; the data weight is the weight corresponding to the physical sign data in the field where this data record acts; the data value is the value of the physical sign data of the patient monitored by this data record; the evaluation value is the final calculation result obtained from the health assessment calculation of the patient participated in by the data record.

4. The data weight clustering and marking method based on a large data set according to claim 3, wherein, Analyzing and statistically calculating the weight set of physical sign data based on the physical sign data record database includes the following sub-steps: Calculate the data value divided by the evaluation value, and mark the calculation result as the data proportion; Mark the data weight and data proportion in a data record as a set of physical sign weight groups; Statistically calculate all the physical sign weight groups in the same physical sign data record database, and integrate them into a weight set; Each physical sign data has a corresponding weight set.

5. The data weight clustering and marking method based on a large data set according to claim 4, characterized in that Analyzing the weight set through a clustering algorithm to obtain the weight clustering points of physical sign data includes the following sub-steps: Establish a plane rectangular coordinate system containing only the first quadrant with the physical sign weight as the X-axis and the data proportion as the Y-axis, and name it the data weight distribution graph. Each weight set constructs an independent data weight distribution graph; Enter all the physical sign weight groups within the same weight set into the data weight distribution graph, name the coordinate points among them as data weight points, use any data weight point as the center of the circle, and use the first correlation distance as the radius to construct a circle, named the correlation circle. Check whether there are data weight points within the correlation circle. If there are, output a data correlation signal; if not, output a data non-correlation signal; If the output data is associated with a signal, mark the data weight points within the associated circle as associated points, then construct an associated circle with the associated points as the center and the first associated distance as the radius, and make a judgment again; When the data association signal is no longer output, mark all the data weight points serving as the center and the associated points as clustering coordinate points, and at the same time, summarize the clustering coordinate points obtained from this analysis into a clustering group, and they will no longer be used as data weight points for subsequent analysis of clustering coordinate points; Reconstruct the associated circle to analyze the clustering group until all data weight points are converted into clustering coordinate points and then stop the analysis. At this time, several clustering groups are obtained. Mark the number of clustering coordinate points in the clustering group as the clustering quantity, and judge whether the clustering quantity is less than or equal to the first error quantity. If so, output a clustering error signal; if not, output a clustering correct signal; If a clustering error signal is output, dissolve this clustering group and mark the clustering coordinate points therein as free coordinate points, and analyze each clustering group to obtain several free coordinate points; Analyze the remaining clustering groups through the mean shift clustering algorithm, and mark the obtained center points as weight clustering points.

6. The method for clustering and marking data weights based on a large dataset according to claim 5, wherein, Calculate the clustering value of the physical sign data based on the weight clustering points, and then cluster and mark the physical sign data according to the clustering values of different physical sign data, including the following sub-steps: Calculate the clustering value of the physical sign data based on the weight clustering points, then sort and number the clustering values of different physical sign data to obtain a clustering queue; Cluster and mark the physical sign data based on the clustering queue.

7. An electronic device, characterized in that, It includes a processor and a memory. The memory stores computer-readable instructions. When the computer-readable instructions are executed by the processor, the steps in the method according to any one of claims 1-6 are run.

8. A storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps in the method according to any one of claims 1-6 are run.

Citation Information

Patent Citations

  • Disease risk factor extraction method based on attribute weight clustering

    CN109978007A

  • Evaluation index weight data processing method based on clustering analysis

    CN117332287A