Multi-center collaborative collection and management method and system for student physical fitness test data

By dividing data into shards and constructing a spatiotemporal knowledge graph in a multi-center environment, and adopting a federated learning framework for collaborative computing, the problems of data surge and privacy protection in student physical fitness test data management are solved, and efficient data analysis and privacy protection are achieved.

CN120528947BActive Publication Date: 2025-09-23ZHEJIANG INT STUDIES UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511025320.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-24
Publication Date
2025-09-23
Estimated Expiration
2045-07-24

AI Technical Summary

Technical Problem

Existing student physical fitness test data management technology is unable to cope with the problems of data surge and load imbalance in multi-center collaborative collection scenarios. It lacks a data integration mechanism, cannot effectively capture the dynamic changes in students' physical fitness development, and data privacy protection is imperfect.

Method used

By dividing physical test data into multiple data shards, dynamically allocating them based on spatiotemporal proximity and storage node capacity information, building a spatiotemporal knowledge graph and using graph neural networks for feature extraction, and utilizing a federated learning framework to build local training models among multiple storage nodes, collaborative computing and analysis can be performed while ensuring data privacy.

Benefits of technology

It achieves efficient management and storage balance of massive student physical fitness test data, improves data access speed and the accuracy of analysis results, protects students' personal privacy and data security, and meets the strict requirements of the education field for data privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120528947B_ABST
    Figure CN120528947B_ABST
Patent Text Reader

Abstract

This invention provides a multi-center collaborative collection and management method and system for student physical fitness test data. This technology involves acquiring data uploaded by multiple physical fitness test centers and allocating it in shards based on storage node capacity information; constructing a spatiotemporal knowledge graph to map test information; extracting physical fitness feature vectors using a graph neural network; and employing a federated learning framework to build a global model, generate analysis reports, and distribute them. This invention enables efficient storage and collaborative analysis of multi-center data, ensuring data privacy and security, and improving data processing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field, and in particular to a method and system for multi-center collaborative collection and management of student physical fitness test data. Background Art

[0002] With the development of educational informatization and the era of big data, the collection, management, and analysis of student physical fitness test data have become an essential component of educational evaluation systems. Traditional student physical fitness testing relies primarily on manual record-keeping and single-center data storage. As testing scale expands and the number of testing centers increases, physical fitness test data has become multi-source, heterogeneous, and large-scale. Currently, schools and physical fitness testing centers across the country generate massive amounts of physical fitness test data annually. This data not only includes students' basic physical indicators but also covers various physical fitness test results, with distinct temporal and spatial distribution characteristics. How to efficiently manage this data, which is dispersed across different testing centers, and to uncover valuable insights into physical fitness development patterns, has become a key issue in the current field of education and health.

[0003] Existing physical fitness test data management technologies have the following shortcomings: First, traditional data storage methods mostly use a centralized architecture, which is difficult to cope with the surge in data volume and unbalanced load in multi-center collaborative collection scenarios. The data of each test center is often stored in isolation, lacking an effective data integration mechanism, resulting in the inability to fully utilize data resources. Second, existing physical fitness data analysis methods lack in-depth exploration of the spatiotemporal characteristics of data, and are unable to effectively capture the dynamic changes in the physical fitness development of students in different regions and different time periods, making it difficult to form a comprehensive and objective physical fitness evaluation system. In addition, during the multi-center data collaborative analysis process, the data privacy protection mechanism is imperfect, and data sharing between test centers poses security risks, making it difficult to maximize data value while protecting personal privacy.

[0004] With the development of artificial intelligence and distributed computing technologies, there is an urgent need for a technical solution that can efficiently manage multi-center physical fitness test data and conduct collaborative analysis while protecting data privacy, so as to support more scientific and comprehensive student physical health evaluation and intervention decisions. Summary of the Invention

[0005] The embodiments of the present invention provide a multi-center collaborative collection and management method and system for student physical fitness test data, which can solve the problems in the prior art.

[0006] A first aspect of an embodiment of the present invention provides a multi-center collaborative collection and management method for student physical fitness test data, comprising:

[0007] Obtain physical fitness test data uploaded by multiple physical fitness test centers, divide the physical fitness test data into multiple data shards, obtain capacity information of each storage node in the distributed storage cluster, calculate the relative load difference value between the storage nodes based on the capacity information of the storage nodes, trigger the shard migration process when the relative load difference value exceeds a preset difference threshold, and distribute the data shards to multiple storage nodes in the distributed storage cluster according to their spatiotemporal proximity;

[0008] Constructing a spatiotemporal knowledge graph, mapping the test subject information and test item information of the physical fitness test data into entity nodes in the spatiotemporal knowledge graph, mapping the test result information into attributes of the entity nodes, and mapping the spatiotemporal association indexes into relationship edges between the entity nodes;

[0009] Using a graph neural network to perform feature extraction and representation learning on the spatiotemporal knowledge graph, obtaining the physical feature vectors of the test subject under different spatiotemporal conditions;

[0010] Based on the physical characteristic vector, a federated learning framework is used to construct a local training model between multiple storage nodes. While ensuring data privacy, the data of different storage nodes are collaboratively calculated to construct a global model. Based on the global model, a comprehensive data analysis report of the physical test data is generated, and the comprehensive data analysis report is distributed to terminal devices of multiple physical test centers through an encrypted transmission channel.

[0011] Dividing the physical fitness test data into multiple data shards, obtaining capacity information of each storage node in the distributed storage cluster, calculating relative load difference values ​​between the storage nodes based on the capacity information of the storage nodes, and triggering a shard migration process when the relative load difference value exceeds a preset difference threshold includes:

[0012] Obtain capacity information for each storage node in the distributed storage cluster, construct a hierarchical consistent hash ring based on the capacity information, and calculate a virtual node hash value for each storage node. The virtual node hash value is obtained by applying a power function to the ratio of the storage node capacity value to the system average capacity value, forming an initial shard distribution plan.

[0013] Collecting access behavior data of data shards according to the initial shard distribution scheme, the access behavior data including access frequency data, data access volume data, and access time distribution data, and calculating relative load difference values ​​between storage nodes based on the access behavior data and the capacity information of the storage nodes;

[0014] When the relative load difference value exceeds a preset difference threshold, the shard migration process is triggered.

[0015] Allocating the data shards to a plurality of storage nodes in a distributed storage cluster according to spatiotemporal proximity includes:

[0016] Calculate the time dimension similarity between data slices, the time dimension similarity is calculated based on the timestamp difference of the data slices, and calculate the space dimension similarity between data slices, the space dimension similarity is calculated based on the coordinate distance of the data slices in the multidimensional space;

[0017] The temporal dimension similarity and the spatial dimension similarity are weighted and combined according to a preset spatiotemporal weight factor to obtain a comprehensive proximity between the data slices; the comprehensive proximity greater than a preset proximity threshold is retained, and the comprehensive proximity less than the preset proximity threshold is set to zero, thereby constructing a proximity matrix; based on the proximity matrix, the boundary positions of the data slices are optimized to maximize the proximity between the internal data points of the data slices and minimize the proximity of the data points between the data slices;

[0018] Calculate the storage capacity utilization rate, computing resource utilization rate, and network bandwidth utilization rate of each storage node respectively, and weight the three utilization rates according to the preset resource weight coefficient to obtain the comprehensive load value of the storage node;

[0019] When it is detected that the comprehensive load value exceeds a preset data migration threshold, the data shards on the storage node are reallocated to other storage nodes with lower loads.

[0020] After constructing the spatiotemporal knowledge graph, the method further includes:

[0021] Obtain a set of neighbor nodes of a node to be processed in a spatiotemporal knowledge graph, wherein the set of neighbor nodes includes nodes that are directly associated with the node to be processed; calculate the similarity between the attribute feature vector of the node to be processed and the attribute feature vector of each neighbor node in the set of neighbor nodes to obtain a node similarity value;

[0022] Performing exponential transformation and normalization on the node similarity values ​​to obtain similarity weights, performing weighted summation on the attribute values ​​of the neighboring nodes based on the similarity weights, and generating missing attribute supplementary values ​​for the node to be processed;

[0023] Obtain historical attribute values ​​of nodes in the spatiotemporal knowledge graph, calculate attribute changes at adjacent time points, and calculate the ratio of the attribute changes to the historical average attribute value, and adjust the update frequency of the spatiotemporal knowledge graph based on the ratio.

[0024] Based on the physical characteristic vector, a local training model is constructed among multiple storage nodes using a federated learning framework. Under the premise of ensuring data privacy, data from different storage nodes are collaboratively calculated to construct a global model. A comprehensive data analysis report of the physical test data generated based on the global model includes the following:

[0025] Adding Laplace noise to the physical feature vector, and homomorphically encrypting the physical feature vector after adding the noise using a public key to generate an encrypted feature vector;

[0026] Building a local training model at each storage node based on the encrypted feature vector, calculating the number of samples and the data quality score of each storage node, and determining the aggregation weight of the storage node according to the number of samples and the data quality score;

[0027] Distributing key pairs among multiple storage nodes, generating a random mask based on the key pairs, and superimposing the random mask with parameters of the local training model in combination with aggregation weights to obtain masked model parameters;

[0028] Collect the masked model parameters of all storage nodes, perform security aggregation operations to obtain global model parameters, and distribute the global model parameters to each storage node; generate a comprehensive data analysis report of the physical fitness test data based on the global model.

[0029] Key pairs are distributed among multiple storage nodes. A random mask is generated based on the key pairs. The random mask is superimposed on the parameters of the local training model in combination with the aggregation weights to obtain the masked model parameters, including:

[0030] Determine the public key of each storage node and distribute the public key among the storage nodes; generate an initial shared key between the storage node pair through modular exponentiation based on the private key of each storage node and the public keys of other storage nodes; perform a cryptographic hash operation on the initial shared key to obtain a final shared key;

[0031] Based on the final shared key between the storage node pair and the current calculation round identifier, a random noise seed is generated through a pseudo-random function. The sum of the noise seed differences between each storage node and other storage nodes is calculated to obtain the noise accumulation value.

[0032] Inputting the noise accumulation value into a deterministic random expansion function to generate a mask matrix of the storage node;

[0033] The local model parameters of each storage node are multiplied by the aggregation weight, and the product is added to the mask matrix of the storage node to obtain the masked model parameters.

[0034] A second aspect of an embodiment of the present invention provides a multi-center collaborative collection and management system for student physical fitness test data, including:

[0035] The first unit is configured to obtain physical fitness test data uploaded by multiple physical fitness test centers, divide the physical fitness test data into multiple data slices, and distribute the data slices to multiple storage nodes in a distributed storage cluster according to spatiotemporal proximity;

[0036] The second unit is used to construct a spatiotemporal knowledge graph, map the test subject information and test item information of the physical test data into entity nodes in the spatiotemporal knowledge graph, map the test result information into attributes of the entity nodes, and map the spatiotemporal association indexes into relationship edges between the entity nodes;

[0037] The third unit is used to obtain the physical characteristic vectors of the test subject under different time and space conditions;

[0038] The fourth unit is used to build a local training model between multiple storage nodes based on the physical feature vector using a federated learning framework, perform collaborative calculations on the data of different storage nodes to generate a comprehensive data analysis report of the physical test data, and distribute the comprehensive data analysis report to terminal devices of multiple physical test centers through an encrypted transmission channel.

[0039] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including:

[0040] processor;

[0041] a memory for storing processor-executable instructions;

[0042] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0043] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0044] The beneficial effects of this application are as follows:

[0045] The present invention achieves efficient management and storage balance of massive student physical fitness test data by dividing physical fitness test data into multiple data slices and storing them in a distributed cluster according to temporal and spatial proximity. The system can dynamically adjust data distribution according to the capacity information of storage nodes, ensure the balanced utilization of storage resources, and improve data access speed and system stability.

[0046] By constructing a spatiotemporal knowledge graph and applying graph neural networks for feature extraction and representation learning, the present invention achieves deep semantic understanding and association analysis of students' physical fitness data, can capture the changing patterns of physical characteristics under different spatiotemporal conditions, and provide rich feature representations for subsequent data mining and analysis, thereby improving the accuracy and interpretability of the analysis results.

[0047] This invention adopts a federated learning framework to collaboratively train models among multiple storage nodes, realizing collaborative analysis of multi-center data without sharing original data, effectively protecting students' personal privacy and data security, and ensuring the secure distribution of analysis reports through encrypted transmission channels, meeting the strict requirements of the education field for data privacy protection and providing reliable technical support for school physical health management. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 This is a flow chart of a multi-center collaborative collection and management method for student physical fitness test data according to an embodiment of the present invention;

[0049] Figure 2 This is a logical block diagram of allocating data shards to a distributed storage cluster based on temporal and spatial proximity according to an embodiment of the present invention. DETAILED DESCRIPTION

[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0051] The following specific embodiments are used to describe the technical solution of the present invention in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0052] Figure 1 FIG. 1 is a flow chart of a multi-center collaborative collection and management method for student physical fitness test data according to an embodiment of the present invention. Figure 1 As shown, the method includes:

[0053] Obtain physical fitness test data uploaded by multiple physical fitness test centers, divide the physical fitness test data into multiple data shards, obtain capacity information of each storage node in the distributed storage cluster, calculate the relative load difference value between the storage nodes based on the capacity information of the storage nodes, trigger the shard migration process when the relative load difference value exceeds a preset difference threshold, and distribute the data shards to multiple storage nodes in the distributed storage cluster according to their spatiotemporal proximity;

[0054] Constructing a spatiotemporal knowledge graph, mapping the test subject information and test item information of the physical fitness test data into entity nodes in the spatiotemporal knowledge graph, mapping the test result information into attributes of the entity nodes, and mapping the spatiotemporal association indexes into relationship edges between the entity nodes;

[0055] Using a graph neural network to perform feature extraction and representation learning on the spatiotemporal knowledge graph, obtaining the physical feature vectors of the test subject under different spatiotemporal conditions;

[0056] Based on the physical characteristic vector, a federated learning framework is used to construct a local training model between multiple storage nodes. While ensuring data privacy, the data of different storage nodes are collaboratively calculated to construct a global model. Based on the global model, a comprehensive data analysis report of the physical test data is generated, and the comprehensive data analysis report is distributed to terminal devices of multiple physical test centers through an encrypted transmission channel.

[0057] In an optional embodiment, the physical fitness test data is divided into multiple data shards, capacity information of each storage node in the distributed storage cluster is obtained, and relative load difference values ​​between the storage nodes are calculated based on the capacity information of the storage nodes. When the relative load difference value exceeds a preset difference threshold, a shard migration process is triggered, including:

[0058] Obtain capacity information for each storage node in the distributed storage cluster, construct a hierarchical consistent hash ring based on the capacity information, and calculate a virtual node hash value for each storage node. The virtual node hash value is obtained by applying a power function to the ratio of the storage node capacity value to the system average capacity value, forming an initial shard distribution plan.

[0059] Collecting access behavior data of data shards according to the initial shard distribution scheme, the access behavior data including access frequency data, data access volume data, and access time distribution data, and calculating relative load difference values ​​between storage nodes based on the access behavior data and the capacity information of the storage nodes;

[0060] When the relative load difference value exceeds a preset difference threshold, the shard migration process is triggered.

[0061] In actual application scenarios, physical fitness test data is usually high-dimensional and large-scale, and includes multiple indicators such as human height, weight, vital capacity, heart rate, etc. In order to effectively manage these data, the present invention first divides the physical fitness test data into multiple data slices. The division of data slices can be based on dimensions such as geographical location, test time, or the age group of the test subject. For example, the physical fitness test data of the 20-30 age group in a certain area can be used as one slice, and the physical fitness test data of the 31-40 age group can be used as another slice. In this embodiment, it is assumed that there are 10 million physical fitness test records, each record contains 30 test indicators, divided into 500 data slices, and each slice contains approximately 20,000 records.

[0062] After data sharding is complete, the system obtains capacity information for each storage node in the distributed storage cluster. This capacity information includes the node's total storage space, used storage space, CPU usage, memory usage, etc. For example, there are five storage nodes in the storage cluster. Node A has a total storage space of 2TB and a used storage space of 0.8TB; Node B has a total storage space of 1.5TB and a used storage space of 0.9TB; Node C has a total storage space of 3TB and a used storage space of 1.2TB; Node D has a total storage space of 2.5TB and a used storage space of 1TB; and Node E has a total storage space of 2TB and a used storage space of 1.1TB.

[0063] Based on the acquired capacity information, the system constructs a hierarchical consistent hash ring. The hierarchical consistent hash ring is an improvement to the traditional consistent hash algorithm and can better adapt to heterogeneous storage environments. The construction process includes calculating the system average capacity value. For the above example, the system average capacity value is (2TB+1.5TB+3TB+2.5TB+2TB) / 5=2.2TB. Subsequently, the virtual node hash value of each storage node is calculated. The virtual node hash value is obtained by performing a power function mapping on the ratio of the capacity value of the storage node to the system average capacity value. In this embodiment, assuming that the square of the ratio is used as the power function mapping, the number of virtual nodes of node A is (2 / 2.2) 2 ×100≈83, the number of virtual nodes of node B is (1.5 / 2.2) 2 ×100≈46, the number of virtual nodes of node C is (3 / 2.2) 2 ×100≈186, the number of virtual nodes of node D is (2.5 / 2.2) 2 ×100≈129, the number of virtual nodes of node E is (2 / 2.2) 2 ×100≈83. These virtual nodes are evenly distributed on the hash ring, forming the initial sharding distribution scheme.

[0064] According to the initial shard distribution plan, the system begins to collect access behavior data for data shards. Access behavior data includes access frequency data, data access volume data, and access time distribution data. Access frequency data records the number of times a shard is accessed per unit time. For example, shard 1 is accessed 20 times per minute and shard 2 is accessed 15 times per minute. Data access volume data records the amount of data read or written to a shard per unit time. For example, shard 1 has 10MB of data read and 5MB of data written per minute. Access time distribution data records the time distribution characteristics of shard accesses. For example, shard 1 is frequently accessed between 9:00 and 11:00 a.m. and shard 2 is frequently accessed between 2:00 and 4:00 p.m.

[0065] Based on the collected access behavior data and storage node capacity information, the relative load difference between storage nodes is calculated. This calculation takes into account three dimensions: storage load, access load, and time-distributed load. Storage load represents the ratio of a node's used storage space to its total storage space; access load represents the ratio of access requests handled by a node to the total access requests in the system; and time-distributed load represents the fluctuation in a node's load over time. In this embodiment, assume that the storage load of node A is 0.8 / 2 = 0.4, the access load is 0.25, and the time-distributed load is 0.3; the storage load of node B is 0.9 / 1.5 = 0.6, the access load is 0.15, and the time-distributed load is 0.2; the storage load of node C is 1.2 / 3 = 0.4, the access load is 0.3, and the time-distributed load is 0.25; the storage load of node D is 1 / 2.5 = 0.4, the access load is 0.2, and the time-distributed load is 0.15; and the storage load of node E is 1.1 / 2 = 0.55, the access load is 0.1, and the time-distributed load is 0.1. Through weighted calculation, the comprehensive loads of each node are 0.33, 0.35, 0.33, 0.27, and 0.28, respectively. The relative load difference value can be obtained by calculating the load difference between the highest-load node and the lowest-load node and dividing it by the average load, that is, (0.35-0.27) / 0.31≈0.26.

[0066] When the relative load difference exceeds the preset difference threshold, the system triggers the shard migration process. Assuming the preset difference threshold is 0.2, since the calculated relative load difference of 0.26 is greater than 0.2, the shard migration process is triggered. The shard migration process includes determining the migration source and target nodes, selecting the appropriate migration data shards, performing the data migration operation, and updating the shard distribution information. In this example, the system selects Node B, which has the highest load, as the migration source node and Node D, which has the lowest load, as the migration target node. A portion of the data shards will be migrated from Node B to Node D to balance the load on each node.

[0067] Through the above technical solution, the present invention realizes the optimization of physical fitness test data distribution based on storage node capacity, which can effectively balance the load of each node in the distributed storage cluster and improve the overall system performance and resource utilization.

[0068] In an optional embodiment, allocating the data shards to multiple storage nodes in a distributed storage cluster according to spatiotemporal proximity includes:

[0069] Calculate the time dimension similarity between data slices, the time dimension similarity is calculated based on the timestamp difference of the data slices, and calculate the space dimension similarity between data slices, the space dimension similarity is calculated based on the coordinate distance of the data slices in the multidimensional space;

[0070] The temporal dimension similarity and the spatial dimension similarity are weighted and combined according to a preset spatiotemporal weight factor to obtain a comprehensive proximity between the data slices; the comprehensive proximity greater than a preset proximity threshold is retained, and the comprehensive proximity less than the preset proximity threshold is set to zero, thereby constructing a proximity matrix; based on the proximity matrix, the boundary positions of the data slices are optimized to maximize the proximity between the internal data points of the data slices and minimize the proximity of the data points between the data slices;

[0071] Calculate the storage capacity utilization rate, computing resource utilization rate, and network bandwidth utilization rate of each storage node respectively, and weight the three utilization rates according to the preset resource weight coefficient to obtain the comprehensive load value of the storage node;

[0072] When it is detected that the comprehensive load value exceeds a preset data migration threshold, the data shards on the storage node are reallocated to other storage nodes with lower loads.

[0073] like Figure 2 As shown, the method includes:

[0074] After receiving the spatiotemporal data to be stored, the data is divided into multiple data shards according to the preset sharding rules. For example, for urban traffic flow data, it can be sharded according to one time window per hour and one spatial unit per region, resulting in several data shards.

[0075] For the divided data shards, the system needs to calculate the time dimension similarity between them. Specifically, for any two data shards A and B, assuming their central timestamps are TA and TB respectively, the system calculates their time difference ΔT = |TA - TB|. The system then converts the time difference into a time dimension similarity ST using a time decay function. For example, when ΔT is 0, ST is 1; when ΔT is 3600 seconds (1 hour), ST is 0.8; and when ΔT is 86400 seconds (24 hours), ST is 0.1. This indicates that the closer the data shards are in time, the higher their time dimension similarity.

[0076] At the same time, the system calculates the spatial dimension similarity between data shards. Assuming that the center coordinates of data shards A and B are (XA, YA) and (XB, YB) respectively, the system calculates their Euclidean distance D = The system then converts the distance into spatial dimension similarity (SS) using a spatial decay function. For example, when D is 0, SS is 1; when D is 1000 meters, SS is 0.7; and when D is 10000 meters, SS is 0.05. This indicates that the closer the data shards are in space, the higher their spatial dimension similarity.

[0077] Next, the system weights the temporal and spatial similarities based on the preset spatiotemporal weighting factors WT and WS, resulting in a comprehensive proximity S = WT × ST + WS × SS. In one specific embodiment, WT = 0.4 and WS = 0.6 can be set, indicating that the spatial dimension is slightly more important than the temporal dimension. For example, for two data shards, if ST = 0.8 and SS = 0.7, then S = 0.4 × 0.8 + 0.6 × 0.7 = 0.74.

[0078] To reduce computational complexity and highlight significant proximity relationships, the system compares the comprehensive proximity with a preset threshold. Assuming the preset proximity threshold is 0.5, the system retains comprehensive proximities greater than 0.5 and sets comprehensive proximities less than 0.5 to 0. In this way, the system constructs a proximity matrix M, where M[i][j] represents the comprehensive proximity between data shards i and j.

[0079] Based on the constructed proximity matrix, the system optimizes the boundary positions of the data shards. Specifically, the system checks the average proximity of each data point to other data points in its own shard, as well as the average proximity to data points in adjacent shards. If a data point's average proximity to adjacent shards is higher than its average proximity to the current shard, the data point is reallocated to the adjacent shard. For example, for a data point P on the boundary, if its average proximity to other data points in the current shard A is 0.6, and its average proximity to data points in the adjacent shard B is 0.8, then P is moved from shard A to shard B.

[0080] Through multiple rounds of iterative optimization, we ultimately achieve optimized data sharding, maximizing the proximity between data points within a shard and minimizing the proximity between data points across shards. In one real-world example, before optimization, the average intra-shard proximity was 0.65, while the average inter-shard proximity was 0.35. After optimization, the average intra-shard proximity increased to 0.82, while the average inter-shard proximity decreased to 0.18.

[0081] To ensure the proper allocation of data shards to the distributed storage cluster, the system monitors the resource usage of each storage node. Specifically, the system calculates each storage node's storage capacity utilization (UC), computing resource utilization (UP), and network bandwidth utilization (UN). For example, for storage node N1, if its total storage capacity is 1TB and 500GB is used, then UC = 50%; if its CPU utilization is 60%, then UP = 60%; and if its network bandwidth utilization is 40%, then UN = 40%.

[0082] The three utilization rates are weighted and combined based on the preset resource weight coefficients WC, WP, and WN to obtain the storage node's overall load value, L = WC × UC + WP × UP + WN × UN. In a specific example, WC = 0.5, WP = 0.3, and WN = 0.2, indicating that storage capacity has the highest weight, followed by computing resources, and finally network bandwidth. For storage node N1, its overall load value, L, is 0.5 × 50% + 0.3 × 60% + 0.2 × 40% = 51%.

[0083] The system continuously monitors the combined load of each storage node. When the combined load of a storage node exceeds a preset data migration threshold (e.g., 80%), the system triggers data migration. For example, if the combined load of storage node N2 reaches 85%, exceeding the preset threshold of 80%, the system will select the appropriate data shard from N2 for migration.

[0084] When selecting a migration target node, the system prioritizes nodes with lower combined load values ​​and higher spatiotemporal proximity to the source data shard. For example, if data shard D1 needs to be migrated from N2, and D1 has a higher spatiotemporal proximity to data shard D2 stored on node N3 (e.g., 0.75), and N3 has a lower combined load value (e.g., 40%), the system will migrate D1 to N3.

[0085] Through the above mechanism, the system realizes distributed storage data allocation based on spatiotemporal proximity, which not only ensures the access locality of spatiotemporal related data, but also realizes the load balancing of the storage cluster, thereby improving the overall system performance and resource utilization efficiency.

[0086] In an optional embodiment, after constructing the spatiotemporal knowledge graph, the method further includes:

[0087] Obtain a set of neighbor nodes of a node to be processed in a spatiotemporal knowledge graph, wherein the set of neighbor nodes includes nodes that are directly associated with the node to be processed; calculate the similarity between the attribute feature vector of the node to be processed and the attribute feature vector of each neighbor node in the set of neighbor nodes to obtain a node similarity value;

[0088] Performing exponential transformation and normalization on the node similarity values ​​to obtain similarity weights, performing weighted summation on the attribute values ​​of the neighboring nodes based on the similarity weights, and generating missing attribute supplementary values ​​for the node to be processed;

[0089] Obtain historical attribute values ​​of nodes in the spatiotemporal knowledge graph, calculate attribute changes at adjacent time points, and calculate the ratio of the attribute changes to the historical average attribute value, and adjust the update frequency of the spatiotemporal knowledge graph based on the ratio.

[0090] After the spatiotemporal knowledge graph is constructed, it is necessary to supplement attributes and adjust the update frequency for the pending nodes in the knowledge graph. For the pending node, the first step is to obtain its neighbor node set. A neighbor node set refers to the set of nodes directly associated with the pending node. For example, in an enterprise knowledge graph, if the pending node is "Enterprise A," its neighbor nodes may include "Enterprise A's Legal Representative," "Enterprise A's Registered Address," "Enterprise A's Partner Enterprise B," and other directly related nodes.

[0091] After obtaining the set of neighboring nodes, the similarity of the attribute feature vectors between the node to be processed and each of its neighboring nodes is calculated. An attribute feature vector is a numerical representation of a node's attributes and can contain multiple dimensions. For example, the attribute feature vector of an enterprise node might include information on multiple dimensions, such as registered capital, years of establishment, and number of employees. For each neighboring node, the similarity between its attribute feature vector and the attribute feature vector of the node to be processed is calculated. This similarity can be calculated using the cosine similarity method, which divides the dot product of the two feature vectors by the product of their norms. For example, if the attribute feature vector of the node to be processed, "Enterprise A," is [1 million RMB, 5 years, 200 employees], and the attribute feature vector of its neighboring node, "Enterprise B," is [1.2 million RMB, 4 years, 180 employees], the similarity between the two can be calculated.

[0092] The calculated node similarity values ​​are subjected to an exponential transformation and normalized to obtain similarity weights. Exponential transformation enhances the weights of nodes with high similarity and suppresses the influence of nodes with low similarity. Normalization ensures that the sum of all weights is 1. For example, if the similarity values ​​of the node to be processed and its three neighbors are 0.8, 0.6, and 0.3, respectively, the exponential transformation (based on the natural logarithm e) yields 2.226, 1.822, and 1.350, and after normalization, the weights are 0.412, 0.337, and 0.251.

[0093] Based on the calculated similarity weights, the attribute values ​​of neighboring nodes are weighted and summed to generate the missing attribute supplement value for the node being processed. If the node being processed is missing an attribute value, but its neighboring nodes have that attribute information, the missing attribute can be supplemented through weighted summation. For example, if the node being processed "Enterprise A" is missing the "Annual Turnover" attribute, and its three neighboring nodes "Enterprise B," "Enterprise C," and "Enterprise D" have annual turnovers of 5 million yuan, 6 million yuan, and 4.5 million yuan, respectively, with weights of 0.412, 0.337, and 0.251, respectively, the supplementary annual turnover value for "Enterprise A" is calculated as: 500 × 0.412 + 600 × 0.337 + 450 × 0.251 = 5.2355 million yuan.

[0094] The spatiotemporal knowledge graph not only needs to supplement missing attributes but also requires a reasonable update frequency. The system obtains the historical attribute values ​​of nodes in the spatiotemporal knowledge graph and calculates the attribute change between adjacent time points. For example, if "Company A" has 150, 170, and 200 employees at time points T1, T2, and T3, respectively, the attribute change between adjacent time points is: 20 employees from T1 to T2, and 30 employees from T2 to T3.

[0095] Calculate the ratio of the attribute change to the historical average attribute value. The historical average attribute value refers to the average value of the attribute at each historical point in time. For example, the average number of employees at "Company A" at time points T1, T2, and T3 is (150 + 170 + 200) / 3 = 173.33. The change ratio from T1 to T2 is 20 / 173.33 = 0.115, and the change ratio from T2 to T3 is 30 / 173.33 = 0.173.

[0096] Based on the calculated ratio, the system can adjust the update frequency of the spatiotemporal knowledge graph. A large attribute change ratio indicates that the node's attributes are changing rapidly, requiring a higher update frequency. A small attribute change ratio indicates that the node's attributes are relatively stable, requiring a lower update frequency. The system can set a threshold to determine the direction of update frequency adjustment. For example, if the change ratio is greater than 0.15, the original update cycle will be shortened to 80%; if the change ratio is between 0.05 and 0.15, the original update cycle will remain unchanged; if the change ratio is less than 0.05, the original update cycle will be extended to 120%. For example, for "Company A," the employee number change ratio from T2 to T3 is 0.173, exceeding the threshold of 0.15. If the original update cycle is 30 days, the adjusted update cycle will be 30 × 80% = 24 days.

[0097] This method effectively supplements missing attributes of nodes in spatiotemporal knowledge graphs and dynamically adjusts the update frequency based on attribute changes, improving the integrity and timeliness of the knowledge graph. This method is particularly suitable for scenarios where attribute values ​​vary significantly over time and there is a certain degree of similarity between neighboring nodes, such as enterprise knowledge graphs and product knowledge graphs.

[0098] In an optional embodiment, based on the physical fitness feature vector, a federated learning framework is used to construct a local training model across multiple storage nodes. While ensuring data privacy, data from different storage nodes are collaboratively calculated to construct a global model. A comprehensive data analysis report for physical fitness test data generated based on the global model includes:

[0099] Adding Laplace noise to the physical feature vector, and homomorphically encrypting the physical feature vector after adding the noise using a public key to generate an encrypted feature vector;

[0100] Building a local training model at each storage node based on the encrypted feature vector, calculating the number of samples and the data quality score of each storage node, and determining the aggregation weight of the storage node according to the number of samples and the data quality score;

[0101] Distributing key pairs among multiple storage nodes, generating a random mask based on the key pairs, and superimposing the random mask with parameters of the local training model in combination with aggregation weights to obtain masked model parameters;

[0102] Collect the masked model parameters of all storage nodes, perform security aggregation operations to obtain global model parameters, and distribute the global model parameters to each storage node; generate a comprehensive data analysis report of the physical fitness test data based on the global model.

[0103] The specific steps for adding Laplace noise to a physical eigenvector are as follows: First, obtain the original physical eigenvector. For example, a student's physical eigenvector is [175.5, 68.3, 3200, 45.6, 10.5, 30.2, 89.3], which represent height (cm), weight (kg), vital capacity (ml), grip strength (kg), 50-meter run (seconds), sit-and-reach (cm), and step test index, respectively. Then, set the scale parameter of the Laplace distribution. The scale parameter directly affects the amount of noise added. For different physical indicators, different scale parameters can be set, such as 1.0 for height, 0.8 for weight, and 100 for vital capacity. Then, generate random noise for each eigenvalue component. The noise follows a Laplace distribution, for example, the generated noise values ​​are [0.7, -0.5, 123, 0.8, 0.02, -0.4, 1.1]. Adding random noise to the original feature vector yields the noisy feature vector [176.2, 67.8, 3323, 46.4, 10.52, 29.8, 90.4]. The noisy feature vector retains the statistical properties of the original data while introducing random perturbations to protect individual privacy.

[0104] The homomorphic encryption implementation process involves generating public and private keys using the Paillier homomorphic encryption algorithm. The public key consists of the product n of two large prime numbers p and q and the generator g, while the private key contains information about p and q. For example, the generated public key is (n=56153, g=34589), and the private key is the corresponding decomposition information. The public key is used to encrypt each feature component after adding noise. The encryption process uses the Paillier algorithm's encryption function E(m,r)=g m ·r nmodn 2, where m is the eigenvalue to be encrypted and r is a random number. For the noisy height data of 176.2, we randomly select r = 15738 and calculate E(176.2, 15738) to obtain the encrypted value 2876543290. This process continues by encrypting all eigenvalues, resulting in the encrypted eigenvector [2876543290, 1873456298, 4529876532, 3678945102, 2345671034, 3987654320, 4023456781]. The encrypted data can be securely transmitted and stored on the network, ensuring that the original data is not leaked.

[0105] The detailed implementation method of building a local training model in each storage node is: receiving the encrypted physical feature vector and using the characteristics of homomorphic encryption to perform calculations in the encrypted domain. For the physical fitness assessment task, the logistic regression model is used as the local training model. For storage node A, assume that the node stores the physical fitness test data of 1,000 students. Set the initial parameters of the model. For the 7 physical characteristics, initialize the weight vector to [0.1, 0.1, 0.1, 0.1, 0.1, 0.1], and the bias term is 0.1. Use the gradient descent method for model training. Since the data is encrypted, the properties of homomorphic encryption can be used to calculate the gradient directly in the encrypted domain. During the calculation process in the encrypted domain, the additive homomorphic property of Paillier encryption is used: E(m1)·E(m2)=E(m1+m2)modn 2 The model parameters are updated through repeated iterations. Training stops when the loss function change is less than the preset threshold of 0.001 or the maximum number of iterations is 500. After training, the parameters of the local model are the weight vector [0.25, 0.32, 0.18, 0.27, -0.35, 0.22, 0.30] and the bias term 0.15.

[0106] The specific method for calculating the number of samples and data quality score for each storage node is to count the number of valid samples for each storage node. For example, storage node A has 1000 samples, storage node B has 1500 samples, and storage node C has 800 samples. The data quality score is calculated based on multiple dimensions: data completeness, data consistency, and data accuracy. Data completeness is determined by calculating the proportion of non-null features in each sample. For example, a data completeness score of 0.95 for storage node A indicates that, on average, 95% of the feature values ​​are valid. Data consistency is assessed by checking whether the data conforms to the expected value range and distribution characteristics. For example, if the height data is within a reasonable range, the consistency score for storage node A is 0.92. Data accuracy is determined by evaluating the measurement accuracy and reliability of the data. For example, the accuracy score for storage node A is 0.88. The weighted average of the scores for the three dimensions, with weights of 0.3, 0.3, and 0.4, yields a total data quality score for storage node A of 0.88 × 0.4 + 0.92 × 0.3 + 0.95 × 0.3 = 0.913. Similarly, the data quality scores of other nodes are calculated, such as storage node B is 0.875, and storage node C is 0.902.

[0107] The method for determining aggregation weight based on sample quantity and data quality score is as follows: Normalize the sample quantity to obtain the sample quantity weight factor. For example, the sample quantity weight of storage node A is 1000 / (1000+1500+800)=0.303, while that of storage node B is 0.455 and that of storage node C is 0.242. Normalize the data quality score to obtain the quality weight factor. For example, the quality weight of storage node A is 0.913 / (0.913+0.875+0.902)=0.340, while that of storage node B is 0.326 and that of storage node C is 0.334. Assume that the importance weights of sample quantity and data quality are 0.6 and 0.4, respectively. Calculate the final aggregation weight, which is the weighted average of the sample quantity weight and the data quality weight. The aggregation weight of storage node A is 0.303 × 0.6 + 0.340 × 0.4 = 0.318, that of storage node B is 0.455 × 0.6 + 0.326 × 0.4 = 0.403, and that of storage node C is 0.242 × 0.6 + 0.334 × 0.4 = 0.279. The aggregation weight ensures that nodes with a large number of samples and high-quality data play a greater role in model aggregation.

[0108] The specific steps for distributing key pairs among multiple storage nodes are as follows: A central server generates key pairs for secure aggregation. For N storage nodes, N sets of public-private key pairs (PKi, SKi) are generated, where i = 1, 2, ... N. For example, for storage nodes A, B, and C, the generated public keys are PKA, PKB, and PKC, respectively, and the generated private keys are SKA, SKB, and SKC, respectively. Using a secure communication channel, such as the TLS protocol, the public keys are distributed to all storage nodes, and each node obtains the public keys of all nodes, including itself. Private keys are only distributed to the corresponding nodes; for example, storage node A only obtains SKA. Each node verifies the validity of the received key and confirms that the key has not been tampered with. This distribution mechanism ensures the reliability of the subsequent secure aggregation process.

[0109] The specific method for generating random masks based on key pairs is as follows: a shared random seed is generated between each pair of storage nodes. The shared random seed Sij between storage nodes i and j is generated through a key exchange protocol, such as Diffie-Hellman key exchange. Specifically, node i uses its private key SKi and node j's public key PKj to calculate the shared seed, and node j uses its private key SKj and node i's public key PKi to calculate the same shared seed. For example, the shared random seed between storage nodes A and B is SAB = 73529483, between A and C it is SAC = 46821937, and between B and C it is SBC = 59274681. Based on the shared random seed, a pseudo-random number generator (PRNG) is used to generate random masks. For storage node i, two sets of random masks are generated: a positive mask and a negative mask that are shared with every other node. For example, storage node A generates a positive mask [2.5, -1.8, 3.2, 0.7, -2.1, 1.5, 0.9] based on SAB and shares it with B. The negative mask is the negative value of the positive mask [-2.5, 1.8, -3.2, -0.7, 2.1, -1.5, -0.9]. Based on SAC, the positive mask [1.3, -0.9, 2.7, -1.5, 0.8, -2.2, 1.1] and the negative mask [-1.3, 0.9, -2.7, 1.5, -0.8, 2.2, -1.1] shared with C are generated. The dimensions of the random mask are the same as those of the model parameter vector, ensuring that the mask can be directly added to the model parameters.

[0110] The specific process of superimposing the random mask with the parameters of the locally trained model combined with the aggregation weights is as follows: storage node i calculates the sum of the masks shared with all other nodes. For example, the sum of the masks calculated by storage node A is the positive mask shared with B plus the positive mask shared with C, that is, [2.5+1.3,-1.8+(-0.9),3.2+2.7,0.7+(-1.5),-2.1+0.8,1.5+(-2.2),0.9+1.1] = [3.8,-2.7,5.9,-0.8,-1.3,-0.7,2.0]. Multiply the local model parameters by the aggregation weight. For example, multiply the local model parameters of storage node A [0.25, 0.32, 0.18, 0.27, -0.35, 0.22, 0.30] by the aggregation weight 0.318 to obtain [0.0795, 0.1018, 0.0572, 0.0859, -0.1113, 0.0700, 0.0954]. Adding the weighted local model parameters to the masked sum yields the masked model parameters [0.0795+3.8, 0.1018+(-2.7), 0.0572+5.9, 0.0859+(-0.8), -0.1113+(-1.3), 0.0700+(-0.7), 0.0954+2.0] = [3.8795, -2.5982, 5.9572, -0.7141, -1.4113, -0.63, 2.0954]. The masked parameters contain valuable model information while protecting the privacy of the original parameters through masking.

[0111] The specific process of collecting and aggregating masked model parameters is as follows: the central server collects masked model parameters from all storage nodes. For example, the collected masked parameters for storage node A are [3.8795, -2.5982, 5.9572, -0.7141, -1.4113, -0.63, 2.0954], the parameters for storage node B are [2.4566, 3.7821, -4.2198, 1.8374, 0.9257, -2.1839, -1.5632], and the parameters for storage node C are [-5.3361, -0.1839, -0.7374, -0.1233, 0.4856, 2.8139, -0.5322]. The central server simply sums the collected parameters and obtains [3.8795+2.4566+(-5.3361),-2.5982+3.7821+(-0.1839),5.9572+(-4.2198)+(-0.7374),-0.7141+1.8374+(-0.1233),-1.4113+0.9257+0.4856,-0.63+(-2.1839)+2.8139,2.0954+(-1.5632)+(-0.5322)]=[1.0,1.0,1.0,1.0,0,0,0]. Due to the characteristics of random masking, the positive and negative masks shared by different nodes cancel each other out during the summation process. The final aggregation result only contains the sum of the weighted model parameters, such as [0.0795+0.1614+0.0837, 0.1018+0.1209+0.0893...]. The aggregated global model parameters are securely distributed back to each storage node.

[0112] The specific steps for generating a comprehensive analysis report of physical fitness test data based on a global model are as follows: The global model is used to predict and analyze the local data of each storage node. For individual physical fitness assessments, scores for each indicator and the total score are calculated. For example, if a student's physical fitness feature vector is [178.5, 65.2, 3450, 48.3, 9.8, 32.5, 92.7], the global model predicts a physical fitness level of "good" with a confidence level of 0.85. The generated individual physical fitness report includes scores for each indicator, a comprehensive rating, and improvement suggestions. Statistical analysis of physical fitness at the group level is performed. For example, the distribution of physical fitness levels by age group shows that for the 15-18 age group, the excellent rate is 15.3%, the good rate is 45.7%, the passing rate is 32.1%, and the failing rate is 6.9%. By gender, the boys' score was 14.2% excellent, 43.8% good, 35.5% qualified, and 6.5% unqualified. For the girls' score, the scores were 16.5%, 47.9% good, 28.4% qualified, and 7.2% unqualified. By region, the urban students' score was 17.8%, while the rural students' score was 12.5%. A chart of physical fitness trends was generated, such as the three-year change in the pass rate for physical fitness: 78.5% in 2023, 80.2% in 2024, and 82.3% in 2025, showing an upward trend. Trends in different physical fitness categories were analyzed, such as the annual improvement in endurance and a slight decline in flexibility. Based on these results, targeted improvement suggestions were provided, such as strengthening flexibility training and optimizing sports facilities in rural schools.

[0113] In an optional embodiment, a key pair is distributed among multiple storage nodes, a random mask is generated based on the key pair, and the random mask is superimposed with the parameters of the local training model in combination with the aggregation weight to obtain the masked model parameters including:

[0114] Determine the public key of each storage node and distribute the public key among the storage nodes; generate an initial shared key between the storage node pair through modular exponentiation based on the private key of each storage node and the public keys of other storage nodes; perform a cryptographic hash operation on the initial shared key to obtain a final shared key;

[0115] Based on the final shared key between the storage node pair and the current calculation round identifier, a random noise seed is generated through a pseudo-random function. The sum of the noise seed differences between each storage node and other storage nodes is calculated to obtain the noise accumulation value.

[0116] Inputting the noise accumulation value into a deterministic random expansion function to generate a mask matrix of the storage node;

[0117] The local model parameters of each storage node are multiplied by the aggregation weight, and the product is added to the mask matrix of the storage node to obtain the masked model parameters.

[0118] To achieve secure federated learning model parameter aggregation, the system distributes key pairs across multiple storage nodes and generates random masks based on these key pairs. Storage nodes can be servers or edge devices participating in federated learning, and each node has its own local training data and model.

[0119] At the beginning of the key distribution phase, the system generates a public-private key pair for each storage node. For example, if there are three storage nodes in the system, A, B, and C, they generate the public-private key pairs (PKA, SKA), (PKB, SKB), and (PKC, SKC), respectively. Public keys are distributed among all storage nodes via a secure channel, ensuring that each node has access to the public keys of all other nodes. For example, node A receives PKB and PKC, node B receives PKA and PKC, and node C receives PKA and PKB.

[0120] After obtaining the public keys of other nodes, each storage node uses its own private key and the public keys of other nodes to generate an initial shared key through modular exponentiation. Taking node A as an example, it uses its own private key SKA and node B's public key PKB to perform modular exponentiation to generate the initial shared key KAB initial Similarly, use SKA and PKC to perform modular exponentiation to generate KAC initial To enhance security, cryptographic hash operations are performed on these initial shared keys, such as using the SHA-256 algorithm, to obtain the final shared keys KAB and KAC. At this point, nodes B and C also perform the same operation to obtain the final shared keys between them and other nodes.

[0121] The mask generation phase uses these shared keys to create random noise. The system introduces the current calculation round identifier round id , ensuring that each round of training generates a different mask. For storage node A, the system uses the final shared key KAB with round id As input, a random noise seed SAB is generated by a pseudo-random function PRF; similarly, KAC and round id Generate SAC. Due to the symmetry of the key, node B uses KBA (equivalent to KAB) and round id The resulting seed SBA is equal to SAB in value but opposite in sign.

[0122] Node A calculates the sum of the noise seed differences: noise sum A =SAB+SAC. Node B calculation: noise sum B =SBA+SBC. Node C calculation: noise sum C=SCA+SCB. Since SBA=-SAB, SCA=-SAC, and SCB=-SBC, the sum of the noise accumulation values ​​of the three nodes is zero, ensuring that the mask can be offset during the aggregation process.

[0123] Each storage node inputs its noise accumulation value into the deterministic random expansion function DRBG to generate a mask matrix with the same dimension as the model parameters. For example, if the model has 1 million parameters, the mask matrix of node A is A is a vector containing 1 million elements, each of which is a sum A Extends the generated random value.

[0124] The mask processing process of local model parameters is as follows: Assume that the local model parameter of node A is model params A , whose aggregation weight is weight A (can be determined based on data quantity or quality), the system will model params A and weight A Multiply and then add the mask matrix mask A , get the masked model parameters params A =model_params A ×weight A +mask A .

[0125] Taking a specific numerical example, assuming that the model has only 3 parameters, the local model parameters of node A are [0.2, 0.3, 0.5], the aggregation weight is 0.4, and the generated mask matrix is ​​[1.2, -0.8, 0.6]. The model parameters after masking are [0.2×0.4+1.2, 0.3×0.4-0.8, 0.5×0.4+0.6]=[1.28, -0.68, 0.8].

[0126] The model parameters after masking are calculated as [-0.61, 0.62, -0.12] for node B (parameters [0.3, 0.4, 0.6], weight 0.3, mask [-0.7, 0.5, -0.3]) and [-0.425, 0.405, -0.135] for node C (parameters [0.25, 0.35, 0.55], weight 0.3, mask [-0.5, 0.3, -0.3]), respectively.

[0127] When the three masked model parameters are aggregated, the masked parts cancel each other out (because 1.2 - 0.7 - 0.5 = 0, -0.8 + 0.5 + 0.3 = 0, and 0.6 - 0.3 - 0.3 = 0), and the final result contains only the weighted sum of the original model parameters: [0.2 × 0.4 + 0.3 × 0.3 + 0.25 × 0.3, 0.3 × 0.4 + 0.4 × 0.3 + 0.35 × 0.3, 0.5 × 0.4 + 0.6 × 0.3 + 0.55 × 0.3] = [0.245, 0.345, 0.545].

[0128] This masking mechanism ensures that during the model aggregation process, nodes do not leak the original model parameters, but only share the masked data. Because the mask is a random value generated based on a secure key, unauthorized parties cannot deduce the original model information from the masked parameters, thereby protecting the data privacy of all participants while ensuring the accuracy of the aggregation results.

[0129] A second aspect of an embodiment of the present invention provides a multi-center collaborative collection and management system for student physical fitness test data, including:

[0130] The first unit is configured to obtain physical fitness test data uploaded by multiple physical fitness test centers, divide the physical fitness test data into multiple data slices, and distribute the data slices to multiple storage nodes in a distributed storage cluster according to spatiotemporal proximity;

[0131] The second unit is used to construct a spatiotemporal knowledge graph, map the test subject information and test item information of the physical test data into entity nodes in the spatiotemporal knowledge graph, map the test result information into attributes of the entity nodes, and map the spatiotemporal association indexes into relationship edges between the entity nodes;

[0132] The third unit is used to obtain the physical characteristic vectors of the test subject under different time and space conditions;

[0133] The fourth unit is used to build a local training model between multiple storage nodes based on the physical feature vector using a federated learning framework, perform collaborative calculations on the data of different storage nodes to generate a comprehensive data analysis report of the physical test data, and distribute the comprehensive data analysis report to terminal devices of multiple physical test centers through an encrypted transmission channel.

[0134] According to a third aspect of an embodiment of the present invention, an electronic device is provided, including:

[0135] processor;

[0136] a memory for storing processor-executable instructions;

[0137] The processor is configured to call the instructions stored in the memory to execute the aforementioned method.

[0138] According to a fourth aspect of an embodiment of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the method described above is implemented.

[0139] The present invention may be a method, an apparatus, a system and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for executing various aspects of the present invention.

[0140] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A multi-center collaborative collection and management method for student physical fitness test data, characterized in that: include: Acquire physical fitness test data uploaded by multiple physical fitness test centers, divide the physical fitness test data into multiple data shards, obtain capacity information of each storage node in a distributed storage cluster, calculate relative load difference values ​​between storage nodes based on the capacity information of the storage nodes, trigger a shard migration process when the relative load difference value exceeds a preset difference threshold, and distribute the data shards to multiple storage nodes in the distributed storage cluster according to spatiotemporal proximity, including: Obtain capacity information for each storage node in the distributed storage cluster, construct a hierarchical consistent hash ring based on the capacity information, and calculate a virtual node hash value for each storage node. The virtual node hash value is obtained by applying a power function to the ratio of the storage node capacity value to the system average capacity value, forming an initial shard distribution plan. Collecting access behavior data of data shards according to the initial shard distribution scheme, the access behavior data including access frequency data, data access volume data, and access time distribution data, and calculating relative load difference values ​​between storage nodes based on the access behavior data and the capacity information of the storage nodes; When the relative load difference value exceeds a preset difference threshold, the shard migration process is triggered; Calculate the time dimension similarity between data slices, the time dimension similarity is calculated based on the timestamp difference of the data slices, and calculate the space dimension similarity between data slices, the space dimension similarity is calculated based on the coordinate distance of the data slices in the multidimensional space; The temporal dimension similarity and the spatial dimension similarity are weighted and combined according to a preset spatiotemporal weight factor to obtain a comprehensive proximity between the data slices; the comprehensive proximity greater than a preset proximity threshold is retained, and the comprehensive proximity less than the preset proximity threshold is set to zero, thereby constructing a proximity matrix; based on the proximity matrix, the boundary positions of the data slices are optimized to maximize the proximity between the internal data points of the data slices and minimize the proximity of the data points between the data slices; Calculate the storage capacity utilization rate, computing resource utilization rate, and network bandwidth utilization rate of each storage node respectively, and weight the three utilization rates according to the preset resource weight coefficient to obtain the comprehensive load value of the storage node; When it is detected that the comprehensive load value exceeds a preset data migration threshold, the data shards on the storage node are reallocated to other storage nodes with lower loads; Constructing a spatiotemporal knowledge graph, mapping the test subject information and test item information of the physical fitness test data into entity nodes in the spatiotemporal knowledge graph, mapping the test result information into attributes of the entity nodes, and mapping the spatiotemporal association indexes into relationship edges between the entity nodes; Using a graph neural network to perform feature extraction and representation learning on the spatiotemporal knowledge graph, obtaining the physical feature vectors of the test subject under different spatiotemporal conditions; Based on the physical characteristic vector, a federated learning framework is used to construct a local training model between multiple storage nodes. While ensuring data privacy, the data of different storage nodes are collaboratively calculated to construct a global model. Based on the global model, a comprehensive data analysis report of the physical test data is generated, and the comprehensive data analysis report is distributed to terminal devices of multiple physical test centers through an encrypted transmission channel.

2. The method according to claim 1, characterized in that After constructing the spatiotemporal knowledge graph, the method further includes: Obtain a set of neighbor nodes of a node to be processed in a spatiotemporal knowledge graph, wherein the set of neighbor nodes includes nodes that are directly associated with the node to be processed; calculate the similarity between the attribute feature vector of the node to be processed and the attribute feature vector of each neighbor node in the set of neighbor nodes to obtain a node similarity value; Performing exponential transformation and normalization on the node similarity values ​​to obtain similarity weights, performing weighted summation on the attribute values ​​of the neighboring nodes based on the similarity weights, and generating missing attribute supplementary values ​​for the node to be processed; Obtain historical attribute values ​​of nodes in the spatiotemporal knowledge graph, calculate attribute changes at adjacent time points, and calculate the ratio of the attribute changes to the historical average attribute value, and adjust the update frequency of the spatiotemporal knowledge graph based on the ratio.

3. The method according to claim 1, characterized in that Based on the physical characteristic vector, a local training model is constructed among multiple storage nodes using a federated learning framework. Under the premise of ensuring data privacy, data from different storage nodes are collaboratively calculated to construct a global model. A comprehensive data analysis report of the physical test data generated based on the global model includes the following: Adding Laplace noise to the physical feature vector, and homomorphically encrypting the physical feature vector after adding the noise using a public key to generate an encrypted feature vector; Building a local training model at each storage node based on the encrypted feature vector, calculating the number of samples and the data quality score of each storage node, and determining the aggregation weight of the storage node according to the number of samples and the data quality score; Distributing key pairs among multiple storage nodes, generating a random mask based on the key pairs, and superimposing the random mask with parameters of the local training model in combination with aggregation weights to obtain masked model parameters; Collect the masked model parameters of all storage nodes, perform security aggregation operations to obtain global model parameters, and distribute the global model parameters to each storage node; generate a comprehensive data analysis report of the physical fitness test data based on the global model.

4. The method according to claim 3, characterized in that Key pairs are distributed among multiple storage nodes. A random mask is generated based on the key pairs. The random mask is superimposed on the parameters of the local training model in combination with the aggregation weights to obtain the masked model parameters, including: Determine the public key of each storage node and distribute the public key among the storage nodes; generate an initial shared key between the storage node pair through modular exponentiation based on the private key of each storage node and the public keys of other storage nodes; perform a cryptographic hash operation on the initial shared key to obtain a final shared key; Based on the final shared key between the storage node pair and the current calculation round identifier, a random noise seed is generated through a pseudo-random function. The sum of the noise seed differences between each storage node and other storage nodes is calculated to obtain the noise accumulation value. Inputting the noise accumulation value into a deterministic random expansion function to generate a mask matrix of the storage node; The local model parameters of each storage node are multiplied by the aggregation weight, and the product is added to the mask matrix of the storage node to obtain the masked model parameters.

5. A multi-center collaborative collection and management system for student physical fitness test data, used to implement the method according to any one of claims 1 to 4, characterized in that: include: The first unit is configured to obtain physical fitness test data uploaded by multiple physical fitness test centers, divide the physical fitness test data into multiple data slices, and distribute the data slices to multiple storage nodes in a distributed storage cluster according to spatiotemporal proximity; The second unit is used to construct a spatiotemporal knowledge graph, map the test subject information and test item information of the physical test data into entity nodes in the spatiotemporal knowledge graph, map the test result information into attributes of the entity nodes, and map the spatiotemporal association indexes into relationship edges between the entity nodes; The third unit is used to obtain the physical characteristic vectors of the test subject under different time and space conditions; The fourth unit is used to build a local training model between multiple storage nodes based on the physical feature vector using a federated learning framework, perform collaborative calculations on the data of different storage nodes to generate a comprehensive data analysis report of the physical test data, and distribute the comprehensive data analysis report to terminal devices of multiple physical test centers through an encrypted transmission channel.

6. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method according to any one of claims 1 to 4.

7. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 4 is implemented.

Citation Information

Patent Citations

  • Statistical management system and method for student learning archives based on space-time database

    CN114462769A

  • Station area intelligent fusion terminal data processing system based on edge calculation

    CN119440800A