Evaluation robot data interaction system based on edge computing
Patent Information
- Application Number
- CN202611039977.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-09-25
AI Technical Summary
[0004]为了解决现有技术中PAF亲和计算未适配幼儿肢体比例、单帧匹配缺少时序约束、重叠场景易误配及固定膨胀率难以适应动作节奏变化的问题,本发明提供了一种基于边缘计算的测评机器人数据交互系统及方法
[0007]本发明引入带幼儿肢体长度先验的椭圆高斯权重矢量,使亲和场计算更契合幼儿骨骼比例特征,并有利于降低背景噪声、遮挡边缘及衣物干扰。通过非均匀采样及位置衰减机制,在关节端点区域提高采样密度,在中段区域减少冗余采样,从而兼顾亲和分数计算稳定性与边缘侧运算效率。在骨架匹配阶段,结合历史代价值形成参考基准,并对异常波动配对施加惩罚,增强骨架序列的时间连续性;针对肢体重叠场景,利用局部热力响应方向构建附加惩罚权重并重新匹配,降低误匹配概率。进一步根据平均关节速度调整时序卷积网络膨胀率,提升了行为测评结果的准确性与稳定性。
Smart Images

Figure CN122821179A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of data interaction, and in particular relates to a data interaction system for an evaluation robot based on edge computing. Background Technology
[0002] Early childhood is a critical period for the development of motor skills and cognitive abilities. Assessing children's behavior helps educators and parents understand their developmental status and identify potential safety risks. Current early childhood behavior assessments generally rely on manual scoring, using task completion as the basis for evaluation. This method is not only susceptible to subjective human factors—different assessors may use different scoring standards for the same behavior—but also exhibits behavioral biases when children are with unfamiliar versus familiar individuals, further reducing accuracy. While video-based behavior recognition and scoring can mitigate these two factors, children's body shapes differ significantly from adults; their limbs are shorter, proportions are unique, and their movements are highly random and variable, often including crawling, rolling, limb overlapping, and rapid running. This makes general human posture estimation and behavior recognition models prone to joint shifts and limb association errors in early childhood scenarios. Existing bottom-up posture estimation methods based on partial affinity fields typically output joint confidence maps and partial affinity field vector maps through deep networks, then generate skeletons through bipartite graph matching, and finally use temporal convolutional networks to classify and score the skeleton sequences.
[0003] However, existing PAF affinity calculations mostly use uniform width and uniform line integral sampling, which do not fully incorporate the prior knowledge of the thickness ratio of children's limbs and bone length. This can easily lead to affinity value distortion in short and thick limbs or joint endpoints. Existing matching strategies mostly rely on single-frame affinity scores and lack temporal constraints on cost fluctuations between consecutive frames. This can easily lead to mismatches in scenarios involving overlapping limbs, hugging, or multi-person interaction. At the same time, temporal convolutional networks with fixed dilation rates are difficult to adapt to changes in the rhythm of children's movements and joint speeds, which may affect the timeliness of behavior recognition and the stability of scoring. Summary of the Invention
[0004] To address the problems in existing technologies such as PAF affinity calculation not adapting to the proportions of young children's limbs, lack of temporal constraints in single-frame matching, easy mismatch in overlapping scenes, and difficulty in adapting to changes in movement rhythm with fixed expansion rates, this invention provides a data interaction system and method for evaluation robots based on edge computing.
[0005] In a first aspect, the present invention provides a data interaction system for an evaluation robot based on edge computing, comprising: The output module is used to evaluate the robot's acquisition of children's videos. It inputs the children's video frames into a dual-branch network and outputs a joint confidence map and a partial affinity field vector map. The summation module is used to calculate the affinity field by weighting the unit vector along the limb direction into an elliptical Gaussian weight vector with the prior of the child's limb length. When sampling the inter-joint line integral, a non-uniform sampling strategy is adopted, with dense sampling in the end-point densified area and sparse sampling in the middle sparse area. The affinity value of each sampling point is multiplied by a position attenuation coefficient that is proportional to the value of the weight vector and then weighted and summed to obtain the affinity score. The update module is used to construct an initial cost matrix based on the affinity score, distance decay factor, and confidence product, and to pair using the Hungarian algorithm. The cost value of the previous frame is used as a reference baseline cost value through exponential moving average, and a penalty is applied to pairings where the increase in matching cost exceeds the fluctuation threshold. When limbs overlap, the local gradient direction of the non-extreme position in the neighborhood of the thermal peak of the joint confidence map is extracted, the local gradient direction is rotated by 90 degrees to obtain the local extension direction, and 1 is subtracted from the absolute value of the cosine of the angle between the local extension direction and the major axis to apply additional penalty weights to update the matrix and rematch. The adjustment module is used to calculate joint features based on the successfully matched skeleton sequences and normalize them according to prior information. The input to the temporal convolutional network is used to obtain the child's behavior category and behavior score. The dilation rate of the temporal convolutional network is adjusted according to the average joint velocity of the input sequence.
[0006] On the other hand, the present invention also provides a data interaction method for an evaluation robot based on edge computing, comprising: The assessment robot acquires videos of young children, inputs the video frames into a dual-branch network, and outputs a joint confidence map and a partial affinity field vector map. When calculating the affinity field, the unit vector along the limb direction is weighted into an elliptical Gaussian weight vector with the prior of the child's limb length; when sampling the inter-joint line integral, a non-uniform sampling strategy is adopted, with dense sampling in the end-point densified area and sparse sampling in the middle section sparse area. The affinity value of each sampling point is multiplied by a position attenuation coefficient that is proportional to the value of the weight vector and then weighted and summed to obtain the affinity score. An initial cost matrix is constructed based on the product of affinity score, distance decay factor, and confidence score, and the Hungarian algorithm is used for pairing. The cost value of the previous frame is used as a reference baseline cost value through exponential moving average, and a penalty is applied to pairings where the increase in matching cost exceeds the fluctuation threshold. When limbs overlap, the local gradient direction of the non-extreme position in the neighborhood of the thermal peak of the joint confidence map is extracted. The local gradient direction is rotated by 90 degrees to obtain the local extension direction. The absolute value of the cosine of the angle between the local extension direction and the major axis is subtracted from 1 to apply an additional penalty weight to update the matrix and rematch. Joint features are calculated based on the successfully matched skeleton sequences and normalized according to prior information. These features are then input into a temporal convolutional network to obtain the child's behavior category and behavior score. The dilation rate of the temporal convolutional network is adjusted according to the average joint velocity of the input sequence.
[0007] This invention introduces an elliptical Gaussian weight vector with prior knowledge of limb length in young children, making affinity field calculation more consistent with the skeletal proportions of children and helping to reduce background noise, occlusion edges, and clothing interference. Through non-uniform sampling and positional attenuation mechanisms, sampling density is increased in the joint endpoint regions and redundant sampling is reduced in the mid-section regions, thus balancing the stability of affinity score calculation with the efficiency of edge-side computation. In the skeleton matching stage, a reference benchmark is formed by combining historical cost values, and penalties are applied to abnormal fluctuation pairings to enhance the temporal continuity of the skeleton sequence. For limb overlap scenarios, additional penalty weights are constructed using local thermal response directions and re-matching is performed to reduce the probability of mismatches. Furthermore, the dilation rate of the temporal convolutional network is adjusted according to the average joint velocity, improving the accuracy and stability of the behavior assessment results. Attached Figure Description
[0008] Figure 1 This is a diagram of the architecture of a two-branch network; Figure 2 This is a structural diagram of a temporal convolutional network; Figure 3 This is a flowchart of Example 2. Detailed Implementation
[0009] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this specification as detailed in the appended claims.
[0010] It should be understood that the terms “comprising” and “having” and any variations thereof in the embodiments of this specification are intended to cover but not exclude inclusion. For example, a product or device that includes a series of components is not necessarily limited to those components that are explicitly listed, but may include other components that are not explicitly listed or that are inherent to such product or device.
[0011] If the technical solution of this application involves personal information, the product using this technical solution has clearly informed the user of the personal information processing rules and obtained the user's voluntary consent before processing the personal information. If the technical solution of this application involves sensitive personal information, the product using this technical solution has obtained the user's separate consent before processing the sensitive personal information, and also meets the requirement of "express consent". For example, at personal information collection devices such as cameras, clear and prominent signs are set up to inform users that they have entered the scope of personal information collection and that personal information will be collected. If an individual voluntarily enters the collection scope, it is deemed that they have agreed to the collection of their personal information; or on the personal information processing device, with clear signs / information informing users of the personal information processing rules, authorization is obtained from the individual through pop-up information or by asking the individual to upload their personal information; wherein, the personal information processing rules may include information such as the personal information processor, the purpose of personal information processing, the processing method, and the types of personal information processed.
[0012] In Embodiment 1 of the present invention, a data interaction system for an evaluation robot based on edge computing is proposed, comprising: The output module is used to evaluate the robot's acquisition of videos of young children. It inputs the video frames into a dual-branch network and outputs a joint confidence map and a partial affinity field vector map.
[0013] The assessment robot acquires videos of young children. Optionally, the assessment robot has a built-in edge AI computing chip, such as... Figure 1 As shown, the video data of young children is read frame by frame and converted into an RGB image matrix. The image resolution is then uniformly adjusted to a specific size and normalized by dividing by 255. A backbone network is constructed, and the HRNet high-resolution network is used to extract spatial feature maps of the images, which are then fed into two parallel convolutional branches. The first branch outputs a two-dimensional Gaussian distribution heatmap encoding the spatial location of each joint, i.e., a joint confidence map. The second branch outputs a two-dimensional vector field encoding the limb orientation and spatial location, i.e., a partial affinity field vector map.
[0014] The backbone network is used to extract multi-scale image features from children's video frames. It can employ a high-resolution convolutional network structure, including basic convolutional layers, multi-scale feature extraction layers, and feature fusion layers. This structure preserves the spatial details needed for joint localization and extracts the semantic features required for limb association. The feature maps output by the backbone network are fed into two parallel convolutional branches. The first branch consists of several two-dimensional convolutional layers and activation layers, outputting confidence maps corresponding to each joint category. The second branch also consists of several two-dimensional convolutional layers and activation layers, outputting partial affinity field vector maps corresponding to each limb connection relationship. Figure 1As shown. During training, joint heatmap labels are generated based on the labeled joint coordinates, and partial affinity field vector labels are generated based on the labeled limb connection relationships. The error between the predicted confidence map and the heatmap labels, and the error between the predicted partial affinity field vector map and the vector labels are weighted to form a loss function, and the parameters of the backbone network and the two branch networks are updated through backpropagation.
[0015] The summation module is used to calculate the affinity field by weighting the unit vector along the limb direction into an elliptical Gaussian weight vector with the prior of the child's limb length. When sampling the inter-joint line integral, a non-uniform sampling strategy is adopted, with dense sampling in the end-point densified area and sparse sampling in the middle section sparse area. The affinity value of each sampling point is multiplied by a position attenuation coefficient that is proportional to the value of the weight vector and then weighted and summed to obtain the affinity score.
[0016] When calculating the affinity field between two connected nodes, a two-dimensional grid coordinate system is constructed, and the unit vector along the limb connection direction is calculated by subtracting coordinates and dividing by Euclidean distance. Statistical minor axis length values for corresponding limb categories are read from a pre-established JSON-formatted prior database of infant skeletal proportions, which is periodically distributed from the cloud server to the local memory cache of edge nodes. These minor axis length values are used as the minor axis scale parameter of the elliptic Gaussian function model, and the length of the endpoint encryption zone is determined based on this parameter. An elliptic Gaussian function model is constructed based on the parent-child joint connection direction and the minor axis scale parameter. The unit vector is multiplied by the elliptic Gaussian weights to obtain the elliptic Gaussian weight vector with prior knowledge.
[0017] When performing inter-joint line integration, non-uniform sampling is achieved by combining the linspace function and the exponential distribution function to generate a sequence of sampling point coordinates. This ensures that the sampling step size in the encrypted region within the minor axis length range of the connection endpoint is smaller than the sampling step size in the sparse region in the middle section.
[0018] Extract the bilinear interpolated two-dimensional vector of each sampling point from the partial affinity field vector diagram, and perform a dot product operation with the unit vector. Read the elliptical Gaussian weight value of the corresponding position of each sampling point as the position attenuation coefficient, multiply the dot product result by the position attenuation coefficient of the corresponding sampling point, sum the product results of all sampling points, and perform scale normalization according to the number of sampling points to output the affinity score of the candidate limb connection.
[0019] In an optional embodiment, the step of weighting the unit vector along the limb direction into an elliptic Gaussian weight vector with prior knowledge of the child's limb length includes: Obtain the two-dimensional coordinates of the parent and child joints, and calculate the Euclidean distance between the parent and child joints; The minor axis prior parameters matching the child's age group are obtained by querying the skeletal proportion database. Calculate the longitudinal and lateral vertical distances of spatial pixels to the line connecting parent and child joints; Based on the longitudinal projection distance, horizontal and vertical distance, Euclidean distance, and minor axis prior parameters, the elliptic Gaussian weight value of the spatial pixel is calculated using a Gaussian function.
[0020] The two-dimensional coordinates of the parent and child joints of the target connection are extracted from the output of the dual-branch network or from the pre-detection. Taking the analysis of the left shoulder and left elbow connection of a young child as an example, the coordinate positions on the feature map of the current frame are obtained respectively. For example, the coordinates of the left shoulder are 120 horizontally and 150 vertically, and the coordinates of the left elbow are 140 horizontally and 210 vertically. Substituting into the Euclidean distance formula, the distance between the two joints is calculated to be 63.2 pixels. At the same time, the backend accesses a skeletal proportion database based on real body measurement data. According to the input age of the child, such as 3 to 4 years old, the short axis prior parameters of the corresponding upper arm part are retrieved. These prior parameters usually refer to the standard width of a specific limb segment and can be mapped to pixel size.
[0021] For any local spatial pixel within the receptive field, such as a pixel located at coordinates 130 horizontally and 180 vertically, the projection point of this pixel on the line connecting the ideal centers of the parent and child joints is calculated using the algebraic method of vector orthogonal projection. The distance between this projection point and the midpoint of the line connecting the parent and child joints is then calculated as the longitudinal projection distance. For example, 30 pixels, and the horizontal and vertical distances from that point to the line connecting the parent and child joints. For example, 2.5 pixels. During this process, The preferred value is 0.5 times the actual calculated Euclidean distance, for example, 31.6, used to define the effective response range in the major axis direction; This is equivalent to retrieving the minor axis prior parameter, such as 8.5 pixels.
[0022] Substitute the example values into the formula. The calculated value is approximately 0.372. This calculation method transforms the rigid binary truncated response into a soft weight distribution that conforms to the proportional characteristics of a child's limbs. The closer the limb distribution area is to the line connecting the father and child joints, the higher the weight; the further away from the line or beyond the reasonable width range of the limb, the weight decreases smoothly, thus helping to reduce unstructured noise interference caused by loose clothing, obscured edges, or background textures.
[0023] In an optional embodiment, the non-uniform sampling strategy employed during inter-joint line integration sampling—with dense sampling in the endpoint enrichment region and sparse sampling in the middle sparse region—includes: Using the parent-child joint as the starting and ending points, the connection is divided into an end-point encrypted region near both ends and a middle-sparse region in the middle. The length of the end-point encrypted region is determined according to the minor axis scale. Within the endpoint encryption zone, non-uniform sampling point coordinates are generated using exponential normalization, resulting in denser sampling points closer to the joint endpoints.
[0024] Based on the extracted skeletal attributes, the virtual parent-child joint connection lines are divided into three segments. Assuming the total length of each segment is 63 pixels and the prior parameter for the minor axis length is 8.5 pixels, segments of 8.5 pixels each are cut from the start and end points as endpoint encryption regions, and the remaining approximately 46 pixels of internal space are designated as the middle sparse region. Since the endpoint encryption region is located in the pivotal area where limbs intersect, it is easily affected by joint bending, occlusion, and local pose changes. Therefore, a higher density of non-uniform sampling points is allocated in this region.
[0025] During the generation of non-uniform points in the endpoint encryption zone, parameters are set to control the variation of sampling density. Its preferred range is 0.5 to 1.5, for example... Simultaneously, the total number of sampling points within the endpoint encryption zone will be... Set to 6. For the 6th endpoint within the encryption zone. Each sampling point has coordinates according to the formula. It is confirmed that, among them, , For the first The coordinates of each sampling point For the corresponding joint endpoints, It is a unit vector pointing from one joint endpoint to another joint endpoint. The length of the endpoint encryption area. The parameters used to control the variation of sampling density, This represents the total number of sampling points within the endpoint encryption zone. For the endpoint encryption zone on the endpoint side, the endpoint joint is used as the corresponding... And set the direction vector to a unit vector pointing from the endpoint of the end joint to the endpoint of another joint.
[0026] The first point closest to the starting point is... And the 6th point farthest from the starting point and the endpoint. For example, when , hour, The corresponding normalized offset ratio is approximately 0.0042. Multiplying this by the 8.5-pixel region indicates that the sampling point is located very close to the joint endpoint. When the normalization ratio is 1.0, it indicates that the sampling point is located at the end boundary of the endpoint encryption zone. Through the above exponential normalization sampling method, the sampling points can maintain a high density near the joint endpoints and gradually increase the sampling interval as they move away from the endpoints. Combined with 3 to 5 equidistant sampling points in the middle section, the ability to perceive some affinity field features in the abrupt change region of the joint angle of young children can be improved while taking into account the computational efficiency on the edge side.
[0027] In an optional embodiment, the step of multiplying the affinity values of each sampling point by a position attenuation coefficient proportional to the weight vector value and then weighting and summing them to obtain the affinity score includes: Extract the original vector values of each sampling point on the partial affinity field vector map; The Gaussian weight value of the ellipse corresponding to each sampling point is obtained as the position attenuation coefficient; Calculate the dot product of the original vector value and the unit vector of the limb connection, multiply it by the position attenuation coefficient, and then perform a weighted summation along the sampling path. Normalize the summation according to the number of sampling points to obtain the corresponding affinity score.
[0028] For the completed overall sampling point array, assuming the entire connection line has a total of M=15 sampling points after endpoint encryption and mid-section sparsification, a dual-threaded data retrieval step is executed. Using a sub-pixel bilinear interpolation algorithm, the original two-dimensional vector features at each sampling point m are located and extracted on a partial affinity field two-dimensional vector map. Assuming the vector extracted at point 3 of the sequence has a horizontal dimension of 0.3 and a vertical dimension of 0.8, the elliptic Gaussian distribution weight values, which are mapped to the spatial location of this sampling point and have been previously calculated and stored, are retrieved simultaneously. Here, it is assumed that the spatial response weight of this local point is 0.85.
[0029] Extract the ideal limb unit vector from the candidate parent joint to the candidate child joint. For example, the horizontal axis is 0.316 and the vertical axis is 0.948. Performing an inner product dot product operation on this unit vector yields approximately 0.8532, representing the calculated consistency score of the angle between the predicted vector and the ideal vector. A weighted suppression coefficient, approximately 0.725, related to the geometric distribution prior, is used to reduce the artificially high weights of outlier prediction points.
[0030] The above calculation paradigm is iterated until all 15 points have been processed and the algebraic sum is calculated. Normalization can then be performed according to the number of sampling points, thereby deriving the comprehensive affinity score of a single candidate skeletal connection. By nesting the Gaussian soft threshold attenuation strategy with the proportion of the torso of a young child outside the dot product operation, it is possible to filter out the background vectors of artifacts that are detached from the periphery of the limbs with a small computational cost. This helps to reduce the interference of background vectors on affinity correlation results in low-quality blurred frame scenes.
[0031] The update module is used to construct an initial cost matrix based on the product of affinity score, distance decay factor, and confidence score, and to pair using the Hungarian algorithm. The cost value of the previous frame is used as a reference baseline cost value through exponential moving average, and a penalty is applied to pairings where the increase in matching cost exceeds the fluctuation threshold. When limbs overlap, the local gradient direction of the non-extreme position in the neighborhood of the thermal peak of the joint confidence map is extracted, the local gradient direction is rotated by 90 degrees to obtain the local extension direction, and 1 is subtracted from the absolute value of the cosine of the angle between the local extension direction and the major axis to apply additional penalty weights to update the matrix and rematch.
[0032] Obtain the spatial Euclidean distance between candidate parent-child joints. Divide this spatial Euclidean distance by the prior length of the corresponding limb segment or a preset scale parameter to obtain a dimensionless normalized distance. Calculate a distance attenuation factor based on this normalized distance, for example, using a distance attenuation factor... ,in, The spatial Euclidean distance between the candidate joints of the father and son is given. This represents the prior length of the corresponding limb segment. The stability constant is used. The product of affinity score, distance decay factor, and confidence score after nonnegation and normalization is used as the connection confidence score. The result of subtracting the connection confidence score from 1 is used as the basic connection cost, which is then amplified and mapped according to a preset scaling factor to obtain the connection cost used to construct the initial cost matrix. The two-dimensional matrix is initialized, and each candidate connection cost is filled into the corresponding row and column to form the initial cost matrix. The Hungarian algorithm is used for optimal matching of the bipartite graph. At the same time, the historical matching state is maintained in memory. The exponential moving average algorithm is used to calculate the weighted sum of the historical cost values of the previous frame and earlier frames according to the set smoothing factor, which is used as the reference baseline cost value.
[0033] The initial matching cost of the current frame is subtracted from the reference baseline cost to obtain the cost difference. When the cost difference is greater than a set fluctuation threshold, the portion exceeding the fluctuation threshold is extracted as an abnormal increment. This abnormal increment is multiplied by a penalty coefficient and then added to the initial matching cost of the current frame to obtain the final cost. When the cost difference is less than or equal to the fluctuation threshold, no penalty is applied, and the final cost is equal to the initial matching cost of the current frame. When overlapping bounding boxes of different toddler limbs are detected by the intersection-union algorithm, the Sobel operator is used to calculate the local pixel gradient direction vector within the neighborhood of a specific pixel around the corresponding thermal peak in the joint confidence map, after removing extreme points. The local pixel gradient direction vector is rotated by 90 degrees to obtain the local extension direction vector. The dot product of this local extension direction vector and the major axis of the limb connection is calculated and divided by the product of their magnitudes to obtain the cosine of the angle. The value obtained by subtracting the absolute value of the cosine from 1 is taken as the additional penalty weight. The additional penalty weight is multiplied by a preset penalty coefficient and then added as a positive penalty term to the cost of the corresponding candidate pair to update the cost matrix; and a rematch is performed to obtain the topological connection relationship after processing the overlap.
[0034] In an optional embodiment, the step of using the exponential moving average of the previous frame's cost value as a reference baseline cost value and penalizing pairings where the increase in matching cost exceeds a fluctuation threshold includes: Record the cost of a successful matching of the same limb in the previous frame, and update the reference cost of the current frame using the exponential moving average algorithm. Calculate the difference between the initial matching cost of the current frame and the reference baseline cost; When the difference is greater than the set fluctuation threshold, the part of the difference that exceeds the fluctuation threshold is extracted as an abnormal increment. The abnormal increment is multiplied by the penalty coefficient and then added to the initial matching cost of the current frame to obtain the final state cost. When the difference is less than or equal to the fluctuation threshold, no penalty is applied, and the final state cost is equal to the initial matching cost of the current frame.
[0035] A sliding state buffer is maintained, and the cost stability characteristics over historical time series are represented using an exponential moving average mechanism. Specifically, the smoothing update factor for the exponential moving average is preferably set to a range of 0.1 to 0.3, for example, 0.2. If the smoothing cost value accumulated for a certain connection, such as the right thigh region, in T-2 and earlier is 4.5, while the actual cost at the time of matching in frame T-1 is 5.0, then the reference cost value for updating the current frame T can be calculated to be 4.6.
[0036] After performing the initial cost calculation using the Hungarian algorithm on the current T frame, extract the initial matching cost of the current frame corresponding to this connection. Assuming the value rises to 8.2 due to sudden changes in illumination or momentary occlusion within the frame, the diagnostic step calculates the incremental difference. In the judgment criteria, the fluctuation threshold is... Calibration is required based on the general attributes of children's behavior, with a preferred configuration range of 1.5 to 3.0, according to the set calibration parameters. For example, because the actual deviation value of 3.6 is higher than the threshold of 2.0, the penalty control module is triggered.
[0037] During the punishment execution phase, the pre-set punishment coefficient is substituted. The preferred range is 0.5 to 1.2, and here we take... The newly added penalty can be solved using the calculation formula. The final cost is then added to the initial cost, resulting in a final cost value of 9.8. This burst penalty calculation model for detecting transient temporal changes can increase the association cost of suspicious pairings that exhibit non-physiological structural mutations in a single frame. This leads to a lower probability of selecting abnormal candidate pairings during rematching and correction, which is beneficial for improving the stability of video associations in multi-objective interactive scenarios for young children.
[0038] In an optional embodiment, when the limbs overlap, the local gradient direction of the non-extreme location in the neighborhood of the thermal peak of the joint confidence map is extracted, the local gradient direction is rotated by 90 degrees to obtain the local extension direction, and the matrix is updated by subtracting the absolute value of the cosine of the angle between the local extension direction and the major axis from 1 as an additional penalty weight and then rematching, including: The non-extreme pixel positions of the location key points in the neighborhood of the thermal peak on the confidence map are calculated using the discrete difference operator to calculate the local two-dimensional spatial gradient vector at that position. Rotating the two-dimensional spatial gradient vector by 90 degrees yields a local extension direction vector used to characterize the extension trend of the local thermal response. Calculate the absolute value of the cosine of the angle between the local extension direction vector and the direction vector of the major axis of the ellipse connecting the limbs; The result of subtracting the absolute value of the cosine from 1 is converted into an additional penalty weight, and the positive penalty term corresponding to the additional penalty weight is superimposed on the cost matrix for rematching.
[0039] For locally congested regions where deep hugging or multiple overlapping of hands and feet leads to false feature interference, the extreme values of the target joint response are identified from the joint confidence scores (heatmap) output by the dual-branch network. These extreme values are typically set at extreme heatmap vertices at a horizontal 160° and a vertical 200° coordinate in a two-dimensional plane. Within a defined small pixel neighborhood, such as a 3×3 grid area extending outward from the center, non-extreme auxiliary pixels with slightly lower response values but rich variation patterns are identified, for example, at a pixel position of horizontal 161° and vertical 200°. Values are then sought by expanding the search to the surrounding neighborhood using a two-dimensional discrete difference calculation formula. For example, if the surrounding response points are measured as follows... , , as well as Through formula The two-dimensional spatial gradient vector representing the steep decay trend of local confidence is generated by parsing. equal The algebraic vector norm, i.e., the modulus, is calculated to be approximately 0.116. Further rotating this gradient vector by 90 degrees yields the local extension direction vector, which is used to characterize the local extension trend near the thermal response contour lines.
[0040] Obtain the estimated major axis vector representing the macroscopic extension direction of skeletal connections. For example, if the measured trend is 0.6 horizontally and 0.8 vertically, and the standard norm is set to 1.0, the local extension direction vector and the estimated major axis vector are subjected to a mathematical model of dot product and modular standardization to extract the absolute cosine value parameter between them. This absolute cosine value is between 0 and 1. The more consistent the local extension direction is with the major axis direction of the candidate limb, the larger the absolute cosine value, indicating that the local thermal response is more consistent with the morphological characteristics of a single limb extending along the skeletal direction, and the smaller the corresponding additional penalty weight; when the local extension direction deviates more significantly from the major axis direction of the candidate limb, the smaller the absolute cosine value, indicating that the region may be affected by overlapping occlusion or false peak interference, and the corresponding additional penalty weight is larger.
[0041] The additional penalty weight coefficient is calculated by inverting the complement of the numerical values. This additional penalty weight coefficient is then multiplied by a preset penalty coefficient and added as a positive penalty term to the cost of the corresponding candidate pair to update the cost matrix. This process increases the cost of candidate pairs with abnormal local morphological orientations, thereby reducing their probability of being selected during rematching.
[0042] The adjustment module is used to calculate joint features based on the successfully matched skeleton sequences and normalize them according to prior information. The input to the temporal convolutional network is used to obtain the child's behavior category and behavior score. The dilation rate of the temporal convolutional network is adjusted according to the average joint velocity of the input sequence.
[0043] Based on the successfully extracted complete 2D skeleton sequence of a young child from multi-frame matching, the angle sequence of each joint and the positional displacement feature sequence relative to the center point of the human body are calculated. Normalization is performed by dividing the positional displacement features by the corresponding standard deviation based on the limb length standard deviation in the prior database of young children's skeletons. The L2 norm of the spatial coordinate difference between each joint in adjacent frames is calculated, and the average value is calculated over the entire time window to obtain the average joint velocity of the current input sequence. A mapping dictionary of multi-level velocity intervals and dilation rate values is established. This mapping dictionary is retrieved based on the obtained average joint velocity, and the dilation parameter of the one-dimensional convolutional layer Conv1d inside the temporal convolutional network is adjusted. When the average velocity increases, the dilation parameter is reduced to detect high-frequency fine-grained movements.
[0044] The normalized feature sequence is input into the adjusted temporal convolutional network, whose structure is as follows: Figure 2 As shown, spatiotemporal context features are extracted; these features are mapped to the target feature space through a fully connected layer, and then output as a probability distribution vector of the action category by a Softmax classifier. The index label corresponding to the element with the highest probability is taken as the child's behavior category, and this highest probability value is multiplied by 100 to convert it into a percentage format, which is then output as the child's behavior score. The assessment robot keeps the original video locally to protect the child's privacy, and only packages the extracted structured behavior categories, behavior scores, and skeleton coordinate data sequences into a JSON payload, which is then uploaded to the cloud management platform via the MQTT IoT protocol, realizing the collaboration between real-time edge perception and large-scale data analysis in the cloud.
[0045] Obtain existing assessment and scoring standards for the aforementioned behavior categories. These standards reference the motor development goals in the health domain of the "Guidelines for Learning and Development of Children Aged 3-6" and the "National Physical Fitness Measurement Standards (Preschool Section)," including quantitative indicators for achievements such as standing long jump, continuous two-footed jumps, and walking on a balance beam. In addition to these general standards, assessment standards can be developed independently based on the assessment content and actual early childhood education scenarios. For example, for specific sensory integration training, such as obstacle course and balance board challenges, customized motor assessment standards can be developed to assess body control and completion time; for tabletop games and fine motor skills, such as building blocks and bead stringing, assessment standards can be developed to assess hand-eye coordination, motor stability, and attention span; and for daily habits and posture, assessment standards can be developed to assess spinal neutrality maintenance rate and visual distance standardization, etc.
[0046] In one embodiment, the score is calculated based on a weighted average of three dimensions: the first dimension is the maximum probability value output by the temporal convolutional network, representing the basic reliability of behavior recognition; the second dimension is the standardization of movement, obtained by calculating the dynamic temporal warping spatial deformation error between the current normalized skeleton sequence and the standard standard movement skeleton template for the same age group distributed from the cloud, with a smaller error indicating higher standardization; the third dimension is the smoothness of movement, calculated based on the variance of the average joint velocity and joint acceleration as previously calculated, with a smaller variance indicating smoother and more coordinated movement. The calculation results of the above three dimensions are weighted and summed according to preset weights, such as confidence at 20%, standardization at 50%, and smoothness at 30%, and converted into a percentage format as the output score for the child's behavior.
[0047] In an optional embodiment, the dilation rate of the temporal convolutional network is adjusted based on the average joint velocity of the input sequence, including: Statistically analyze the displacement of all joints in the input skeleton sequence within the time window, and calculate the average joint velocity of the current input sequence; Divide the preset baseline motion speed by the average joint speed to obtain the relative speed ratio; Round the relative velocity ratio up and set the minimum expansion rate to 1 to obtain the expansion rate parameter; The dilation rate parameter is applied to the dilated convolutional layer of the temporal convolutional network to adjust the receptive field size.
[0048] The temporal convolutional network takes as input the spatial coordinate features of each joint in a multi-frame input skeleton sequence, and outputs the child's action category or temporal action score based on the temporal skeleton feature mapping. The network structure consists of multiple cascaded dilated convolutional layers combined with activation function layers. The parameter adjustment step in the network's forward computation mechanism begins with measuring the intensity of the macroscopic sequence. Within a predefined time span window, preferably 30 to 60 consecutive video frames covering the minimum cycle of the action, the algorithm calculates the pixel Euclidean displacement of the spatial position differences between adjacent frames for all paired child skeleton systems' main activity nodes, including 14 main joint coordinate points encompassing the head, shoulders, elbows, hips, knees, and the entire body. The average joint velocity, representing the core variable of the current motion attribute, is then calculated by averaging the accumulated values. Taking the identification of a toddler's slow climbing or building block activity as an example, the measured movement speed was weak, and the average displacement parameter remained at a low level of about 2.5 pixels per frame.
[0049] Read the preset baseline motion speed This prior benchmark represents the optimal reference scale for recognizing typical normal running and jumping motion features. It is typically set within the range of 5.0 to 8.0 pixels per frame, based on the statistical distribution or calibration results of the standard test set; assuming it is set to 6.0. The dimensionless relative velocity ratio is obtained by performing a reverse division of the two values; for the above example, the ratio is calculated to be 2.4. Using an upper bound approximation strategy, an integer-wise stretching process is applied to obtain a positive integer base of 3.
[0050] By using the defense lower bound comparison operation max[1,3]=3, the inflation rate parameter under the response configuration is officially set. The parameter is set to 3. This parameter is then pushed and loaded into the hyperparameter list of each dilated convolutional layer within the temporal convolutional network responsible for behavior recognition, replacing the initial fixed setting (e.g., a setting of 1). This extends the observation span of a single convolution mapping on the time series proportionally by a factor of 3. The mechanism of adjusting the dilation rate based on the average joint velocity can expand the temporal receptive field in slow-motion scenarios to enhance the coverage of long-term dependent features; and reduce the temporal sampling span in fast-motion scenarios to improve the response to high-frequency action detail changes, thereby improving the stability of action recognition and behavior scoring under different behavioral rhythms.
[0051] Embodiment 2 of the present invention proposes a data interaction method for an evaluation robot based on edge computing, such as... Figure 3 As shown, it includes: S1, the evaluation robot acquires videos of children, inputs the video frames into a dual-branch network, and outputs a joint confidence map and a partial affinity field vector map; S2, When calculating the affinity field, the unit vector along the limb direction is weighted into an elliptical Gaussian weight vector with the prior of the child's limb length; when sampling the inter-joint line integral, a non-uniform sampling strategy is adopted, with dense sampling in the end-point dense area and sparse sampling in the middle sparse area, and the affinity value of each sampling point is multiplied by the position attenuation coefficient that is proportional to the value of the weight vector and then weighted and summed to obtain the affinity score. S3: Based on the product of affinity score, distance decay factor, and confidence score, an initial cost matrix is constructed, and the Hungarian algorithm is used for pairing. The cost value of the previous frame is used as a reference baseline cost value through exponential moving average. Pairings with matching cost increases exceeding the fluctuation threshold are penalized. When limbs overlap, the local gradient direction of the non-extreme position in the neighborhood of the thermal peak of the joint confidence map is extracted. The local gradient direction is rotated by 90 degrees to obtain the local extension direction. The absolute value of the cosine of the angle between the local extension direction and the major axis is subtracted from 1 to apply additional penalty weights to update the matrix and rematch. S4: Calculate joint features based on the successfully matched skeleton sequence and normalize them according to prior information. Input the features into the temporal convolutional network to obtain the child's behavior category and behavior score. The dilation rate of the temporal convolutional network is adjusted according to the average joint velocity of the input sequence.
[0052] It should be understood that in the foregoing description of the embodiments in this specification, various features are combined in a single embodiment, drawing, or description for the purpose of simplifying the description and to aid in understanding a feature. However, this does not mean that the combination of these features is necessary, and those skilled in the art, upon reading this specification, may readily identify some of the devices as separate embodiments. That is, the embodiments in this specification can also be understood as an integration of multiple secondary embodiments. And the content of each secondary embodiment is valid even if it contains fewer than all the features of a single foregoing disclosed embodiment.
Claims
1. A data interaction system for an evaluation robot based on edge computing, characterized in that, Includes the following modules: The output module is used to evaluate the robot's acquisition of children's videos. It inputs the children's video frames into a dual-branch network and outputs a joint confidence map and a partial affinity field vector map. The summation module is used to calculate the affinity field by weighting the unit vector along the limb direction into an elliptical Gaussian weight vector with the prior of the child's limb length. When sampling the inter-joint line integral, a non-uniform sampling strategy is adopted, with dense sampling in the end-point densified area and sparse sampling in the middle sparse area. The affinity value of each sampling point is multiplied by a position attenuation coefficient that is proportional to the value of the weight vector and then weighted and summed to obtain the affinity score. The update module is used to construct an initial cost matrix based on the affinity score, distance decay factor, and confidence product, and then pair them using the Hungarian algorithm. The previous frame's cost value is used as a reference baseline value through exponential moving average. Pairings where the increase in matching cost exceeds the fluctuation threshold are penalized. When limbs overlap, the local gradient direction of the non-extreme position in the neighborhood of the thermal peak of the joint confidence map is extracted. The local gradient direction is rotated by 90 degrees to obtain the local extension direction. The absolute value of the cosine of the angle between the local extension direction and the major axis is subtracted from 1 to apply additional penalty weights to update the matrix and rematch. The adjustment module is used to calculate joint features based on the successfully matched skeleton sequences and normalize them according to prior information. The input to the temporal convolutional network is used to obtain the child's behavior category and behavior score. The dilation rate of the temporal convolutional network is adjusted according to the average joint velocity of the input sequence.
2. The system according to claim 1, characterized in that, The step of weighting the unit vector along the limb direction into an elliptical Gaussian weight vector with prior knowledge of the child's limb length includes: Obtain the two-dimensional coordinates of the parent and child joints, and calculate the Euclidean distance between the parent and child joints; The minor axis prior parameters matching the child's age group are obtained by querying the skeletal proportion database. Calculate the longitudinal and lateral vertical distances of spatial pixels to the line connecting parent and child joints; Based on the longitudinal projection distance, horizontal and vertical distance, Euclidean distance, and minor axis prior parameters, the elliptic Gaussian weight value of the spatial pixel is calculated using a Gaussian function.
3. The system according to claim 1, characterized in that, The non-uniform sampling strategy employed during inter-joint line integration sampling, involving dense sampling in the endpoint enrichment region and sparse sampling in the middle sparse region, includes: Using the parent-child joint as the starting and ending points, the connection is divided into an end-point encrypted region near both ends and a middle-sparse region in the middle. The length of the end-point encrypted region is determined according to the minor axis scale. Within the endpoint encryption zone, non-uniform sampling point coordinates are generated using exponential normalization, resulting in denser sampling points closer to the joint endpoints.
4. The system according to claim 2 or 3, characterized in that, The step of multiplying the affinity values of each sampling point by a position attenuation coefficient proportional to the weight vector value and then weighting and summing them to obtain the affinity score includes: Extract the original vector values of each sampling point on the partial affinity field vector map; The Gaussian weight value of the ellipse corresponding to each sampling point is obtained as the position attenuation coefficient; Calculate the dot product of the original vector value and the unit vector of the limb connection, multiply it by the position attenuation coefficient, and then perform a weighted summation along the sampling path. Normalize the summation according to the number of sampling points to obtain the corresponding affinity score.
5. The system according to claim 1, characterized in that, The step of using the exponential moving average of the previous frame's cost value as a reference baseline cost value, and penalizing pairings where the increase in matching cost exceeds a fluctuation threshold, includes: Record the cost of a successful matching of the same limb in the previous frame, and update the reference cost of the current frame using the exponential moving average algorithm. Calculate the difference between the initial matching cost of the current frame and the reference baseline cost; When the difference is greater than the set fluctuation threshold, the part of the difference that exceeds the fluctuation threshold is extracted as an abnormal increment. The abnormal increment is multiplied by the penalty coefficient and then added to the initial matching cost of the current frame to obtain the final state cost. When the difference is less than or equal to the fluctuation threshold, no penalty is applied, and the final state cost is equal to the initial matching cost of the current frame.
6. The system according to claim 1, characterized in that, When the limbs overlap, the local gradient direction of the non-extreme position in the neighborhood of the thermal peak of the joint confidence map is extracted. The local gradient direction is rotated by 90 degrees to obtain the local extension direction. The matrix is updated by subtracting the absolute value of the cosine of the angle between the local extension direction and the major axis from 1 as an additional penalty weight, and then re-matched, including: The non-extreme pixel positions of the location key points in the neighborhood of the thermal peak on the confidence map are calculated using the discrete difference operator to calculate the local two-dimensional spatial gradient vector at that position. Rotating the two-dimensional spatial gradient vector by 90 degrees yields a local extension direction vector used to characterize the extension trend of the local thermal response. Calculate the absolute value of the cosine of the angle between the local extension direction vector and the direction vector of the major axis of the ellipse connecting the limbs; The result of subtracting the absolute value of the cosine from 1 is converted into an additional penalty weight, and the positive penalty term corresponding to the additional penalty weight is superimposed on the cost matrix for rematching.
7. The system according to claim 1, characterized in that, The dilation rate of the temporal convolutional network is adjusted based on the average joint velocity of the input sequence, including: Statistically analyze the displacement of all joints in the input skeleton sequence within the time window, and calculate the average joint velocity of the current input sequence; Divide the preset baseline motion speed by the average joint speed to obtain the relative speed ratio; Round the relative velocity ratio up and set the minimum expansion rate to 1 to obtain the expansion rate parameter; The dilation rate parameter is applied to the dilated convolutional layer of the temporal convolutional network to adjust the receptive field size.
8. A data interaction method for an evaluation robot based on edge computing, characterized in that, include: The assessment robot acquires videos of young children, inputs the video frames into a dual-branch network, and outputs a joint confidence map and a partial affinity field vector map. When calculating the affinity field, the unit vector along the limb direction is weighted into an elliptical Gaussian weight vector with the prior of the child's limb length; when sampling the inter-joint line integral, a non-uniform sampling strategy is adopted, with dense sampling in the end-point densified area and sparse sampling in the middle section sparse area. The affinity value of each sampling point is multiplied by a position attenuation coefficient that is proportional to the value of the weight vector and then weighted and summed to obtain the affinity score. An initial cost matrix is constructed based on the product of affinity score, distance decay factor, and confidence score, and the Hungarian algorithm is used for pairing. The cost value of the previous frame is used as a reference baseline cost value through exponential moving average, and a penalty is applied to pairings where the increase in matching cost exceeds the fluctuation threshold. When limbs overlap, the local gradient direction of the non-extreme position in the neighborhood of the thermal peak of the joint confidence map is extracted. The local gradient direction is rotated by 90 degrees to obtain the local extension direction. The absolute value of the cosine of the angle between the local extension direction and the major axis is subtracted from 1 to apply an additional penalty weight to update the matrix and rematch. Joint features are calculated based on the successfully matched skeleton sequences and normalized according to prior information. These features are then input into a temporal convolutional network to obtain the child's behavior category and behavior score. The dilation rate of the temporal convolutional network is adjusted according to the average joint velocity of the input sequence.
9. The method according to claim 8, characterized in that, The step of weighting the unit vector along the limb direction into an elliptical Gaussian weight vector with prior knowledge of the child's limb length includes: Obtain the two-dimensional coordinates of the parent and child joints, and calculate the Euclidean distance between the parent and child joints; The minor axis prior parameters matching the child's age group are obtained by querying the skeletal proportion database. Calculate the longitudinal and lateral vertical distances of spatial pixels to the line connecting parent and child joints; Based on the longitudinal projection distance, horizontal and vertical distance, Euclidean distance, and minor axis prior parameters, the elliptic Gaussian weight value of the spatial pixel is calculated using a Gaussian function.
10. The method according to claim 8, characterized in that, The non-uniform sampling strategy employed during inter-joint line integration sampling, involving dense sampling in the endpoint enrichment region and sparse sampling in the middle sparse region, includes: Using the parent-child joint as the starting and ending points, the connection is divided into an end-point encrypted region near both ends and a middle-sparse region in the middle. The length of the end-point encrypted region is determined according to the minor axis scale. Within the endpoint encryption zone, non-uniform sampling point coordinates are generated using exponential normalization, resulting in denser sampling points closer to the joint endpoints.