Target state identification method and device

By grouping and learning features from sparse radar point clouds, and optimizing deep neural networks, the problem of high data labeling requirements in human keypoint detection is solved, achieving efficient and accurate detection while reducing data labeling costs.

CN120932284APending Publication Date: 2025-11-11FUJITSU LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410564662.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-08
Publication Date
2025-11-11

AI Technical Summary

Technical Problem

Existing methods for human keypoint detection based on wireless signals require a large amount of labeled data. The labeling process is complex and costly, making it difficult to learn general feature representations in sparse radar point clouds.

Method used

By grouping the sparse point cloud information output by radar equipment, two sets of point cloud features are constructed. The similarity between the two sets of point cloud features is learned using a machine learning model, the deep neural network is optimized, and a small amount of labeled data is used for training.

Benefits of technology

It reduces data annotation costs, improves detection accuracy, and is easy to implement, simple to operate, has strong noise resistance, and high privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932284A_ABST
    Figure CN120932284A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a target state recognition method and device. The method comprises the following steps: grouping input point clouds to obtain two groups of point clouds, and inputting the two groups of point clouds into a machine learning model to obtain two groups of features corresponding to the two groups of point clouds; updating the parameters of the machine learning model according to the similarity of the two groups of features corresponding to the two groups of point clouds; and identifying a target state by using the updated machine learning model. Therefore, the problem of low sparse point cloud recognition precision is effectively solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this application relate to the fields of human key point detection, motion detection, and behavior analysis. Background Technology

[0002] Human keypoint detection is crucial for describing human posture and predicting human movements, and plays a fundamental role in research related to behavior (actions, states, etc.) recognition, person tracking, and gait recognition. Currently, human keypoint detection is widely used in patient monitoring systems, smart homes, and smart elderly care, providing health monitoring services for relevant groups (such as patients and the elderly). Furthermore, in noisy environments such as airports and factories, gesture and movement recognition based on human keypoint detection can transmit information more quickly and accurately.

[0003] It should be noted that the above introduction to the technical background is only for the purpose of providing a clear and complete explanation of the technical solutions of this application and for the convenience of those skilled in the art to understand them. It should not be assumed that the above technical solutions are known to those skilled in the art simply because these solutions have been described in the background section of this application. Summary of the Invention

[0004] The inventors discovered that common methods for human keypoint detection based on wireless signals mainly utilize supervised deep learning. These methods require large amounts of labeled data, and the labeling process is highly complex, consuming significant time and resources. Self-supervised learning methods, in particular, rely on pretext tasks to extract supervisory information from large-scale unsupervised data, then train the network to learn representations valuable for downstream tasks. However, learning a universal feature representation for sparse radar point clouds is extremely difficult.

[0005] To address at least one of the aforementioned technical problems or other similar issues, embodiments of this application provide a target state recognition method and apparatus.

[0006] According to one aspect of the embodiments of this application, a target state recognition method is provided, including:

[0007] The input point cloud is grouped to obtain two groups of point clouds, which are then input into a machine learning model to obtain two sets of features corresponding to the two groups of point clouds.

[0008] The parameters of the machine learning model are updated based on the similarity of the two sets of features corresponding to the two sets of point clouds.

[0009] The updated machine learning model is used to identify the target state.

[0010] According to another aspect of the embodiments of this application, a target state identification device is provided, comprising:

[0011] The first processing unit groups the input point cloud into two groups of point clouds, which are then input into a machine learning model to obtain two sets of features corresponding to the two groups of point clouds.

[0012] The update unit updates the parameters of the machine learning model based on the similarity of the two sets of features corresponding to the two sets of point clouds.

[0013] The second processing unit uses the updated machine learning model to identify the target state.

[0014] One of the beneficial effects of this application's embodiments is that it utilizes the sparse point cloud information output by radar equipment to construct two different point cloud self-learning features. The deep neural network is optimized by learning the similarity between the two point cloud features, and a small amount of labeled data is used to guide the network's learning, thereby achieving human keypoint recognition based on a small amount of labeled data. Furthermore, this semi-supervised learning-based target state recognition (e.g., human keypoint recognition) method requires little labeled data, greatly reducing the cost of data annotation, and achieves high detection accuracy. It also features ease of implementation, simple operation, strong noise resistance, and high privacy protection.

[0015] Referring to the following description and accompanying drawings, specific implementation methods of the embodiments of this application are disclosed in detail, indicating how the principles of the embodiments of this application can be adopted. It should be understood that the implementation methods of this application are not limited in scope. Within the spirit and scope of the appended claims, the implementation methods of this application include many changes, modifications, and equivalents. Attached Figure Description

[0016] The accompanying drawings, which form part of the specification, are used to provide a further understanding of the embodiments of this application and illustrate the implementation methods of this application, together with the textual description, to explain the principles of this application. Obviously, the drawings described below are merely some embodiments of this application, and those skilled in the art can obtain other implementation methods based on these drawings without creative effort. In the drawings:

[0017] Figure 1 This is a schematic diagram of a target state recognition method according to an embodiment of this application;

[0018] Figure 2 This is a schematic diagram illustrating a specific example of the grouping method according to an embodiment of this application;

[0019] Figure 3 yes Figure 2 The example shown is a schematic diagram of an example of the original radar point cloud;

[0020] Figure 4 yes Figure 2 A schematic diagram of an example of the point cloud obtained in group 1 from the examples;

[0021] Figure 5 yes Figure 2 The example obtained Figure 2 A schematic diagram of an example point cloud.

[0022] Figure 6 This is a schematic diagram showing the two sets of features obtained after inputting two sets of point clouds into a machine learning model.

[0023] Figure 7 This is a schematic diagram of updating the parameters of the machine learning model based on the similarity of the two sets of features corresponding to the two sets of point clouds.

[0024] Figure 8 This is a schematic diagram illustrating the calculation of the loss function for each group of point cloud data, including both labeled and unlabeled data.

[0025] Figure 9 This is a schematic diagram illustrating the updating of the total loss function of a machine learning model according to the method of an embodiment of this application;

[0026] Figure 10 This is a schematic diagram of an implementation scenario of an embodiment of this application;

[0027] Figure 11 This is a schematic diagram of the comparison results;

[0028] Figure 12 This is a schematic diagram of a target state recognition device according to an embodiment of this application;

[0029] Figure 13 This is a schematic diagram of an electronic device according to an embodiment of this application. Detailed Implementation

[0030] Referring to the accompanying drawings, the foregoing and other features of the embodiments of this application will become apparent from the following description. Specific embodiments of this application are specifically disclosed in the description and drawings, illustrating partial implementations in which the principles of the embodiments of this application can be adopted. It should be understood that this application is not limited to the described embodiments; rather, the embodiments of this application include all modifications, variations, and equivalents falling within the scope of the appended claims.

[0031] In the embodiments of this application, the terms "first," "second," etc., are used to distinguish different elements by name, but do not indicate the spatial arrangement or chronological order of these elements, and these elements should not be limited by these terms. The term "and / or" includes any one or more of the terms listed in association and all combinations thereof. The terms "comprising," "including," "having," etc., refer to the presence of the stated features, elements, components, or assemblies, but do not exclude the presence or addition of one or more other features, elements, components, or assemblies.

[0032] In the embodiments of this application, the singular forms "a," "the," etc., including the plural forms, should be broadly understood as "a kind" or "a class" rather than limited to the meaning of "an." Furthermore, the term "the" should be understood to include both the singular and plural forms, unless the context explicitly indicates otherwise. Additionally, the term "according to" should be understood as "at least partially based on…," and the term "based on" should be understood as "at least partially based on…," unless the context explicitly indicates otherwise.

[0033] Features described and / or illustrated for one embodiment may be used in the same or similar manner in one or more other embodiments, combined with features in other embodiments, or substituted for features in other embodiments. The term "comprising / including" as used herein means the presence of a feature, integral, step, or component, but does not exclude the presence or addition of one or more other features, integrals, steps, or components.

[0034] In this embodiment, the radar can be a millimeter-wave (mmWave) radar, but is not limited to this. The radar transmits electromagnetic waves through a transmitting antenna, and after reflection from different objects, receives the corresponding reflected waves (which can be called radar echo information or radar echo signals). By processing and analyzing the radar echo information, information such as the physiological signals of the detected target can be extracted, which can meet the needs of many application scenarios.

[0035] In the embodiments of this application, the detection target can be people of various ages, such as the elderly, children, or elderly people and / or caregivers, children and / or guardians. This application is not limited to these; the detection target can also be animals with vital characteristics. The following explanation uses the human body as an example.

[0036] First aspect of the embodiments

[0037] This application provides a target state recognition method. Figure 1 This is a schematic diagram of a target state recognition method according to an embodiment of this application, as shown below. Figure 1 As shown, the method includes:

[0038] 110. The input point cloud is grouped to obtain two sets of point clouds, which are then input into the machine learning model to obtain two sets of features corresponding to the two sets of point clouds.

[0039] 120. Update the parameters of the machine learning model based on the similarity of the two sets of features corresponding to the two sets of point clouds.

[0040] 130. Use the updated machine learning model described above to identify the target state.

[0041] It is worth noting that the above appendix Figure 1 The embodiments of this application have only been illustrated schematically, and the application is not limited thereto. For example, the execution order between various operations can be appropriately adjusted, and other operations can be added or some operations can be removed. Those skilled in the art can make appropriate modifications based on the above description, and are not limited to the above-described embodiments. Figure 1 The records.

[0042] According to the above embodiments, two different self-learning point cloud features are constructed using sparse point cloud information output by radar equipment. The deep neural network is optimized by learning the similarity between the two point cloud features, and a small amount of labeled data is used to guide the network's learning, thereby achieving human keypoint recognition based on a small amount of labeled data. Furthermore, this semi-supervised learning-based target state recognition (e.g., human keypoint recognition) method requires little labeled data, greatly reducing the cost of data annotation, and has high detection accuracy. It also features ease of implementation, simple operation, strong noise resistance, and high privacy protection.

[0043] In the embodiments of this application, the input point cloud can be the original point cloud output by the radar device, or it can be the point cloud obtained after preprocessing the original point cloud output by the radar device. The preprocessing here is, for example, filtering, noise reduction, etc.

[0044] According to the above embodiment, by grouping the input point cloud, two sets of point clouds are obtained, which are then input into the machine learning model as the input point clouds of the machine learning model.

[0045] In some embodiments, the two sets of point clouds can be transformed first, for example, through random transformation, and the transformed point clouds can be used as the input point clouds for a machine learning model, thereby enhancing the diversity of the data. This application does not limit the transformation method. In some possible implementations, the point cloud velocity information of the two sets of point clouds can be mirrored to obtain the transformed point cloud; in other possible implementations, the point cloud position information of the two sets of point clouds can be mirrored to obtain the transformed point cloud.

[0046] In the above embodiment, by inputting the two transformed point clouds into the machine learning model respectively, two sets of features can be obtained. In operation 120, the parameters of the machine learning model can be updated according to the similarity of the last layer features corresponding to the two transformed point clouds.

[0047] In this application embodiment, the network architecture of the machine learning model is not limited. For example, the machine learning model can be a neural network model, a clustering model, a dimensionality reduction model, a Naive Bayes model, a support vector machine model, a random forest model, a decision tree model, a logistic regression model, a linear regression model, etc.

[0048] In some embodiments, the grouping principle is to minimize the number of point clouds belonging to both groups, i.e., to minimize the overlap, for example, to be 0 or M1 + M2 - M. Here, M is the number of point clouds in the input point cloud, and M1 and M2 are the number of point clouds in the two groups, respectively. For example, if the input point cloud M = 300, according to the above embodiment, these 300 point clouds need to be grouped first to obtain two groups of point clouds, one group containing M1 point clouds and the other group containing M2 point clouds.

[0049] In some possible implementations, when the number of points (M) in the input point cloud is greater than or equal to the sum of the number of points in the two groups of point clouds to be grouped (M1+M2), a point cloud extraction algorithm is used to extract two groups of point clouds from the input point cloud according to the number of points required for each group of point clouds to be grouped, and the overlap between the two groups of point clouds is 0.

[0050] For example, if the number of points in the input point cloud is M, the number of points required for the first group of point clouds is M1, and the number of points required for the second group of point clouds is M2, and M1 < M, M2 < M, and M1 + M2 ≤ M, then a point cloud extraction algorithm can be used to extract M1 points from the input point cloud as the points in the first group, and then the point cloud extraction algorithm can be used to extract M2 points from the remaining M-M1 points as the points in the second group.

[0051] Taking M=300 as an example again, assuming M1=150 and M2=150, we can directly use the point cloud extraction algorithm to extract 150 points from these 300 point clouds as the first group of point clouds, and use the remaining 150 point clouds as the second group of point clouds. For another example, assuming M=300, M1=150 and M2=100, we can directly use the point cloud extraction algorithm to extract 150 points from these 300 point clouds as the first group of point clouds, and use the point cloud extraction algorithm to extract 100 points from the remaining 150 point clouds as the second group of point clouds. Of course, we can also first extract 100 points from the 300 point clouds as the second group of point clouds, and then extract 150 points from the remaining 200 point clouds as the first group of point clouds, and so on.

[0052] In some other possible implementations, when the number of points (M) in the input point cloud is less than the sum of the number of points in the two groups of point clouds to be grouped (M1+M2), a point cloud extraction algorithm is used to extract two groups of point clouds from the input point cloud according to the number of points required for each group of point clouds to be grouped. The overlap between the two groups of point clouds is greater than 0, which is M1+M2-M.

[0053] For example, if the number of points in the input point cloud is M, the number of points required for the first group of point clouds is M1, and the number of points required for the second group of point clouds is M2, and M1 < M, M2 < M, M1 + M2 > M, and M1 < M2, then first divide the input M point clouds into two groups, with each group containing 2 / M points; use a point cloud extraction algorithm to extract M2 - 2 / M points from one group and merge them with the points from the other group to form the point cloud in the second group; use a point cloud extraction algorithm to extract M1 points from the remaining points in one group and / or the points from the other group to form the point cloud in the first group.

[0054] Taking M=300 as an example again, assuming M1=150 and M2=200, we can first divide these 300 point clouds into two groups. Then, we can extract 200-150=50 point clouds from one group and merge them with 150 point clouds from the other group to form the point cloud in the second group. Then, we can extract 150 point clouds from the remaining 150-50=100 point clouds in one group and / or the 150 point clouds from the other group to form the point cloud in the first group.

[0055] For example, if the number of points in the input point cloud is M, the number of points required for the first group of point clouds is M1, and the number of points required for the second group of point clouds is M2, and M1 < M, M2 < M, M1 + M2 > M, and M1 < M2, then firstly, a point cloud extraction algorithm is used to extract M2 points from the input M point clouds as the point clouds in the second group. Then, a point cloud extraction algorithm is used to extract M1 - (M - M2) points from the second group of point clouds and merge them with the remaining M - M2 points from the input M point clouds as the point clouds in the first group.

[0056] Taking M=300 as an example again, assuming M1=150 and M2=200, we can first extract 200 point clouds from these 300 point clouds as the point clouds in the second group, and then extract 150-(300-200)=50 point clouds from these 200 point clouds and merge them with the remaining 300-200=100 point clouds from these 300 point clouds as the point clouds in the first group.

[0057] In the above embodiments, there are no restrictions on the point cloud extraction algorithm. The point cloud extraction algorithm may be, for example, a grid sampling method, a random sampling method, a farthest point sampling method, a uniform sampling method, a voxel sampling method, or any combination of the above sampling methods.

[0058] Grid sampling utilizes the three-dimensional spatial information of point clouds to discretize the original point cloud into small grids. Each small grid contains several points, and the point closest to the center of the grid is taken as the sampling point. This sampling point can then be used as the extracted point cloud information.

[0059] Random sampling is the simplest and most efficient sampling method. It generates sampling points, i.e., the extracted point cloud information, by randomly selecting seed points and adding them to the sampling point set.

[0060] The farthest point sampling method is an iterative method for selecting the farthest point in a point cloud. During each sampling, the point closest to the set is found and added to the sampling point set, which is then used as the extracted point cloud information.

[0061] The uniform sampling method samples at fixed intervals to extract the required point cloud information.

[0062] Voxel sampling is a method to reduce the number of point clouds while preserving their shape features and spatial structure information. It involves meshing the point cloud space (i.e., voxelizing), then sampling one point from each voxel to obtain a downsampled point cloud, which is then used as the extracted point cloud information.

[0063] The above is only a brief explanation of the point cloud extraction algorithm. For details, please refer to the relevant technologies. It will not be elaborated here.

[0064] Figure 2 This is a schematic diagram illustrating a specific example of the grouping method in an embodiment of this application. In this example, it is assumed that the number of original radar point clouds (input point clouds) is M. Figure 2 As shown, the method includes:

[0065] 210: Let M1 be the number of point clouds in one group and M2 be the number of point clouds in another group, and satisfy M1 < M2 < M;

[0066] 220: Use a point cloud extraction algorithm to extract M² point clouds from M point clouds;

[0067] 230: The extracted M2 point clouds are used as the point clouds in group 1 (corresponding to the second group mentioned above), thus obtaining group 1;

[0068] 240: The remaining point cloud has a point cloud count of M-M2;

[0069] 250: Determine if the sum of M1 and M2 is less than or equal to M; if yes, execute 260, otherwise execute 270;

[0070] 260: Extract point cloud M1 from the remaining point cloud and use it as the point cloud in group 2 (corresponding to the first group mentioned above);

[0071] 270: Use a point cloud extraction algorithm to extract M2+M1-M point clouds from group 1;

[0072] 280: Merge the point cloud extracted in operation 270 with the remaining point cloud in operation 240 to form the point cloud in group 2 (corresponding to the first group mentioned above), thus obtaining group 2.

[0073] It is worth noting that the above appendix Figure 2 The embodiments of this application have only been illustrated schematically, and the application is not limited thereto. For example, the execution order between various operations can be appropriately adjusted, and other operations can be added or some operations can be removed. Those skilled in the art can make appropriate modifications based on the above description, and are not limited to the above-described embodiments. Figure 2 The records.

[0074] According to the above embodiment, the principle of grouping is to minimize the overlap between different groups. The point cloud extraction algorithm is used to first obtain group 1 with the largest number of points. Then, the sum of the number of point clouds in different groups is compared with the total number of the original radar point cloud. If M1+M2≤M, then points in group 2 are extracted from the remaining points. If M1+M2>M, then some points in group 2 are extracted from group 1. Then, the extracted points are combined with the remaining points to obtain group 2. Figure 3 yes Figure 2 The example shown is a schematic diagram of an example of the original radar point cloud. Figure 4 yes Figure 2 A schematic diagram of an example of the point cloud obtained in group 1 from the examples. Figure 5 yes Figure 2 The example obtained Figure 2 A schematic diagram of an example point cloud.

[0075] In this embodiment of the application, in operation 110, by inputting two sets of point clouds (which may be obtained by grouping the input point clouds, or by transforming the two grouped point clouds, and for ease of explanation, they are collectively referred to as "two sets of point clouds") into the machine learning model, two sets of features corresponding to the two sets of point clouds can be obtained.

[0076] Figure 6 This is a schematic diagram showing the two sets of features obtained after inputting two sets of point clouds into a machine learning model. For example... Figure 6 As shown, when the original point cloud x (e.g. Figure 3After grouping, at least two sets of network input point cloud information x1 and x2 (as shown) will be obtained. Figure 4 and Figure 5 (As shown). The original point cloud x and the two resulting point clouds x1 and x2 each have K information features, where K=6. The information features are the point cloud velocity v, energy p, position x, position y, position z, and frame number fid, represented as [v, p, x, y, z, fid]. After these two point clouds x1 and x2 are input into the machine learning model, intermediate layer features y1′ and y2′ are obtained, as well as the final layer features (outputs) y1 and y2. It is expected that the outputs (y1 and y2) corresponding to these two point clouds x1 and x2, or their corresponding intermediate layer features (y1′ and y2′), are consistent.

[0077] According to the above embodiment, in operation 120, the parameters of the above machine learning model can be updated based on the similarity of the two sets of features (y1′ and y2′, or y1 and y2) corresponding to the two sets of point clouds (x1 and x2).

[0078] Figure 7 This is a schematic diagram illustrating how the parameters of the machine learning model are updated based on the similarity of the two sets of features corresponding to these two sets of point clouds, as shown below. Figure 7 As shown, in some embodiments, the method includes:

[0079] 710: Calculate the loss function for each group of point clouds, including both labeled data (point clouds with ground truth values) and unlabeled data (point clouds without ground truth values);

[0080] 720: The calculated loss functions are weighted and summed to obtain the total loss function of the machine learning model;

[0081] 730: Update the parameters of the machine learning model using the total loss function described above.

[0082] It is worth noting that the above appendix Figure 7 The embodiments of this application have only been illustrated schematically, and the application is not limited thereto. For example, the execution order between various operations can be appropriately adjusted, and other operations can be added or some operations can be removed. Those skilled in the art can make appropriate modifications based on the above description, and are not limited to the above-described embodiments. Figure 7 The records.

[0083] In some embodiments, for labeled data, a first loss function is calculated, and optionally, a second loss function may also be calculated. The first loss function is the loss function between the predicted value and the true value of the labeled data corresponding to a set of point clouds after the machine learning model, and the second loss function is the loss function between the features (intermediate layer features and / or the last layer features) of the labeled data corresponding to two sets of point clouds after the machine learning model. For unlabeled data, a third loss function is calculated, which is the loss function between the features (intermediate layer features and / or the last layer features) of the unlabeled data corresponding to two sets of point clouds after the machine learning model.

[0084] Figure 8 This is a schematic diagram illustrating the calculation of the loss function for each group of point clouds in the above embodiments, specifically for labeled data (point clouds with truth values) and unlabeled data (point clouds without truth values).

[0085] like Figure 8 As shown, taking the input point cloud with 300 points as an example, after grouping, two groups of point clouds are obtained, namely x1 and x2. After being input into the machine learning model, the intermediate layer features are y1' and y2', and the final layer features are y1 and y2.

[0086] In the above embodiments, in some possible implementations, the first loss function and the second loss function can be weighted and summed to obtain the loss function for labeled data. label The third loss function is weighted and summed to obtain the loss function for unlabeled data. unlabel Based on the loss function for labeled data label loss function for unlabeled data unlabel This yields the total loss function of the machine learning model.

[0087] Among them, the loss function for labeled data is loss. label It can be represented as: loss label =α1*loss1+β1*loss2+γ1*loss3+ρ1*loss4; Loss function for unlabeled data unlabel It can be represented as: loss unlabel =β²*loss₅ + γ²*loss₆ + ρ²*loss₇; The total loss function of a machine learning model can be expressed as: loss = loss label +loss unlabel .

[0088] Where α1, β1, γ1, ρ1, β2, γ2, and ρ2 are weighting coefficients, α1 > 0, β1 > 0, β2 > 0, γ1 ≥ 0, ρ1 ≥ 0, γ2 ≥ 0, and ρ2 ≥ 0, respectively.

[0089] like Figure 8 As shown, loss1 is a set of point clouds ( Figure 8 Taking point cloud group x1 as an example, loss1 is the loss function between the predicted value (y1) and the true value (true) of the labeled data; loss2 is the loss function between the last layer features (y1 and y2) of the two labeled point cloud groups; loss5 is the loss function between the last layer features (y1 and y2) of the two unlabeled point cloud groups; loss3 is the loss function between the intermediate layer features (y1' and y2') of the two labeled point cloud groups; loss6 is the loss function between the intermediate layer features (y1' and y2') of the two unlabeled point cloud groups. The loss function is between y1' and y2'; loss4 is the loss function between the mixed features (including intermediate layer features y1' and y2' and the last layer features y1 and y2) corresponding to the two sets of point clouds with labeled data; loss7 is the loss function between the mixed features (including intermediate layer features y1' and y2' and the last layer features y1 and y2) corresponding to the two sets of point clouds with unlabeled data. loss4 and loss7 can be obtained by comparing and analyzing the weighted result of the last layer features of the two point clouds with the weighted result of a certain intermediate layer feature.

[0090] For example, loss4 can be calculated using the following formula (1):

[0091] loss4 = Mse(mixed) y ,λ*y1+(1-λ)*y2) (1)

[0092] Where Mse (Mean Squared Error) represents the root mean square error loss function, mixed y It is obtained by weighted summation of the intermediate feature values ​​of two sets of point clouds, and then feeding the weighted result into the network of the subsequent layers for recognition. λ is a random number given a beta distribution, used to represent the weights of the feature values ​​y1 and y2 of the two output layers.

[0093] The above example uses the root mean square error loss function, but this application is not limited to this. The loss functions mentioned above can also be the absolute value error loss function or others.

[0094] Figure 9 This is a schematic diagram illustrating the updating of the total loss function of a machine learning model according to the method of an embodiment of this application.

[0095] like Figure 9 As shown, by grouping the input point cloud (original point cloud or transformed point cloud), the two point cloud groups x1 and x2 with the least overlap are obtained. These two point cloud groups x1 and x2 are then input into the machine learning model to obtain two sets of features. The first set of features includes the label data y. 1l and unlabeled data y 1u The second set of features also includes labeled data y. 2l and unlabeled data y 2u This feature y 1l y 1u and y 2l y 2u These can be intermediate features of a machine learning model or features from the final layer of a machine learning model. For a set of point cloud features, the labeled data y represents the features. 1l Labeled data y of features corresponding to another set of point clouds and / or other point clouds. 2l The loss function between the predicted and true values ​​can be calculated, denoted as Ll; for a set of unlabeled data y corresponding to the features of a point cloud. 1u Unlabeled data y of features corresponding to another set of point clouds 2u We can calculate the loss function between the predicted values ​​of the two sets of unlabeled data corresponding to the two sets of point clouds, denoted as Lu; by weighted summing of these two loss functions, we obtain the total loss function of the machine learning model, L = αLl + βLu.

[0096] exist Figure 9 In the example, taking labeled data as an example, the loss function between the predicted value and the true value of the corresponding feature is calculated. This application is not limited to this, and applies to the labeled data y corresponding to these two sets of point clouds. 1l and y 2l It can also calculate the loss function between its predicted values.

[0097] According to the above embodiments, the parameters of the machine learning model can be updated using the total loss function, i.e., the machine learning model can be optimized. Then, the updated (optimized) machine learning model can be used to identify the target state. This improves the accuracy and reliability of the identification process.

[0098] The effects of the embodiments of this application will be illustrated below with a specific example.

[0099] Figure 10 This is a schematic diagram of an implementation scenario of an embodiment of this application, as shown below. Figure 10 As shown in the table below, in this scenario, the walking area is 3.7m × 4.4m, and the collected behaviors are shown in Table 1. Among them, there are 20 targets with labels and 68 targets without labels.

[0100] Table 1:

[0101] ID Behavior 1 Walk → Station 2 Walk → Sit 3 Walk → Lie down on the bed 4 Walk → Lie down on the floor 5 Sitting → Falling 6 Lie down on the bed → fall down 7 Standing → Falling 8 Walk → Fall face down towards the radar 9 Walk → Fall backwards towards the radar, face down

[0102] Table 2 below illustrates the application of existing supervised methods, self-supervised methods, and semi-supervised methods using embodiments of this application. Figure 10 The comparison results of target state recognition in the scenario shown.

[0103] Table 2:

[0104] Loss Epoch Person117 Person128 Person049 Mean Supervision 80 0.0689 0.0481 0.0457 0.0542 Self-monitoring 30 0.0667 0.0459 0.0434 0.0520 Semi-supervised 70 0.0628 0.0441 0.0441 0.0503

[0105] Figure 11 This is a schematic diagram of the comparison results above. See Table 2 and... Figure 11 As shown, compared with existing supervised and self-supervised methods, the semi-supervised method of this application for target state recognition has higher accuracy.

[0106] The method described in this application can be applied to scenarios such as smart homes, smart elderly care, and police stations, and can be used for key point detection, posture detection, human behavior analysis, fall detection, etc. According to the method described in this application, the cost of data collection and labeling is reduced because labeling data to obtain truth values ​​is both difficult and time-consuming, and unlabeled data is usually abundant and can be easily or cheaply obtained. Furthermore, it improves the versatility of the model, making it easy to transfer to different application scenarios, such as from homes to nursing homes.

[0107] The above embodiments are merely illustrative examples of embodiments of this application, but this application is not limited thereto, and appropriate modifications can be made based on the above embodiments. For example, the above embodiments can be used alone, or one or more of the above embodiments can be combined.

[0108] According to the above embodiments, two different self-learning point cloud features are constructed using sparse point cloud information output by radar equipment. The deep neural network is optimized by learning the similarity between the two point cloud features, and a small amount of labeled data is used to guide the network's learning, thereby achieving human keypoint recognition based on a small amount of labeled data. Furthermore, this semi-supervised learning-based target state recognition (e.g., human keypoint recognition) method requires little labeled data, greatly reducing the cost of data annotation, and has high detection accuracy. It also features ease of implementation, simple operation, strong noise resistance, and high privacy protection.

[0109] Second aspect of the embodiments

[0110] This application provides a target state recognition device, and the contents that are the same as those in the first aspect of the embodiment will not be repeated.

[0111] Figure 12 This is a schematic diagram of a target state recognition device according to an embodiment of this application, as shown below. Figure 12 As shown, the target state identification device 1200 of this application embodiment includes:

[0112] The first processing unit 1210 groups the input point cloud into two groups of point clouds, which are then input into the machine learning model to obtain two sets of features corresponding to the two groups of point clouds.

[0113] The update unit 1220 updates the parameters of the machine learning model based on the similarity of the two sets of features corresponding to the two sets of point clouds.

[0114] The second processing unit 1230 uses the updated machine learning model described above to identify the target state.

[0115] In some embodiments, the number of point clouds belonging to both of the above two point cloud groups is minimized, for example, 0 or M1+M2-M, where M is the number of point clouds in the input point cloud, and M1 and M2 are the number of point clouds in the two above point cloud groups, respectively.

[0116] In some embodiments, when the number of points in the input point cloud is greater than or equal to the sum of the number of points in the two groups of point clouds to be grouped, the first processing unit 1210 extracts two groups of point clouds from the input point cloud using a point cloud extraction algorithm based on the number of points required for each group of point clouds to be grouped, and the overlap degree of the two groups of point clouds is 0; when the number of points in the input point cloud is less than the sum of the number of points in the two groups of point clouds to be grouped, the first processing unit 1210 extracts two groups of point clouds from the input point cloud using a point cloud extraction algorithm based on the number of points required for each group of point clouds to be grouped, and the overlap degree of the two groups of point clouds is greater than 0.

[0117] In the above embodiments, in some possible implementations, the first processing unit 1210 first uses a point cloud extraction algorithm to extract M1 point clouds from the input point cloud as the point clouds in the first group, and then uses a point cloud extraction algorithm to extract M2 point clouds from the remaining M-M1 point clouds as the point clouds in the second group; wherein, the number of point clouds in the input point cloud is M, the number of point clouds required for the first group of point clouds is M1, the number of point clouds required for the second group of point clouds is M2, and M1 < M, M2 < M, M1 + M2 ≤ M.

[0118] In the above embodiments, in some other possible implementations, the first processing unit 1210 first divides the input point cloud into two groups on average; uses a point cloud extraction algorithm to extract M2-2 / M point clouds from one group and merges them with the point clouds in the other group to form the point cloud in the second group; uses a point cloud extraction algorithm to extract M1 point clouds from the remaining point clouds in one group and / or the point clouds in the other group to form the point cloud in the first group; or, the first processing unit 1210 first uses a point cloud extraction algorithm to extract M2 point clouds from the input point cloud to form the point cloud in the second group, and then uses a point cloud extraction algorithm to extract M1-(M-M2) point clouds from the point clouds in the second group and merges them with the remaining M-M2 point clouds in the input point cloud to form the point cloud in the first group; wherein, the number of point clouds in the input point cloud is M, the number of point clouds required for the first group of point clouds is M1, the number of point clouds required for the second group of point clouds is M2, and M1 < M, M2 < M, M1+M2 > M, M1 < M2.

[0119] In some embodiments, the first processing unit 1210 further transforms the two sets of point clouds to obtain two transformed sets of point clouds, and inputs the two transformed sets of point clouds into a machine learning model to obtain two sets of features or inputs; the update unit 1220 updates the parameters of the machine learning model according to the similarity of the last layer features corresponding to the two transformed sets of point clouds.

[0120] In the above embodiments, the first processing unit 1210 transforms the two sets of point clouds, which may include:

[0121] Perform a mirror transformation on the point cloud velocity information of the two sets of point clouds mentioned above; or,

[0122] The point cloud position information of the two sets of point clouds above is mirrored.

[0123] In some embodiments, the updating unit 1220 updates the parameters of the machine learning model, including:

[0124] For each group of point cloud data, the loss function is calculated separately for both labeled and unlabeled data.

[0125] The calculated loss functions are weighted and summed to obtain the total loss function of the above machine learning model.

[0126] The parameters of the machine learning model are updated using the total loss function described above.

[0127] In the above embodiments, the update unit 1220 calculates the loss function for the labeled and unlabeled data of each group of point clouds, which may include:

[0128] For labeled data, update unit 1220 calculates a first loss function and a second loss function. The first loss function is the loss function between the predicted value and the true value of the labeled data corresponding to a set of point clouds after the machine learning model. The second loss function is the loss function between the features of the labeled data corresponding to two sets of point clouds after the machine learning model.

[0129] For unlabeled data, update unit 1220 calculates a third loss function, which is the loss function between the features of the unlabeled data corresponding to the two sets of point clouds after the above machine learning model.

[0130] In the above embodiments, the update unit 1220 performs a weighted summation of the calculated loss function to obtain the total loss function of the machine learning model, which may include:

[0131] The first loss function and the second loss function are weighted and summed to obtain the loss function for labeled data. label The third loss function is then weighted and summed to obtain the loss function for unlabeled data. unlabel And based on the loss function of labeled data label loss function for unlabeled data unlabel This yields the total loss function of the machine learning model.

[0132] in,

[0133] loss function for labeled data label Represented as: loss label =α1*loss1+β1*loss2+γ1*loss3+ρ1*loss4;

[0134] loss function for unlabeled data unlabel Represented as: loss unlabel =β2*loss5+γ2*loss6+ρ2*loss7;

[0135] The total loss function of a machine learning model is expressed as: loss = loss label +loss unlabel ;

[0136] Where α1, β1, γ1, ρ1, β2, γ2, and ρ2 are weighting coefficients, α1 > 0, β1 > 0, β2 > 0, γ1 ≥ 0, ρ1 ≥ 0, γ2 ≥ 0, and ρ2 ≥ 0;

[0137] loss1 is the loss function between the predicted and actual values ​​of a set of labeled data corresponding to point clouds;

[0138] loss2 and loss5 are the loss functions between the last layer features corresponding to the two sets of point clouds mentioned above.

[0139] loss3 and loss6 are the loss functions between the intermediate layer features corresponding to the two sets of point clouds mentioned above.

[0140] loss4 and loss7 are the loss functions between the mixed features corresponding to the two sets of point clouds mentioned above. The mixed features include the intermediate layer features and the last layer features.

[0141] It is worth noting that the above description only covers the components or modules relevant to this application, but this application is not limited thereto. The target state identification device 1200 may also include other components or modules, and for details regarding these components or modules, please refer to related technologies.

[0142] For the sake of simplicity, Figure 12 The diagram only exemplifies the connection relationships or signal flow between various components or modules; however, those skilled in the art should understand that various related technologies, such as bus connections, can be employed. The aforementioned components or modules can be implemented using hardware facilities such as processors and memory; this application does not limit the scope of the embodiments.

[0143] The above embodiments are merely illustrative examples of embodiments of this application, but this application is not limited thereto, and appropriate modifications can be made based on the above embodiments. For example, the above embodiments can be used alone, or one or more of the above embodiments can be combined.

[0144] According to the above embodiments, two different self-learning point cloud features are constructed using sparse point cloud information output by radar equipment. The deep neural network is optimized by learning the similarity between the two point cloud features, and a small amount of labeled data is used to guide the network's learning, thereby achieving human keypoint recognition based on a small amount of labeled data. Furthermore, this semi-supervised learning-based target state recognition (e.g., human keypoint recognition) method requires little labeled data, greatly reducing the cost of data annotation, and has high detection accuracy. It also features ease of implementation, simple operation, strong noise resistance, and high privacy protection.

[0145] Third aspect of the embodiments

[0146] This application provides an electronic device including a target state identification device 1200 as described in the second aspect of the embodiment, the contents of which are incorporated herein by reference. This electronic device may be, for example, a computer, server, workstation, laptop computer, smartphone, etc.; however, this application is not limited thereto.

[0147] Figure 13This is a schematic diagram of an electronic device according to an embodiment of this application. Figure 13 As shown, the electronic device 1300 may include a processor (e.g., a central processing unit, CPU) 1310 and a memory 1320; the memory 1320 is coupled to the central processing unit 1310. The memory 1320 may store various types of data; it also stores information processing programs and executes these programs under the control of the processor 1310.

[0148] In some embodiments, the functionality of the target state recognition device 1200 is integrated into the processor 1310. The processor 1310 is configured to implement the target state recognition method as described in the first aspect of the embodiment.

[0149] In some embodiments, the target state recognition device 1200 is configured separately from the processor 1310. For example, the target state recognition device 1200 can be configured as a chip connected to the processor 1310, and the function of the target state recognition device 1200 can be realized through the control of the processor 1310.

[0150] For example, processor 1310 is configured to perform the following control: grouping the input point cloud into two groups of point clouds, inputting them into a machine learning model to obtain two sets of features or outputs corresponding to the two groups of point clouds; updating the parameters of the machine learning model based on the similarity of the two sets of features or outputs corresponding to the two groups of point clouds; and using the updated machine learning model to identify the target state.

[0151] In addition, such as Figure 13 As shown, the electronic device 1300 may further include: an input / output (I / O) device 1330 and a display 1340, etc.; the functions of the above components are similar to those in the prior art, and will not be described in detail here. It is worth noting that the electronic device 1300 is not necessarily required to include... Figure 13 All components shown; in addition, the electronic device 1300 may also include Figure 13 For components not shown, please refer to relevant technologies.

[0152] This application also provides a computer-readable program, wherein when the program is executed in an electronic device, the program causes the computer in the electronic device to perform the target state recognition method as described in the first aspect embodiment.

[0153] This application also provides a storage medium storing a computer-readable program, wherein the computer-readable program causes a computer in an electronic device to perform the target state recognition method as described in the first aspect embodiment.

[0154] The apparatus and methods described above in this application can be implemented in hardware or in combination with software. This application relates to a computer-readable program that, when executed by a logic component, enables the logic component to implement the apparatus or components described above, or to implement the various methods or steps described above. This application also relates to storage media for storing the above programs, such as hard disks, magnetic disks, optical disks, DVDs, flash memory, etc.

[0155] The methods / apparatus described in conjunction with the embodiments of this application can be directly embodied in hardware, software modules executed by a processor, or a combination of both. For example, one or more and / or combinations of one or more functional block diagrams shown in the figures can correspond to various software modules in a computer program flow, or to various hardware modules. These software modules can correspond to the various steps shown in the figures, respectively. These hardware modules can be implemented, for example, using a field-programmable gate array (FPGA) to embed these software modules.

[0156] The software module can reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art. A storage medium can be coupled to the processor, enabling the processor to read information from and write information to the storage medium; or the storage medium can be an integral part of the processor. The processor and storage medium can reside in an ASIC. The software module can be stored in the memory of a mobile terminal or in a memory card that can be inserted into the mobile terminal. For example, if the device (such as a mobile terminal) uses a high-capacity MEGA-SIM card or a high-capacity flash memory device, the software module can be stored in the MEGA-SIM card or the high-capacity flash memory device.

[0157] One or more and / or one or more combinations of functional blocks described in the accompanying drawings can be implemented as a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, or any suitable combination thereof for performing the functions described herein. One or more and / or one or more combinations of functional blocks described in the accompanying drawings can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors in communication with a DSP, or any other such configuration.

[0158] The present application has been described above with reference to specific embodiments. However, those skilled in the art should understand that these descriptions are exemplary and not intended to limit the scope of protection of the present application. Those skilled in the art can make various modifications and variations to the present application based on the principles thereof, and these modifications and variations are also within the scope of the present application.

[0159] Regarding the implementation methods including the above embodiments, the following notes are also disclosed:

[0160] 1. A target state recognition method, wherein the method comprises:

[0161] The input point cloud is grouped to obtain two groups of point clouds, which are then input into a machine learning model to obtain two sets of features corresponding to the two groups of point clouds.

[0162] The parameters of the machine learning model are updated based on the similarity of the two sets of features corresponding to the two sets of point clouds.

[0163] The updated machine learning model is used to identify the target state.

[0164] 2. According to the method described in Appendix 1, wherein,

[0165] The number of points belonging to both sets of point clouds is either 0 or M1+M2-M, where M is the number of points in the input point cloud, and M1 and M2 are the number of points in the two sets of point clouds, respectively.

[0166] 3. According to the method described in Appendix 1, wherein,

[0167] When the number of points in the input point cloud is greater than or equal to the sum of the number of points in the two groups of point clouds to be grouped, a point cloud extraction algorithm is used to extract two groups of point clouds from the input point cloud according to the number of points required for each group of point clouds to be grouped, and the overlap of the two groups of point clouds is 0.

[0168] When the number of points in the input point cloud is less than the sum of the number of points in the two groups of point clouds to be grouped, a point cloud extraction algorithm is used to extract two groups of point clouds from the input point cloud according to the number of points required for each group of point clouds to be grouped, and the overlap of the two groups of point clouds is greater than 0.

[0169] 4. According to the method described in Appendix 3, wherein,

[0170] First, a point cloud extraction algorithm is used to extract M1 point clouds from the input point cloud as the first group of point clouds. Then, a point cloud extraction algorithm is used to extract M2 point clouds from the remaining M-M1 point clouds as the second group of point clouds.

[0171] The input point cloud has a point cloud number of M, the first group of point clouds requires a point cloud number of M1, the second group of point clouds requires a point cloud number of M2, and M1 < M, M2 < M, M1 + M2 ≤ M.

[0172] 5. According to the method described in Appendix 3, wherein,

[0173] First, divide the input point cloud into two equal groups; use a point cloud extraction algorithm to extract M²-2 / M point clouds from one group and merge them with the point clouds from the other group to form the point cloud in the second group; use a point cloud extraction algorithm to extract M1 point clouds from the remaining point clouds in one group and / or the point clouds from the other group to form the point cloud in the first group; or,

[0174] First, a point cloud extraction algorithm is used to extract M2 point clouds from the input point cloud as the point clouds in the second group. Then, a point cloud extraction algorithm is used to extract M1-(M-M2) point clouds from the second group of point clouds and merge them with the remaining M-M2 point clouds in the input point cloud as the point clouds in the first group.

[0175] The input point cloud has a point cloud number of M, the first group of point clouds requires a point cloud number of M1, the second group of point clouds requires a point cloud number of M2, and M1 < M, M2 < M, M1 + M2 > M, M1 < M2.

[0176] 6. The method according to Appendix 1, wherein the method further comprises:

[0177] The two sets of point clouds are transformed to obtain two transformed sets of point clouds, and the two transformed sets of point clouds are input into the machine learning model to obtain two sets of features or inputs.

[0178] The parameters of the machine learning model are updated based on the similarity of the last layer features corresponding to the two transformed point clouds.

[0179] 7. The method according to Appendix 6, wherein transforming the two sets of point clouds includes:

[0180] Perform a mirror transformation on the point cloud velocity information of the two sets of point clouds; or,

[0181] The point cloud position information of the two sets of point clouds is mirrored.

[0182] 8. The method according to Appendix 1, wherein updating the parameters of the machine learning model includes:

[0183] For each group of point cloud data, the loss function is calculated separately for both labeled and unlabeled data.

[0184] The calculated loss functions are weighted and summed to obtain the total loss function of the machine learning model.

[0185] The parameters of the machine learning model are updated using the total loss function.

[0186] 9. The method described in Appendix 8, wherein the loss function is calculated for each group of point cloud labeled and unlabeled data, including:

[0187] For labeled data, a first loss function and a second loss function are calculated. The first loss function is the loss function between the predicted value and the true value of the labeled data corresponding to a set of point clouds after the machine learning model. The second loss function is the loss function between the features of the labeled data corresponding to two sets of point clouds after the machine learning model.

[0188] For unlabeled data, a third loss function is calculated, which is the loss function between the features of the unlabeled data corresponding to the two sets of point clouds after the machine learning model.

[0189] 10. The method according to Appendix 9, wherein the weighted summation of the calculated loss functions to obtain the total loss function of the machine learning model includes:

[0190] The first loss function and the second loss function are weighted and summed to obtain the loss function for labeled data. label The third loss function is then weighted and summed to obtain the loss function for unlabeled data. unlabel And based on the loss function of the labeled data label and the loss function of the unlabeled data. unlabel The total loss function of the machine learning model is obtained.

[0191] in,

[0192] The loss function for labeled data label Represented as: loss label =α1*loss1+β1*loss2+γ1*loss3+ρ1*loss4;

[0193] The loss function of the unlabeled data unlabel Represented as: loss unlabel =β2*loss5+γ2*loss6+ρ2*loss7;

[0194] The total loss function of the machine learning model is expressed as: loss = loss label +loss unlabel ;

[0195] Where α1, β1, γ1, ρ1, β2, γ2, and ρ2 are weighting coefficients, α1 > 0, β1 > 0, β2 > 0, γ1 ≥ 0, ρ1 ≥ 0, γ2 ≥ 0, and ρ2 ≥ 0;

[0196] The loss1 is a loss function between the predicted and actual values ​​of a set of labeled data corresponding to point clouds.

[0197] The loss2 and loss5 are the loss functions between the last layer features corresponding to the two sets of point clouds;

[0198] The loss3 and loss6 are the loss functions between the intermediate layer features corresponding to the two sets of point clouds;

[0199] The loss4 and loss7 are loss functions between the mixed features corresponding to the two sets of point clouds, and the mixed features include the intermediate layer features and the last layer features.

[0200] 11. An electronic device comprising a memory and a processor, the memory storing a computer program, characterized in that the processor is configured to execute the computer program to implement the method as described in any one of Appendices 1 to 10.

[0201] 12. A storage medium storing a computer-readable program, wherein the computer-readable program causes a computer in an electronic device to perform the method described in any one of Appendices 1 to 10.

Claims

1. A target state recognition device, wherein, The device includes: The first processing unit groups the input point cloud into two groups of point clouds, which are then input into a machine learning model to obtain two sets of features corresponding to the two groups of point clouds. The update unit updates the parameters of the machine learning model based on the similarity of the two sets of features corresponding to the two sets of point clouds. The second processing unit uses the updated machine learning model to identify the target state.

2. The apparatus according to claim 1, wherein, The number of points belonging to both sets of point clouds is either 0 or M1+M2-M, where M is the number of points in the input point cloud, and M1 and M2 are the number of points in the two sets of point clouds, respectively.

3. The apparatus according to claim 1, wherein, When the number of points in the input point cloud is greater than or equal to the sum of the number of points in the two groups of point clouds to be grouped, the first processing unit extracts two groups of point clouds from the input point cloud using a point cloud extraction algorithm based on the number of points required for each group of point clouds to be grouped, and the overlap of the two groups of point clouds is 0. When the number of points in the input point cloud is less than the sum of the number of points in the two groups of point clouds to be grouped, the first processing unit extracts two groups of point clouds from the input point cloud using a point cloud extraction algorithm based on the number of points required for each group of point clouds to be grouped, and the overlap of the two groups of point clouds is greater than 0.

4. The apparatus according to claim 3, wherein, The first processing unit first uses a point cloud extraction algorithm to extract M1 point clouds from the input point cloud as the point clouds in the first group, and then uses a point cloud extraction algorithm to extract M2 point clouds from the remaining M-M1 point clouds as the point clouds in the second group. The input point cloud has a point cloud number of M, the first group of point clouds requires a point cloud number of M1, the second group of point clouds requires a point cloud number of M2, and M1 < M, M2 < M, M1 + M2 ≤ M.

5. The apparatus according to claim 3, wherein, The first processing unit first divides the input point cloud into two equal groups; it then uses a point cloud extraction algorithm to extract M²-2 / M point clouds from one group and merges them with the point clouds from the other group to form the point cloud in the second group; finally, it uses the same point cloud extraction algorithm to extract M¹ point clouds from the remaining point clouds in one group and / or the point clouds from the other group to form the point cloud in the first group; or... The first processing unit first uses a point cloud extraction algorithm to extract M2 point clouds from the input point cloud as the point clouds in the second group, and then uses a point cloud extraction algorithm to extract M1-(M-M2) point clouds from the second group of point clouds and merges them with the remaining M-M2 point clouds in the input point cloud as the point clouds in the first group. The input point cloud has a point cloud number of M, the first group of point clouds requires a point cloud number of M1, the second group of point clouds requires a point cloud number of M2, and M1 < M, M2 < M, M1 + M2 > M, M1 < M2.

6. The apparatus according to claim 1, wherein, The first processing unit further transforms the two sets of point clouds to obtain two transformed sets of point clouds, and inputs the two transformed sets of point clouds into the machine learning model to obtain two sets of features or inputs; The updating unit updates the parameters of the machine learning model based on the similarity of the last layer features corresponding to the two transformed point clouds.

7. The apparatus according to claim 6, wherein, The first processing unit transforms the two sets of point clouds, including: Perform a mirror transformation on the point cloud velocity information of the two sets of point clouds; or, The point cloud position information of the two sets of point clouds is mirrored.

8. The apparatus according to claim 1, wherein, The updating unit updates the parameters of the machine learning model, including: For each group of point cloud data, the loss function is calculated separately for both labeled and unlabeled data. The calculated loss functions are weighted and summed to obtain the total loss function of the machine learning model. The parameters of the machine learning model are updated using the total loss function.

9. The apparatus according to claim 8, wherein, The update unit calculates the loss function for each group of point cloud labeled and unlabeled data, including: For labeled data, the update unit calculates a first loss function and a second loss function. The first loss function is the loss function between the predicted value and the true value of the labeled data corresponding to a set of point clouds after the machine learning model. The second loss function is the loss function between the features of the labeled data corresponding to two sets of point clouds after the machine learning model. For unlabeled data, the update unit calculates a third loss function, which is the loss function between the features of the unlabeled data corresponding to the two sets of point clouds after passing through the machine learning model.

10. The apparatus according to claim 9, wherein, The update unit performs a weighted summation of the calculated loss functions to obtain the total loss function of the machine learning model, including: The first loss function and the second loss function are weighted and summed to obtain the loss function for labeled data. label The third loss function is then weighted and summed to obtain the loss function for unlabeled data. unlabel And based on the loss function of the labeled data label and the loss function of the unlabeled data unlabel The total loss function of the machine learning model is obtained. in, The loss function for labeled data label Represented as: loss label =α1*loss1+β1*loss2+γ1*loss3+ρ1*loss4; The loss function of the unlabeled data unlabel Represented as: loss unlabel =β2*loss5+γ2*loss6+ρ2*loss7; The total loss function of the machine learning model is expressed as: loss = loss label +loss unlabel ; Where α1, β1, γ1, ρ1, β2, γ2, and ρ2 are weighting coefficients, α1 > 0, β1 > 0, β2 > 0, γ1 ≥ 0, ρ1 ≥ 0, γ2 ≥ 0, and ρ2 ≥ 0; The loss1 is a loss function between the predicted and actual values ​​of a set of labeled data corresponding to point clouds. The loss2 and loss5 are the loss functions between the last layer features corresponding to the two sets of point clouds; The loss3 and loss6 are the loss functions between the intermediate layer features corresponding to the two sets of point clouds; The loss4 and loss7 are loss functions between the mixed features corresponding to the two sets of point clouds, and the mixed features include the intermediate layer features and the last layer features.