An AI and multi-sensor based force action intelligent recognition method and device

CN121009334BActive Publication Date: 2026-09-18SHENZHEN BREO TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511034483.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-25
Publication Date
2026-09-18
Estimated Expiration
2045-07-25

AI Technical Summary

Technical Problem

[0004]然而,由于多个施力动作类别对应的压力特征存在相似的情况,因此,采用上述方式,可能会导致识别结果错误,即施力动作识别准确性较低

Benefits of technology

[0107] In the AI- and multi-sensor-based intelligent force application method provided in this application embodiment, the target massage operation is comprehensively quantified based on the multi-dimensional data corresponding to each equally spaced array node, resulting in high-precision quantification results. Furthermore, feature processing of the multi-dimensional data of each array node maximizes the discriminative feature space that distinguishes different force applications. Embedded vectors fuse temporal and spatial information to capture the temporal and spatial information of the target massage operation. Therefore, it can accurately quantify and identify force applications, improving the accuracy of force application recognition and providing a replicable and scalable technical path for traditional physiotherapy, rehabilitation training, and intelligent teaching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121009334B_ABST
    Figure CN121009334B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a kind of based on AI and multi-sensor's force action intelligent identification method and device, it is related to computer technical field.The method is: obtaining in target massage operation, the multidimensional data corresponding to each array node respectively included in contact surface array, wherein, each array node is equidistant arrangement, each array node except boundary node at least has the neighbor node of preset number in each array node;The multidimensional data of each array node is extracted respectively, obtains the initial feature vector of each array node respectively, and according to the space adjacency relation of each array node, the initial feature vector obtained is updated, obtains the embedding vector of target massage operation;When sequence information and space information are fused to embedding vector, obtain target modeling vector, and determine the force action recognition result of target massage operation based on target modeling vector.The force action can be accurately quantified and identified by multidimensional data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to an intelligent recognition method and device for force application actions based on AI and multiple sensors. Background Technology

[0002] Massage is an effective way to adjust one's physical and mental state, and it involves many different types of pressure movements.

[0003] In existing technologies, when identifying force application actions, the operator's current force application action category is usually determined based on the operator's current pressure-time curve data.

[0004] However, since the pressure characteristics corresponding to multiple force application action categories are similar, the above method may lead to incorrect identification results, that is, the accuracy of force application action identification is low. Summary of the Invention

[0005] This application provides an intelligent recognition method and device for force application actions based on AI and multiple sensors, in order to improve the accuracy of force application action recognition.

[0006] In a first aspect, embodiments of this application provide an intelligent recognition method for force application actions based on AI and multiple sensors, the method comprising:

[0007] In the target massage operation, acquire the multidimensional data corresponding to each array node in the contact surface array, wherein each array node is arranged at equal intervals, and each array node except for the boundary node has at least a preset number of neighbor nodes.

[0008] Feature extraction is performed on the multidimensional data of each array node to obtain the initial feature vector of each array node. Based on the spatial adjacency relationship of each array node, the obtained initial feature vectors are updated to obtain the embedding vector of the target massage operation.

[0009] By fusing temporal and spatial information into the embedded vector, a target modeling vector is obtained. Based on the target modeling vector, the force application action recognition result of the target massage operation is determined.

[0010] In one alternative embodiment, the multidimensional data includes force dimension data, attitude dimension data, and temperature dimension data.

[0011] In one optional embodiment, acquiring multidimensional data corresponding to each array node in the contact surface array during the target massage operation includes:

[0012] Initial data for each array node during the target massage operation is acquired by deploying a contact surface array on the surface of the first object.

[0013] Second initial data is collected by deploying a hand multimodal sensor on the second object's hand;

[0014] Based on the second initial data and the first initial data of each array node, the multidimensional data corresponding to each array node is determined.

[0015] In one optional embodiment, each array node includes a pressure sensor, a force sensor, and a temperature sensor, and the first initial data includes normal pressure data, tangential force data, and contact surface temperature data.

[0016] The hand multimodal sensing includes an inertial measurement unit, a depth camera, and a temperature electrode. The second initial data includes inertial data, hand point cloud data, and back of hand temperature data.

[0017] In one optional embodiment, based on the second initial data and the first initial data of each array node, the multidimensional data corresponding to each array node is determined, including:

[0018] Based on inertial data and hand point cloud data, the attitude time series is determined and used as the attitude dimension data corresponding to each array node.

[0019] Based on the attitude time series, the force time series and torque time series of each hand joint are determined, and the normal pressure data, tangential force data and the force time series and torque time series of each hand joint of each array node are used as the force dimension data corresponding to each array node.

[0020] The contact surface temperature data and back of hand temperature data of each array node are used as the temperature dimension data of each array node.

[0021] In one optional embodiment, determining the attitude time series based on inertial data and hand point cloud data includes:

[0022] An adaptive extended Kalman filter is applied to the inertial data and hand point cloud data to obtain the attitude time series.

[0023] In one optional embodiment, determining the force time series and torque time series of each hand joint based on the posture time series includes:

[0024] Based on the posture time series and the DH parameter set of each hand joint, the end pose of each finger is determined;

[0025] Based on the end-effector pose, inverse dynamics calculations are performed to obtain the force time series and torque time series of each hand joint.

[0026] In one optional embodiment, feature extraction is performed on the multidimensional data of each array node to obtain an initial feature vector for each array node, including:

[0027] Feature extraction of different domains is performed on the multidimensional data of each array node to obtain multiple sub-initial feature vectors for each array node.

[0028] The multiple sub-initial feature vectors of each array node are standardized and concatenated to obtain the initial feature vector of each array node.

[0029] In one optional embodiment, feature extraction of different domains is performed on the multidimensional data of each array node to obtain multiple sub-initial feature vectors for each array node, including:

[0030] Temporal features are extracted from the force dimension data of each array node to obtain the first sub-initial feature vector of each array node.

[0031] Frequency domain features are extracted from the attitude dimension data of each array node to obtain the second sub-initial feature vector of each array node.

[0032] Topological domain features are extracted from the temperature dimension data of each array node to obtain the third sub-initial feature vector of each array node.

[0033] In one optional embodiment, the obtained initial feature vectors are updated according to the spatial adjacency relationship of each array node to obtain the embedding vector of the target massage operation, including:

[0034] A graph is constructed based on the spatial adjacency relationship of each array node, where the vertices of the graph are each array node, and the edge set of the graph is the spatial adjacency relationship between each array node.

[0035] Based on the graph, graph convolution and attention aggregation are performed on each initial feature vector to obtain the embedding vector.

[0036] In one optional embodiment, the embedded vector is fused with temporal and spatial information to obtain a modeling vector for the target massage operation, including:

[0037] Temporal dynamic feature extraction is performed on the embedded vector to obtain temporal feature vector, and spatial distribution feature extraction is performed on the embedded vector to obtain spatial feature vector. The temporal feature vector represents the temporal dynamic semantics of the target massage operation, and the spatial feature vector represents the spatial distribution semantics of the target massage operation.

[0038] The temporal feature vector and the spatial feature vector are fused to obtain the target modeling vector.

[0039] In one optional embodiment, the force application action recognition result of the target massage operation is determined based on the target modeling vector, including:

[0040] Determine the distance between the target modeling vector and the center vector of each force application action category, where the center vector is determined based on the sample modeling vectors of the corresponding force application action category;

[0041] Based on the obtained distances, the force application action recognition results of the target massage operation are determined.

[0042] In one optional embodiment, based on the obtained distances, the force application action recognition result of the target massage operation is determined, including:

[0043] When there is a target distance greater than the distance threshold among all distances, the force application action category of the target massage operation is determined to be the force application action category corresponding to the target distance;

[0044] When there is no target distance greater than the distance threshold among the various distances, the force application action recognition result of the target massage operation is determined based on the distance between the target modeling vector and the sample modeling vectors of each force application action category.

[0045] In one optional embodiment, the force application action recognition result is obtained by inputting the target modeling vector into the target force application action recognition model, wherein the target force application action recognition model is trained in the following manner:

[0046] Obtain a training sample set, wherein each training sample in the training sample set includes at least: the sample embedding vector of the sample massage operation;

[0047] Based on the training sample set, the force application action recognition model to be trained is iteratively trained to obtain the target force application action recognition model. During each iteration of training, the following operations are performed:

[0048] The temporal and spatial information of the sample embedding vectors of the selected target training samples are fused to obtain the sample modeling vectors of the target training samples.

[0049] Determine the sample distance between the sample modeling vector of the target training sample and the center vector of each force application action category;

[0050] Based on the obtained distances between samples, the sample force action recognition results of the target training sample are determined, and the parameters are tuned based on the loss value corresponding to the sample force action recognition results.

[0051] In an optional embodiment, after determining the force application action recognition result of the target training sample based on the obtained sample distances, and after parameter tuning based on the loss value corresponding to the sample force application action recognition result, the method further includes:

[0052] When the target training sample is a high-confidence sample, the center vector of the force application action category to which the target training sample belongs is updated based on the sample modeling vector of the target training sample.

[0053] Secondly, embodiments of this application also provide an intelligent recognition device for force application actions based on AI and multiple sensors, the device comprising:

[0054] The acquisition module is used to acquire the multi-dimensional data corresponding to each array node in the contact surface array during the target massage operation. The array nodes are arranged at equal intervals, and each array node, except for the boundary nodes, has at least a preset number of neighbor nodes.

[0055] The data fusion module is used to extract features from the multidimensional data of each array node, obtain the initial feature vector of each array node, and update the obtained initial feature vectors according to the spatial adjacency relationship of each array node to obtain the embedding vector of the target massage operation.

[0056] The force application action recognition module is used to fuse temporal and spatial information into embedded vectors to obtain target modeling vectors, and based on the target modeling vectors, to determine the force application action recognition result of the target massage operation.

[0057] In one alternative embodiment, the multidimensional data includes force dimension data, attitude dimension data, and temperature dimension data.

[0058] In an optional embodiment, when acquiring the multidimensional data corresponding to each array node in the contact surface array during the target massage operation, the acquisition module is further configured to:

[0059] Initial data for each array node during the target massage operation is acquired by deploying a contact surface array on the surface of the first object.

[0060] Second initial data is collected by deploying a hand multimodal sensor on the second object's hand;

[0061] Based on the second initial data and the first initial data of each array node, the multidimensional data corresponding to each array node is determined.

[0062] In one optional embodiment, each array node includes a pressure sensor, a force sensor, and a temperature sensor, and the first initial data includes normal pressure data, tangential force data, and contact surface temperature data.

[0063] The hand multimodal sensing includes an inertial measurement unit, a depth camera, and a temperature electrode. The second initial data includes inertial data, hand point cloud data, and back of hand temperature data.

[0064] In an optional embodiment, when determining the multidimensional data corresponding to each array node based on the second initial data and the first initial data of each array node, the acquisition module is further configured to:

[0065] Based on inertial data and hand point cloud data, the attitude time series is determined and used as the attitude dimension data corresponding to each array node.

[0066] Based on the attitude time series, the force time series and torque time series of each hand joint are determined, and the normal pressure data, tangential force data and the force time series and torque time series of each hand joint of each array node are used as the force dimension data corresponding to each array node.

[0067] The contact surface temperature data and back of hand temperature data of each array node are used as the temperature dimension data of each array node.

[0068] In an optional embodiment, when determining the attitude time series based on inertial data and hand point cloud data, the acquisition module is further configured to:

[0069] An adaptive extended Kalman filter is applied to the inertial data and hand point cloud data to obtain the attitude time series.

[0070] In an optional embodiment, when determining the force time series and torque time series of each hand joint based on the posture time series, the acquisition module is further configured to:

[0071] Based on the posture time series and the DH parameter set of each hand joint, the end pose of each finger is determined;

[0072] Based on the end-effector pose, inverse dynamics calculations are performed to obtain the force time series and torque time series of each hand joint.

[0073] In an optional embodiment, when performing feature extraction on the multidimensional data of each array node to obtain the initial feature vector of each array node, the data fusion module is further used to:

[0074] Feature extraction of different domains is performed on the multidimensional data of each array node to obtain multiple sub-initial feature vectors for each array node.

[0075] The multiple sub-initial feature vectors of each array node are standardized and concatenated to obtain the initial feature vector of each array node.

[0076] In an optional embodiment, when performing feature extraction from different domains on the multidimensional data of each array node to obtain multiple sub-initial feature vectors for each array node, the data fusion module is further used to:

[0077] Temporal features are extracted from the force dimension data of each array node to obtain the first sub-initial feature vector of each array node.

[0078] Frequency domain features are extracted from the attitude dimension data of each array node to obtain the second sub-initial feature vector of each array node.

[0079] Topological domain features are extracted from the temperature dimension data of each array node to obtain the third sub-initial feature vector of each array node.

[0080] In an optional embodiment, when updating the obtained initial feature vectors based on the spatial adjacency relationships of each array node to obtain the embedding vector of the target massage operation, the data fusion module is further used to:

[0081] A graph is constructed based on the spatial adjacency relationship of each array node, where the vertices of the graph are each array node, and the edge set of the graph is the spatial adjacency relationship between each array node.

[0082] Based on the graph, graph convolution and attention aggregation are performed on each initial feature vector to obtain the embedding vector.

[0083] In an optional embodiment, when obtaining the modeling vector of the target massage operation by fusing temporal and spatial information from the embedded vector, the force application action recognition module is further used to:

[0084] Temporal dynamic feature extraction is performed on the embedded vector to obtain temporal feature vector, and spatial distribution feature extraction is performed on the embedded vector to obtain spatial feature vector. The temporal feature vector represents the temporal dynamic semantics of the target massage operation, and the spatial feature vector represents the spatial distribution semantics of the target massage operation.

[0085] The temporal feature vector and the spatial feature vector are fused to obtain the target modeling vector.

[0086] In an optional embodiment, when determining the force application action recognition result of the target massage operation based on the target modeling vector, the force application action recognition module is further configured to:

[0087] Determine the distance between the target modeling vector and the center vector of each force application action category, where the center vector is determined based on the sample modeling vectors of the corresponding force application action category;

[0088] Based on the obtained distances, the force application action recognition results of the target massage operation are determined.

[0089] In an optional embodiment, when determining the force application action recognition result of the target massage operation based on the obtained distances, the force application action recognition module is further configured to:

[0090] When there is a target distance greater than the distance threshold among all distances, the force application action category of the target massage operation is determined to be the force application action category corresponding to the target distance;

[0091] When there is no target distance greater than the distance threshold among the various distances, the force application action recognition result of the target massage operation is determined based on the distance between the target modeling vector and the sample modeling vectors of each force application action category.

[0092] In an optional embodiment, the device further includes a training module, which is used for:

[0093] Obtain a training sample set, wherein each training sample in the training sample set includes at least: the sample embedding vector of the sample massage operation;

[0094] Based on the training sample set, the force application action recognition model to be trained is iteratively trained to obtain the target force application action recognition model. During each iteration of training, the following operations are performed:

[0095] The temporal and spatial information of the sample embedding vectors of the selected target training samples are fused to obtain the sample modeling vectors of the target training samples.

[0096] Determine the sample distance between the sample modeling vector of the target training sample and the center vector of each force application action category;

[0097] Based on the obtained distances between samples, the sample force action recognition results of the target training sample are determined, and the parameters are tuned based on the loss value corresponding to the sample force action recognition results.

[0098] In one optional embodiment, based on the obtained distances between samples, the force application action recognition result of the target training sample is determined, and after parameter tuning based on the loss value corresponding to the force application action recognition result of the sample, the training module is further used for:

[0099] When the target training sample is a high-confidence sample, the center vector of the force application action category to which the target training sample belongs is updated based on the sample modeling vector of the target training sample.

[0100] Thirdly, embodiments of this application also provide an electronic device, including:

[0101] Processor; and

[0102] Stored program memory,

[0103] The program includes instructions that, when executed by the processor, cause the processor to perform the intelligent recognition method for force application based on AI and multiple sensors as described in the first aspect.

[0104] Fourthly, embodiments of this application also provide a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to execute the intelligent recognition method for force application actions based on AI and multiple sensors as described in the first aspect.

[0105] Fifthly, this application provides a computer program product that, when invoked by a computer, causes the computer to execute the intelligent recognition method steps for force application actions based on AI and multiple sensors as described in the first aspect.

[0106] The beneficial effects of this application are as follows:

[0107] In the AI- and multi-sensor-based intelligent force application method provided in this application embodiment, the target massage operation is comprehensively quantified based on the multi-dimensional data corresponding to each equally spaced array node, resulting in high-precision quantification results. Furthermore, feature processing of the multi-dimensional data of each array node maximizes the discriminative feature space that distinguishes different force applications. Embedded vectors fuse temporal and spatial information to capture the temporal and spatial information of the target massage operation. Therefore, it can accurately quantify and identify force applications, improving the accuracy of force application recognition and providing a replicable and scalable technical path for traditional physiotherapy, rehabilitation training, and intelligent teaching.

[0108] Furthermore, other features and advantages of this application will be set forth in the following description and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0109] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described herein are used to provide a further understanding of this application, constitute a part of this application, and do not constitute an improper limitation of this application. In the accompanying drawings:

[0110] Figure 1 This is a schematic diagram of an optional system architecture applicable to the embodiments of this application;

[0111] Figure 2 A schematic diagram illustrating the implementation process of an intelligent recognition method for force application actions based on AI and multiple sensors, provided in an embodiment of this application;

[0112] Figure 3 This is a first schematic diagram of a contact surface array provided in an embodiment of this application;

[0113] Figure 4 This is a second schematic diagram of a contact surface array provided in an embodiment of this application;

[0114] Figure 5 A logical schematic diagram for obtaining an embedding vector provided in an embodiment of this application;

[0115] Figure 6 A logical diagram illustrating the determination of force application action recognition results provided in an embodiment of this application;

[0116] Figure 7 A schematic diagram illustrating the training process of a target force application action recognition model provided in an embodiment of this application;

[0117] Figure 8 A schematic diagram of the structure of an intelligent recognition device for force application based on AI and multiple sensors, provided in an embodiment of this application;

[0118] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0119] Embodiments of this application will now be described in more detail with reference to the accompanying drawings. While some embodiments of this application are shown in the drawings, it should be understood that this application can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this application. It should be understood that the drawings and embodiments of this application are for illustrative purposes only and are not intended to limit the scope of protection of this application.

[0120] It should be understood that the steps described in the method embodiments of this application may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this application is not limited in this respect.

[0121] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc., mentioned in this application are used only to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.

[0122] It should be noted that the terms "a" and "a plurality of" used in this application are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".

[0123] The names of the messages or information exchanged between multiple devices in the embodiments of this application are for illustrative purposes only and are not intended to limit the scope of these messages or information.

[0124] The following explanations of some terms used in the embodiments of this application are provided to facilitate understanding by those skilled in the art.

[0125] (1) The first object: is the object being massaged.

[0126] (2) The second object: is the object to which the massage is performed.

[0127] (3) Nine-Degree of Freedom Inertial Measurement Unit (9-DoF IMU): It includes a three-axis accelerometer, a three-axis gyroscope and a three-axis magnetometer, which are used to detect linear acceleration, angular velocity and magnetic field strength, respectively.

[0128] (4) Time of Flight (TOF) Depth Camera: This is an advanced imaging device that uses light pulse sensing technology to measure the distance and depth between an object and the camera, enabling fast and accurate acquisition of depth information.

[0129] (5) Adaptive Extended Kalman Filter (AdaptiveEKF): This is an improved Kalman filter algorithm that can automatically adjust the filter parameters to adapt to changes in the variance of system noise and observation noise when these variances are unknown or change over time.

[0130] (6) DH parameters. A standardized method is provided to describe the geometric relationship between joints and links. The DH parameter model defines the relative position and orientation between two adjacent links through four key parameters: joint angle (θ), link length (a), link offset (d), and link deflection angle (α).

[0131] (7) Graph Convolutional Network (GCN): By defining convolution operations on the graph to extract features and learn representations of nodes, it can process data with generalized topological graph structures and deeply explore data features and patterns.

[0132] (8) Attention Mechanism: It is an algorithm that simulates human attention allocation to improve the model’s ability to handle complex tasks. It can dynamically focus attention on key parts of the input data, enabling the network to efficiently handle tasks with large input length or structural variations.

[0133] (9) Objective function: In machine learning, it is a function used to measure and optimize the distance between the model's predicted value and the actual sample labeled value.

[0134] (10) Embodied Intelligence: This refers to the integration of artificial intelligence into physical entities such as robots, endowing them with the ability to perceive, learn, and dynamically interact with their environment like humans. Embodied intelligence emphasizes that the behavior of intelligent agents does not solely rely on abstract computation and reasoning abilities, but is deeply rooted in their interactions with their environment. Embodied intelligence is often associated with the intelligent behaviors of robots, animals, and humans, reflecting the ability to make decisions, learn, and reason through perception, action, and adaptation to the environment.

[0135] Based on the above explanations of terms and related terminology, the design concept of the embodiments of this application will be briefly introduced below:

[0136] Massage is an effective way to adjust one's physical and mental state, and it involves many different types of pressure movements.

[0137] In existing technologies, when identifying force application actions, the operator's current force application action category is usually determined based on the operator's current pressure-time curve data.

[0138] However, since the pressure characteristics corresponding to multiple force application action categories are similar, the above method may lead to incorrect identification results, that is, the accuracy of force application action identification is low.

[0139] In one optional implementation, this application proposes an intelligent recognition method for force application actions based on AI and multiple sensors. Specifically, it may include: acquiring multi-dimensional data corresponding to each array node in the contact surface array during a target massage operation, wherein each array node is arranged at equal intervals, and each array node, except for boundary nodes, has at least a preset number of neighbor nodes; then extracting features from the multi-dimensional data of each array node to obtain an initial feature vector for each array node; updating the obtained initial feature vectors according to the spatial adjacency relationship of each array node to obtain an embedding vector for the target massage operation; finally, fusing temporal and spatial information into the embedding vector to obtain a target modeling vector; and determining the force application action recognition result of the target massage operation based on the target modeling vector.

[0140] By employing the above method, the target massage operation is comprehensively quantified based on the multidimensional data corresponding to each equally spaced array node, resulting in high-precision quantification. Furthermore, feature processing of the multidimensional data of each array node maximizes the discriminative feature space that distinguishes different force application actions. Embedded vectors are fused with spatiotemporal information to capture the temporal and spatial information of the target massage operation. Therefore, force application actions can be accurately quantified and identified, improving the accuracy of force application action recognition and providing a replicable and scalable technical path for traditional physiotherapy, rehabilitation training, and intelligent teaching. In addition, it provides data and model support for future embodied intelligence and AI-automated massage.

[0141] In particular, the preferred embodiments of this application will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit this application. Furthermore, the embodiments of this application and the features in the embodiments can be combined with each other without conflict.

[0142] See Figure 1 The diagram illustrates an optional system architecture applicable to an embodiment of this application. This system architecture may include an edge device 101 and a cloud device 102. The edge device 101 and the cloud device 102 can interact via a communication network, where the communication network employs wireless and wired communication methods. For example, the edge device 101 can access the network via Wi-Fi 6E or cellular mobile communication technology (such as 5G) to perform real-time data transmission with the cloud device 102; it also supports BLE 5.3 short-range wireless communication for low-power interconnection of local sensor networks.

[0143] This application embodiment does not impose any limitation on the number of communication devices involved in the above system architecture. For example, the above system architecture may include more edge devices, or it may include fewer edge devices, or it may also include other network devices. Figure 1 As shown, only the edge device 101 and the cloud device 102 are described as examples. The following is a brief introduction to the above communication devices and their respective functions.

[0144] Edge terminal 101 is a data acquisition and preprocessing terminal, used to acquire multi-dimensional data corresponding to each array node in the contact surface array during the target massage operation. It extracts features from the multi-dimensional data of each array node to obtain initial feature vectors, and updates these initial feature vectors based on the spatial adjacency relationships of the array nodes to obtain the embedding vector of the target massage operation. It then fuses temporal and spatial information into the embedding vector to obtain the target modeling vector, and determines the force application action recognition result of the target massage operation based on the target modeling vector. Additionally, it is used to acquire a training sample set and send it to the cloud 102. Each training sample in the training sample set includes at least one sample embedding vector of the sample massage operation.

[0145] Optionally, a pre-trained force application action recognition model can be deployed on the edge device 101. In this way, after the edge device 101 obtains the embedding vector of the target massage operation, it can input the embedding vector into the pre-trained force application action recognition model to determine the force application action recognition result of the target massage operation, thereby realizing offline force application action recognition in the offline scenario.

[0146] Cloud 102 is the model training and data management terminal, used to receive training sample sets and iteratively train the force action recognition model to be trained based on the training sample sets to obtain the target force action recognition model. During each iteration, the following operations are performed: the spatiotemporal information of the sample embedding vectors of the selected target training samples is fused to obtain the sample modeling vector of the target training samples; the sample distance between the sample modeling vector of the target training samples and the center vector of each force action category is determined; based on the obtained sample distances, the sample force action recognition result of the target training samples is determined, and parameter tuning is performed based on the loss value corresponding to the sample force action recognition result. Alternatively, the latest trained force action recognition model can be sent to edge terminal 101, or a lightweight version of the latest force action recognition model can be sent to edge terminal 101; this embodiment does not impose any restrictions on this.

[0147] Additionally, it is worth noting that the trained target force application action recognition model can also be deployed on the cloud 102. In this way, after the edge terminal 101 obtains the embedding vector of the target massage operation, it sends the embedding vector of the target massage operation to the cloud 102. The cloud 102 inputs the embedding vector into the pre-trained target force application action recognition model to determine the force application action recognition result of the target massage operation.

[0148] In the embodiments of this application, the target force action recognition model involves artificial intelligence (AI) technology and is designed based on speech technology, natural language processing technology and machine learning (ML) in artificial intelligence.

[0149] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence.

[0150] Artificial intelligence (AI) studies the design principles and implementation methods of various intelligent machines, enabling them to perceive, reason, and make decisions. AI technology mainly includes computer vision, natural language processing, and machine learning / deep learning. With the research and advancement of AI technology, it is being researched and applied in multiple fields, such as smart homes, intelligent customer service, virtual assistants, smart speakers, intelligent marketing, autonomous driving, robotics, and smart healthcare. It is believed that with further technological development, AI will be applied in even more fields and play an increasingly important role.

[0151] Machine learning is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory, among others. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Compared to data mining, which focuses on finding patterns in large datasets, machine learning emphasizes algorithm design, enabling computers to automatically "learn" patterns from data and use these patterns to predict unknown data.

[0152] Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence. Its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, and inductive learning. Reinforcement learning (RL), also known as reward learning, evaluation learning, or enhancement learning, is one of the paradigms and methodologies of machine learning. It is used to describe and solve problems where an agent learns strategies to maximize rewards or achieve specific goals during its interaction with the environment.

[0153] The following describes the intelligent recognition method for force application based on AI and multiple sensors provided by exemplary embodiments of this application, in conjunction with the above-described system architecture and with reference to the accompanying drawings. It should be noted that the above-described system architecture is only shown to facilitate understanding of the spirit and principles of this application, and the embodiments of this application are not limited in any way in this respect.

[0154] See Figure 2 The diagram shown illustrates the implementation flow of an intelligent force recognition method based on AI and multiple sensors provided in this application. The specific implementation flow of this method is as follows:

[0155] S20: Obtain the multi-dimensional data corresponding to each array node contained in the contact surface array during the target massage operation.

[0156] The array nodes are arranged at equal intervals, and each array node, except for the boundary nodes, has at least a preset number of neighboring nodes. The multidimensional data includes force dimension data, attitude dimension data, and temperature dimension data.

[0157] In this embodiment of the application, when the second object performs a target massage operation on the first object, the multi-dimensional data corresponding to each array node is obtained.

[0158] Optionally, in this application embodiment, a possible implementation is provided for obtaining the multidimensional data corresponding to each array node, specifically by performing the following operations:

[0159] S200: Acquire initial data from each array node during the target massage operation by deploying a contact surface array on the surface of the first object.

[0160] The contact surface array includes each array node.

[0161] For example, see Figure 3 The diagram shown is a first schematic diagram of the contact surface array in an embodiment of this application. The contact surface array is a hexagonal array, including 27 array nodes. The 27 array nodes are arranged in an equally spaced "staggered row" manner (3 rows × 9 columns), so that each array node (except the boundary) has at least six neighboring nodes (top / bottom / left / right / diagonal), thereby achieving optimal spatial coverage and uniform reconstruction accuracy.

[0162] The spacing can be 4.2mm, but this embodiment does not impose any restrictions on it.

[0163] For example, see Figure 4The diagram shown is a second schematic diagram of the contact surface array in the embodiment of this application. The contact surface array is a hexagonal array, including 27 equally spaced array nodes, so that each array node (except for the boundary) has at least six neighboring nodes. There is one array node in the center, six array nodes in the first adjacent layer, twelve array nodes in the second adjacent layer, and eight array nodes attached to the periphery, thereby achieving optimal spatial coverage and uniform reconstruction accuracy.

[0164] In this embodiment of the application, each array node is a sensing unit, and each array node includes: a pressure sensor, a force sensor and a temperature sensor. The first initial data includes: normal pressure data, tangential force data and contact surface temperature data.

[0165] The pressure sensor can be a flexible strain pressure sensor (based on a micro-dome structure of MXene / PDMS, measuring normal pressure and surface micro-strain in real time), the force sensor can be a 6-dimensional force sensor (measuring force and torque in the X / Y / Z directions), and the temperature sensor is a CNT-PDMS temperature sensor (32×32 pixel array). Normal pressure data includes normal pressure and bending strain; tangential force data includes shear force along the X and Y directions; and contact surface temperature data includes temperature readings from the temperature sensor, though this embodiment does not impose limitations on these data. Thus, each array node simultaneously acquires five physical quantities, enabling a 5-dimensional, 27-channel synchronous acquisition of the contact surface array with 27 array nodes.

[0166] S201: Acquire second initial data by deploying a hand multimodal sensor on the hand of the second object.

[0167] In this embodiment of the application, the hand multimodal sensing includes: an inertial measurement unit, a depth camera, and a temperature electrode; the second initial data includes: inertial data, hand point cloud data, and hand back temperature data.

[0168] The inertial measurement unit can be a wearable 9-DoF IMU, the depth camera can be a ToF depth camera, and the temperature electrode can be a hand temperature electrode (4×4 array, measuring the surface temperature of the back of the hand).

[0169] In this embodiment, a hand temperature electrode is deployed on the back of the second object's hand to detect the skin temperature. This hand temperature electrode differs from the temperature sensor integrated with the pressure and force sensors in each array node of the contact surface array. The temperature sensor monitors the intensity of palm heat application and frictional heat generation. The hand temperature electrode does not interfere with the contact surface array, providing a "static reference" that differentiates from the palm temperature, separating the "external heat source" from the "internal heating." Furthermore, the hand temperature electrode cannot be integrated into each array node because each node simultaneously collects normal pressure and tangential force data. Adding a hand temperature electrode would affect flexibility and pressure sensitivity. Moreover, placing the hand temperature electrode within an array node would significantly disrupt its readings due to heat application, rendering it meaningless as a "reference body temperature." Additionally, fatigue and DC current can be estimated based on the hand temperature data collected by the hand temperature electrode, and a safety protection threshold can be established during intense heat application. The heat therapy power can be adjusted in real-time based on the contact surface temperature data collected by the temperature sensor integrated with the pressure and force sensors in each array node.

[0170] In addition, it is worth noting that in the embodiments of this application, the sensor is calibrated before the sensor collects data to improve the accuracy of the collected data. Specifically, the 6-dimensional force sensor is calibrated with matrix calibration and temperature drift compensation, the CNT-PDMS temperature sensor is calibrated with three-point water bath calibration and online drift compensation, and the 9-DoF IMU is calibrated with internal parameters, external parameters and clock.

[0171] S202: Based on the second initial data and the first initial data of each array node, determine the multidimensional data corresponding to each array node.

[0172] Optionally, in this embodiment of the application, a possible implementation is provided for determining the multidimensional data corresponding to each array node based on the second initial data and the first initial data of each array node, specifically by performing the following operations:

[0173] S2020: Based on inertial data and hand point cloud data, determine the attitude time series, and use the attitude time series as the attitude dimension data corresponding to each array node.

[0174] For example, the attitude dimension data corresponding to an array node includes: attitude time series.

[0175] In this embodiment, adaptive extended Kalman filtering is performed on inertial data and hand point cloud data to obtain attitude time series.

[0176] The attitude time series includes quaternion time series and position time series.

[0177] This avoids the problem of insufficient estimation accuracy in fast-moving or occluded situations using traditional methods, and enables accurate real-time estimation of hand posture data during massage operations.

[0178] Specifically, in this embodiment, a state vector is constructed based on inertial data, and then the state vector is updated based on Kalman gain and hand point cloud data to obtain a quaternion time series and a position time series.

[0179] The state vector includes position, velocity, quaternion, accelerometer bias, and gyroscope bias, and the quaternion time series includes the posterior quaternion at each time point.

[0180] In this embodiment of the application, a state vector update includes: predicting the prior state vector at the current time based on the posterior state vector and the local state increment at the previous time, and then determining the posterior state vector at the current time based on the Kalman gain, the actual observation, the predicted observation, and the prior state vector at the current time.

[0181] Here, the prior state vector at the current moment is the initial estimate of the state vector at the current moment, and the posterior state vector at the current moment is the optimized state vector obtained after correcting the error of the prior state vector at the current moment, which is closer to the true state.

[0182] For example, the above state vector can be specifically represented as follows: x = [p, v, q, b] a b g ] T Where p is position, v is velocity, q is quaternion, and b is... a For accelerometer bias, b g This is for gyroscope bias.

[0183] The formula for calculating the prior state vector at the current moment can be specifically expressed as follows:

[0184]

[0185] Where, x k-1 Let u be the posterior state vector at time k-1. k-1 This is the IMU measurement input at time k-1; u k-1 =[a k-1 ,ω k-1 ] T a k-1 Let ω be the "pure inertial" acceleration after removing gravity at time k-1, representing the linear acceleration of the wrist reference point in the X, Y, and Z directions. k-1Let f(·) be the rotational speed of the wrist reference point around the X, Y, and Z axes at time k-1; f(·) is the state transition model, describing the dynamic change of the state over time. Short-time prediction is achieved based on the IMU pre-integration principle, and the quaternion increment Δq = ∏ is calculated through IMU pre-integration. i exp((ω i -b g )Δt), velocity increment Δv=∫(a i -b a )dt, position increment Δp=∫v i dt, ω i It is the instantaneous angular velocity reading of the gyroscope, a i It is the instantaneous linear acceleration from the accelerometer, w i Not measured by the IMU, it is an intermediate estimate of the system state, updated using pre-integrated results combined with gravity g, and the estimated velocity of the wrist reference point from the previous keyframe; w k-1 The process noise coupling matrix describes the mapping matrix of random disturbances from independent noise sources such as friction, mechanical clearance, and IMU quantization error onto state variables. It acts on each state dimension and follows a zero-mean Gaussian distribution. Q k-1 This is the covariance matrix of the process noise. The larger the value, the lower the model's "confidence" in the noise source, which is used to adjust the uncertainty of EKF prediction.

[0186] The actual predicted values ​​mentioned above can be specifically represented as follows: Where h(·) is the observation model function, the joint mapping function from state to observation, which converts the predicted hand pose and position into observable joint coordinates; v k For observation noise; R k To observe the noise covariance matrix. k =[j1,…,j M ], containing M joints, M=20, but this is not limited in the embodiments of this application.

[0187] The Kalman gain at the current moment can be specifically expressed as follows:

[0188] K k =P k|k-1 H T HP k|k-1 H T +R k ) -1

[0189] Among them, P k|k-1 Let P be the prior covariance matrix at time k, and let P be the uncertainty of the state vector X-1. k|k-1 =FP k-1 F T+Q, F is the state transition Jacobian, which maps small perturbations in the old state to the sensitivity matrix of the predicted state perturbation; Q is the process noise covariance matrix, describing the uncertainty of the model error plus the sensor's inherent noise injection within one prediction step; P k-1 Let be the posterior covariance matrix at time k-1. A larger value indicates less accurate predictions, meaning the filter becomes more dependent on observations; a smaller value indicates greater confidence in the model. H is the observation matrix, representing the partial derivatives of the state vector with respect to the observation vectors, describing how state variables are mapped to the observation space. P k To observe the noise covariance matrix, the magnitude and cross-correlation of sensor measurement errors are quantified. The variance of the basic observation noise characterizes the inherent noise of the ToF depth camera under ideal conditions, where λ is the attenuation factor and N is the noise level. p The number of matched depth points reflects the number of detectable hand joints in the current ToF frame and is used to dynamically adjust the observation confidence. I is the identity matrix.

[0190] The posterior state vector at the current moment can be specifically represented as follows:

[0191] x k =x k|k-1 +K k (z k -h(x k|k-1 ))

[0192] Where, x k|k-1 Let K be the prior state vector at time k. k Let z be the Kalman gain at time k. k h(x) is the actual observed value at time k. k|k-1 ) represents the predicted observation value at time k.

[0193] Furthermore, after obtaining the posterior state vector at the current time, the posterior covariance matrix at the current time is determined. The covariance matrix at the current time can be specifically represented as follows: P k =(IK k H)P k|k-1 I is the identity mapping, and a "residual projection" is constructed in the covariance update formula.

[0194] In this embodiment, the quaternions in the state vector at each time point are used as the quaternion time series, and the positions in the state vector at each time point are used as the position time series, thereby obtaining the attitude time series.

[0195] Optionally, multiple depth cameras can be deployed, each with a corresponding quaternion time series. The final quaternion time series is obtained by fusing the time series from each quaternion. Similarly, the position time series can be obtained by fusing the position time series from multiple depth cameras. In this way, the accuracy of the pose dimension data is improved by compensating for single-viewpoint occlusion through multi-viewpoint compensation.

[0196] The final quaternion time series described above can be specifically represented as follows:

[0197]

[0198] Where L is the total number of depth cameras, q i Let w be the quaternion time series corresponding to the i-th depth camera. i Let i be the weight of the i-th depth camera. Let be the covariance of the pose estimation of the i-th depth camera.

[0199] Additionally, it is worth noting that in the embodiments of this application, when |a k -g|<∈ a And |ω k |<∈ ω If a preset number of frames are met consecutively, the zero count is reset to zero and the biases of the accelerometer and gyroscope are reassessed. Where a k Let g be the linear acceleration (including gravity) in the body coordinate system at the current moment k, used to determine whether the linear motion is sufficiently "static", where g is the gravity vector and ω is the linear acceleration (including gravity). k Given the IMU gyroscope readings (debiased / filtered), the angular velocity vectors of the body around the X / Y / Z axes at the current moment are used to determine whether the rotation has almost stopped. a For acceleration threshold, ∈ ω This is the angular velocity threshold.

[0200] In addition, it is worth noting that in the embodiments of this application, cubic B-splines can also be fitted to the quaternion sequence to alleviate high-frequency jumps.

[0201] S2021: Based on the attitude time series, determine the force time series and torque time series of each hand joint, and use the normal pressure data, tangential force data and the force time series and torque time series of each hand joint of each array node as the force dimension data corresponding to each array node.

[0202] For example, the force dimension data corresponding to an array node includes: the force time series and torque time series of each hand joint, as well as the normal pressure data and tangential force data of the array node.

[0203] Additionally, it is worth noting that in this embodiment, each finger has 4 joints, for a total of 20 joints across the five fingers.

[0204] In this embodiment, the end pose of each finger is determined based on the posture time series and the DH parameter set of each hand joint. Then, based on the end pose, inverse dynamics calculation is performed to obtain the force time series and torque time series of each hand joint.

[0205] Specifically, the DH parameter set of each hand joint is defined, and then the transformation matrix is ​​constructed to obtain the total fingertip transformation including the end pose. Then, the mapping of joint angular velocity and end velocity is defined through the Jacobian matrix. Finally, based on the end loading, the force time series and torque time series of each finger joint are obtained by reverse recursion (k=4 to 1).

[0206] For example, the DH parameter set of the j-th joint of the i-th finger can be represented as: {a i,j ,α i,j ,d i,j ,θ i,j}

[0207] Among them, a i,j Let α be the length of the link. i,j d is the deflection angle of the connecting rod. i,j For the link offset, θ i,j This refers to the joint angle.

[0208] The transformation matrix can be represented as:

[0209]

[0210] The terminal velocity mapping can be expressed as:

[0211] Among them, the linear velocity series J v,i,k =z k-1 ×(p e,i -p k-1 ), angular velocity train J ω,i,k =z k-1 p e,i The orientation of the fingertip relative to the base, z k-1 Let p be the joint rotation axis of the k-th joint. k-1 v is the origin of the k-th joint; e,i ω represents the instantaneous translational velocity of the origin of the end frame in 3D space, describing how many millimeters the fingertip moves per second in the X, Y, and Z directions; e,i This represents the instantaneous rotation rate of the end frame around its own X, Y, and Z axes; it describes the instantaneous rotation of the fingertip. This represents the joint angular velocity.

[0212] The formula for calculating the total fingertip transformation can be expressed as:

[0213]

[0214] The force at the k-th joint of the i-th finger can be expressed as:

[0215]

[0216] Where, m k,i Let a be the mass of the k-th joint of the i-th finger. k,i Let be the link length of the k-th joint of the i-th finger, be the external force / torque return stage of the k-th joint of the i-th finger, and be the attitude matrix of the link k+1 coordinate system relative to the inertial or base coordinate system. It transforms any vector (force, torque, velocity, acceleration) represented in the local coordinate system of the link to the inertial / world system, or inversely transforms it (takes the transpose).

[0217] The torque of the k-th joint of the i-th finger can be expressed as:

[0218]

[0219] Among them, I k,i Let a be the inertia of the k-th joint of the i-th finger. k,i Let ω be the deflection angle of the link at the k-th joint of the i-th finger. k,i Let p be the angular velocity of the k-th joint of the i-th finger. k,i Let p be the position vector of the k-th joint of the i-th finger. k+1,i Let be the position vector of the (k+1)th joint of the i-th finger.

[0220] Furthermore, regarding τ k,i Take a sample to obtain the torque values ​​of each joint of each finger.

[0221] In this way, the fine movements and mechanical characteristics of the fingers during the massage process are accurately quantified, improving the accuracy of force application recognition.

[0222] S2022: Use the contact surface temperature data and back of hand temperature data of each array node as the temperature dimension data of each array node.

[0223] For example, the temperature dimension data corresponding to an array node includes: hand back temperature data and contact surface temperature data of the array node.

[0224] In this way, high-dimensional discriminative features that reflect the dynamic changes of the applied force and characterize the spectral features and spatial distribution are extracted from multi-source sensor signals, providing input rich in the semantics of the applied force for subsequent force recognition. Force dimension data, posture dimension data, and temperature dimension data each reflect different dimensions of the applied force. Force dimension data reflects the strength of the contact force, posture dimension data describes the trajectory of the hand movement, and temperature dimension data reveals the temperature changes of the touch during the applied force. By combining the features of temporal evolution (time domain), spectral components (frequency domain), and spatial topology (array arrangement), the three elements of the applied force—"when force is applied," "where force is applied," and "frequency of force change"—can be fully reproduced.

[0225] S21: Extract features from the multidimensional data of each array node to obtain the initial feature vector of each array node, and update the obtained initial feature vectors according to the spatial adjacency relationship of each array node to obtain the embedding vector of the target massage operation.

[0226] Optionally, in this embodiment of the application, a possible implementation is provided for extracting features from the multidimensional data of each array node to obtain the initial feature vector of each array node, specifically by performing the following operations:

[0227] S210: Extract features from different domains of the multidimensional data of each array node to obtain multiple sub-initial feature vectors for each array node.

[0228] The different domains include: time domain, frequency domain, and topological domain.

[0229] In this embodiment of the application, temporal features are extracted from the force dimension data of each array node to obtain the first sub-initial feature vector of each array node.

[0230] Specifically, the following operations are performed on each array node: wavelet packet decomposition is performed on the force dimension data of an array node to obtain the quantized energy spectrum, which is used as the first sub-initial feature vector of that array node.

[0231] For example, Daubechies-5 wavelet packet decomposition is performed every 128ms sliding window to obtain the quantized energy spectrum, which can be represented as:

[0232]

[0233] Among them, W p,l [n] represents the coefficient of the l-th subband, where l = 1, 2, ..., L p L p To determine the number of decomposition layers, the dimension of the first sub-initial feature vector is L. p .

[0234] In this embodiment of the application, frequency domain features are extracted from the attitude dimension data of each array node to obtain the second sub-initial feature vector of each array node.

[0235] Specifically, the following operations are performed on each array node: autoregressive modeling is performed on the attitude dimension data of an array node, and the autoregressive coefficients are extracted as the second sub-initial feature vector of that array node.

[0236] For example, in this embodiment of the application, the attitude dimension data (e.g., Euler angles θ(t), obtained through quaternion transformation) is fitted with an AR(p) model to obtain the coefficients {φ}. k}

[0237] in, With φ=[φ1,…,φ p ] is the second sub-initial feature vector, and the dimension of the second sub-initial feature vector is p.

[0238] In this embodiment of the application, topological domain features are extracted from the temperature dimension data of each array node to obtain the third sub-initial feature vector of each array node.

[0239] Specifically, the following operations are performed on each array node: the temperature dimension data of an array node is used to perform gradient calculation to obtain the temperature gradient spectrum, which is used as the third sub-initial feature vector of that array node.

[0240] For example, in this embodiment of the application, the gradient tensor is calculated for the temperature dimension data T(i,j), and the temperature gradient spectrum can be expressed as:

[0241]

[0242] S211: Standardize and concatenate the multiple sub-initial feature vectors of each array node to obtain the initial feature vector of each array node.

[0243] In this embodiment of the application, the mean and variance of the multiple sub-initial feature vectors of each array node are normalized, and then the normalized vectors are concatenated to obtain the initial feature vectors of each array node.

[0244] For example, the initial feature vector of an array node can be represented as: This is the first sub-initial eigenvector after mean and variance normalization. This is the second sub-initial eigenvector after mean and variance normalization. This is the third initial eigenvector after mean and variance normalization.

[0245] In this way, the first sub-initial feature vector represents the time-varying force at different scales, the second sub-initial feature vector captures the rhythm of the gesture and the oscillation frequency, and the third sub-initial feature vector reveals the spatial texture of the heat distribution, which can improve the accuracy of force application action recognition.

[0246] Optionally, in this embodiment of the application, a possible implementation is provided for updating each initial feature vector obtained according to the spatial adjacency relationship of each array node to obtain the embedding vector of the target massage operation, specifically by performing the following operations:

[0247] S212: Construct a graph based on the spatial adjacency relationships of each array node.

[0248] In this graph, the vertices are the array nodes, and the edge set is the spatial adjacency relationship between the array nodes.

[0249] In this embodiment of the application, a graph G = (V, E) is constructed based on the spatial adjacency relationship of each array node.

[0250] S213: Based on the graph, perform graph convolution and attention aggregation on each initial feature vector to obtain the embedding vector.

[0251] In this embodiment, each initial feature vector is used as an attribute of a graph node, an L-layer graph convolution is performed, and an attention mechanism is introduced to calculate edge weights. After weighted aggregation, the vectors are projected through a fully connected layer and activated to obtain the embedded vectors.

[0252] For example, the graph convolution of the (l+1)th layer can be represented as:

[0253]

[0254] Where A is the adjacency matrix, generated from the spatial adjacency relationships of each array node: if node v i With v j Adjacent, then A {ij} =1, otherwise A {ij} =0; I is the adjacency matrix, the identity matrix, which adds self-loops to each node to preserve its own characteristics; D is the degree matrix, i.e., the connectivity degree of node i (including self-loops), used to normalize and suppress degree differences; H (l) Let W be the input feature matrix of the l-th layer. (l) These are learnable weights.

[0255] For example, attention can be represented as:

[0256]

[0257] For example, an embedding vector can be represented as:

[0258]

[0259] Among them, H j (L) The node representation output by the Lth layer of the graph convolution is the high-order feature vector of node j after the Lth layer (which has been fused with neighborhood information). When the graph has 27 array nodes, there are 27 such vectors. For neighborhood indexing, in the attention mechanism, the neighbor node numbers aggregated by node i are used. If six-way adjacency is adopted, then... At most 6; W o Output the projection weight matrix, which linearly maps the features of the stitched whole image to the embedding space.

[0260] Additionally, it is worth noting that in this embodiment, the embedding vector is 512-dimensional.

[0261] See Figure 5 The diagram shown is a logical schematic diagram of obtaining an embedding vector provided in an embodiment of this application.

[0262] S22: The embedded vector is fused with temporal and spatial information to obtain the target modeling vector, and the force application action recognition result of the target massage operation is determined based on the target modeling vector.

[0263] Optionally, in this embodiment of the application, a possible implementation is provided for fusing temporal and spatial information of the embedded vector to obtain the target modeling vector, specifically by performing the following operations:

[0264] S220: Perform temporal dynamic feature extraction on the embedded vector to obtain a temporal feature vector, and perform spatial distribution feature extraction on the embedded vector to obtain a spatial feature vector.

[0265] Among them, the temporal feature vector represents the temporal dynamic semantics of the target massage operation, and the spatial feature vector represents the spatial distribution semantics of the target massage operation.

[0266] In this embodiment, temporal and spatial location encodings are added to the embedded vectors, and the vectors are fed into two parallel Transformer towers (temporal tower and spatial tower) to obtain temporal feature vectors and spatial feature vectors.

[0267] For example, the input to the time-domain tower is E t =f+P t The input to the airspace tower is E. s =f+P s Where f is the embedding vector, P is the temporal code, and P s Spatial location encoding. Each tower contains L layers of identical structure, and the calculation for the l-th layer is: H (l+1) =LN(FFN(MHA(H (l) ))+H (l)), where MHA is multi-head self-attention, FFN is a two-layer feedforward network, and LN is layer normalization.

[0268] S221: The temporal feature vector and the spatial feature vector are fused to obtain the target modeling vector.

[0269] In this embodiment of the application, after obtaining the temporal feature vector and the spatial feature vector, the temporal feature vector and the spatial feature vector are fused by linear gating to obtain the target modeling vector.

[0270] For example, the target modeling vector can be represented as: z = σ(W t z t +W s z s ), where z t z is the temporal feature vector. s Let W be a spatial feature vector, σ be an activation function (e.g., ReLU), and W be a spatial feature vector. t and W s As weight.

[0271] Additionally, it is worth noting that in this embodiment, the target modeling vector is 256-dimensional.

[0272] In this way, by fusion of time and space, a target modeling vector can be obtained, which improves the accuracy of force action recognition and also accurately quantifies the target massage operation.

[0273] Optionally, in this application embodiment, a possible implementation is provided for determining the force application action recognition result of the target massage operation based on the target modeling vector, specifically by performing the following operations:

[0274] S222: Determine the distance between the target modeling vector and the center vector of each force application action category.

[0275] The center vector for each force application action category is determined based on the sample modeling vectors for each force application action category.

[0276] In this embodiment of the application, the prototype library pre-stores the center vectors of each force application action category. The center vector of a force application action category in the prototype library is obtained by clustering or finding the centroid of a large number of sample modeling vectors labeled with that force application action category in the offline stage. There are 32 categories of each force application action category. A unified objective label is formulated to support large-scale intelligent training and feedback.

[0277] For example, the formula for calculating the distance between the target modeling vector and the center vector of each force application action category is: z is the target modeling vector, p j Let be the center vector of the j-th force application action category.

[0278] S223: Based on the obtained distances, determine the force application action recognition result of the target massage operation.

[0279] In this embodiment of the application, the force application recognition result of the target massage operation is determined based on the obtained distances, including the following two cases:

[0280] Case 1: When there is a target distance greater than the distance threshold among the various distances, the force application action category of the target massage operation is determined to be the force application action category corresponding to the target distance.

[0281] In this embodiment of the application, the probability corresponding to each force application action category can be obtained by converting each distance into a probability distribution. When the highest probability among all probabilities is greater than the probability threshold, the force application action category of the target massage operation is determined to be the force application action category corresponding to the highest probability.

[0282] The probability threshold can be 0.6, but this application does not impose any restrictions on it.

[0283] For example, the probability of the j-th force application action category can be expressed as:

[0284]

[0285] Case 2: When there is no target distance greater than the distance threshold among the distances, the force application action recognition result of the massage operation is determined based on the distance between the target modeling vector and the sample modeling vectors of each force application action category.

[0286] In this embodiment of the application, when the largest probability among all probabilities is not greater than the probability threshold, the force application action recognition result of the target massage operation is determined based on the distance between the target modeling vector and the sample modeling vectors of each force application action category.

[0287] Specifically, the k sample modeling vectors that are closest to the target modeling vector among the sample modeling vectors are determined. Then, the proportion of the target category that appears most frequently among the k sample modeling vectors is determined, i.e., the consistency rate. If the consistency rate is greater than the consistency rate threshold, then the force application action category of the target massage operation is determined as the target category.

[0288] The consistency rate threshold can be 0.8, but this application embodiment does not impose any restrictions on it.

[0289] In this way, k-NN backups are performed for low-confidence classes to improve the robustness of classification and ensure high reliability.

[0290] In summary, by mapping the embedding vectors obtained after multimodal fusion to the standard 32 categories of force application actions, the system can automatically identify and label the force application action categories of the target massage operation. Through prototype measurement, it can quickly locate the best matching category in the force application action model library, and enable k-NN backup when the confidence is insufficient to ensure real-time performance and robustness. It also supports continuous cloud updates of the prototype library and model parameters, providing data support for personalized application needs.

[0291] Furthermore, in the embodiments of this application, after determining the force application action recognition result of the target massage operation, a control sequence for the target massage operation can be generated based on the standard trajectory template of the target force application action category of the target massage operation. The control sequence is used to drive the actuator to perform the target massage operation, and / or to determine whether the target massage operation is qualified based on the standard trajectory template of the target force application action category of the target massage operation.

[0292] The actuators are massage robots and other massage equipment.

[0293] In this embodiment of the application, the standard trajectory template can be represented as:

[0294] Among them, u i For the control sequence to be optimized, g i w is the target trajectory point i To track the weights, Δu is the control increment, and λ is the regularization coefficient.

[0295] See Figure 6 The diagram shown is a logical schematic diagram of a method for determining the recognition result of a force application action according to an embodiment of this application.

[0296] The following is a brief introduction to the process of training a target force action recognition model. (See also...) Figure 7 The diagram shown is a flowchart illustrating the training process of the target force action recognition model.

[0297] S70: Obtain the training sample set, where each training sample in the training sample set includes at least: the sample embedding vector of the sample massage operation.

[0298] In this embodiment of the application, the above-described method for obtaining embedding vectors is used to obtain the sample embedding vectors of the sample massage operation and construct a training sample set.

[0299] Additionally, it is worth noting that in this embodiment of the application, the training sample set can be obtained from the edge device and then sent to the cloud.

[0300] S71: Based on the training sample set, iteratively train the force application action recognition model to be trained to obtain the target force application action recognition model. During one iteration of training, the following operations are performed:

[0301] The sample embedding vectors of the selected target training samples are fused with temporal and spatial information to obtain the sample modeling vectors of the target training samples; the sample distances between the sample modeling vectors of the target training samples and the center vectors of each force application action category are determined; based on the obtained sample distances, the sample force application action recognition results of the target training samples are determined, and the parameters are tuned based on the loss values ​​corresponding to the sample force application action recognition results.

[0302] In this embodiment of the application, the above-mentioned spatiotemporal information fusion method is similarly used to obtain the sample modeling vector of the target training sample, determine the sample distance between the sample modeling vector of the target training sample and the center vector of each force application action category, when there is a target sample distance greater than the sample distance threshold among the sample distances, the sample force application action category of the sample massage operation is determined as the force application action category corresponding to the target sample distance, when there is no target sample distance greater than the distance threshold among the sample distances, the sample force application action recognition result of the sample massage operation is determined based on the sample modeling vector and the sample modeling vector of each force application action category.

[0303] Specifically, when determining the force application action recognition result of the sample massage operation based on the sample modeling vector and the sample modeling vector of each force application action category, the top k sample modeling vectors with the closest sample distance to the sample modeling vector of the target training sample are determined. Then, the proportion of the target category that appears most frequently among the k sample modeling vectors is determined, i.e., the sample consistency rate. If the sample consistency rate is greater than the sample consistency rate threshold, the force application action category of the sample massage operation is determined as the target category. Otherwise, the force application action category of the sample massage operation is determined by manual annotation, and the training sample is re-added to the training sample set for training.

[0304] Specifically, in the embodiments of this application, when tuning parameters based on the loss value corresponding to the sample force application action recognition result, the objective function includes, but is not limited to: cross-entropy loss, center loss, and regularization term loss.

[0305] Furthermore, in the embodiments of this application, after adjusting the parameters based on the loss value corresponding to the sample force application action recognition result, when the target training sample is a high-confidence sample, the center vector of the force application action category to which the target training sample belongs can be updated based on the sample modeling vector of the target training sample.

[0306] Specifically, the sample distance between the sample modeling vector of the target training sample and the center vector of each force application action category is determined. When there is a target sample distance among the sample distances that is greater than the sample distance threshold, the target training sample is a high-confidence sample. At this time, momentum update is used to update the center vector of the force application action category to which the target training sample belongs.

[0307] For example, the center vector of the applied action category to which the updated target training sample belongs can be represented as: c j ←mc j +(1-m)z, where c j Let m be the center vector of the force application action category to which the target training sample belongs, m be the momentum coefficient that determines the weight of the old center vector, and z be the sample modeling vector of the target training sample.

[0308] Furthermore, based on the same technical concept, embodiments of this application provide an intelligent recognition device for force application actions based on AI and multiple sensors. This intelligent recognition device for force application actions based on AI and multiple sensors is used to implement the above-described method flow of embodiments of this application. For example, see [link to relevant documentation]. Figure 8 As shown, the intelligent force recognition device 800 based on AI and multiple sensors may include: an acquisition module 801, a data fusion module 802, a force recognition module 803, and a training module 804, wherein:

[0309] The acquisition module 801 is used to acquire the multi-dimensional data corresponding to each array node in the contact surface array during the target massage operation. The array nodes are arranged at equal intervals, and each array node, except for the boundary nodes, has at least a preset number of neighbor nodes.

[0310] The data fusion module 802 is used to extract features from the multidimensional data of each array node, obtain the initial feature vector of each array node, and update the obtained initial feature vectors according to the spatial adjacency relationship of each array node to obtain the embedding vector of the target massage operation.

[0311] The force application action recognition module 803 is used to fuse temporal and spatial information of the embedded vector to obtain the target modeling vector, and to determine the force application action recognition result of the target massage operation based on the target modeling vector.

[0312] In one alternative embodiment, the multidimensional data includes force dimension data, attitude dimension data, and temperature dimension data.

[0313] In an optional embodiment, when acquiring the multidimensional data corresponding to each array node in the contact surface array during the target massage operation, the acquisition module 801 is further configured to:

[0314] Initial data for each array node during the target massage operation is acquired by deploying a contact surface array on the surface of the first object.

[0315] Second initial data is collected by deploying a hand multimodal sensor on the second object's hand;

[0316] Based on the second initial data and the first initial data of each array node, the multidimensional data corresponding to each array node is determined.

[0317] In one optional embodiment, each array node includes a pressure sensor, a force sensor, and a temperature sensor, and the first initial data includes normal pressure data, tangential force data, and contact surface temperature data.

[0318] The hand multimodal sensing includes an inertial measurement unit, a depth camera, and a temperature electrode. The second initial data includes inertial data, hand point cloud data, and back of hand temperature data.

[0319] In an optional embodiment, when determining the multidimensional data corresponding to each array node based on the second initial data and the first initial data of each array node, the acquisition module 801 is further configured to:

[0320] Based on inertial data and hand point cloud data, the attitude time series is determined and used as the attitude dimension data corresponding to each array node.

[0321] Based on the attitude time series, the force time series and torque time series of each hand joint are determined, and the normal pressure data, tangential force data and the force time series and torque time series of each hand joint of each array node are used as the force dimension data corresponding to each array node.

[0322] The contact surface temperature data and back of hand temperature data of each array node are used as the temperature dimension data of each array node.

[0323] In an optional embodiment, when determining the attitude time series based on inertial data and hand point cloud data, the acquisition module 801 is further configured to:

[0324] An adaptive extended Kalman filter is applied to the inertial data and hand point cloud data to obtain the attitude time series.

[0325] In an optional embodiment, when determining the force time series and torque time series of each hand joint based on the posture time series, the acquisition module 801 is further configured to:

[0326] Based on the posture time series and the DH parameter set of each hand joint, the end pose of each finger is determined;

[0327] Based on the end-effector pose, inverse dynamics calculations are performed to obtain the force time series and torque time series of each hand joint.

[0328] In an optional embodiment, when performing feature extraction on the multidimensional data of each array node to obtain the initial feature vector of each array node, the data fusion module 802 is further configured to:

[0329] Feature extraction of different domains is performed on the multidimensional data of each array node to obtain multiple sub-initial feature vectors for each array node.

[0330] The multiple sub-initial feature vectors of each array node are standardized and concatenated to obtain the initial feature vector of each array node.

[0331] In an optional embodiment, when performing feature extraction from different domains on the multidimensional data of each array node to obtain multiple sub-initial feature vectors for each array node, the data fusion module 802 is further used to:

[0332] Temporal features are extracted from the force dimension data of each array node to obtain the first sub-initial feature vector of each array node.

[0333] Frequency domain features are extracted from the attitude dimension data of each array node to obtain the second sub-initial feature vector of each array node.

[0334] Topological domain features are extracted from the temperature dimension data of each array node to obtain the third sub-initial feature vector of each array node.

[0335] In an optional embodiment, when updating the obtained initial feature vectors according to the spatial adjacency relationship of each array node to obtain the embedding vector of the target massage operation, the data fusion module 802 is further configured to:

[0336] A graph is constructed based on the spatial adjacency relationship of each array node, where the vertices of the graph are each array node, and the edge set of the graph is the spatial adjacency relationship between each array node.

[0337] Based on the graph, graph convolution and attention aggregation are performed on each initial feature vector to obtain the embedding vector.

[0338] In an optional embodiment, when obtaining the modeling vector of the target massage operation by fusing temporal and spatial information from the embedded vector, the force application action recognition module 802 is further configured to:

[0339] Temporal dynamic feature extraction is performed on the embedded vector to obtain temporal feature vector, and spatial distribution feature extraction is performed on the embedded vector to obtain spatial feature vector. The temporal feature vector represents the temporal dynamic semantics of the target massage operation, and the spatial feature vector represents the spatial distribution semantics of the target massage operation.

[0340] The temporal feature vector and the spatial feature vector are fused to obtain the target modeling vector.

[0341] In an optional embodiment, when determining the force application action recognition result of the target massage operation based on the target modeling vector, the force application action recognition module 803 is further configured to:

[0342] Determine the distance between the target modeling vector and the center vector of each force application action category, where the center vector is determined based on the sample modeling vectors of the corresponding force application action category;

[0343] Based on the obtained distances, the force application action recognition results of the target massage operation are determined.

[0344] In an optional embodiment, when determining the force application action recognition result of the target massage operation based on the obtained distances, the force application action recognition module 803 is further configured to:

[0345] When there is a target distance greater than the distance threshold among all distances, the force application action category of the target massage operation is determined to be the force application action category corresponding to the target distance;

[0346] When there is no target distance greater than the distance threshold among the various distances, the force application action recognition result of the target massage operation is determined based on the distance between the target modeling vector and the sample modeling vectors of each force application action category.

[0347] In an optional embodiment, the device further includes a training module 804, which is used for:

[0348] Obtain a training sample set, wherein each training sample in the training sample set includes at least: the sample embedding vector of the sample massage operation;

[0349] Based on the training sample set, the force application action recognition model to be trained is iteratively trained to obtain the target force application action recognition model. During each iteration of training, the following operations are performed:

[0350] The temporal and spatial information of the sample embedding vectors of the selected target training samples are fused to obtain the sample modeling vectors of the target training samples.

[0351] Determine the sample distance between the sample modeling vector of the target training sample and the center vector of each force application action category;

[0352] Based on the obtained distances between samples, the sample force action recognition results of the target training sample are determined, and the parameters are tuned based on the loss value corresponding to the sample force action recognition results.

[0353] In an optional embodiment, based on the obtained distances between samples, the force application action recognition result of the target training sample is determined, and after parameter tuning based on the loss value corresponding to the sample force application action recognition result, the training module 804 is further configured to:

[0354] When the target training sample is a high-confidence sample, the center vector of the force application action category to which the target training sample belongs is updated based on the sample modeling vector of the target training sample.

[0355] Based on the description of the method and apparatus embodiments above, an exemplary embodiment of the present invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, which, when executed by the at least one processor, causes the electronic device to perform the method according to an embodiment of the present invention.

[0356] This application also provides a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to an embodiment of this application.

[0357] This application also provides a computer program product, including a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to an embodiment of this application.

[0358] See Figure 9 The diagram shown below illustrates a structural block diagram of an electronic device 900 that can serve as a server or client of this application, which is an example of a hardware device that can be applied to various aspects of this application. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the application described and / or claimed herein.

[0359] like Figure 9As shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 902 or a computer program loaded from a storage unit 908 into a random access memory (RAM) 903. The RAM 903 may also store various programs and data required for the operation of the device 900. The computing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0360] Multiple components in electronic device 900 are connected to I / O interface 905, including: input unit 906, output unit 907, storage unit 908, and communication unit 909. Input unit 906 can be any type of device capable of inputting information to electronic device 900. Input unit 906 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 907 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 908 may include, but is not limited to, disk and optical disk. Communication unit 909 allows electronic device 900 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers and / or chipsets, such as Bluetooth devices, WiFi devices, worldwide interoperability for microwave access (WiMax) devices, cellular communication devices, and / or the like.

[0361] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 901 performs the various methods and processes described above. For example, in some embodiments, the above-described intelligent recognition method for force application based on AI and multiple sensors can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 900 via ROM 902 and / or communication unit 909. In some embodiments, the computing unit 901 can be configured to perform the above-described intelligent recognition method for force application based on AI and multiple sensors by any other suitable means (e.g., by means of firmware).

[0362] The program code used to implement the methods of this application may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing device, such that when executed by the processor or controller, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0363] In the context of this application, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, RAM, ROM, erasable programmable read-only memory (EPROM) or flash memory, optical fibers, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0364] As used in this application, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device, PLD) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0365] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0366] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0367] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.

[0368] Furthermore, it should be understood that the above-disclosed embodiments are merely preferred embodiments of this application and should not be construed as limiting the scope of the invention. Therefore, any equivalent variations made in accordance with the claims of this invention are still within the scope of this application.

Claims

1. An AI and multi-sensor-based intelligent recognition method for force application actions, characterized by, include: In the target massage operation, acquire the multidimensional data corresponding to each array node in the contact surface array, wherein the multidimensional data includes force dimension data, posture dimension data and temperature dimension data, the array nodes are arranged at equal intervals, and each array node except the boundary node has at least a preset number of neighbor nodes. Temporal feature extraction is performed on the force dimension data of each array node to obtain the first sub-initial feature vector of each array node. Frequency domain features are extracted from the attitude dimension data of each array node to obtain the second sub-initial feature vector of each array node. Topological domain features are extracted from the temperature dimension data of each array node to obtain the third sub-initial feature vector of each array node. The first sub-initial feature vector, the second sub-initial feature vector, and the third sub-initial feature vector of each array node are standardized and concatenated to obtain the initial feature vector of each array node. Based on the spatial adjacency relationship of each array node, a graph is constructed, and based on the graph, graph convolution and attention aggregation are performed on each obtained initial feature vector to obtain the embedding vector of the target massage operation, wherein the vertices of the graph are each array node, and the edge set of the graph is the spatial adjacency relationship between each array node; The embedded vector is fused with temporal and spatial information to obtain a target modeling vector, and the force application action recognition result of the target massage operation is determined based on the target modeling vector.

2. The method of claim 1, wherein, The acquisition of multi-dimensional data corresponding to each array node in the contact surface array during the target massage operation includes: First initial data of each array node during the target massage operation is acquired by deploying the contact surface array on the surface of the first object. Second initial data is collected by deploying a hand multimodal sensor on the second object's hand; Based on the second initial data and the first initial data of each array node, the multidimensional data corresponding to each array node is determined.

3. The method of claim 2, wherein, Each array node includes a pressure sensor, a force sensor, and a temperature sensor. The first initial data includes normal pressure data, tangential force data, and contact surface temperature data. The hand multimodal sensing includes an inertial measurement unit, a depth camera, and a temperature electrode. The second initial data includes inertial data, hand point cloud data, and hand back temperature data.

4. The method of claim 3, wherein, The step of determining the multidimensional data corresponding to each array node based on the second initial data and the first initial data of each array node includes: Based on the inertial data and the hand point cloud data, the attitude time series is determined, and the attitude time series is used as the attitude dimension data corresponding to each array node. Based on the posture time series, the force time series and torque time series of each hand joint are determined, and the normal pressure data and tangential force data of each array node, as well as the force time series and torque time series of each hand joint, are used as the force dimension data corresponding to each array node. The contact surface temperature data of each array node and the back of the hand temperature data are used as the temperature dimension data corresponding to each array node.

5. The method of claim 4, wherein, The determination of the attitude time series based on the inertial data and the hand point cloud data includes: The attitude time series is obtained by performing adaptive extended Kalman filtering on the inertial data and the hand point cloud data.

6. The method of claim 4, wherein, The determination of the force time series and torque time series of each hand joint based on the posture time series includes: Based on the posture time series and the DH parameter set of each hand joint, the end pose of each finger is determined; Based on the end-effector pose, inverse dynamics calculations are performed to obtain the force time series and torque time series of each hand joint.

7. The method of claim 1, wherein, The process of fusing temporal and spatial information into the embedded vector to obtain the target modeling vector includes: Temporal dynamic feature extraction is performed on the embedded vector to obtain a temporal feature vector, and spatial distribution feature extraction is performed on the embedded vector to obtain a spatial feature vector, wherein the temporal feature vector represents the temporal dynamic semantics of the target massage operation, and the spatial feature vector represents the spatial distribution semantics of the target massage operation; The temporal feature vector and the spatial feature vector are fused to obtain the target modeling vector.

8. The method of claim 1, wherein, The determination of the force application action recognition result of the target massage operation based on the target modeling vector includes: Determine the distance between the target modeling vector and the center vector of each force application action category, wherein the center vector is determined based on the sample modeling vectors of the corresponding force application action category; Based on the obtained distances, the force application action recognition result of the target massage operation is determined.

9. The method as described in claim 8, characterized in that, The determination of the force application action recognition result of the target massage operation based on the obtained distances includes: When there is a target distance greater than the distance threshold among the distances, the force application action category of the target massage operation is determined to be the force application action category corresponding to the target distance; When there is no target distance greater than the distance threshold among the distances, the force application action recognition result of the target massage operation is determined based on the distance between the target modeling vector and the sample modeling vectors of each force application action category.

10. The method according to any one of claims 8-9, characterized in that, The force application action recognition result is obtained by inputting the target modeling vector into the target force application action recognition model, wherein the target force application action recognition model is trained in the following manner: Obtain a training sample set, wherein each training sample in the training sample set includes at least: a sample embedding vector of a sample massage operation; Based on the training sample set, the force application action recognition model to be trained is iteratively trained to obtain the target force application action recognition model. During one iteration of training, the following operations are performed: The temporal and spatial information of the sample embedding vector of the selected target training sample are fused to obtain the sample modeling vector of the target training sample. Determine the sample distance between the sample modeling vector of the target training sample and the center vector of each force application action category; Based on the obtained distances between the samples, the sample force application action recognition result of the target training sample is determined, and the parameters are tuned based on the loss value corresponding to the sample force application action recognition result.

11. The method as described in claim 10, characterized in that, After determining the sample force application action recognition result of the target training sample based on the obtained sample distances, and tuning the parameters based on the loss value corresponding to the sample force application action recognition result, the method further includes: When the target training sample is a high-confidence sample, the center vector of the force application action category to which the target training sample belongs is updated based on the sample modeling vector of the target training sample.

12. An intelligent recognition device for force application actions based on AI and multiple sensors, characterized in that, include: The acquisition module is used to acquire the multidimensional data corresponding to each array node in the contact surface array during the target massage operation. The multidimensional data includes force dimension data, posture dimension data and temperature dimension data. The array nodes are arranged at equal intervals, and each array node, except for the boundary nodes, has at least a preset number of neighbor nodes. The data fusion module is used to extract temporal features from the force dimension data of each array node to obtain a first sub-initial feature vector for each array node; extract frequency domain features from the attitude dimension data of each array node to obtain a second sub-initial feature vector for each array node; extract topological domain features from the temperature dimension data of each array node to obtain a third sub-initial feature vector for each array node; standardize and concatenate the first, second, and third sub-initial feature vectors of each array node to obtain an initial feature vector for each array node; construct a graph based on the spatial adjacency relationship of each array node; and perform graph convolution and attention aggregation on the obtained initial feature vectors based on the graph to obtain the embedding vector of the target massage operation, wherein the vertices of the graph are each array node, and the edge set of the graph is the spatial adjacency relationship between each array node; The force application action recognition module is used to fuse temporal and spatial information into the embedded vector to obtain a target modeling vector, and to determine the force application action recognition result of the target massage operation based on the target modeling vector.

Citation Information

Patent Citations

  • Massage robot control method and system, medium, program product and terminal

    CN119235637A

  • Smart vision sensor system and method

    US20200097706A1