Multi-modal fusion personnel intention recognition method, system and equipment and medium
Through the multimodal fusion personnel intention recognition method, combined with RGB video and inertial sensing data, the decision-level and feature-level fusion strategies are used to solve the problem of poor single-modal recognition effect, and achieve higher recognition accuracy and robustness.
Patent Information
- Application Number
- CN202510219472.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2045-02-26
AI Technical Summary
In the prior art, the single-modal personnel intention recognition method cannot fully understand the operator's intention, especially in the human-machine manufacturing cooperation scenario, there are defects of poor recognition effects.
A multimodal fusion personnel intention recognition method is adopted, and through decision-level fusion and feature-level fusion strategies, combined with RGB video and inertial sensing data, a personnel intention recognition model is built to improve identification accuracy.
Through multimodal fusion, the limitations of single-modal recognition are overcome, and the accuracy and robustness of operator intentions are significantly improved in the human-machine manufacturing collaboration scenario.
Smart Images

Figure CN120029463A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of human-machine collaboration and intelligent manufacturing, and more specifically, relates to a multi-modal fusion personnel intention recognition method, system, device and medium. Background Art
[0002] Human-machine collaboration, with its advantages of high flexibility, efficiency, adaptability, personalization and safety, has become an important driving force for industrial transformation towards intelligent manufacturing. In human-machine collaboration, the high precision, strength and repeatability of robots are combined with the high flexibility and adaptability of operators to achieve optimal productivity. One of the keys to achieving this goal is that robots can dynamically plan safe responses based on the intentions of operators. Therefore, human intention recognition, as a prerequisite for dynamic decision-making by robots, is crucial to achieving efficient human-machine manufacturing collaboration, including human-machine collaborative assembly, human-machine collaborative welding, etc.
[0003] At present, person intention recognition is mainly achieved through single modality. In terms of single modality, especially RGB video-based person intention recognition and inertial sensor-based person intention recognition, they are the two mainstreams for current person intention recognition. RGB video data can be applied to visually sensitive person recognition, and inertial sensor data can be applied to fine-grained person recognition. However, although these two single-modality methods have made some considerable progress in person intention recognition, they each have limitations that cannot be ignored.
[0004] 1. Person intention recognition based on RGB video: Video can capture complex movements and contextual information of people, but this method is limited by physical factors such as limited field of view (occlusions, viewing angle limitations), lighting environment, and complex background. In particular, if the person is not within the visible range of the camera, the intention recognition cannot be determined.
[0005] 2. HAR based on wearable inertial sensors: The wearability, diversity, and independence of inertial sensors make them not restricted by specific fields of view and lighting conditions, overcoming the limitations of video-based methods. However, this method mainly captures the movement and posture information of people and cannot provide rich contextual information.
[0006] It can be seen that single-mode sensing cannot fully cover all the information related to personnel intention recognition. For human-machine manufacturing collaboration, operators not only present complex actions, but sometimes the actions have subtle and similar motion patterns (such as wedging pins, tightening bolts) or the same action pattern is in different operation sequences (from a manufacturing perspective, it belongs to two types of action intentions). In this application scenario, personnel intention recognition based on single-mode data cannot fully understand personnel intention recognition, and there is a defect that the personnel intention recognition effect is not good. Summary of the invention
[0007] In view of the above problems, the purpose of the present invention is to provide a multi-modal fusion personnel intention recognition method, system, device and medium, which utilizes decision-level fusion and feature-level fusion strategies, integrates the advantages of RGB video and inertial sensing modalities, overcomes the limitations of single modality in personnel intention recognition, and improves the accuracy of personnel intention recognition in human-machine manufacturing collaboration scenarios.
[0008] In order to achieve the above object, the present invention is implemented through the following technical solutions: In a first aspect, an embodiment of the present application provides a multi-modal fusion method for identifying a person's intention, comprising: In the human-machine collaboration scenario, the operator's RGB video data is collected through the visual sensor, and the inertial sensor data is collected through the wearable inertial sensor as the operator's intention data; the inertial sensor data includes but is not limited to: acceleration, angular velocity, joint quaternion and center of mass displacement; The operator's intention data is selected and, after preprocessing, an action sample training set and an action sample test set are generated; Based on a decision-level fusion strategy or a feature-level fusion strategy, a personnel intention recognition model integrating RGB video and inertial sensing modalities is constructed, and the personnel intention recognition model is trained using an action sample training set; The action sample test set is input into the personnel intention recognition model to perform personnel intention recognition and obtain the personnel intention recognition result.
[0009] In an optional implementation, the data selection of the operator's intention data is performed, and after preprocessing, an action sample training set and an action sample test set are generated, including: Selecting RGB video data of a preset length, dividing the selected RGB video data into a plurality of video sequences of equal length, selecting a frame of image from each video sequence, and cropping the image into a fixed size to generate an image sequence as preprocessed RGB video data; Select the inertial sensor data collected within a preset time period, and organize the selected inertial sensor data into an action sequence X={x 1 , x 2, ..., x T}, as the preprocessed inertial sensor data; where T represents the cycle length of the action sequence, x t is the behavior state of the personnel at time step t, each x t Contains m basic features, x t =(x t1 , x t2 , …,x tm ); The preprocessed RGB video data and the preprocessed inertial sensor data are taken as an action sample; Select M action samples as the action sample training set, and select N action samples as the action sample test set.
[0010] In an optional implementation, the construction of a personnel intention recognition model based on RGB video and inertial sensing modality fusion based on a decision-level fusion strategy or a feature-level fusion strategy includes: The decision-level fusion strategy is used to build a personnel intention recognition model that integrates RGB video and inertial sensing modalities. The specific implementation process is as follows: Modeling the temporal and spatial features of the preprocessed RGB video data based on a three-dimensional convolutional network to construct a first personnel intention recognition sub-model; the first personnel intention recognition sub-model is used to generate a first operator intention recognition result based on the preprocessed RGB video data; Based on the bidirectional long short-term memory-bidirectional gated loop-converter network, the temporal characteristics and global characteristics of the preprocessed inertial sensor data are modeled to construct a second person intention recognition sub-model; the second person intention recognition sub-model is used to generate a second operator intention recognition result according to the preprocessed inertial sensor data; Constructing a calculation module based on a decision fusion algorithm; the calculation module is used to obtain the first operator intention recognition result and the second operator intention recognition result, and use the decision fusion algorithm to calculate the personnel intention recognition result; The decision fusion algorithm includes:
[0011] Where, RGB_output[i] is the first operator intention recognition result of the i-th action sample, IS_output[i] is the second operator intention recognition result of the i-th action sample, W RGB is the weight of the first operator's intention recognition result, W IS is the weight of the first operator's intention recognition result, is the person intention recognition result of the i-th action sample.
[0012] In an optional implementation, the construction of a personnel intention recognition model based on RGB video and inertial sensing modality fusion based on a decision-level fusion strategy or a feature-level fusion strategy includes: A human intention recognition model based on RGB video and inertial sensor feature fusion is constructed using a feature-level fusion strategy based on a three-dimensional convolutional network, a bidirectional long short-term memory-bidirectional gated loop-converter network, and a feature analysis network. The feature-level fusion strategy includes: The pre-processed RGB video data is read through a three-dimensional convolutional network, and after feature extraction, a video feature vector is extracted; Read the preprocessed inertial sensing data through a bidirectional long short-term memory - bidirectional gated recurrent unit - Transformer network. After feature extraction, a sensing feature vector is extracted. Concatenate the video feature vector and the sensing feature vector into a fused feature vector through a feature analysis network, and generate a personnel intention recognition result based on the fused feature vector.
[0013] In an optional embodiment, the feature analysis network includes at least two fully connected layers, and introduces non-linearity through the ReLU activation function. The feature analysis network is used to implement the following feature fusion algorithm:
[0014]
[0015] where \(Z\) is the fused feature, \([;]\) represents the concatenation operation, \(W\) 1 and \(W\) 2 respectively represent the weight matrices of two fully connected layers, \(b\) 1 and \(b\) 2 are bias vectors, is the video feature vector, is the sensing feature vector, is the personnel intention recognition result.
[0016] In an optional embodiment, the first personnel intention recognition sub-model includes: a C3D core module and an output module; The C3D core module is used to capture the spatio-temporal features characterizing the operator's activities in the preprocessed RGB video information by performing 3D convolution and 3D pooling operations; the C3D core module includes five 3D sub-modules, and each 3D sub-module includes one or two 3D convolutional layers and one 3D max pooling layer. The output module is used to map the spatio-temporal features to the classification space by deploying fully connected layers, and generate a score distribution covering the overall personnel action categories; the output module includes two fully connected layers and one dropout layer.
[0017] In an optional embodiment, the second personnel intention recognition sub-model includes a BiLSTM-BiGRU module, a Transformer module, and an output module; The BiLSTM-BiGRU module is used to capture the temporal features characterizing the personnel's behavior in the preprocessed inertial sensing data; the BiLSTM-BiGRU module is provided with a bidirectional long short-term memory network and a bidirectional gated recurrent unit stacked together. The Transformer module is provided with two stacked Transformer encoders for capturing the long-term dependencies of the temporal features and enhancing the global features to generate enhanced features. A fully connected layer is deployed in the output module to map the enhanced features generated by the Transformer module to the classification space and generate scores covering the overall personnel action categories.
[0018] In a second aspect, the embodiment of the present application further provides a multi-modal fusion personnel intention recognition system, including: A multimodal data acquisition module is used to collect RGB video data of operators through visual sensors and inertial sensor data through wearable inertial sensors in human-machine collaboration scenarios as operator intention data; the inertial sensor data includes but is not limited to: acceleration, angular velocity, joint quaternion and center of mass displacement; The multimodal data preprocessing module is used to select the operator's intention data and generate an action sample training set and an action sample test set after preprocessing; A personnel intention recognition model building module is used to build a personnel intention recognition model that integrates RGB video and inertial sensing modalities based on a decision-level fusion strategy or a feature-level fusion strategy, and train the personnel intention recognition model using an action sample training set; The personnel intention recognition module is used to input the action sample test set into the personnel intention recognition model to perform personnel intention recognition and obtain the personnel intention recognition result.
[0019] In a third aspect, an embodiment of the present application further provides an electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the steps of the multimodal fusion personnel intention recognition method as described in any one of the above items are implemented.
[0020] In a fourth aspect, an embodiment of the present application further provides a storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of the method for identifying personnel intentions by multimodal fusion as described in any one of the above items are implemented.
[0021] It can be seen from the above technical solutions that the present invention has the following advantages: In the multi-modal fusion personnel intention recognition method provided in this application, two fusion strategies are given: decision-level fusion and feature-level fusion. By integrating the recognition advantages of the two modalities, the limitations of single modality in personnel intention recognition (personnel image recognition based on RGB video or personnel intention recognition based on inertial sensing) can be overcome, and accurate recognition of operator intentions in human-machine manufacturing collaboration scenarios can be achieved.
[0022] This application integrates the probabilistic decision results of the two modal branch models from the perspective of decision-level fusion through RGB video and inertial sensing decision-level fusion to improve the accuracy of personnel intention recognition. RGB video and inertial sensing modalities are carried out along two independent paths, involving data acquisition, data preprocessing, feature extraction, and decision / classification recognition. Finally, the decision results of the two modalities are fused together using probabilistic decision fusion technology to generate more accurate recognition results for the operator's intention.
[0023] This application can also fuse the feature vectors of the two modal branch models through RGB video and inertial sensing feature level fusion to improve the accuracy of personnel intention recognition. The data collection, preprocessing and feature extraction of the two branches, as well as their feature fusion and feature analysis (decision recognition), constitute a complete feature fusion model to generate more accurate recognition results for operator intention.
[0024] This application uses a combination of visual sensors and wearable inertial sensors to collect RGB video data and inertial sensing data of operators, respectively, as key information sources for intent recognition. In the data processing stage, the method cleverly selects and preprocesses these data to generate action sample sets for training and testing. Subsequently, based on decision-level or feature-level fusion strategies, a personnel intention recognition model that integrates RGB video and inertial sensing modalities is constructed. This model not only makes full use of the complementary advantages of the two modal data, but also accurately identifies the operator's intention through advanced deep learning algorithms such as three-dimensional convolutional networks, bidirectional long short-term memory-bidirectional gated loop-converter networks, etc. This application significantly improves the accuracy and robustness of intent recognition, and provides strong support for the intelligence and efficiency of human-machine collaboration. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] In order to more clearly illustrate the technical solution of the present invention, the accompanying drawings required for use in the description will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0026] Figure 1 A flowchart of the multimodal fusion method for identifying human intention provided in this application.
[0027] Figure 2 A flowchart of the decision-level fusion strategy provided for this application.
[0028] Figure 3 A flowchart of the feature-level fusion strategy provided in this application.
[0029] Figure 4A flowchart of another multimodal fusion method for identifying human intention provided in this application.
[0030] Figure 5 A schematic diagram of the structure of the first person intention recognition sub-model provided in this application.
[0031] Figure 6 A schematic diagram of the structure of the second person intention recognition sub-model provided in this application.
[0032] Figure 7 A schematic diagram of the structure of the BiLSTM-BiGRU module provided in this application.
[0033] Figure 8 A schematic diagram of the structure of the human intention recognition model constructed based on the feature fusion strategy provided in this application.
[0034] Fig. 9 A comparison chart of the recognition results of different models (single-modal branch model and fusion model) provided in this application on the test data set.
[0035] Fig.10 A schematic diagram of the structure of the multimodal fusion personnel intention recognition system provided in this application.
[0036] Fig.11 A schematic diagram of the structure of the electronic device provided in this application. DETAILED DESCRIPTION
[0037] In the specific steps of the multimodal fusion personnel intention recognition method described in detail below, various embodiments of the present disclosure will be described in more detail. The present disclosure may have various embodiments, and adjustments and changes may be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but the present disclosure should be understood to cover all adjustments, equivalents and / or alternatives that fall within the spirit and scope of the various embodiments of the present disclosure.
[0038] Hereinafter, the terms "include" or "may include" used in various embodiments of the present disclosure indicate the presence of disclosed functions, operations, or elements, and do not limit the addition of one or more functions, operations, or elements. In addition, as used in various embodiments of the present disclosure, the terms "include", "have", and their cognates are intended only to indicate specific features, numbers, steps, operations, elements, components, or a combination of the foregoing, and should not be understood as first excluding the presence of one or more other features, numbers, steps, operations, elements, components, or a combination of the foregoing or the possibility of adding one or more features, numbers, steps, operations, elements, components, or a combination of the foregoing.
[0039] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0040] See also Figure 1 The figure is a flowchart of a method for multimodal fusion personnel intention recognition in a specific embodiment, the method comprising: S101: In a human-machine collaboration scenario, the operator's RGB video data is collected through a visual sensor, and inertial sensing data is collected through a wearable inertial sensor as the operator's intention data; the inertial sensing data includes but is not limited to: acceleration, angular velocity, joint quaternion and center of mass displacement.
[0041] In this step, inertial sensing data refers to the operator posture information that can be captured by inertial sensors in human-machine manufacturing collaboration scenarios, including acceleration, angular velocity, skeleton quaternion / Euler angle, joint quaternion / Euler angle, joint vector angle and center of mass displacement.
[0042] S102: Select the operator's intention data, and generate an action sample training set and an action sample test set after preprocessing.
[0043] In a specific implementation, the purpose of this step is to preprocess the operator's intention data, including: sampling, image scaling, cropping and normalization of the original RGB video data to ensure that the input RGB data meets the training requirements of the model; normalization and data division of the original inertial sensor data to ensure that the input inertial sensor data meets the training requirements of the model; since the duration of each operator's operation action is different, the data length of the RGB video and inertial sensor data is uniformly intercepted to ensure that the data length of all actions input to the model is consistent.
[0044] Specifically, for the preprocessing of RGB video data, firstly, RGB video data of a preset length is selected, the selected RGB video data is divided into multiple video sequences of equal length, a frame of image is selected from each video sequence, and the image is cropped to a fixed size to generate an image sequence as the preprocessed RGB video data.
[0045] For the preprocessing of inertial sensor data, firstly, the inertial sensor data collected within a preset time period is selected, and the selected inertial sensor data is organized into an action sequence X={x 1 , x 2, ..., x T}, as the preprocessed inertial sensor data; where T represents the cycle length of the action sequence, x t is the behavior state of the personnel at time step t, each x t Contains m basic features, x t =(x t1 , x t2 , …,x tm ).
[0046] Then, the pre-processed RGB video data and the pre-processed inertial sensor data are used as an action sample, thereby forming a plurality of different action samples according to different operators.
[0047] At this time, M action samples are selected as the action sample training set, and N action samples are selected as the action sample test set.
[0048] S103: Based on a decision-level fusion strategy or a feature-level fusion strategy, a personnel intention recognition model integrating RGB video and inertial sensing modalities is constructed, and the personnel intention recognition model is trained using an action sample training set.
[0049] In this step, two fusion strategies can be used according to different needs to build a personnel intention recognition model that integrates RGB video and inertial sensing modalities. The details are as follows: 1. The decision-level fusion strategy refers to the use of probabilistic decision fusion technology to build a decision-fusion personnel intention recognition model, which is used to fuse the personnel intention recognition results of two single-modal branch models. Among them, the two single-modal branch models refer to the personnel intention recognition models constructed for the two single modalities of RGB video and inertial sensing, including modeling the temporal and spatial features of the pre-processed video data in the RGB video branch based on the three-dimensional convolutional network (C3D); modeling the temporal and global features of the pre-processed inertial sensing data in the inertial sensing branch based on the bidirectional long short-term memory-bidirectional gated recurrent-transformer network (BiLSTM-BiGRU-Transformer).
[0050] It should be noted that the probability fusion technology includes majority voting, sum rule, naive-Bayes, weighted average and DS evidence theory. Whichever method is selected, the decision fusion model of RGB video and inertial sensing can be obtained.
[0051] As an example, the specific process of building a personnel intention recognition model using the decision-level fusion strategy is as follows: On the one hand, the temporal and spatial characteristics of the preprocessed RGB video data are modeled based on a three-dimensional convolutional network to construct a first person intention recognition sub-model; the first person intention recognition sub-model is used to generate a first operator intention recognition result based on the preprocessed RGB video data.
[0052] Another method is to model the temporal characteristics and global characteristics of the preprocessed inertial sensor data based on a bidirectional long short-term memory-bidirectional gated loop-converter network, and construct a second person intention recognition sub-model; the second person intention recognition sub-model is used to generate a second operator intention recognition result based on the preprocessed inertial sensor data.
[0053] At the same time, a calculation module is constructed based on the decision fusion algorithm; the calculation module is used to obtain the first operator intention recognition result and the second operator intention recognition result, and calculate the personnel intention recognition result using the decision fusion algorithm.
[0054] Wherein, the decision fusion algorithm includes:
[0055] Where RGB_output[i] is the first operator intention recognition result of the i-th action sample, IS_output[i] is the second operator intention recognition result of the i-th action sample, W RGB is the weight of the first operator's intention recognition result, W IS is the weight of the first operator's intention recognition result, is the person intention recognition result of the i-th action sample.
[0056] It can be seen that the personnel intention recognition model constructed by using the decision-level fusion strategy includes: the first personnel intention recognition sub-model, the second personnel intention recognition sub-model and the calculation module. The first personnel intention recognition sub-model and the second personnel intention recognition sub-model respectively generate a single-dimensional operator intention recognition result, and the calculation module uses the probability decision fusion technology to fuse the recognition results of the two branch models to obtain the final personnel intention recognition result.
[0057] 2. The feature-level fusion strategy refers to splicing the feature vectors output by two single-modal branch networks on the basis of a single-modal branch network, and connecting the feature analysis network to build a feature-fused personnel intention recognition model. Among them, the feature analysis network can be a complex sub-neural network or a simple sub-neural network, such as two fully connected layers. The first fully connected layer is used to accept the deeply fused feature vectors, and then the activation function ReLU is used to introduce nonlinearity. The second fully connected layer generates the final personnel intention recognition result.
[0058] It can be seen that the personnel intention recognition model constructed by the feature-level fusion strategy includes a three-dimensional convolutional network, a bidirectional long short-term memory-bidirectional gated loop-converter network and a feature analysis network.
[0059] As an example, the feature-level fusion strategy includes: First, the pre-processed RGB video data is read through the three-dimensional convolutional network, and after feature extraction, the video feature vector is extracted. At the same time, the pre-processed inertial sensor data is read through the bidirectional long short-term memory-bidirectional gated loop-converter network, and after feature extraction, the sensor feature vector is extracted.
[0060] Then, the video feature vector and the sensor feature vector are spliced into a fused feature vector through the feature analysis network, and the person intention recognition result is generated based on the fused feature vector.
[0061] Wherein, the feature analysis network includes at least two fully connected layers, and nonlinearity is introduced through a ReLU activation function; The feature analysis network is used to implement the following feature fusion algorithm:
[0062]
[0063] In the formula, Z is the fusion feature, [;] represents the concatenation operation, and W 1 and W 2 Represent the weight matrices of the two fully connected layers, b 1 and b 2 is the bias vector, is the video feature vector, is the sensing feature vector, Identify results for human intent.
[0064] S104: Inputting the action sample test set into the personnel intention recognition model to perform personnel intention recognition and obtain a personnel intention recognition result.
[0065] In a specific implementation, an action sample test set or real-time data including preprocessed RGB video data and inertial sensor data is input into the two modal branches of the personnel intention recognition model, and personnel intention recognition is performed through the constructed decision-fusion personnel intention recognition model or feature-fusion personnel intention recognition model to obtain the final personnel intention recognition result.
[0066] In this embodiment, by fusing the RGB video data collected by the visual sensor and the inertial sensing data collected by the wearable inertial sensor, a decision-level or feature-level fusion model is constructed, which effectively improves the accuracy and reliability of operator intention recognition in human-machine collaboration scenarios.
[0067] In an embodiment of the present invention, based on the decision-level fusion strategy disclosed in step S103, a possible embodiment is given below to illustrate its specific implementation scheme in a non-limiting manner.
[0068] like Figure 2As shown in the figure, decision-level fusion involves three independent models, including the first person intent recognition sub-model for intent recognition on the RGB video branch, the second person intent recognition sub-model for intent recognition on the inertial sensing branch, and a computational module for fusing the decisions of the two branches. The RGB video and inertial sensing modes proceed along two independent paths, involving data acquisition, data preprocessing, feature extraction, and decision / classification recognition. Finally, the decision results of the two modes are fused together to generate a more accurate recognition result for person intent. The final result of the decision fusion model depends only on the decision results of the single-modal branch model and is not affected by its internal state.
[0069] Specifically, the decision-level fusion of RGB video and inertial sensing is used to realize personnel intention recognition, including: Based on visual sensors (cameras) and wearable inertial sensors, operators’ intention data in human-machine manufacturing collaboration scenarios is collected, including parts assembly actions, welding actions, etc. Preprocessing the operator intention data to generate feasible preprocessed data; Based on the three-dimensional convolutional network (C3D), a person intention recognition model for RGB video is constructed, that is, the first person intention recognition sub-model; based on the BiLSTM-BiGRU-Transformer, a person intention recognition model for inertial sensing is constructed, that is, the second person intention recognition sub-model.
[0070] Input the preprocessed RGB data into the first personnel intention recognition sub-model to obtain the video branch's understanding of the operator's intention, i.e., the operator's intention recognition result; input the preprocessed inertial sensor data into the second personnel intention recognition sub-model to obtain the inertial sensor branch's understanding of the operator's intention, i.e., the operator's intention recognition result; The probabilistic decision fusion technology is used to integrate the recognition results of the two branch models to obtain the final person intention recognition result.
[0071] In an embodiment of the present invention, based on the feature-level fusion strategy disclosed in step S103, a possible embodiment is given below to illustrate its specific implementation scheme in a non-limiting manner.
[0072] like Figure 3 As shown in Figure 1, feature-level fusion involves an associated fusion model. The final result of the model is not only affected by the fused features, but also by the internal state of the branch model. The data acquisition, preprocessing and feature extraction of the two branches, as well as their feature fusion and feature analysis (decision recognition), constitute a complete feature fusion model to generate more accurate recognition results for operator behavior. The branch model can be fine-tuned based on the gradient feedback information of the entire model, thereby improving the performance of the entire model.
[0073] Specifically, the feature-level fusion of RGB video and inertial sensing is used to realize personnel intention recognition, including: Based on visual sensors (cameras) and wearable inertial sensors, operators’ intention data in human-machine manufacturing collaboration scenarios is collected, including parts assembly actions, welding actions, etc. Preprocessing the operator intention data to generate feasible preprocessed data; A human intention recognition model based on RGB video and inertial sensing feature fusion is constructed based on three-dimensional convolutional network (C3D), BiLSTM-BiGRU-Transformer and feature analysis network; The pre-processed RGB data and inertial sensor data are input into the personnel intention recognition model fused with the RGB video and inertial sensor features to perform personnel intention recognition and obtain a final personnel intention recognition result.
[0074] The process of obtaining the result of personnel intention recognition based on the feature-level fusion model of RGB video and inertial sensing includes: Inputting the preprocessed RGB data into the C3D model to obtain an output vector of the video branch, which is a feature vector of the video branch's understanding of the operator's intention; At the same time, the pre-processed inertial sensing data is input into the BiLSTM-BiGRU- Transformer model to obtain an output vector of the inertial sensing branch, which is a feature vector of the inertial sensing branch's understanding of the operator's intention; The feature vectors of the two branches are concatenated or combined to form a feature vector of deep fusion of the two modalities; The deeply fused feature vector is passed to the feature analysis network to generate the final person intention recognition result.
[0075] Further, as a refinement and extension of the specific implementation of the above embodiment, in order to fully illustrate the specific implementation process in this embodiment, taking the RV reducer assembly as an example, another modal fusion method for identifying personnel intentions is provided, such as Figure 4 As shown, the method comprises the following steps: S201: Determine the assembly action category of the operator according to the RV reducer assembly task.
[0076] Specifically, the RV reducer assembly actions are divided into 13 categories according to the assembly order, namely actions 1-13: install large gasket; lubricate needle bearing; install small gasket; install retaining ring; lubricate main bearing; assemble output shaft and main bearing; install large gasket; assemble support flange and main bearing; install tapered pin; lubricate needle bearing; install limit cover; install small bolt; install large bolt.
[0077] S202: Configure the video sensor and the inertial sensor according to the assembly task and the assembly layout.
[0078] Specifically, the operator's assembly activities are within the field of view of the visual sensor, and the operator's assembly action RGB video is collected. At the same time, 7 wearable inertial sensors are selected to be fixed on the operator's upper limbs to collect the operator's assembly posture information, including acceleration, angular velocity and joint vector angle. The video sensor is sampled at 30 frames / s, and the inertial sensor sampling frequency is 100hz. This embodiment collects the actions of 7 people, each person repeats each action 12 times, 84 groups of data in each category, and a total of 1092 samples.
[0079] S203: Collect RGB video data and inertial sensor data, and construct a motion sample training set and a motion sample test set after preprocessing.
[0080] Specifically, the original RGB video data is sampled and the image is cropped; the original inertial sensor data is normalized and the data is divided; the time length of the RGB video and inertial sensor data is unified, and the first 2s of data are selected for subsequent model training and testing. The data collected by 5 people are selected as the action sample training set, and the data of the other 2 people are selected as the action sample test set; In this step, sampling refers to preprocessing the original RGB video data using a frame skipping strategy. Specifically, the original video is divided into a series of equal sequences. Then, a frame is selected from each sequence. Finally, each original video can be divided into non-overlapping T-frame segments as the input of the network. In this embodiment, each original video is sampled into a 15-frame video sequence, and the image size is uniformly cropped to 112×112. For data partitioning, this embodiment selects data collected from 5 people as the action sample training set, and the data from another 2 people as the action sample test set.
[0081] In addition, since the duration of each person's operation is different, about 6-10 seconds, the time length of the RGB video and inertial sensor data is unified, and the first 2 seconds of data are selected for subsequent model training and testing. Therefore, within the set 2 seconds, 15 frames of RGB video data of each sample are selected by sampling, and each frame image is cropped to 112×112. The inertial sensor data intercepts the first 2 seconds of data, that is, each sample contains 200 data points.
[0082] S204: Build a personnel intention recognition model based on the decision fusion strategy or the feature fusion strategy, and perform model training.
[0083] Specifically, a personnel intention recognition model of RGB video and inertial sensor modality decision fusion / feature fusion is constructed, and the action sample training set data is input into the personnel intention recognition model of decision level fusion / feature level fusion to perform model training.
[0084] In this step, the personnel intention recognition model constructed based on the decision fusion strategy includes a first personnel intention recognition sub-model constructed based on the C3D network and a second personnel intention recognition sub-model constructed based on the BiLSTM-BiGRU- Transformer network.
[0085] See also Figure 5 As shown in FIG. 1 , it is a schematic diagram of the structure of the first personnel intention recognition sub-model constructed by the implementation of the present invention. The model consists of two modules: a C3D core module and an output module. The C3D core module mainly captures the spatiotemporal features representing the operator's activities in the RGB video by performing 3D convolution and 3D pooling operations. The output module maps the extracted spatiotemporal features to the classification space by deploying a fully connected layer, thereby obtaining a score distribution covering the overall personnel action categories.
[0086] The specific parameter configuration of the first personnel intention recognition sub-model is shown in Table 1. The dimension of the input RGB video data is C×L×H×W, where C is the number of RGB image channels, L is the number of frames taken out of each video clip, H is the height of the image, and W is the width of the image. In this embodiment, it is set to 3×15×112×112. The core module consists of five 3D sub-modules, each of which contains one or two 3D convolutional layers and a 3D maximum pooling layer. The output module consists of two fully connected layers and a dropout layer. The output of the last 3D maximum pooling layer is flattened and passed to the first fully connected layer. Based on ReLU activation, the output of the first fully connected layer is passed to a 50% dropout layer, and then connected to the final fully connected layer to generate the recognition result of the operator's behavior. The number of filters of the 3D convolutional layer in the five 3D sub-modules is 64, 128, 256, 512, and 512, respectively. All 3D convolutional filters are 3×3×3 with a stride of 1×1×1. Except for the first layer, all 3D max pooling layers are 2×2×2 with a stride of 2×2×2. The first pooling layer kernel is 1×2×2 with a stride of 1×1×1. The two fully connected layers have 4096 output units.
[0087] Table 1: Specific parameter information table of the first person intention recognition sub-model
[0088] See also Figure 6As shown, it is a schematic diagram of the structure of the second personnel intention recognition sub-model constructed according to an embodiment of the present invention. The model consists of three core modules: a BiLSTM-BiGRU module, a Transformer module, and an output module. In the BiLSTM-BiGRU module, BiLSTM and BiGRU are stacked together to capture the temporal features that characterize personnel behavior in inertial sensing signals. In the Transformer module, Transformer encoders are stacked to capture long-term dependencies and enhance global features. At the same time, in order to reduce the risk of overfitting during model training, a 50% dropout layer is introduced after this module. The output module maps the Transformer-enhanced features to the classification space by deploying a fully connected layer, thereby obtaining a score that covers the overall personnel action category.
[0089] In the BiLSTM-BiGRU module, since the inertial sensor signals that characterize human behavior are typical time series data, BiLSTM and BiGRU are composed of two memory networks with opposite propagation directions, which can capture the forward (historical) and backward (future) information of the sequence data.
[0090] Therefore, in order to fully extract the temporal features of inertial sensor data, the BiLSTM-BiGRU module stacks two layers of BiLSTM and two layers of BiGRU, see Figure 7 , as follows: Assume that the preprocessed action sequence X = {x 1 , x 2 , ..., x T}, where T represents the period length of the sequence. t (tϵ{1, 2, ..., T}) is the behavior state of the person at time step t. t Contains m basic features, including acceleration, angular velocity and joint vector angle, specifically x t =(x t1 , x t2 , …,x tm ). Accordingly, after the BiLSTM-BiGRU model processes the entire sequence X, the overall feature F={f1, f2, ..., fT} is obtained. In the BiLSTM-BiGRU model of this embodiment, the number of layers of BiLSTM and BiGRU are both 2, and the hidden dimension is 128. The first layer of BiLSTM receives the preprocessed inertial sensor data, and the output of the second layer of BiLSTM is the input of the first layer of BiGRU.
[0091] In the Transformer module, the core of the Transformer module is the self-attention mechanism, which can process input sequences in parallel, effectively capture long-term dependencies and global information, and has significant advantages in processing long sequences and large-scale data. Therefore, in order to capture long-term dependencies and enhance global features, two Transformer encoder modules are stacked, and the specific implementation process is as follows: Taking the first layer Transformer encoder as an example, its embedding dimension (d_model) is 128, which corresponds to the output dimension of the second layer BiGRU described above. The temporal features generated by the BiLSTM BiGRU module are passed to the multi-head self-attention mechanism of the first Transformer encoder layer. In this mechanism, these features are divided into four parts, each of which is processed independently by an attention head. Finally, the outputs of the four attention heads are concatenated and linearly transformed to produce enhanced temporal features with the same dimensions as the output of the second layer BiGRU.
[0092] In particular, when implementing the decision fusion strategy, the present invention selects to use weighted average to perform decision-level fusion of RGB video and inertial sensor modalities, that is, the weighted average method is a decision fusion model of the two modalities. For each sample, the result calculation formula of the decision fusion is:
[0093] Where IS_output[i] and RGB_output[i] represent the decision results of the i-th sample of the RGB video and inertial sensing branches, respectively. IS and W RGB are the weights of the two branch results. In the embodiment of the present invention, both weights are 0.5.
[0094] In this step, the personnel intention recognition model constructed based on the feature fusion strategy includes a three-dimensional convolutional network, a bidirectional long short-term memory-bidirectional gated loop-converter network and a feature analysis network.
[0095] The feature analysis network can be a complex sub-neural network or a simple sub-neural network, such as two fully connected layers, see Figure 8 As shown, a personnel intention recognition model constructed based on a feature fusion strategy provided by an embodiment of the present invention is shown. First, the decision results of the RGB video and inertial sensing branches are obtained as feature information. Then the output features of the two branches are concatenated to form a fused feature. Subsequently, the fused feature is passed to the first fully connected layer with a hidden dimension of 128, and the ReLU activation function is applied to introduce nonlinearity. Finally, the second fully connected layer is connected to generate the final personnel recognition result.
[0096] The calculation formula for feature fusion is as follows.
[0097]
[0098]
[0099] Where Z is the fusion feature. [;] indicates the concatenation operation. W 1 and W 2 Represent the weight matrices of the two fully connected layers, b 1 and b 2 is the bias vector.
[0100] In general, in this step, for decision-level fusion, first, the decision recognition results of human actions are obtained through the BiLSTM-BiGRU-Transformer model in the inertial sensing branch and the C3D model in the RGB video branch. Then, the recognition results of the two branches are fused by the weighted average method to achieve the final personnel action recognition. For feature-level fusion, the branch model, the feature fusion of the two branches, and the feature analysis (decision recognition) constitute a feature fusion model integrating RGB video and inertial sensing modalities. Therefore, the personnel intention recognition model of decision fusion has no training parameters, and it is a weighted average calculation module.
[0101] In this step, the training parameters of the two branch sub-models for personnel assembly action recognition and the personnel intention recognition model with feature fusion are shown in Table 2.
[0102] Table 2: Comparison of training parameters of different models for personnel assembly action recognition
[0103] S205: Input the data of the action sample test set into the trained personnel intention recognition model to perform personnel assembly action recognition to obtain a recognition result.
[0104] It can be seen that the present invention compares the proposed personnel intention recognition method of fusion of RGB video and inertial sensor data with the single modal intention recognition method, that is, the recognition results of the decision-level fusion or feature-level fusion of the two modalities are compared with the recognition results of the C3D model of the RGB video branch and the BiLSTM-BiGRU- Transformer model of the inertial sensor branch.
[0105] Fig. 9The recognition results of different models (single-modal branch model and fusion model) in the test data set when the data length is 2s. It can be seen that the fusion model integrating two modalities adopted in the present invention produces better recognition accuracy than the single-modal branch model. When the data length is 2s, the best recognition accuracy of the two fusion models is as high as 94%: decision fusion (95.51%) and feature fusion (94.23%), both of which are higher than the accuracy of the RGB branch model (91.35%) and the inertial sensing branch model (79.49%). Therefore, the RGB video and inertial sensing fusion method for personnel intention recognition proposed in the present invention can overcome the limitations of single modality and significantly improve the robustness of personnel intention recognition. The RGB video modality has the ability to capture the contextual information of the manufacturing scene, and the inertial sensing modality is good at capturing detailed operator postures. The two complement each other and can provide strong support for personnel intention recognition in human-machine manufacturing collaboration scenarios.
[0106] like Fig.10 As shown, the following is an embodiment of the multimodal fusion personnel intention recognition system provided by the embodiment of the present disclosure. The system and the multimodal fusion personnel intention recognition method of the above-mentioned embodiments belong to the same inventive concept. For details not described in detail in the embodiment of the multimodal fusion personnel intention recognition system, please refer to the embodiment of the above-mentioned multimodal fusion personnel intention recognition method.
[0107] A multimodal fusion personnel intention recognition system includes: a multimodal data acquisition module, a multimodal data preprocessing module, a personnel intention recognition model construction module and a personnel intention recognition module.
[0108] The multimodal data acquisition module is used to collect the operator's RGB video data through visual sensors and inertial sensor data through wearable inertial sensors in human-machine collaboration scenarios as the operator's intention data; the inertial sensor data includes but is not limited to: acceleration, angular velocity, joint quaternion and center of mass displacement. The intention data may include part assembly actions, welding actions, etc.
[0109] The multimodal data preprocessing module is used to select the operator's intention data and generate an action sample training set and an action sample test set after preprocessing.
[0110] The personnel intention recognition model building module is used to build a personnel intention recognition model that fuses RGB video and inertial sensing modalities based on a decision-level fusion strategy or a feature-level fusion strategy, and train the personnel intention recognition model using an action sample training set.
[0111] The personnel intention recognition module is used to input the action sample test set into the personnel intention recognition model to perform personnel intention recognition and obtain the personnel intention recognition result.
[0112] The multimodal fusion personnel intention recognition system provided in this embodiment fully utilizes the complementary advantages of visual information and motion information by integrating RGB video data and inertial sensor data. Based on the decision-level or feature-level fusion strategy, an accurate personnel intention recognition model is constructed, which effectively improves the accuracy and real-time performance of operator intention recognition in human-machine collaboration scenarios, and provides strong support for intelligent human-machine interaction.
[0113] Fig.11 A schematic diagram of the hardware structure of an electronic device for implementing various embodiments of the present invention.
[0114] The multimodal fusion personnel intention recognition method provided in the embodiment of the present application can be applied to electronic devices. It can be understood by those skilled in the art that the electronic device structure involved in the embodiment of the present invention does not constitute a limitation on the electronic device, and the electronic device may include more or less components than shown, or combine certain components, or arrange different components. In an embodiment of the present invention, the electronic device includes but is not limited to a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples, and are not intended to limit the implementation of the embodiments of the present application described and / or required herein.
[0115] The electronic device may include a processor, an external memory interface, an internal memory, a universal serial bus (USB) interface, a charging management module, a power management module, a battery, a wireless communication module, an audio module, a speaker, a microphone, a sensor module, buttons, a camera, a display, and a SIM card interface, etc.
[0116] The processor may include one or more processing units, for example, the processor may include a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated into one or more processors.
[0117] The processor can be the nerve center and command center of the electronic device. The controller can generate an operation control signal according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.
[0118] A memory may also be provided in the processor for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. The memory may store instructions or data that the processor has just used or is cyclically used. If the processor needs to use the instruction or data again, it may be directly called from the memory. This avoids repeated access, reduces the waiting time of the processor, and thus improves system efficiency.
[0119] The external memory interface can be used to connect an external memory card, such as a MicroSD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor through the external memory interface to implement data storage functions. For example, files such as music and videos can be saved in the external memory card.
[0120] The internal memory can be used to store computer executable program codes, which include instructions. The processor executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory. The internal memory may include a program storage area and a data storage area. The internal memory may include a high-speed random access memory and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0121] The wireless communication function of an electronic device can be realized through an antenna, a wireless communication module, a modem processor, and a baseband processor.
[0122] The wireless communication module can provide solutions for wireless communications applied to electronic devices, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc.
[0123] The electronic device can implement audio functions through an audio module, speaker, receiver, microphone, headphone jack, application processor, etc.
[0124] The electronic device can implement shooting functions through an ISP, camera, video codec, GPU, display screen, and application processor, etc.
[0125] The electronic device can implement display functions through a GPU, display screen, and application processor, etc.
[0126] The GPU is a microprocessor for image processing, connecting the display screen and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor may include one or more GPUs, which execute program instructions to generate or change display information.
[0127] The display screen is used to display images, videos, etc. The display screen includes a display panel.
[0128] The above-mentioned electronic device realizes the method for recognizing human intention with multi-modal fusion in this application. By combining RGB video data and inertial sensing data and using a decision-level or feature-level fusion strategy to construct a model, the accurate recognition of the operator's intention is achieved. This method has significant advantages in the human-machine collaboration scenario, achieving the beneficial effect of greatly improving the accuracy and efficiency of intention recognition.
[0129] In the storage medium provided in this application, there is a program product capable of realizing the method for recognizing human intention with multi-modal fusion.
[0130] The method for identifying human intention through multimodal fusion includes: in a human-machine collaboration scenario, collecting RGB video data of an operator through a visual sensor and collecting inertial sensing data through a wearable inertial sensor as the intention data of the operator; selecting the intention data of the operator, and generating an action sample training set and an action sample test set after preprocessing; constructing a human intention recognition model that fuses RGB video and inertial sensing modalities based on a decision-level fusion strategy or a feature-level fusion strategy, and training the human intention recognition model by using the action sample training set; inputting the action sample test set into the human intention recognition model to perform human intention recognition, and obtaining a human intention recognition result.
[0131] In some possible implementation manners, the method for identifying human intention through multimodal fusion of the present disclosure may be implemented in the form of a program product, which includes program code. When the program product runs on a terminal device, the program code is used to cause the terminal device to execute the steps according to various exemplary embodiments of the present disclosure described in the above "Exemplary Method" section of this specification.
[0132] The storage medium of the present disclosure may adopt any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0133] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but will be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A multimodal fusion method for identifying personnel intentions, characterized in that: include: In the human-machine collaboration scenario, the operator's RGB video data is collected through the visual sensor, and the inertial sensor data is collected through the wearable inertial sensor as the operator's intention data; the inertial sensor data includes but is not limited to: acceleration, angular velocity, joint quaternion and center of mass displacement; The operator's intention data is selected and, after preprocessing, an action sample training set and an action sample test set are generated; Based on a decision-level fusion strategy or a feature-level fusion strategy, a personnel intention recognition model integrating RGB video and inertial sensing modalities is constructed, and the personnel intention recognition model is trained using an action sample training set; The action sample test set is input into the personnel intention recognition model to perform personnel intention recognition and obtain the personnel intention recognition result.
2. The multimodal fusion personnel intention recognition method according to claim 1 is characterized in that: The operator's intention data is selected and preprocessed to generate an action sample training set and an action sample test set, including: Selecting RGB video data of a preset length, dividing the selected RGB video data into a plurality of video sequences of equal length, selecting a frame of image from each video sequence, and cropping the image into a fixed size to generate an image sequence as preprocessed RGB video data; Select the inertial sensor data collected within a preset time period and organize the selected inertial sensor data into an action sequence X={x1,x 2, ..., x T }, as the preprocessed inertial sensor data; where T represents the cycle length of the action sequence, x t is the behavior state of the personnel at time step t, each x t Contains m basic features, x t =(x t1 , x t2 , …,x tm ); The preprocessed RGB video data and the preprocessed inertial sensor data are taken as an action sample; Select M action samples as the action sample training set, and select N action samples as the action sample test set.
3. The multimodal fusion personnel intention recognition method according to claim 2 is characterized in that: The personnel intention recognition model based on the fusion strategy of the decision level or the feature level is constructed by fusion of RGB video and inertial sensing modality, including: The decision-level fusion strategy is used to build a personnel intention recognition model that integrates RGB video and inertial sensing modalities. The specific implementation process is as follows: Modeling the temporal and spatial features of the preprocessed RGB video data based on a three-dimensional convolutional network to construct a first personnel intention recognition sub-model; the first personnel intention recognition sub-model is used to generate a first operator intention recognition result based on the preprocessed RGB video data; Based on the bidirectional long short-term memory-bidirectional gated loop-converter network, the temporal characteristics and global characteristics of the preprocessed inertial sensor data are modeled to construct a second person intention recognition sub-model; the second person intention recognition sub-model is used to generate a second operator intention recognition result according to the preprocessed inertial sensor data; Constructing a calculation module based on a decision fusion algorithm; the calculation module is used to obtain the first operator intention recognition result and the second operator intention recognition result, and use the decision fusion algorithm to calculate the personnel intention recognition result; The decision fusion algorithm includes: Where, RGB_output[i] is the first operator intention recognition result of the i-th action sample, IS_output[i] is the second operator intention recognition result of the i-th action sample, W RGB is the weight of the first operator's intention recognition result, W IS is the weight of the first operator's intention recognition result, is the person intention recognition result of the i-th action sample.
4. The multimodal fusion personnel intention recognition method according to claim 2 is characterized in that: The personnel intention recognition model based on the fusion strategy of the decision level or the feature level is constructed by fusion of RGB video and inertial sensing modality, including: A human intention recognition model based on RGB video and inertial sensor feature fusion is constructed using a feature-level fusion strategy based on a three-dimensional convolutional network, a bidirectional long short-term memory-bidirectional gated loop-converter network, and a feature analysis network. The feature-level fusion strategy includes: The pre-processed RGB video data is read through a three-dimensional convolutional network, and after feature extraction, a video feature vector is extracted; The pre-processed inertial sensor data is read through a bidirectional long short-term memory-bidirectional gated loop-converter network, and after feature extraction, a sensor feature vector is extracted; The video feature vector and the sensor feature vector are spliced into a fused feature vector through the feature analysis network, and the person intention recognition result is generated based on the fused feature vector.
5. The multimodal fusion personnel intention recognition method according to claim 4 is characterized in that: The feature analysis network includes at least two fully connected layers, and nonlinearity is introduced through a ReLU activation function; The feature analysis network is used to implement the following feature fusion algorithm: Among them, Z is the fusion feature, [;] represents the concatenation operation, W1 and W2 represent the weight matrices of the two fully connected layers, b1 and b2 are bias vectors, is the video feature vector, is the sensing feature vector, Identify results for human intent.
6. The multimodal fusion personnel intention recognition method according to claim 3 is characterized in that: The first personnel intention recognition sub-model includes: a C3D core module and an output module; The C3D core module is used to capture the spatiotemporal features characterizing the operator's activities in the preprocessed RGB video information by performing 3D convolution and 3D pooling operations; the C3D core module includes five 3D sub-modules, each of which includes one or two 3D convolution layers and one 3D maximum pooling layer; The output module is used to map the spatiotemporal features to the classification space by deploying a fully connected layer to generate a score distribution covering the entire personnel action category; the output module includes two fully connected layers and one discard layer.
7. The multimodal fusion personnel intention recognition method according to claim 3 is characterized in that: The second person intention recognition sub-model includes a BiLSTM-BiGRU module, a Transformer module and an output module; The BiLSTM-BiGRU module is used to capture the temporal features that characterize human behavior in the preprocessed inertial sensor data; the BiLSTM-BiGRU module is equipped with a stacked bidirectional long short-term memory network and a bidirectional gated recurrent unit; The Transformer module is provided with two stacked Transformer encoders for capturing the long-term dependencies of the temporal features and enhancing the global features to generate enhanced features. A fully connected layer is deployed in the output module to map the enhanced features generated by the Transformer module to the classification space and generate scores covering the overall personnel action categories.
8. A multi-modal fusion personnel intention recognition system, characterized in that: The system adopts the multimodal fusion personnel intention recognition method as described in any one of claims 1 to 7; The system comprises: A multimodal data acquisition module is used to collect RGB video data of operators through visual sensors and inertial sensor data through wearable inertial sensors in human-machine collaboration scenarios as operator intention data; the inertial sensor data includes but is not limited to: acceleration, angular velocity, joint quaternion and center of mass displacement; The multimodal data preprocessing module is used to select the operator's intention data and generate an action sample training set and an action sample test set after preprocessing; A personnel intention recognition model building module is used to build a personnel intention recognition model that integrates RGB video and inertial sensing modalities based on a decision-level fusion strategy or a feature-level fusion strategy, and train the personnel intention recognition model using an action sample training set; The personnel intention recognition module is used to input the action sample test set into the personnel intention recognition model to perform personnel intention recognition and obtain the personnel intention recognition result.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the multimodal fusion personnel intention recognition method as described in any one of claims 1 to 7 are implemented.
10. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the multimodal fusion person intention recognition method as described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Special personnel emotion recognition method and system based on multi-modal data fusion
CN113887365A
Behavior recognition method and system based on multi-dimensional sensing data and monitoring video multimode heterogeneous fusion
CN114973120A
Human-machine cooperation intention understanding method, system and equipment facing open environment
CN116561702A
Driving intention recognition method based on multi-modal data
CN118470673A
Multi-modal deep learning power generation device anomaly integrated identification method and device
WO2023087525A1