A first-view action recognition method, device and electronic equipment

CN122618686APending Publication Date: 2026-08-21CHINA TELECOM NETWORK SECURITY TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610648087.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-12
Publication Date
2026-08-21

AI Technical Summary

Technical Problem

[0003]然而,现有的动作识别方法大多是基于第三视角视频数据进行开发,应用于第一视角视频数据中表现不佳

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122618686A_ABST
    Figure CN122618686A_ABST
Patent Text Reader

Abstract

The application discloses a first-view action recognition method and device and electronic equipment, and relates to the technical field of image processing. The method comprises the following steps: firstly, performing feature extraction on original video frames to obtain first verb features and first noun features; then, performing local feature alignment on the first verb features and the first noun features to obtain second verb features, and performing gate fusion on the second verb features to obtain target verb features; finally, based on the target verb features and the first noun features, predicting the contact probability of the hand action of a target object and a candidate object, and obtaining an action recognition result. Through the above method, the interaction between noun features and verb features is realized by using a gate mechanism, the problem of semantic association in action recognition can be significantly improved, the hand movement trajectory is predicted, the reliable interactive object is filtered from the physical layer, and irrelevant object objects are excluded, so that the accuracy and robustness of action recognition are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to a first-person perspective action recognition method, apparatus and electronic device. Background Technology

[0002] With technological advancements, smart glasses, action cameras, and other devices are becoming increasingly prevalent in daily life. Consequently, the volume of self-centered first-person perspective video data is rapidly increasing. Unlike traditional third-person perspective video data, first-person perspective video data contains rich information about hand interactions and object manipulation details, which are crucial for motion recognition.

[0003] However, most existing action recognition methods are developed based on third-person view video data, and perform poorly when applied to first-person view video data. Actions in similar scenes depend on subtle differences in hand / object interactions, which existing algorithms struggle to capture accurately, leading to difficulties in distinguishing fine-grained actions. Furthermore, in first-person view, the moving object is surrounded by a large number of irrelevant objects, resulting in significant interference from noun classification.

[0004] Therefore, improving the accuracy and robustness of action recognition in first-person perspective is a pressing technical problem that needs to be solved. Summary of the Invention

[0005] This application provides a first-view action recognition method, device, and electronic device to improve the accuracy and robustness of action recognition in a first-view perspective.

[0006] In a first aspect, this application provides a first-person perspective action recognition method, the method comprising: Feature extraction is performed on the original video frames to obtain the first verb feature and the first noun feature; the first verb feature is used to determine the hand action corresponding to the target object, and the first noun feature is used to determine the candidate object. Local feature alignment is performed on the first verb feature and the first noun feature to obtain the second verb feature, and gated fusion is performed on the second verb feature to obtain the target verb feature; Based on the target verb features and the first noun features, the probability of contact between the target object's hand movements and candidate objects is predicted to obtain the action recognition results.

[0007] By employing the above methods and using local feature alignment and gating mechanisms to process features, we can focus on real interactive objects, significantly improve the classification accuracy of verbs and nouns, and predict hand movement trajectories to determine the contact probability with candidate objects, thereby filtering out the most relevant candidate objects and improving the accuracy and robustness of action recognition.

[0008] In one optional implementation, local feature alignment is performed on the first verb features and the first noun features to obtain the second verb features, including: Based on the first term features, the spatial position of the candidate object is determined; where the spatial position represents the bounding box coordinates of the candidate object in 2D space. The second verb features are obtained by aligning the first verb features with the spatial location using local features.

[0009] By employing the above method, global motion information is decomposed into local details centered on the target object through local feature alignment, reducing the semantic gap between noun features and verb features, thereby enabling better alignment between noun features and verb features.

[0010] In one optional implementation, gating fusion is performed on the first verb features and the second verb features to obtain the target verb features, including: A gating mechanism is used to process the first verb features to obtain the processed first verb features; The processed first verb features and second verb features are then gated and fused to obtain the target verb features.

[0011] By employing the above method, a gating mechanism is used to process verb features, reducing background noise interference and making the action features focus more on the actual object being manipulated by the hand, thus significantly reducing the impact of noise.

[0012] In one optional implementation, based on the target verb features and the first noun features, the probability of contact between the target object's hand and the candidate object is predicted to obtain the action recognition result, including: Based on the target verb features, the verb probability distribution corresponding to the target verb features is obtained; wherein, the verb probability distribution is used to characterize the probability of the target object's hand performing different actions; Based on the verb probability distribution, determine the target action corresponding to the target object's hand; Based on the target action, the hand movement trajectory of the target object is predicted to obtain the hand movement trajectory prediction result; Action recognition results are obtained based on the hand motion trajectory prediction results and the contact probability of candidate objects.

[0013] The above method uses LSTM to predict the trajectory of hand movements, thereby predicting the probability of contact between the hand and each candidate object, and determining the object most likely to be contacted. This improves the accuracy of word classification and action recognition.

[0014] In one optional implementation, based on the target action, the hand movement trajectory of the target object is predicted to obtain the hand movement trajectory prediction result, including: Determine the first centroid position of the target object's hand and the corresponding second centroid position of the candidate object; Based on the first centroid position and the second centroid position, the target position vector is obtained; wherein, the target position vector represents the relative position of the target object's hand and the candidate object. The hand movement trajectory of the target object is predicted based on the target position vector, and the prediction result of the hand movement trajectory is obtained.

[0015] Secondly, this application provides a first-view motion recognition device, the device comprising: The extraction module is used to extract features from the original video frames to obtain the first verb feature and the first noun feature respectively; wherein, the first verb feature is used to determine the hand action corresponding to the target object, and the first noun feature is used to determine the candidate object; The processing module is used to perform local feature alignment on the first verb features and the first noun features to obtain the second verb features, and to perform gated fusion on the second verb features to obtain the target verb features; The recognition module is used to predict the probability of contact between the target object's hand action and the candidate object based on the target verb features and the first noun features, and to obtain the action recognition result.

[0016] In one optional implementation, when performing local feature alignment on the first verb features and the first noun features to obtain the second verb features, the processing module is specifically used for: Based on the first term features, the spatial position of the candidate object is determined; where the spatial position represents the bounding box coordinates of the candidate object in 2D space. The second verb features are obtained by aligning the first verb features with the spatial location using local features.

[0017] In one optional implementation, when gating and fusing the first verb features and the second verb features to obtain the target verb features, the processing module is specifically used for: A gating mechanism is used to process the first verb features to obtain the processed first verb features; The processed first verb features and second verb features are then gated and fused to obtain the target verb features.

[0018] In an optional implementation, when predicting the contact probability between the target object's hand and the candidate object based on the target verb features and the first noun features to obtain the action recognition result, the recognition module is specifically used for: Based on the target verb features, the verb probability distribution corresponding to the target verb features is obtained; wherein, the verb probability distribution is used to characterize the probability of the target object's hand performing different actions; Based on the verb probability distribution, determine the target action corresponding to the target object's hand; Based on the target action, the hand movement trajectory of the target object is predicted to obtain the hand movement trajectory prediction result; Based on the hand motion trajectory prediction results, the contact probability between the target object's hand and the candidate object is determined, and the action recognition result is obtained.

[0019] In one optional implementation, when predicting the hand movement trajectory of a target object based on the target action to obtain the hand movement trajectory prediction result, the recognition module is specifically used for: Determine the first centroid position of the target object's hand and the corresponding second centroid position of the candidate object; Based on the first centroid position and the second centroid position, the target position vector is obtained; wherein, the target position vector represents the relative position of the target object's hand and the candidate object. The hand movement trajectory of the target object is predicted based on the target position vector, and the prediction result of the hand movement trajectory is obtained.

[0020] Thirdly, this application provides an electronic device including a processor and a memory, wherein the memory stores program code that, when executed by the processor, causes the processor to perform the steps of the first-view motion recognition method described in the first aspect.

[0021] Fourthly, this application provides a computer-readable storage medium including program code that, when executed on an electronic device, causes the electronic device to perform the steps of the first-view action recognition method described in the first aspect.

[0022] Fifthly, this application provides a computer program product that, when invoked by a computer, causes the computer to execute the steps of the first-view action recognition method as described in the first aspect.

[0023] Furthermore, other features and advantages of this application will be set forth in the following description and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings: Figure 1 A schematic diagram of a suitable system architecture is provided for an embodiment of this application; Figure 2 A schematic diagram illustrating the implementation process of a first-view action recognition method provided in this application embodiment; Figure 3 A schematic diagram of the structure of a Mask R-CNN network provided in an embodiment of this application; Figure 4 A schematic diagram of the structure of a ResNet-50 backbone network provided in an embodiment of this application; Figure 5 A schematic diagram of the structure of a first-view motion recognition device provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this application. Obviously, the described embodiments are only some embodiments of the technical solutions of this application, and not all embodiments. Based on the embodiments recorded in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the technical solutions of this application.

[0026] It should be noted that in the description of this application, "multiple" is understood as "at least two". "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. A connected to B can represent: A and B directly connected, or A and B connected through C. Furthermore, in the description of this application, terms such as "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or order.

[0027] Furthermore, the data collection, dissemination, and use in the technical solution of this application all comply with the requirements of relevant national laws and regulations.

[0028] The design concept of the embodiments of this application is briefly introduced below: With technological advancements, smart glasses, action cameras, and other devices are becoming increasingly prevalent in daily life. Consequently, the volume of self-centered first-person perspective video data is rapidly increasing. Unlike traditional third-person perspective video data, first-person perspective video data contains rich information about hand interactions and object manipulation details, which are crucial for motion recognition.

[0029] However, most existing action recognition methods are developed based on third-person view video data, and their performance is poor when applied to first-person view video data. Actions in similar scenes rely on subtle differences in hand / object interactions, which existing algorithms struggle to capture accurately, leading to difficulties in distinguishing fine-grained actions. Furthermore, in first-person view, the moving object is surrounded by numerous irrelevant objects, resulting in significant interference from noun classification. Therefore, action recognition accuracy on first-person view video data is insufficient.

[0030] In view of this, this application provides a first-view action recognition method, which includes: first, extracting features from the original video frames to obtain first verb features and first noun features; wherein, the first verb features are used to determine the hand action corresponding to the target object, and the first noun features are used to determine the candidate object; then, performing local feature alignment on the first verb features and the first noun features to obtain second verb features, and performing gated fusion on the second verb features to obtain target verb features; finally, based on the target verb features and the first noun features, predicting the contact probability between the target object's hand action and the candidate object to obtain the action recognition result.

[0031] By using the above method, the interaction between noun features and verb features is realized through a gating mechanism, which can significantly improve the problem of weak semantic correlation in action recognition, and perform hand movement trajectory prediction. It can also screen credible interaction objects from a physical level and exclude irrelevant objects, thereby improving the accuracy and robustness of action recognition.

[0032] It should be noted that the acquisition, storage, use, and processing of data in this application embodiment all comply with the relevant provisions of national laws and regulations.

[0033] See Figure 1 The diagram shown illustrates a system architecture according to an embodiment of this application. This system architecture includes a target terminal 101 and a server 102. The target terminal 101 and the server 102 can interact via a communication network. The communication network can employ wireless communication or wired communication methods.

[0034] For example, the target terminal 101 can access the network and communicate with the server 102 through cellular mobile communication technology, wherein the cellular mobile communication technology includes, for example, 5th generation mobile networks (5G) technology.

[0035] Optionally, the target terminal 101 can access the network and communicate with the server 102 via short-range wireless communication, wherein the short-range wireless communication method includes, for example, Wireless Fidelity (Wi-Fi) technology.

[0036] This application embodiment does not impose any limitation on the number of communication devices involved in the above system architecture. For example, there may be more target terminals, or no target terminals, or other network devices may be included, such as... Figure 1 As shown, only the target terminal 101 and server 102 are described as examples. The following is a brief introduction to each of the above devices and their respective functions.

[0037] The target terminal 101 is a device that can provide voice and / or data connectivity to a user, and may be a device that supports limited and / or wireless connectivity.

[0038] For example, the target terminal 101 includes, but is not limited to, wearable devices such as smart glasses and action cameras, which are capable of acquiring raw video data from a first-person perspective.

[0039] Server 102 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.

[0040] It is worth noting that, in the embodiments of this application, the methods in these embodiments can be executed by an electronic device, which can be the target terminal 101 or the server 102. That is, the method can be executed by the terminal device 101 or the server 102 alone, or by the target terminal device 101 and the server 102 together. Furthermore, the executing entities for each method can be the same or different, and this embodiment of the application does not impose any restrictions on this.

[0041] For example, when the target terminal 101 executes the first-view action recognition method provided in this application alone, the target terminal 101 can directly extract features from the original video frame to obtain the first verb feature and the first noun feature respectively; then, it performs local feature alignment on the first verb feature and the first noun feature to obtain the second verb feature, and performs gated fusion on the second verb feature to obtain the target verb feature; finally, based on the target verb feature and the first noun feature, it predicts the contact probability between the target object's hand action and the candidate object to obtain the action recognition result.

[0042] For example, when the target terminal 101 and the server 102 jointly execute the first-view action recognition method of this application, the target terminal 101 can collect raw video frame data in real time and send the raw video frames to the server 102. The server 102 can extract features from the raw video frames to obtain first verb features and first noun features respectively. Then, the first verb features and first noun features are aligned locally to obtain second verb features, and the second verb features are gated and fused to obtain target verb features. Finally, based on the target verb features and the first noun features, the contact probability between the target object's hand movements and the candidate object is predicted to obtain the action recognition result.

[0043] The first-view action recognition method provided by an exemplary embodiment of this application will now be described with reference to the accompanying drawings.

[0044] See Figure 2 The diagram shown illustrates the implementation flow of a first-view action recognition method provided in this application. The specific implementation flow of this method is as follows: S1: Extract features from the original video frames to obtain the first verb feature and the first noun feature respectively.

[0045] Firstly, in this embodiment, a large amount of raw video data can be obtained from the first-person perspective of the target object through a wearable device. For example, this can be achieved using a motion camera worn on the target object's head or chest, or through smart glasses worn by the target object. This raw video data clearly shows the interaction between the target object's hands and objects or the environment.

[0046] Furthermore, after obtaining the original video data, feature extraction can be performed on the original video frames to obtain the first verb feature and the first noun feature respectively.

[0047] It should be noted that, in the embodiments of this application, the first verb feature is used to determine the hand action of the target object, such as taking, cutting, pouring, pushing, etc., and the first noun feature is used to determine the candidate object, that is, the object in the first viewpoint, such as cup, apple, mobile phone, etc.

[0048] Furthermore, the feature extraction network used in this embodiment consists of two parts. First, a 3D CNN is used to process the original video frames to obtain the first verb features. The acquired original video frames are preprocessed; the input original video frame size can be 224×224, and the original video frames are downsampled to 30 FPS. Then, the R3D backbone network is used to extract features from the original video frames to obtain the first verb features. The R3D backbone network is shown in the table below.

[0049] Table 1

[0050] It is worth noting that if the size of the original input video frame is 8×224×224×3, then the temporal dimension L of the output size is the same as that of the original input video frame, i.e., L=8. The network convolutional residual blocks are shown in parentheses in Table 1. The original video frame undergoes one spatial downsampling using a 1×2×2 convolutional step on conv1, and three spatiotemporal downsamplings using a 2×2×2 convolutional step on conv31, conv41, and conv51. After passing through conv5_X, the obtained features... Perform global average pooling to obtain the first verb feature. The formula for obtaining the first verb feature is as follows:

[0051] Second, Mask R-CNN is used to extract features from the original video frames to obtain the first noun features. See also Figure 3 As shown, Mask R-CNN uses ResNet-50 as the backbone network for feature extraction. This network contains 50 convolutional layers and introduces residual connections, including multiple residual modules. Each residual module contains multiple convolutional layers and identity mappings. Through direct connections across layers, it can solve the problems of vanishing and exploding gradients. (See also...) Figure 4 As shown, ResNet-50 includes two basic blocks (Conv Block (CB)). Figure 4 The left side and the Identity Block (IB) Figure 4 (Right side) In this part, the input and output dimensions of CB are different and cannot be continuously connected. It can be used to change the dimensions of the network. The output dimension of IB is the same as the output dimension and can be continuously connected. It can be used to deepen the network.

[0052] By processing the original video frames using the aforementioned 3D CNN and Mask R-CNN, we can obtain the first verb features representing the hand action corresponding to the target object and the first noun features representing the candidate object.

[0053] S2: Perform local feature alignment on the first verb feature and the first noun feature to obtain the second verb feature, and perform gated fusion on the second verb feature to obtain the target verb feature.

[0054] Directly associating detected candidate objects with actions can lead to irrelevant objects interfering with action recognition. For example, in a kitchen scene, a "table" contributes nothing to the "chopping vegetables" action. Therefore, in this embodiment, after obtaining the first verb feature and the first noun feature, it is necessary to strengthen the spatiotemporal correlation between hand actions and candidate objects. This involves locally aligning the first verb feature and the first noun feature in spatial location.

[0055] Because of the significant semantic differences between verb features and noun features, directly aligning the first verb feature and the first noun feature globally would not effectively integrate them. Therefore, this application proposes a local alignment method.

[0056] In this embodiment, local alignment can decompose global motion information into local details centered on the target object. Specifically, after determining the first noun feature, the detection network can also determine the spatial location corresponding to the candidate object. This spatial location can be represented as... This spatial location allows us to identify the bounding box coordinates of the candidate object in 2D space.

[0057] Then, after determining the spatial location corresponding to the first noun feature, bilinear interpolation (ROIAlign) can be used to perform local feature alignment between the spatial locations of the first verb feature and the first noun feature, i.e. The final second verb feature can be represented as: ; in, , , , , , Indicates learnable weights, This represents the second verb feature after local feature alignment.

[0058] Furthermore, since there are inaccurate detection areas throughout the detection area, there is a lot of background noise interference in the extracted features. Therefore, in order to ensure that the extracted verb features only focus on "candidate objects that are actually manipulated by the target object's hand" and ignore irrelevant details in the background, this application embodiment also needs to perform gating mechanism processing on the second verb features after local feature alignment, in order to emphasize the action-related information in the second verb features.

[0059] Specifically, a gating mechanism is first used to process the first verb features to obtain the processed first verb features. Optionally, in this embodiment, the gating mechanism used to process the first verb features to obtain the processed first verb features can be expressed by the following formula:

[0060] in, Indicates learnable parameters, Indicates the gating weight, Indicates the characteristics of the first verb. This indicates the first verb feature after being processed by the gating mechanism.

[0061] Furthermore, after obtaining the processed first verb features, the processed first verb features and the second verb features can be gated and fused to obtain the target verb features. In this embodiment, the target verb features can be obtained using the following formula:

[0062] in, Indicates the characteristics of the target verb. This indicates the characteristics of the first verb after processing by the gating mechanism. This indicates element-wise multiplication. It indicates the characteristics of the second verb.

[0063] By introducing a gating mechanism through the above method, verb features and noun features can interact dynamically, and the weight of noun features on verb classification can be calculated to enhance the influence of related actions and suppress irrelevant noise.

[0064] S3: Based on the target verb features and the first noun features, predict the probability of contact between the target object's hand movements and the candidate object to obtain the action recognition result.

[0065] In this embodiment of the application, after obtaining the target verb features processed by the gating mechanism, the obtained target verb features can be input into the fully connected layer, and the verb probability distribution corresponding to the target verb features can be obtained through the softmax function.

[0066] It should be noted that these verb probabilities are used to represent the probability of the target object's hand performing different actions. For example, based on the obtained target verb features, action classification can be performed to determine that the target object's hand action could be "take" (30% probability), "cut" (60% probability), or "push" (10% probability). Based on the above probability distribution, the probability of the target object's hand action being "cut" is the highest. Therefore, the target action corresponding to the target object's hand can be determined as "cut," and the corresponding action label can be output to identify the target action currently corresponding to the target object's hand.

[0067] Once the target action corresponding to the target object's hand is determined, the hand region detection uses a Long Short-Term Memory (LSTM) network to predict the target object's hand movement trajectory. This can predict the temporal movement trend of the target object's hand, thereby determining the current movement trajectory of the target object's hand action. Furthermore, it can determine the contact probability between the target object and candidate objects. In this way, it can be determined which candidate object the target object's hand action actually acts on, that is, the final candidate object that is operated on.

[0068] In one alternative implementation, by predicting the hand movement trajectory of the target object, it is possible to determine the target object that the target object's hand actually contacts, thereby filtering out irrelevant candidate objects. Specifically, in this embodiment, to reduce computational complexity, the centroid positions of the target object's hand and the candidate object's centroid positions are used as inputs to the LSTM to predict the hand movement trajectory.

[0069] In the aforementioned extraction of the first verb and first noun features, the detection network can directly extract the first centroid position corresponding to the hand of the target object and the second centroid position corresponding to the candidate object. Optionally, the first centroid position can be obtained by calculating the mean of the pixel coordinates of the hand, as shown in the following formula:

[0070] in( () represents the coordinates of the i-th hand pixel in the hand mask, and N is the total pixel value of the hand mask.

[0071] Using the above formula, the first centroid position corresponding to the hand of the target object in N consecutive video frames can be determined. For example, the first centroid position corresponding to the hand of the target object in 8 consecutive video frames can be determined.

[0072] Then, the target position vector is obtained by subtracting the first centroid position from the second centroid position corresponding to each candidate object. That is, the target position vector = first centroid position - second centroid position. In this way, the target position vector can represent the relative position of the target object's hand and the candidate object. The obtained multiple target position vectors are input into the LSTM to predict where the target object's hand will move to in the next M frames (e.g., 5 frames), thereby determining the predicted trajectory of the target object's hand.

[0073] Furthermore, after obtaining the predicted trajectory of the target object's hand, the average distance from the target object's hand to the candidate object in the next M frames can be used as the input to a multi-layer perceptron (MLP) to determine the contact probability between the target object and the candidate object. In this embodiment, the average distance from the target object's hand to the candidate object can be, for example, taken as the L2 norm, and calculated using the following formula:

[0074] Where M represents the number of future video frames predicted. This indicates the predicted results of hand movements.

[0075] The output layer of the MLP has two nodes, used to output the non-contact probability and the contact probability of the target object's hand with a candidate object, respectively. The output layer also has a softmax function for activation, which normalizes the sum of the non-contact and contact probabilities to 1. For example, if the non-contact probability is 0.1 and the contact probability is 0.9, then the contact probability of the target object's hand with the candidate object is 90%. Simultaneously, the aforementioned noun feature extraction determines the category of the candidate object (e.g., cup), and the target object's hand action (e.g., take) is determined based on the target verb features. Therefore, based on the predicted probability, it is determined that the target object's hand will contact the candidate object, thus determining the final action recognition result. That is, the final action recognition result is verb + noun, for example, take + cup = take cup. The action recognition result can not only determine the target object's hand action but also the object the hand contactes. By predicting the hand's movement trend and the contact probability with candidate objects, the most relevant object can be selected, thereby improving the accuracy of noun classification. This improves the accuracy of action recognition.

[0076] Furthermore, based on the same technical concept, embodiments of this application provide a first-view motion recognition device, which is used to implement the above-described method flow of embodiments of this application. See also... Figure 5As shown, the device includes: an extraction module 501, a processing module 502, and an identification module 503, wherein, The extraction module 501 is used to extract features from the original video frames to obtain first verb features and first noun features respectively; wherein, the first verb features are used to determine the hand action corresponding to the target object, and the first noun features are used to determine the candidate object; The processing module 502 is used to perform local feature alignment on the first verb feature and the first noun feature to obtain the second verb feature, and to perform gated fusion on the second verb feature to obtain the target verb feature; The recognition module 503 is used to predict the probability of contact between the hand action of the target object and the candidate object based on the target verb features and the first noun features, and to obtain the action recognition result.

[0077] In one optional implementation, when performing local feature alignment on the first verb features and the first noun features to obtain the second verb features, the processing module 502 is specifically used for: Based on the first term features, the spatial position of the candidate object is determined; where the spatial position represents the bounding box coordinates of the candidate object in 2D space. The second verb features are obtained by aligning the first verb features with the spatial location using local features.

[0078] In one optional implementation, when gating and fusing the first verb features and the second verb features to obtain the target verb features, the processing module 502 is specifically used for: A gating mechanism is used to process the first verb features to obtain the processed first verb features; The processed first verb features and second verb features are then gated and fused to obtain the target verb features.

[0079] In an optional implementation, when predicting the contact probability between the target object's hand and the candidate object based on the target verb features and the first noun features to obtain the action recognition result, the recognition module 503 is specifically used for: Based on the target verb features, the verb probability distribution corresponding to the target verb features is obtained; wherein, the verb probability distribution is used to characterize the probability of the target object's hand performing different actions; Based on the verb probability distribution, determine the target action corresponding to the target object's hand; Based on the target action, the hand movement trajectory of the target object is predicted to obtain the hand movement trajectory prediction result; Based on the hand motion trajectory prediction results, the contact probability between the target object's hand and the candidate object is determined, and the action recognition result is obtained.

[0080] In one optional implementation, when predicting the hand movement trajectory of a target object based on the target action to obtain the hand movement trajectory prediction result, the recognition module 503 is specifically used for: Determine the first centroid position of the target object's hand and the corresponding second centroid position of the candidate object; Based on the first centroid position and the second centroid position, the target position vector is obtained; wherein, the target position vector represents the relative position of the target object's hand and the candidate object. The hand movement trajectory of the target object is predicted based on the target position vector, and the prediction result of the hand movement trajectory is obtained.

[0081] Based on the same technical concept, embodiments of this application also provide an electronic device that can implement the first-view action recognition method flow provided in the above embodiments of this application. In one embodiment, the electronic device may be a server, a terminal device, or other electronic devices. See also... Figure 6 As shown, the electronic device may include: At least one processor 601 and a memory 602 connected to at least one processor 601. In this embodiment, the specific connection medium between the processor 601 and the memory 602 is not limited. Figure 6 The example shown is the connection between processor 601 and memory 602 via bus 600. Bus 600 is... Figure 6 The connections between other components are indicated by thick lines and are for illustrative purposes only, not as limiting information. The 600 bus can be divided into address bus, data bus, control bus, etc., for ease of representation. Figure 6 The term is represented by a single thick line, but this does not imply that there is only one bus or one type of bus. Alternatively, the processor 601 can also be called a controller; there is no restriction on the name.

[0082] In this embodiment, memory 602 stores instructions executable by at least one processor 601. By executing the instructions stored in memory 602, at least one processor 601 can execute a first-person perspective action recognition method described above. Processor 601 can implement... Figure 5 The functions of each module in the device shown.

[0083] The processor 601 is the control center of the device. It can connect to various parts of the control device through various interfaces and lines. By running or executing instructions stored in memory 602 and calling data stored in memory 602, the processor can perform various functions and process data, thereby monitoring the device as a whole.

[0084] In one possible design, processor 601 may include one or more processing units. Processor 601 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into processor 601. In some embodiments, processor 601 and memory 602 may be implemented on the same chip; in some embodiments, they may also be implemented on separate chips.

[0085] The processor 601 can be a general-purpose processor, such as a CPU, digital signal processor, application-specific integrated circuit, field-programmable gate array or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the first-view action recognition method disclosed in the embodiments of this application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.

[0086] Memory 602, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory 602 may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic memory, magnetic disk, optical disk, etc. Memory 602 can be any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. In the embodiments of this application, memory 602 can also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.

[0087] By designing and programming the processor 601, the code corresponding to the first-view action recognition method described in the foregoing embodiments can be embedded into the chip, thereby enabling the chip to execute the code during operation. Figure 2The illustrated embodiment presents the steps of a first-view action recognition method. How to design and program the processor 601 is a technique well-known to those skilled in the art and will not be described further here.

[0088] Based on the same inventive concept, embodiments of this application also provide a storage medium storing computer instructions that, when executed on a computer, cause the computer to perform a first-person action recognition method as described above.

[0089] In some possible implementations, this application also provides that various aspects of a first-view motion recognition method can also be implemented in the form of a program product, which includes program code that, when the program product is run on a device, causes the control device to perform the steps in a first-view motion recognition method according to various exemplary embodiments of this application as described above.

[0090] It should be noted that although several units or sub-units of the device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.

[0091] Furthermore, although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0092] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0093] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable first-view motion recognition device to produce a server, such that the instructions, which execute via the processor of the computer or other programmable first-view motion recognition device, generate instructions for implementing the process... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0094] Program code for performing the operations of this application can be written using any combination of one or more programming languages, including object-oriented programming languages ​​such as Java and C++, as well as conventional procedural programming languages ​​such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0095] These computer program instructions can also be loaded onto a computer or other programmable first-person perspective motion recognition device, causing a series of operational steps to be executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable device for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0096] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.

Claims

1. A first-person perspective action recognition method, characterized in that, The method includes: Feature extraction is performed on the original video frames to obtain first verb features and first noun features respectively; wherein, the first verb features are used to determine the hand action corresponding to the target object, and the first noun features are used to determine the candidate object; Local feature alignment is performed on the first verb feature and the first noun feature to obtain the second verb feature, and gated fusion is performed on the second verb feature to obtain the target verb feature; Based on the target verb features and the first noun features, the probability of contact between the target object's hand movements and the candidate object is predicted, and the action recognition result is obtained.

2. The method as described in claim 1, characterized in that, The step of performing local feature alignment on the first verb features and the first noun features to obtain the second verb features includes: Based on the first noun feature, the spatial position corresponding to the candidate object is determined; wherein, the spatial position represents the rectangular coordinates of the candidate object in 2D space; The second verb feature is obtained by aligning the first verb feature with the spatial location using local features.

3. The method as described in claim 1, characterized in that, The step of gating and fusing the first verb features and the second verb features to obtain the target verb features includes: The first verb features are processed using a gating mechanism to obtain the processed first verb features; The processed first verb features and the second verb features are then gated and fused to obtain the target verb features.

4. The method as described in claim 1, characterized in that, The step of predicting the contact probability between the target object's hand and the candidate object based on the target verb features and the first noun features, and obtaining the action recognition result, includes: Based on the target verb features, a verb probability distribution corresponding to the target verb features is obtained; wherein, the verb probability distribution is used to characterize the probability that the target object's hand will perform different actions; Based on the verb probability distribution, the target action corresponding to the hand of the target object is determined; Based on the target action, the hand movement trajectory of the target object is predicted to obtain the hand movement trajectory prediction result; Based on the hand motion trajectory prediction results, the contact probability between the target object's hand and the candidate object is determined, and the action recognition result is obtained.

5. The method as described in claim 4, characterized in that, The step of predicting the hand movement trajectory of the target object based on the target action to obtain the hand movement trajectory prediction result includes: Determine the first centroid position of the hand of the target object and the second centroid position corresponding to the candidate object; Based on the first centroid position and the second centroid position, a target position vector is obtained; wherein, the target position vector represents the relative position of the hand of the target object and the candidate object. Based on the target position vector, the hand movement trajectory of the target object is predicted to obtain the hand movement trajectory prediction result.

6. A first-person perspective motion recognition device, characterized in that, The device includes: The extraction module is used to extract features from the original video frames to obtain first verb features and first noun features respectively; wherein, the first verb features are used to determine the hand action corresponding to the target object, and the first noun features are used to determine the candidate object; The processing module is used to perform local feature alignment on the first verb feature and the first noun feature to obtain the second verb feature, and to perform gated fusion on the second verb feature to obtain the target verb feature; The recognition module is used to predict the contact probability between the hand action of the target object and the candidate object based on the target verb features and the first noun features, and to obtain the action recognition result.

7. The apparatus as claimed in claim 6, characterized in that, When performing local feature alignment on the first verb feature and the first noun feature to obtain the second verb feature, the processing module is specifically used for: Based on the first noun feature, the spatial position corresponding to the candidate object is determined; wherein, the spatial position represents the rectangular coordinates of the candidate object in 2D space; The second verb feature is obtained by aligning local features based on the first verb feature and the spatial location.

8. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1-5.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-5.

10. A computer program product, characterized in that, The computer program product includes: computer program code, which, when run on a computer, causes the computer to perform the method described in any one of claims 1-5.