Behavior detection method and device, computer device and storage medium

CN115966020BActive Publication Date: 2026-09-15BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211710538.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-29
Publication Date
2026-09-15
Estimated Expiration
2042-12-29

AI Technical Summary

Technical Problem

[0003]一般的,可以利用训练得到的神经网络,对三维场景中虚拟对象的交互动作进行识别,但是,基于深度学习的动作识别技术依赖大量的标注数据,而对于刚刚兴起的三维场景数字内容,用于训练角色交互动作识别的深度学习网络的数据量不足,且需要消耗较多的时间成本和人力成本进行数据标注,使得神经网络的训练效率和精度较低,进一步降低了交互动作识别的效率和精度

Benefits of technology

[0038] The behavior detection method, apparatus, computer device, and storage medium provided in this disclosure are as follows: The method acquires three-dimensional keypoint sequences corresponding to multiple virtual objects in a target three-dimensional space; based on the three-dimensional keypoint sequences of the virtual objects, it generates two-dimensional keypoint sequences of the virtual objects from multiple viewpoints; for each viewpoint, based on the two-dimensional keypoint sequences of the multiple virtual objects from that viewpoint and a trained behavior detection network, it determines the interaction behavior features of the multiple virtual objects from that viewpoint; based on the interaction behavior features of the multiple virtual objects from each viewpoint and the behavior detection network, it determines the interaction behavior detection results of the multiple virtual objects. In this disclosure, the input to the behavior detection network is a two-dimensional keypoint sequence, therefore, the behavior detection network can be trained from the keypoint sequences of training samples in a two-dimensional scene, alleviating the problems of high annotation costs and limited samples caused by directly using three-dimensional spatial data to train a neural network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115966020B_ABST
    Figure CN115966020B_ABST
Patent Text Reader

Abstract

The present disclosure provides a behavior detection method and device, computer equipment and a storage medium, wherein the method comprises: acquiring a plurality of three-dimensional key point sequences corresponding to a plurality of virtual objects in a target three-dimensional space; the three-dimensional key point sequence comprises three-dimensional key point information collected at different time points; based on the three-dimensional key point sequence of the virtual object, a two-dimensional key point sequence of the virtual object under multiple perspectives is generated; for each perspective, based on the two-dimensional key point sequence of the plurality of virtual objects under the perspective and a trained behavior detection network, the interactive behavior features of the plurality of virtual objects under the perspective are determined; based on the interactive behavior features of the plurality of virtual objects under each perspective and the behavior detection network, an interactive behavior detection result of the plurality of virtual objects is determined.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of deep learning technology, and more specifically, to a behavior detection method, apparatus, computer device, and storage medium. Background Technology

[0002] With the continuous development of multimedia technology, digital content space is gradually evolving towards three-dimensional environments. For example, applications of Virtual Reality (VR) and Augmented Reality (AR) technologies in social and entertainment formats are increasingly gaining public attention. Therefore, there is a need to recognize the interactive actions between characters in user-controlled, highly free, and open three-dimensional scenes.

[0003] Generally, trained neural networks can be used to recognize the interactive actions of virtual objects in a 3D scene. However, deep learning-based action recognition technology relies on a large amount of labeled data. For the emerging 3D scene digital content, the amount of data used to train deep learning networks for character interaction action recognition is insufficient, and data labeling requires a lot of time and manpower, resulting in low training efficiency and accuracy of the neural network, which further reduces the efficiency and accuracy of interactive action recognition. Summary of the Invention

[0004] This disclosure provides at least one behavior detection method, apparatus, computer device, and storage medium.

[0005] In a first aspect, embodiments of this disclosure provide a behavior detection method, including:

[0006] Obtain a sequence of 3D key points corresponding to multiple virtual objects in the target 3D space; the sequence of 3D key points includes 3D key point information collected at different time points;

[0007] Based on the three-dimensional key point sequence of the virtual object, generate a two-dimensional key point sequence of the virtual object from multiple perspectives;

[0008] For each viewpoint, based on the two-dimensional keypoint sequence of the multiple virtual objects under the viewpoint and the trained behavior detection network, the interaction behavior features of the multiple virtual objects under the viewpoint are determined;

[0009] Based on the interaction behavior features of the multiple virtual objects under each of the aforementioned perspectives and the behavior detection network, the interaction behavior detection results of the multiple virtual objects are determined.

[0010] In one optional implementation, determining the interaction behavior detection results of the plurality of virtual objects based on the interaction behavior features of the plurality of virtual objects under each of the aforementioned viewpoints and the behavior detection network includes:

[0011] For each viewpoint, based on the interaction behavior features of the multiple virtual objects under that viewpoint and the behavior detection network, intermediate detection results are generated under that viewpoint;

[0012] Based on the intermediate detection results corresponding to each of the aforementioned perspectives, the interaction behavior detection results of the multiple virtual objects are determined.

[0013] In one optional implementation, determining the interaction behavior detection results of the plurality of virtual objects based on the intermediate detection results corresponding to each of the aforementioned viewpoints includes:

[0014] Based on the intermediate detection results corresponding to each of the aforementioned viewpoints and the weights corresponding to each of the aforementioned viewpoints obtained by training the behavior detection network, the interaction behavior detection results of the multiple virtual objects are determined.

[0015] In one optional implementation, determining the interaction behavior detection results of the plurality of virtual objects based on the interaction behavior features of the plurality of virtual objects under each of the aforementioned viewpoints and the behavior detection network includes:

[0016] The interactive behavior features of the multiple virtual objects under each of the various perspectives are fused to generate fused feature data;

[0017] Based on the fused feature data and the behavior detection network, the interaction behavior detection results of the multiple virtual objects are determined.

[0018] In one optional implementation, obtaining the sequence of 3D key points corresponding to multiple virtual objects in the target 3D space includes:

[0019] Obtain the position information of the plurality of virtual objects in the target three-dimensional space;

[0020] Based on the location information corresponding to the plurality of virtual objects, if it is determined that the target distance between at least two of the virtual objects is less than a preset distance value, the three-dimensional key point sequence corresponding to the at least two virtual objects is obtained.

[0021] In one optional implementation, the step of training the behavior detection network includes:

[0022] Acquire multiple sample video data carrying labeled tags; the labeled tags are used to indicate the interaction behavior between different sample objects in the sample video data;

[0023] Determine the two-dimensional keypoint sequence of the sample objects included in each of the sample video data;

[0024] The two-dimensional keypoint sequence corresponding to each of the sample video data is input into the neural network to be trained to generate the predicted behavior result corresponding to each of the sample video data.

[0025] Based on the predicted behavior result and the labeled label corresponding to each sample video data, the parameters of the neural network to be trained are adjusted until the adjusted neural network meets the training cutoff condition, thus obtaining the behavior detection network.

[0026] In one optional implementation, after obtaining the behavior detection network, the method further includes:

[0027] Obtain three-dimensional sample data of the target three-dimensional space, wherein the three-dimensional sample data includes behavioral annotation labels of multiple sample objects, and a sequence of three-dimensional key points of each sample object that matches the behavioral annotation labels;

[0028] Using the three-dimensional sample data, the parameters of the behavior detection network are adjusted to generate an adjusted behavior detection network.

[0029] In one optional implementation, the method further includes:

[0030] If the interaction behavior detection result indicates that the interaction behavior between the multiple virtual objects matches the preset behavior, a warning message is generated and displayed.

[0031] Secondly, embodiments of this disclosure also provide a behavior detection device, comprising:

[0032] The acquisition module is used to acquire a sequence of three-dimensional key points corresponding to multiple virtual objects in the target three-dimensional space; the sequence of three-dimensional key points includes three-dimensional key point information collected at different time points;

[0033] A generation module is used to generate a two-dimensional keypoint sequence of the virtual object from multiple perspectives based on the three-dimensional keypoint sequence of the virtual object.

[0034] The first determining module is used to determine the interactive behavior features of the multiple virtual objects in the viewpoint based on the two-dimensional keypoint sequence of the multiple virtual objects in the viewpoint and the trained behavior detection network for each viewpoint.

[0035] The second determining module is used to determine the interaction behavior detection results of the multiple virtual objects based on the interaction behavior features of the multiple virtual objects under each of the aforementioned perspectives and the behavior detection network.

[0036] Thirdly, embodiments of this disclosure also provide a computer device, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the computer device is running, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, the steps of the first aspect above, or any possible implementation of the first aspect, are performed.

[0037] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the first aspect or any possible implementation of the first aspect.

[0038] The behavior detection method, apparatus, computer device, and storage medium provided in this disclosure are as follows: The method acquires three-dimensional keypoint sequences corresponding to multiple virtual objects in a target three-dimensional space; based on the three-dimensional keypoint sequences of the virtual objects, it generates two-dimensional keypoint sequences of the virtual objects from multiple viewpoints; for each viewpoint, based on the two-dimensional keypoint sequences of the multiple virtual objects from that viewpoint and a trained behavior detection network, it determines the interaction behavior features of the multiple virtual objects from that viewpoint; based on the interaction behavior features of the multiple virtual objects from each viewpoint and the behavior detection network, it determines the interaction behavior detection results of the multiple virtual objects. In this disclosure, the input to the behavior detection network is a two-dimensional keypoint sequence, therefore, the behavior detection network can be trained from the keypoint sequences of training samples in a two-dimensional scene, alleviating the problems of high annotation costs and limited samples caused by directly using three-dimensional spatial data to train a neural network.

[0039] Meanwhile, the embodiments of this disclosure obtain a three-dimensional keypoint sequence. Since the three-dimensional keypoint sequence includes three-dimensional keypoint information collected at different time points, and the feature distribution between the keypoint information of real objects in real scenes and the keypoint information of virtual objects in virtual scenes is consistent, it can alleviate the accuracy problem caused by data inconsistency after neural network transfer. Furthermore, the method obtains a two-dimensional keypoint sequence of virtual objects from multiple perspectives. Through the two-dimensional keypoint sequence from multiple perspectives, interactive behavior features can be obtained more comprehensively and richly from multiple angles. Then, by utilizing the interactive behavior features of multiple virtual objects from various perspectives and the behavior detection network, the interactive behavior detection results of multiple virtual objects can be determined more accurately, thereby improving the accuracy of interactive behavior detection.

[0040] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description

[0041] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.

[0042] Figure 1 A flowchart of a behavior detection method provided in an embodiment of this disclosure is shown;

[0043] Figure 2 A flowchart of another behavior detection method provided by an embodiment of this disclosure is shown;

[0044] Figure 3 A schematic diagram of the architecture of a behavior detection device provided in an embodiment of this disclosure is shown;

[0045] Figure 4 A schematic diagram of the structure of a computer device provided in an embodiment of this disclosure is shown. Detailed Implementation

[0046] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.

[0047] In related technologies, trained neural networks can be used to identify the interactive actions of virtual objects in a 3D scene. However, deep learning-based action recognition technology relies on a large amount of labeled data. For the emerging 3D scene digital content, the amount of data used to train deep learning networks for character interaction action recognition is insufficient, and it requires a lot of time and manpower to label the data, resulting in low training efficiency and accuracy of the neural network, which further reduces the efficiency and accuracy of interactive action recognition.

[0048] In another approach, a neural network for interactive behavior detection can be trained using a publicly available image dataset from a 2D scene. This neural network can then be transferred to a 3D scene, and the images to be detected in the 3D scene can be input into the trained neural network for detection to obtain the results. However, since the image dataset in the 2D scene contains images of real objects, while the 3D scene contains images of various complex virtual objects, there is an inconsistency in data distribution between the real object images in the 2D scene and the virtual object images in the 3D scene. If the neural network trained in the 2D scene is used to detect the images to be detected in the 3D scene, the recognition accuracy of the neural network transferred to the 3D scene will be low.

[0049] Based on the above research, this disclosure provides a behavior detection method. This method obtains 3D keypoint sequences corresponding to multiple virtual objects in a target 3D space; generates 2D keypoint sequences of the virtual objects from multiple viewpoints based on the 3D keypoint sequences of the virtual objects; for each viewpoint, determines the interactive behavior features of the multiple virtual objects from that viewpoint based on the 2D keypoint sequences of the multiple virtual objects from that viewpoint and a trained behavior detection network; and determines the interactive behavior detection results of the multiple virtual objects based on the interactive behavior features of the multiple virtual objects from each viewpoint and the behavior detection network. In this disclosure, the input to the behavior detection network is a 2D keypoint sequence, therefore, the behavior detection network can be trained from the keypoint sequences of training samples in a 2D scene, alleviating the problems of high annotation costs and limited sample size caused by directly training neural networks using 3D spatial data.

[0050] Meanwhile, this disclosure obtains a 3D keypoint sequence. Since the 3D keypoint sequence includes 3D keypoint information collected at different time points, and the feature distribution between the keypoint information of real objects in real scenes and the keypoint information of virtual objects in virtual scenes is consistent, it can alleviate the accuracy problem caused by data inconsistency after neural network transfer. Furthermore, this method obtains a 2D keypoint sequence of virtual objects from multiple perspectives. Through the 2D keypoint sequence from multiple perspectives, interactive behavior features can be obtained more comprehensively and richly from multiple angles. Then, by utilizing the interactive behavior features of multiple virtual objects from various perspectives and the behavior detection network, the interactive behavior detection results of multiple virtual objects can be determined more accurately, improving the accuracy of interactive behavior detection.

[0051] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.

[0052] In this document, the term "and / or" merely describes a relationship, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.

[0053] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.

[0054] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.

[0055] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.

[0056] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.

[0057] To facilitate understanding of this embodiment, a behavior detection method disclosed in this disclosure will first be described in detail. This behavior detection method can be applied to computing devices with certain computing capabilities, such as servers and terminal devices. The following description uses a server as the executing entity to illustrate the behavior detection method.

[0058] See Figure 1 The diagram shows a flowchart of a behavior detection method provided in an embodiment of this disclosure. The method includes steps S101 to S104, wherein:

[0059] S101, Obtain the sequence of three-dimensional key points corresponding to multiple virtual objects in the target three-dimensional space; the sequence of three-dimensional key points includes three-dimensional key point information collected at different time points;

[0060] S102, Based on the three-dimensional key point sequence of the virtual object, generate a two-dimensional key point sequence of the virtual object from multiple perspectives;

[0061] S103, for each viewpoint, based on the two-dimensional keypoint sequence of the multiple virtual objects under the viewpoint and the trained behavior detection network, determine the interactive behavior features of the multiple virtual objects under the viewpoint;

[0062] S104, based on the interaction behavior features of the multiple virtual objects under each of the aforementioned perspectives and the behavior detection network, determine the interaction behavior detection results of the multiple virtual objects.

[0063] The following provides a detailed explanation of S101-S104.

[0064] In S101, the target 3D space can be any virtual 3D space, such as the 3D space in a 3D game scene, the 3D space in a virtual reality scene, etc. Multiple virtual objects are pre-constructed in this target 3D space, and each virtual object has 3D keypoints. The number and location of the 3D keypoints of the virtual objects can be determined as needed.

[0065] The system acquires sequences of 3D keypoints corresponding to multiple virtual objects in the target 3D space. These sequences include 3D keypoint information collected at different time points. The executing entity can perform real-time detection of the target 3D space to obtain the 3D keypoint information of each virtual object located within that space at different time points; that is, the 3D keypoint information of any virtual object in the target 3D space is known data.

[0066] If the 3D key point information of a virtual object in the target 3D space is unknown, video data and image data of the virtual object from multiple acquisition angles can be obtained. Based on the video data and image data of the virtual object from multiple angles, the 3D key point information of the virtual object can be determined. For example, a key point detection network can be used to detect each video frame in the image data or video data to obtain 2D key point information. Then, the 3D key point information of the virtual object can be determined through the 2D key information from multiple angles.

[0067] In one optional embodiment, obtaining the sequence of three-dimensional key points corresponding to multiple virtual objects in the target three-dimensional space includes: obtaining the position information corresponding to the multiple virtual objects in the target three-dimensional space; and, based on the position information corresponding to the multiple virtual objects, determining that there is a target distance between at least two of the virtual objects that is less than a preset distance value, obtaining the sequence of three-dimensional key points corresponding to the at least two virtual objects.

[0068] This disclosure allows for the setting of detection trigger logic, such as detecting interactive behavior when multiple virtual objects are close to each other. In implementation, the location information of multiple virtual objects in the target 3D space can be obtained, and based on the location information, the target distance between any two virtual objects can be determined. When it is determined that the target distance between at least two virtual objects is less than a preset distance value, the 3D keypoint sequence corresponding to those at least two virtual objects is obtained.

[0069] Here, by acquiring the position information of multiple virtual objects, the location information is used to conveniently and efficiently determine whether multiple virtual objects have triggered the interaction behavior detection logic. For example, when it is determined that the target distance between at least two virtual objects is less than a preset distance value, it is determined that the at least two virtual objects have triggered the interaction behavior detection logic, and the three-dimensional key point sequences corresponding to the at least two virtual objects are obtained to provide data support for subsequent interaction behavior detection.

[0070] In S102, for each virtual object, the 3D keypoint sequence of the virtual object can be mapped from multiple perspectives to generate a 2D keypoint sequence of the virtual object from each perspective. The number and angle information of the perspectives can be set as needed; for example, the number of perspectives can be 3, with perspective 1 being the frontal view of the virtual object, perspective 2 being the left side of the virtual object, and perspective 3 being the right side of the virtual object.

[0071] In practice, projection can be used to generate two-dimensional keypoint sequences from multiple perspectives based on the three-dimensional keypoint sequence of the virtual object. For example, if the 3D keypoint sequence includes five frames of 3D keypoint information, namely 3D keypoint information 1, 3D keypoint information 2, 3D keypoint information 3, 3D keypoint information 4, and 3D keypoint information 5; using a projection method, 3D keypoint information 1 is projected onto multiple viewpoints to obtain 2D keypoint information 1-1 under viewpoint 1, 2D keypoint information 2-1 under viewpoint 2, and so on; 3D keypoint information 2 is projected onto multiple viewpoints to obtain 2D keypoint information 1-2 under viewpoint 1, 2D keypoint information 2-2 under viewpoint 2, and so on; similarly, 3D keypoint information 3 can be obtained as 2D keypoint information 1-3 under viewpoint 1, 2D keypoint information 2-3 under viewpoint 2, etc.; 3D keypoint information 4 can be obtained as 2D keypoint information 1-4 under viewpoint 1, 2D keypoint information 2-4 under viewpoint 2, etc.; and 3D keypoint information 5 can be obtained as 2D keypoint information 1-5 under viewpoint 1, 2D keypoint information 2-5 under viewpoint 2, etc. Therefore, the two-dimensional keypoint sequence under viewpoint 1 includes: two-dimensional keypoint information 1-1, two-dimensional keypoint information 1-2, two-dimensional keypoint information 1-3, two-dimensional keypoint information 1-4, and two-dimensional keypoint information 1-5; the two-dimensional keypoint sequence under viewpoint 2 includes: two-dimensional keypoint information 2-1, two-dimensional keypoint information 2-2, two-dimensional keypoint information 2-3, two-dimensional keypoint information 2-4, and two-dimensional keypoint information 2-5.

[0072] In S103, each viewpoint can invoke a pre-trained behavior detection network to detect the 2D keypoint sequence under that viewpoint. The 2D keypoint sequences under different viewpoints can be detected in parallel by invoking multiple behavior detection networks to obtain the interaction behavior features of multiple virtual objects under various viewpoints. Alternatively, a pre-trained behavior detection network can be invoked, and the 2D keypoint sequences under each viewpoint can be input into the behavior detection network in sequence to obtain the interaction behavior features of multiple virtual objects under various viewpoints.

[0073] During implementation, for each viewpoint, the two-dimensional keypoint sequence of multiple virtual objects under that viewpoint is input into the behavior detection network. For example, when multiple virtual objects include virtual object 1 and virtual object 2, the two-dimensional keypoint sequence of virtual object 1 under that viewpoint and the two-dimensional keypoint sequence of virtual object 2 under that viewpoint are input into the behavior detection network to generate the interaction behavior features of multiple virtual objects under that viewpoint. These interaction behavior features can be the feature data output by the behavior detection network.

[0074] In S104, the interaction behavior detection results of multiple virtual objects can be determined based on the interaction behavior features of multiple virtual objects from various perspectives and the behavior detection network. For example, the classifier in the behavior detection network can be used to detect the interaction behavior features of multiple virtual objects from each perspective to obtain the intermediate detection results of multiple virtual objects from that perspective. The intermediate detection results of multiple virtual objects from various perspectives are combined, and the interaction behavior detection results of multiple virtual objects are determined according to the voting method.

[0075] Alternatively, the interactive behavior features of multiple virtual objects from various perspectives can be fused. For example, the interactive behavior features of multiple virtual objects from various perspectives can be concatenated to obtain concatenated feature data. Then, the concatenated feature data can be convolved to generate fused feature data. Finally, a classifier in the behavior detection network can be used to detect the fused feature data to obtain the interactive behavior detection results of multiple virtual objects.

[0076] In one optional implementation, determining the interaction behavior detection results of the plurality of virtual objects based on the interaction behavior features of the plurality of virtual objects under each of the aforementioned viewpoints and the behavior detection network includes:

[0077] Step a1: For each viewpoint, based on the interaction behavior features of the multiple virtual objects under the viewpoint and the behavior detection network, generate intermediate detection results under the viewpoint;

[0078] Step a2: Based on the intermediate detection results corresponding to each of the aforementioned viewpoints, determine the interaction behavior detection results of the multiple virtual objects.

[0079] In step a1, for each viewpoint, a classifier in the behavior detection network can be used to detect the interactive behavior features of multiple virtual objects under that viewpoint, generating intermediate detection results for that viewpoint. These intermediate detection results can include the confidence scores of multiple virtual objects belonging to various interactive behaviors. For example, pre-defined interactive behaviors may include: greeting, handshake, hug, and fighting. The intermediate detection results could then include: greeting - confidence score 0.9, handshake - confidence score 0.1, hug - confidence score 0, and fighting - confidence score 0. Alternatively, the intermediate detection results can also include the target interactive behavior to which multiple virtual objects belong, and the confidence score of belonging to that target interactive behavior. For example, the intermediate detection results could include: greeting - confidence score 0.9.

[0080] In step a2, the interaction behavior detection results of multiple virtual objects can be determined based on the intermediate detection results corresponding to each viewpoint. For example, the interaction behaviors with higher confidence and more frequency can be determined based on the intermediate detection results corresponding to each viewpoint and thus identified as the interaction behavior detection results. For example, if the intermediate detection results from viewpoint 1 include: greeting interaction - confidence 0.9, handshake interaction - confidence 0.1, hug interaction - confidence 0, and fighting interaction - confidence 0; the intermediate detection results from viewpoint 2 include: greeting interaction - confidence 0.85, handshake interaction - confidence 0.05, hug interaction - confidence 0.1, and fighting interaction - confidence 0; and the intermediate detection results from viewpoint 3 include: greeting interaction - confidence 0.75, handshake interaction - confidence 0.1, hug interaction - confidence 0.1, and fighting interaction - confidence 0.05; then the interaction behavior detection result can be determined as: greeting interaction. Alternatively, a voting method can be used to determine the interaction behavior detection results of multiple virtual objects based on the intermediate detection results corresponding to each viewpoint.

[0081] Here, for each viewpoint, the intermediate detection results of multiple virtual objects under that viewpoint are determined. Then, by comprehensively considering the intermediate detection results corresponding to each viewpoint, the interaction behavior detection results of multiple virtual objects can be determined more accurately.

[0082] In one optional implementation, determining the interaction behavior detection results of the plurality of virtual objects based on the intermediate detection results corresponding to each of the respective viewpoints includes: determining the interaction behavior detection results of the plurality of virtual objects based on the intermediate detection results corresponding to each of the respective viewpoints and the weights corresponding to each of the respective viewpoints obtained by training the behavior detection network.

[0083] Considering that the influence of interactive behavior features from different perspectives on the interactive behavior detection results may vary, in order to more accurately determine the interactive behavior detection results, learnable weight parameters can be set for each perspective. These weight parameters can be used as network parameters of the behavior detection network. When training the behavior detection network, the weights corresponding to each perspective can be obtained more accurately. Then, based on the intermediate detection results corresponding to each perspective and the weights corresponding to each perspective obtained from the training of the behavior detection network, the interactive behavior detection results of multiple virtual objects can be determined more accurately.

[0084] After obtaining the intermediate detection results corresponding to each viewpoint, the interaction behavior detection results of multiple virtual objects can be determined based on the weights corresponding to each viewpoint obtained by the behavior detection network and the intermediate detection results corresponding to each viewpoint. For example, if the intermediate detection results from perspective 1 include: greeting interaction - confidence 0.7, handshake interaction - confidence 0.3, hug interaction - confidence 0, and fighting interaction - confidence 0; the intermediate detection results from perspective 2 include: greeting interaction - confidence 0.5, handshake interaction - confidence 0.5, hug interaction - confidence 0, and fighting interaction - confidence 0; and the intermediate detection results from perspective 3 include: greeting interaction - confidence 0.3, handshake interaction - confidence 0.6, hug interaction - confidence 0.1, and fighting interaction - confidence 0; and the weight of perspective 1 is 0.3, the weight of perspective 2 is 0.2, and the weight of perspective 3 is 0.5, then the interaction detection result can be determined as: handshake interaction. Alternatively, a voting method can be used, combining the weights corresponding to each perspective and the intermediate detection results, to determine the interaction detection results for multiple virtual objects.

[0085] In another optional implementation, determining the interaction behavior detection results of the multiple virtual objects based on the interaction behavior features of the multiple virtual objects under each of the aforementioned viewpoints and the behavior detection network includes:

[0086] Step b1: The interactive behavior features of the multiple virtual objects under each of the various viewpoints are fused to generate fused feature data.

[0087] Step b2: Based on the fused feature data and the behavior detection network, determine the interaction behavior detection results of the multiple virtual objects.

[0088] In implementation, the interactive behavior features of multiple virtual objects from various perspectives can be fused to generate fused feature data. There are several feature fusion methods, which can be set according to needs during implementation; for example, the interactive behavior features of multiple virtual objects from various perspectives can be concatenated, and the concatenated features can be used as fused feature data; alternatively, the interactive behavior features of multiple virtual objects from various perspectives can be concatenated first, and the concatenated features can be subjected to feature extraction (such as convolution processing, fully connected processing, etc.) to generate fused feature data.

[0089] The classifier in the behavior detection network is then used to process the fused feature data to determine the interaction behavior detection results of multiple virtual objects.

[0090] Here, the interaction behavior features of multiple virtual objects from various perspectives are fused to generate fused feature data with rich feature information. This allows the behavior detection network to accurately determine the interaction behavior detection results based on the fused feature data with rich feature information, resulting in high detection accuracy.

[0091] In one optional implementation, the method further includes: generating and displaying a warning message when the interaction behavior detection result indicates that the interaction behavior between the plurality of virtual objects matches a preset behavior.

[0092] During implementation, the interaction behavior detection results can be used to determine whether the interaction behavior between multiple virtual objects matches preset behaviors. Preset behaviors can be pre-set as needed, and may include, but are not limited to, fighting or pushing behaviors. If a match is found, a warning message can be generated and displayed.

[0093] If the interaction behavior detection results indicate that the interaction behavior between multiple virtual objects does not match the preset behavior, it can be determined that the interaction behavior between multiple virtual objects is a normal interaction behavior. Then, different virtual effects can be displayed according to the interaction behavior detection results. For example, when the interaction behavior detection results indicate that the interaction behavior between multiple virtual objects is an hugging interaction behavior, a virtual effect of playing fireworks can be displayed.

[0094] This disclosure also includes the process of training the behavior detection network, which is described below.

[0095] The steps for training the behavior detection network include: acquiring multiple sample video data carrying labeled tags; the labeled tags are used to indicate the interaction behavior between different sample objects in the sample video data; determining the two-dimensional keypoint sequence of the sample objects included in each sample video data; inputting the two-dimensional keypoint sequence corresponding to each sample video data into the neural network to be trained to generate a predicted behavior result corresponding to each sample video data; and adjusting the parameters of the neural network to be trained based on the predicted behavior result and the labeled tags for each sample video data until the adjusted neural network meets the training cutoff condition, thereby obtaining the behavior detection network.

[0096] Acquire multiple sample video data sets carrying labeled tags. These sample video data can be video data from open-source datasets or real-time generated video data including the interactive behaviors of multiple sample objects. The sample video data includes interactive behaviors between different sample objects, and the labeled tags are used to indicate these interactive behaviors. That is, the labeled tags can be labels representing the interactive behaviors of different sample objects, such as labels for greetings, fighting, etc.

[0097] For each sample video data set, multiple sample video frames are acquired. A keypoint detection network is used to determine the two-dimensional keypoint information of each sample object in the multiple sample video frames; alternatively, the two-dimensional keypoint information of each sample object in the multiple sample video frames can be determined in response to an annotation operation. Then, based on the two-dimensional keypoint information corresponding to each sample object in the multiple sample video frames, a two-dimensional keypoint sequence of the sample objects included in the sample video data is generated.

[0098] The two-dimensional keypoint sequences corresponding to each sample video data are input into the neural network to be trained, generating the predicted behavior result for each sample video data. The network structure of the neural network to be trained can be set as needed, such as a graph convolutional neural network or a spatiotemporal graph convolutional neural network. Then, based on the corresponding keypoint sequences of each sample video data...

[0099] The process involves predicting behavioral outcomes and assigning labels, determining a loss value, and then adjusting the parameters of the neural network (5) to be trained based on this loss value until the adjusted neural network meets the training cutoff criteria, thus obtaining the behavior detection network. Training cutoff criteria may include, for example, a training iteration count exceeding a set threshold, a loss value less than a set threshold, neural network convergence, and detection accuracy exceeding a set accuracy threshold.

[0100] The training process in this disclosure uses sample video data, which is data from a two-dimensional 0 scene. The sample video data has a large amount of data and a small amount of annotation, which ensures the required amount of sample data and reduces the time cost of annotation operations.

[0101] In one optional embodiment, after obtaining the behavior detection network, the method further includes: acquiring three-dimensional sample data of the target three-dimensional space, wherein the three-dimensional sample data includes behavior annotation labels of multiple sample objects, and three-dimensional key point sequences of each of the sample objects that match the behavior annotation labels; and using the three-dimensional sample data to adjust the parameters of the behavior detection network to generate an adjusted behavior detection network.

[0102] During implementation, three-dimensional sample data of the target three-dimensional space can also be acquired. This three-dimensional sample data includes behavioral annotation labels for multiple sample objects, as well as individual samples that match the behavioral annotation labels.

[0103] The sequence of 3D key points for this object. Then, the parameters of the behavior detection network are fine-tuned using 3D sample data from the target's 3D space, generating the adjusted behavior detection network.

[0104] For example, based on the 3D keypoint sequences of each sample object in the 3D sample data, multiple 2D keypoint sequences of the sample objects from multiple perspectives can be generated. These 2D keypoint sequences from each perspective are then input into a behavior detection network to obtain predicted behavior labels. Finally, based on the predicted behavior labels...

[0105] The behavior detection network is fine-tuned by adjusting parameters 5 based on the loss value, and the adjusted behavior detection network is generated by adding and subtracting labels and behavior annotations.

[0106] This allows us to subsequently utilize the adjusted behavior detection network to determine the interactive behavior features of multiple virtual objects from each perspective, based on the two-dimensional keypoint sequences of multiple virtual objects in each viewpoint; and to use the adjusted behavior detection network to determine the interactive behavior detection results of multiple virtual objects based on the interactive behavior features of multiple virtual objects from each viewpoint.

[0107] Here, by using 3D sample data from the target 3D space, the parameters of the behavior detection network are adjusted so that the adjusted behavior detection network is more compatible with the target 3D space. That is, the adjusted behavior detection network can more accurately detect the interactive behavior of multiple virtual objects in the target 3D space, thereby improving the detection accuracy.

[0108] See Figure 2 As shown, combined with Figure 2 The behavior detection method proposed in this disclosure is illustrated by example. The method specifically includes the following steps:

[0109] Step 1: Obtain multiple sample video data.

[0110] Step 2: For each sample video data, perform pose estimation on each sample video frame within the sample video data to obtain the two-dimensional key point information of each sample object in the sample video frame, and then obtain the two-dimensional key point sequence of each sample object in the sample video data.

[0111] Step 3: Using the two-dimensional keypoint sequences corresponding to each sample video data, train the neural network to obtain the behavior detection network. For example, the two-dimensional keypoint sequences corresponding to each sample video data can be input into the neural network to obtain classification results. Based on the classification results and the labeled data of the sample video data, adjust the parameters of the neural network to be trained until the training cutoff condition is met, thus obtaining the behavior detection network. Then, transfer the behavior detection network to the target 3D space to detect the interaction behavior of multiple virtual objects in the target 3D space.

[0112] Neural network to be trained, for example Figure 2 The graph convolutional network shown is an example.

[0113] Here, three-dimensional sample data of the target three-dimensional space can also be obtained, including behavioral annotation labels of multiple sample objects, and three-dimensional key point sequences of each sample object that match the behavioral annotation labels; and the behavior detection network can be adjusted (fine-tuned) using the three-dimensional sample data to generate an adjusted behavior detection network, so as to perform network migration on the adjusted behavior detection network.

[0114] Step 4: Obtain the position information of multiple virtual objects in the target 3D space.

[0115] Step 5: Based on the position information corresponding to multiple virtual objects, determine whether there are at least two virtual objects whose target distance is less than a preset distance value. If so, obtain the 3D keypoint sequences corresponding to at least two virtual objects. Figure 2 The distance shown is too close, triggering the logic.

[0116] Step 6: Based on the 3D keypoint sequence of the virtual object, generate a 2D keypoint sequence of the virtual object from multiple viewpoints. That is... Figure 2 Two-dimensional multi-view keypoint mapping in the model.

[0117] Step 7: For each viewpoint, input the two-dimensional keypoint sequence of multiple virtual objects in the viewpoint into the trained behavior detection network to generate the interactive behavior features of multiple virtual objects in the viewpoint.

[0118] Step 8: Determine the interaction behavior detection results of multiple virtual objects based on the interaction behavior characteristics of multiple virtual objects from various perspectives and the classifier of the behavior detection network.

[0119] For example, for each viewpoint, intermediate detection results can be generated based on the interaction behavior features of multiple virtual objects under that viewpoint and the classifier of the behavior detection network; then, based on the intermediate detection results corresponding to each viewpoint, the interaction behavior detection results of multiple virtual objects can be determined. Alternatively, the interaction behavior features of multiple virtual objects under various viewpoints can be fused to generate fused feature data; then, based on the fused feature data and the behavior detection network, the interaction behavior detection results of multiple virtual objects can be determined.

[0120] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.

[0121] Based on the same inventive concept, this disclosure also provides a behavior detection device corresponding to the behavior detection method. Since the principle of the device in this disclosure for solving the problem is similar to that of the behavior detection method described above, the implementation of the device can refer to the implementation of the method, and the repeated parts will not be described again.

[0122] Reference Figure 3 The diagram shown is an architectural schematic of a behavior detection device provided in an embodiment of this disclosure. The device includes: an acquisition module 301, a generation module 302, a first determination module 303, and a second determination module 304; wherein,

[0123] The acquisition module 301 is used to acquire a sequence of three-dimensional key points corresponding to multiple virtual objects in the target three-dimensional space; the sequence of three-dimensional key points includes three-dimensional key point information collected at different time points;

[0124] The generation module 302 is used to generate a two-dimensional key point sequence of the virtual object from multiple perspectives based on the three-dimensional key point sequence of the virtual object;

[0125] The first determining module 303 is used to determine the interactive behavior features of the multiple virtual objects in the viewpoint based on the two-dimensional keypoint sequence of the multiple virtual objects in the viewpoint and the trained behavior detection network for each viewpoint.

[0126] The second determining module 304 is used to determine the interaction behavior detection results of the multiple virtual objects based on the interaction behavior features of the multiple virtual objects under each of the aforementioned perspectives and the behavior detection network.

[0127] In an optional implementation, the second determining module 304, when determining the interaction behavior detection results of the plurality of virtual objects based on the interaction behavior features of the plurality of virtual objects under each of the aforementioned viewpoints and the behavior detection network, is configured to:

[0128] For each viewpoint, based on the interaction behavior features of the multiple virtual objects under that viewpoint and the behavior detection network, intermediate detection results are generated under that viewpoint;

[0129] Based on the intermediate detection results corresponding to each of the aforementioned perspectives, the interaction behavior detection results of the multiple virtual objects are determined.

[0130] In an optional implementation, the second determining module 304, when determining the interaction behavior detection results of the plurality of virtual objects based on the intermediate detection results corresponding to each of the aforementioned viewpoints, is configured to:

[0131] Based on the intermediate detection results corresponding to each of the aforementioned viewpoints and the weights corresponding to each of the aforementioned viewpoints obtained by training the behavior detection network, the interaction behavior detection results of the multiple virtual objects are determined.

[0132] In an optional implementation, the second determining module 304, when determining the interaction behavior detection results of the plurality of virtual objects based on the interaction behavior features of the plurality of virtual objects under each of the aforementioned viewpoints and the behavior detection network, is configured to:

[0133] The interactive behavior features of the multiple virtual objects under each of the various perspectives are fused to generate fused feature data;

[0134] Based on the fused feature data and the behavior detection network, the interaction behavior detection results of the multiple virtual objects are determined.

[0135] In an optional implementation, the acquisition module 301, when acquiring the sequence of 3D key points corresponding to multiple virtual objects in the target 3D space, is used to:

[0136] Obtain the position information of the plurality of virtual objects in the target three-dimensional space;

[0137] Based on the location information corresponding to the plurality of virtual objects, if it is determined that the target distance between at least two of the virtual objects is less than a preset distance value, the three-dimensional key point sequence corresponding to the at least two virtual objects is obtained.

[0138] In an optional embodiment, the apparatus further includes a training module 305 for training the behavior detection network according to the following steps:

[0139] Acquire multiple sample video data carrying labeled tags; the labeled tags are used to indicate the interaction behavior between different sample objects in the sample video data;

[0140] Determine the two-dimensional keypoint sequence of the sample objects included in each of the sample video data;

[0141] The two-dimensional keypoint sequence corresponding to each of the sample video data is input into the neural network to be trained to generate the predicted behavior result corresponding to each of the sample video data.

[0142] Based on the predicted behavior result and the labeled label corresponding to each sample video data, the parameters of the neural network to be trained are adjusted until the adjusted neural network meets the training cutoff condition, thus obtaining the behavior detection network.

[0143] In an optional implementation, after obtaining the behavior detection network, the training module 305 is further configured to:

[0144] Obtain three-dimensional sample data of the target three-dimensional space, wherein the three-dimensional sample data includes behavioral annotation labels of multiple sample objects, and a sequence of three-dimensional key points of each sample object that matches the behavioral annotation labels;

[0145] Using the three-dimensional sample data, the parameters of the behavior detection network are adjusted to generate an adjusted behavior detection network.

[0146] In an optional embodiment, the device further includes: a warning module 306, used for:

[0147] If the interaction behavior detection result indicates that the interaction behavior between the multiple virtual objects matches the preset behavior, a warning message is generated and displayed.

[0148] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.

[0149] Based on the same technical concept, this disclosure also provides a computer device. (See also...) Figure 4The diagram shows the structure of a computer device 400 provided in this embodiment, including a processor 401, a memory 402, and a bus 403. The memory 402 stores execution instructions and includes main memory 4021 and external memory 4022. The main memory 4021, also called internal memory, is used to temporarily store computational data in the processor 401 and data exchanged with external memory 4022 such as a hard disk. The processor 401 exchanges data with the external memory 4022 through the main memory 4021. When the computer device 400 is running, the processor 401 and the memory 402 communicate through the bus 403, causing the processor 401 to execute the following instructions:

[0150] Obtain a sequence of 3D key points corresponding to multiple virtual objects in the target 3D space; the sequence of 3D key points includes 3D key point information collected at different time points;

[0151] Based on the three-dimensional key point sequence of the virtual object, generate a two-dimensional key point sequence of the virtual object from multiple perspectives;

[0152] For each viewpoint, based on the two-dimensional keypoint sequence of the multiple virtual objects under the viewpoint and the trained behavior detection network, the interaction behavior features of the multiple virtual objects under the viewpoint are determined;

[0153] Based on the interaction behavior features of the multiple virtual objects under each of the aforementioned perspectives and the behavior detection network, the interaction behavior detection results of the multiple virtual objects are determined.

[0154] In one possible design, the instructions executed by processor 401, which include determining the interaction behavior detection results of the multiple virtual objects based on the interaction behavior features of the multiple virtual objects from various perspectives and the behavior detection network, include:

[0155] For each viewpoint, based on the interaction behavior features of the multiple virtual objects under that viewpoint and the behavior detection network, intermediate detection results are generated under that viewpoint;

[0156] Based on the intermediate detection results corresponding to each of the aforementioned perspectives, the interaction behavior detection results of the multiple virtual objects are determined.

[0157] In one possible design, the instructions executed by processor 401, which include determining the interaction behavior detection results of the multiple virtual objects based on the intermediate detection results corresponding to each of the aforementioned viewpoints, include:

[0158] Based on the intermediate detection results corresponding to each of the aforementioned viewpoints and the weights corresponding to each of the aforementioned viewpoints obtained by training the behavior detection network, the interaction behavior detection results of the multiple virtual objects are determined.

[0159] In one possible design, the instructions executed by processor 401, which include determining the interaction behavior detection results of the multiple virtual objects based on the interaction behavior features of the multiple virtual objects from various perspectives and the behavior detection network, include:

[0160] The interactive behavior features of the multiple virtual objects under each of the various perspectives are fused to generate fused feature data;

[0161] Based on the fused feature data and the behavior detection network, the interaction behavior detection results of the multiple virtual objects are determined.

[0162] In one possible design, the instruction executed by processor 401, which involves obtaining the sequence of 3D key points corresponding to multiple virtual objects in the target 3D space, includes:

[0163] Obtain the position information of the plurality of virtual objects in the target three-dimensional space;

[0164] Based on the location information corresponding to the plurality of virtual objects, if it is determined that the target distance between at least two of the virtual objects is less than a preset distance value, the three-dimensional key point sequence corresponding to the at least two virtual objects is obtained.

[0165] In one possible design, the steps for training the behavior detection network in the instructions executed by processor 401 include:

[0166] Acquire multiple sample video data carrying labeled tags; the labeled tags are used to indicate the interaction behavior between different sample objects in the sample video data;

[0167] Determine the two-dimensional keypoint sequence of the sample objects included in each of the sample video data;

[0168] The two-dimensional keypoint sequence corresponding to each of the sample video data is input into the neural network to be trained to generate the predicted behavior result corresponding to each of the sample video data.

[0169] Based on the predicted behavior result and the labeled label corresponding to each sample video data, the parameters of the neural network to be trained are adjusted until the adjusted neural network meets the training cutoff condition, thus obtaining the behavior detection network.

[0170] In one possible design, after obtaining the behavior detection network, the method further includes the following instructions executed by processor 401:

[0171] Obtain three-dimensional sample data of the target three-dimensional space, wherein the three-dimensional sample data includes behavioral annotation labels of multiple sample objects, and a sequence of three-dimensional key points of each sample object that matches the behavioral annotation labels;

[0172] Using the three-dimensional sample data, the parameters of the behavior detection network are adjusted to generate an adjusted behavior detection network.

[0173] In one possible design, the method further includes the following instructions executed by processor 401:

[0174] If the interaction behavior detection result indicates that the interaction behavior between the multiple virtual objects matches the preset behavior, a warning message is generated and displayed.

[0175] This disclosure also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the behavior detection method described in the above method embodiments. The storage medium can be a volatile or non-volatile computer-readable storage medium.

[0176] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the behavior detection method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.

[0177] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.

[0178] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.

[0179] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0180] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0181] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0182] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.

Claims

1. A behavior detection method, comprising: Obtain the sequence of 3D key points corresponding to multiple virtual objects in the target 3D space; The sequence of three-dimensional key points corresponding to each virtual object includes three-dimensional key point information collected at different time points; For each virtual object, a two-dimensional keypoint sequence of the virtual object is generated from multiple perspectives based on the three-dimensional keypoint sequence of the virtual object; For each viewpoint, based on the two-dimensional keypoint sequence and behavior detection network of the multiple virtual objects under the viewpoint, the interactive behavior features of the multiple virtual objects under the viewpoint are determined; And based on the interaction behavior features of the multiple virtual objects under each of the aforementioned perspectives and the classifier in the behavior detection network, the interaction behavior detection results between the multiple virtual objects are determined; The step of determining the interaction behavior detection results between the multiple virtual objects based on the interaction behavior features of the multiple virtual objects under each of the aforementioned viewpoints and the classifier in the behavior detection network includes: For each viewpoint, based on the interaction behavior features of the plurality of virtual objects under that viewpoint and the classifier of the behavior detection network, intermediate detection results are generated for that viewpoint; based on the intermediate detection results corresponding to each of the viewpoints, the interaction behavior detection results between the plurality of virtual objects are determined; or The interactive behavior features of the multiple virtual objects under each of the aforementioned perspectives are fused to generate fused feature data; based on the fused feature data and the classifier of the behavior detection network, the interactive behavior detection results between the multiple virtual objects are determined. The steps for training the behavior detection network include: Multiple sample video data carrying labeled tags are acquired; the labeled tags of each sample video data are used to indicate the interaction behavior between different sample objects in the sample video data; a two-dimensional keypoint sequence of the sample objects included in each sample video data is determined; the two-dimensional keypoint sequence corresponding to each sample video data is input into the neural network to be trained to generate a predicted behavior result corresponding to each sample video data; based on the predicted behavior result and the labeled tags corresponding to each sample video data, the parameters of the neural network to be trained are adjusted until the adjusted neural network to be trained meets the training cutoff condition, thereby obtaining the behavior detection network.

2. The method of claim 1, wherein, The determination of interaction behavior detection results between the multiple virtual objects based on the intermediate detection results corresponding to each of the aforementioned viewpoints includes: Based on the intermediate detection results corresponding to each of the aforementioned viewpoints and the weights corresponding to each of the aforementioned viewpoints obtained by the behavior detection network, the interaction behavior detection results between the multiple virtual objects are determined.

3. The method of claim 1, wherein, The step of obtaining the sequence of 3D key points corresponding to multiple virtual objects in the target 3D space includes: Obtain the position information of the plurality of virtual objects in the target three-dimensional space; Based on the location information corresponding to the plurality of virtual objects, if it is determined that the target distance between at least two of the virtual objects is less than a preset distance value, the three-dimensional key point sequence corresponding to the at least two virtual objects is obtained.

4. The method according to claim 1, wherein, After obtaining the behavior detection network, the method further includes: Obtain three-dimensional sample data of the target three-dimensional space, wherein the three-dimensional sample data includes behavioral annotation labels of multiple sample objects, and a sequence of three-dimensional key points of each sample object included in the three-dimensional sample data that matches the behavioral annotation labels; Using the three-dimensional sample data, the parameters of the behavior detection network are adjusted to generate an adjusted behavior detection network.

5. The method according to any one of claims 1-4, further comprising: If the interaction behavior detection result indicates that the interaction behavior between the multiple virtual objects matches the preset behavior, a warning message is generated and displayed.

6. A behavior detection device, comprising: The acquisition module is used to acquire the sequence of 3D key points corresponding to multiple virtual objects in the target 3D space. The sequence of three-dimensional key points corresponding to each virtual object includes three-dimensional key point information collected at different time points; The generation module is used to generate a two-dimensional keypoint sequence of the virtual object from multiple perspectives, based on the three-dimensional keypoint sequence of the virtual object; The first determining module is used to determine the interactive behavior features of the multiple virtual objects in the viewpoint based on the two-dimensional keypoint sequence and behavior detection network of the multiple virtual objects in the viewpoint for each viewpoint; The second determining module is used to determine the interaction behavior detection results between the multiple virtual objects based on the interaction behavior features of the multiple virtual objects under each of the views and the classifier in the behavior detection network; Specifically, when performing the step of determining the interaction behavior detection results between the multiple virtual objects based on the interaction behavior features of the multiple virtual objects under each of the aforementioned viewpoints and the classifier in the behavior detection network, the second determining module is used to: For each viewpoint, based on the interaction behavior features of the plurality of virtual objects under that viewpoint and the classifier of the behavior detection network, intermediate detection results are generated for that viewpoint; based on the intermediate detection results corresponding to each of the viewpoints, the interaction behavior detection results between the plurality of virtual objects are determined; or The interactive behavior features of the multiple virtual objects under each of the aforementioned perspectives are fused to generate fused feature data; based on the fused feature data and the classifier of the behavior detection network, the interactive behavior detection results between the multiple virtual objects are determined. The steps for training the behavior detection network include: Multiple sample video data carrying labeled tags are acquired; the labeled tags of each sample video data are used to indicate the interaction behavior between different sample objects in the sample video data; a two-dimensional keypoint sequence of the sample objects included in each sample video data is determined; the two-dimensional keypoint sequence corresponding to each sample video data is input into the neural network to be trained to generate a predicted behavior result corresponding to each sample video data; based on the predicted behavior result and the labeled tags corresponding to each sample video data, the parameters of the neural network to be trained are adjusted until the adjusted neural network to be trained meets the training cutoff condition, thereby obtaining the behavior detection network.

7. A computer device, comprising: The computer device includes a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the computer device is running, the processor communicates with the memory via the bus, and the machine-readable instructions, when executed by the processor, perform the steps of the behavior detection method as described in any one of claims 1 to 5.

8. A computer-readable storage medium on which a computer program is stored, wherein, The computer program is executed by the processor to perform the steps of the behavior detection method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Training data generation method and device, terminal equipment and storage medium

    CN114359658A

  • 2D human body pose generation method and device for strong interaction human body motion

    CN114548224A