Multi-modal human skeleton behavior recognition method and device, equipment and storage medium
Through multimodal data processing, including feature encoding of skeleton sequences and prompt words, the similarity is calculated to determine the behavior recognition results, and the problems of insufficient recognition ability and low accuracy caused by relying on single-modal data in the prior art are solved, and high-accuracy human behavior recognition is achieved in the case of few samples.
Patent Information
- Application Number
- CN202510006338.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-09
AI Technical Summary
The existing human skeleton behavior recognition methods mainly rely on skeleton single-modal data, making it difficult to use other data modes to improve model recognition capabilities, and the recognition results are not accurate enough.
A multimodal human body skeleton behavior recognition method is proposed. By obtaining the sequence of skeletons to be identified and the set of prompt words, performing feature coding and text encoding, calculating the similarity between the spatiotemporal embedding feature vector and the text embedding feature matrix, and then determining the target behavior recognition result.
Quickly realize business needs with few samples or even zero samples. Multimodal data improves the accuracy of the method and can more accurately identify human behavior.
Smart Images

Figure CN119964235A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer vision technology, and in particular to a multimodal human skeleton behavior recognition method, device, equipment and storage medium. Background Art
[0002] Human behavior recognition is an extremely important and active research direction in the field of computer vision, with a wide range of applications, such as intelligent monitoring systems, human-computer interaction, virtual reality, robotics, etc. The behavior recognition method based on skeleton data has better robustness in the face of complex backgrounds, viewpoint changes, light and dark changes, etc. In addition, skeleton data is more concise, and its calculation method occupies less computing resources.
[0003] In existing human skeleton behavior recognition methods, usually only skeleton unimodal data is used for training and inference.
[0004] In this way, it is difficult to use data modalities other than skeleton data to improve the model recognition capability, and the accuracy of the recognition results cannot be guaranteed. Summary of the invention
[0005] The present disclosure aims to solve one of the technical problems in the related art at least to some extent.
[0006] To this end, the purpose of the present invention is to propose a multimodal human skeleton behavior recognition method, device, computer equipment and storage medium, which can quickly meet business needs with few or even zero samples, and multimodal data can further improve the accuracy of the method.
[0007] To achieve the above-mentioned purpose, the multimodal human skeleton behavior recognition method proposed in the first aspect of the present disclosure includes:
[0008] Obtaining a skeleton sequence to be identified of a target human skeleton and a prompt word set, wherein the prompt word set includes a plurality of candidate prompt texts, and the candidate prompt texts are associated with candidate behavior recognition results;
[0009] Performing feature encoding on the skeleton sequence to be identified to obtain a spatiotemporal embedding feature vector;
[0010] Performing text encoding on the candidate prompt text to obtain a text embedding feature matrix;
[0011] Determining the similarity between the spatiotemporal embedding feature vector and the text embedding feature matrix;
[0012] A target behavior recognition result matching the skeleton sequence to be recognized is determined from the plurality of candidate behavior recognition results according to the similarity.
[0013] To achieve the above-mentioned purpose, a multimodal human skeleton behavior recognition device proposed in a second aspect of the present disclosure includes:
[0014] An acquisition module, used to acquire a skeleton sequence to be identified of a target human skeleton and a prompt word set, wherein the prompt word set includes a plurality of candidate prompt texts, and the candidate prompt texts are associated with candidate behavior recognition results;
[0015] A first encoding module is used to perform feature encoding on the skeleton sequence to be identified to obtain a spatiotemporal embedding feature vector;
[0016] A second encoding module is used to perform text encoding on the candidate prompt text to obtain a text embedding feature matrix;
[0017] A first determination module is used to determine the similarity between the spatiotemporal embedding feature vector and the text embedding feature matrix;
[0018] The second determination module is used to determine a target behavior recognition result that matches the skeleton sequence to be recognized from the plurality of candidate behavior recognition results according to the similarity.
[0019] The computer device proposed in the third aspect embodiment of the present disclosure includes: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the multimodal human skeleton behavior recognition method proposed in the first aspect embodiment of the present disclosure is implemented.
[0020] The fourth aspect embodiment of the present disclosure proposes a non-temporary computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the multimodal human skeleton behavior recognition method proposed in the first aspect embodiment of the present disclosure.
[0021] The fifth aspect embodiment of the present disclosure proposes a computer program product. When the instructions in the computer program product are executed by a processor, the multimodal human skeleton behavior recognition method proposed in the first aspect embodiment of the present disclosure is executed.
[0022] The multimodal human skeleton behavior recognition method, device, computer equipment and storage medium provided by the present disclosure obtain the skeleton sequence to be recognized and the prompt word set of the target human skeleton, wherein the prompt word set includes multiple candidate prompt texts, and the candidate prompt texts are associated with the candidate behavior recognition results; feature encoding is performed on the skeleton sequence to be recognized to obtain a spatiotemporal embedding feature vector; text encoding is performed on the candidate prompt text to obtain a text embedding feature matrix; the similarity between the spatiotemporal embedding feature vector and the text embedding feature matrix is determined; and the target behavior recognition result matching the skeleton sequence to be recognized is determined from multiple candidate behavior recognition results according to the similarity. In this way, business needs can be quickly met with few or even zero samples, and multimodal data can further improve the accuracy of the method.
[0023] Additional aspects and advantages of the present disclosure will be given in part in the following description and in part will be obvious from the following description or learned through practice of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The above and / or additional aspects and advantages of the present disclosure will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which:
[0025] Figure 1 is a flow chart of a multimodal human skeleton behavior recognition method proposed in an embodiment of the present disclosure;
[0026] Figure 2 is a flowchart of a multimodal human skeleton behavior recognition method proposed in another embodiment of the present disclosure;
[0027] Figure 3 It is a schematic diagram of the process of the multimodal human skeleton behavior recognition method proposed in the present disclosure;
[0028] Figure 4 is a reasoning flow chart of a multimodal human skeleton behavior recognition method proposed in the present disclosure;
[0029] Figure 5 It is a joint point map proposed according to the present disclosure;
[0030] Figure 6 is a schematic diagram of the structure of a multimodal human skeleton behavior recognition device proposed in an embodiment of the present disclosure;
[0031] Figure 7 A block diagram of an exemplary computer device suitable for implementing embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0032] Embodiments of the present disclosure are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present disclosure, and are not to be construed as limitations of the present disclosure. On the contrary, the embodiments of the present disclosure include all changes, modifications, and equivalents that fall within the spirit and connotation of the appended claims.
[0033] Figure 1 It is a flowchart of a multimodal human skeleton behavior recognition method proposed in an embodiment of the present disclosure.
[0034] Among them, it should be noted that the executor of the multimodal human skeleton behavior recognition method of this embodiment is a multimodal human skeleton behavior recognition device, which can be implemented by software and / or hardware. The device can be configured in a computer device, and the computer device can include but is not limited to a terminal, a server, etc. For example, the terminal can be a mobile phone, a handheld computer, etc.
[0035] like Figure 1 As shown, the multimodal human skeleton behavior recognition method includes:
[0036] S101: Obtain a skeleton sequence to be identified of a target human skeleton and a prompt word set, wherein the prompt word set includes a plurality of candidate prompt texts, and the candidate prompt texts are associated with candidate behavior recognition results.
[0037] The target human skeleton may refer to a human skeleton for which behavior recognition is to be performed.
[0038] Among them, the skeleton sequence can refer to a data sequence that captures and represents the positions of various key parts (usually joints) of a human body or object in the time dimension. These joint point data are usually used to describe the dynamic changes of human body posture, motion trajectory or certain actions. The skeleton sequence to be identified refers to the skeleton sequence to be used for behavior recognition in the embodiment of the present disclosure. The skeleton sequence to be identified may include multiple frames of skeleton data, and each frame of skeleton data is used to indicate the position information of different joint points in the skeleton.
[0039] The prompt word set refers to a set consisting of multiple candidate prompt texts. The candidate prompt word texts can be used to describe the behavior to be identified in the embodiment of the present disclosure. For example, the candidate prompt word text can be "This is a fighting action."
[0040] The candidate behavior recognition result refers to the behavior that may be recognized in the embodiment of the present disclosure, such as "fighting".
[0041] Optionally, in some embodiments, the candidate prompt text is generated based on the following method: obtaining business scenario requirement information; determining multiple behavior tags to be identified based on the business scenario requirement information; if the behavior tag to be identified belongs to a historical tag set, then generating a corresponding candidate prompt text based on the behavior tag to be identified, wherein the historical tag set includes multiple historical behavior tags, and the historical behavior tags are obtained through training of a skeleton behavior recognition model; if the behavior tag to be identified does not belong to the historical tag set, then obtaining a behavior description text corresponding to the behavior tag to be identified, and generating a candidate prompt text based on the behavior description text. In this way, it can be ensured that the obtained candidate prompt text is suitable for personalized application scenarios, and the accuracy of the description of user behavior by the candidate prompt text is improved.
[0042] The skeleton behavior recognition model refers to a model for skeleton behavior recognition obtained based on training data in the embodiment of the present disclosure. In the embodiment of the present disclosure, the historical behavior label can also be obtained by training a skeleton behavior recognition encoder, which is not limited.
[0043] Among them, the business scenario requirement information can be used to describe the relevant requirements for behavior recognition in the application scenarios of the embodiments of the present disclosure.
[0044] The to-be-identified behavior label may be used to identify the to-be-identified behavior in the embodiments of the present disclosure.
[0045] For example, when the business scenario requirement information indicates that the current scenario needs to identify behaviors such as smoking and fighting, "smoking" and "fighting" can be used as labels of the behaviors to be identified.
[0046] The historical tag set refers to a set of historical behavior tags that have been trained by the behavior skeleton recognition model in the embodiment of the present disclosure.
[0047] The behavior description text can be used to describe the relevant features of the behavior. For example, when the behavior tag to be identified is "wandering", the corresponding behavior description text can be "this is the action of a person walking back and forth in an area".
[0048] In the disclosed embodiment, when the skeleton sequence to be identified and the prompt word set of the target human skeleton are obtained, multi-modal data support can be provided for the skeleton behavior recognition process.
[0049] S102: Feature encoding is performed on the skeleton sequence to be identified to obtain a spatiotemporal embedding feature vector.
[0050] The spatiotemporal embedding feature vector refers to a vector obtained by feature encoding the skeleton sequence to be identified and used to describe the relevant features of the skeleton sequence to be identified.
[0051] In the embodiment of the present disclosure, when feature encoding is performed on the skeleton sequence to be identified to obtain a spatiotemporal embedding feature vector, any human skeleton behavior recognition network can be selected as a feature encoder according to the needs of the application scenario.
[0052] S103: Perform text encoding on the candidate prompt text to obtain a text embedding feature matrix.
[0053] The text embedding feature matrix refers to the feature matrix obtained by text encoding the candidate prompt text in the embodiment of the present disclosure.
[0054] In the embodiment of the present disclosure, when text encoding is performed on the candidate prompt text to obtain a text embedding feature matrix, any applicable text encoder may be selected for encoding, for example, BERT may be selected as the text encoder.
[0055] It is understandable that, in the embodiment of the present disclosure, when encoding the skeleton sequence to be recognized and the candidate prompt text, the dimensional consistency of the output spatiotemporal embedding feature vector and the text embedding feature matrix should be ensured to facilitate subsequent similarity calculation.
[0056] S104: Determine the similarity between the spatiotemporal embedding feature vector and the text embedding feature matrix.
[0057] Among them, similarity can be used to describe the similarity between the spatiotemporal embedding feature vector and the text embedding feature matrix.
[0058] For example, in the embodiments of the present disclosure, methods such as cosine similarity or Euclidean distance can be used to measure the similarity between the spatiotemporal embedding feature vector and the text embedding feature matrix, and there is no limitation to this.
[0059] In the embodiment of the present disclosure, when determining the similarity between the spatiotemporal embedding feature vector and the text embedding feature matrix, a reliable judgment basis can be provided for the subsequent determination of the target behavior recognition result that matches the skeleton sequence to be recognized.
[0060] S105: Determine a target behavior recognition result that matches the skeleton sequence to be recognized from the multiple candidate behavior recognition results according to the similarity.
[0061] Among them, the target behavior recognition result can be used to indicate the behavior related to the skeleton sequence to be recognized.
[0062] That is, in the embodiment of the present disclosure, after determining the similarity between the spatiotemporal embedding feature vector and the text embedding feature matrix, a target behavior recognition result that matches the skeleton sequence to be recognized can be determined from multiple candidate behavior recognition results based on the similarity.
[0063] In this embodiment, by obtaining the skeleton sequence to be identified and the prompt word set of the target human skeleton, wherein the prompt word set includes multiple candidate prompt texts, and the candidate prompt texts are associated with the candidate behavior recognition results; feature encoding is performed on the skeleton sequence to be identified to obtain a spatiotemporal embedding feature vector; text encoding is performed on the candidate prompt text to obtain a text embedding feature matrix; the similarity between the spatiotemporal embedding feature vector and the text embedding feature matrix is determined; and the target behavior recognition result matching the skeleton sequence to be identified is determined from multiple candidate behavior recognition results according to the similarity. In this way, business needs can be quickly met with few or even zero samples, and multimodal data can further improve the accuracy of the method.
[0064] Figure 2 It is a flowchart of a multimodal human skeleton behavior recognition method proposed in another embodiment of the present disclosure.
[0065] like Figure 2 As shown, the multimodal human skeleton behavior recognition method includes:
[0066] S201: Obtain a skeleton sequence to be identified of a target human skeleton and a prompt word set, wherein the prompt word set includes a plurality of candidate prompt texts, and the candidate prompt texts are associated with candidate behavior recognition results.
[0067] The description of S201 can be specifically referred to the above embodiment, which will not be repeated here.
[0068] S202: Normalize the skeleton sequence to be identified to obtain a target skeleton sequence.
[0069] The target skeleton sequence refers to a skeleton sequence obtained after normalizing multiple skeleton structures in the skeleton sequence to be identified.
[0070] In the embodiment of the present disclosure, when normalizing the skeleton sequence to be identified, the skeleton sequence to be identified can be input into a pre-trained machine learning model to obtain a corresponding target skeleton sequence, or the skeleton sequence to be identified can be normalized based on a third-party normalization processing device to obtain a target skeleton sequence, and there is no limitation on this.
[0071] Optionally, in some embodiments, when normalizing the skeleton sequence to be identified to obtain the target skeleton sequence, a skeleton behavior identification coordinate system may be constructed; target joints in the target human skeleton may be determined; and the skeleton sequence to be identified may be normalized based on the skeleton behavior identification coordinate system and the target joints to obtain the target skeleton sequence. Thus, the normalization of the skeleton sequence to be identified may be accurately and quickly achieved based on the skeleton behavior identification coordinate system and the target joints, so as to analyze and compare the position information of the joints corresponding to different time sequences in the target skeleton sequence.
[0072] The skeleton behavior recognition coordinate system refers to a coordinate system used to recognize skeleton behaviors in the embodiments of the present disclosure. The skeleton behavior recognition coordinate system can be a two-dimensional coordinate system or a three-dimensional coordinate system, which is not limited.
[0073] Among them, the target joint point refers to the joint point used for normalizing the skeletons of different time sequences in the implementation of the present disclosure. For example, the spinal joints in the human body trunk can be selected as the target joint points, so that the spinal joints in the skeletons corresponding to different time sequences are located at the origin of the coordinate system; the hip joints and spinal joints can be selected as the target joint points, so that the hip joints and spinal joints in the skeletons corresponding to different time sequences are on the same coordinate axis in the coordinate system; and the right shoulder joint and the left shoulder joint can be selected as the target joint points, so that the right shoulder joint and the left shoulder joint in the skeletons corresponding to different time sequences are also on the same coordinate axis in the coordinate system.
[0074] In the embodiment of the present disclosure, when the skeleton sequence to be identified is normalized to obtain a target skeleton sequence, the clarity of the indication of the user behavior characteristics by the obtained target skeleton sequence can be effectively improved.
[0075] S203: Perform feature encoding on the target skeleton sequence to obtain a spatiotemporal embedding feature vector.
[0076] That is to say, in the embodiment of the present disclosure, after obtaining the skeleton sequence to be identified of the target human skeleton, the skeleton sequence to be identified can be normalized to obtain the target skeleton sequence; the target skeleton sequence can be feature encoded to obtain the spatiotemporal embedding feature vector. Thus, the standardization of the target skeleton sequence can be effectively improved by normalizing the skeleton sequence to be identified, thereby effectively improving the indication effect of the obtained spatiotemporal embedding feature vector on the user behavior characteristics.
[0077] S204: Perform text encoding on the candidate prompt text to obtain a text embedding feature matrix.
[0078] The description of S204 can be specifically referred to the above embodiment, which will not be repeated here.
[0079] S205: Normalize the spatiotemporal embedding feature vector.
[0080] That is, in the embodiments of the present disclosure, the spatiotemporal feature vector may be normalized (such as L2 normalization or Min-Max normalization) so that the vector is more stable in the subsequent similarity calculation and the weight of each feature is relatively balanced.
[0081] S206: Normalize the text embedding feature matrix.
[0082] That is, in the embodiments of the present disclosure, the text embedding matrix can be normalized similar to the processing process of the above-mentioned spatiotemporal features to ensure that the text features and spatiotemporal features are comparable in the subsequent similarity calculations.
[0083] S207: Calculate and determine similarity based on the normalized spatiotemporal embedding feature vector and the normalized text embedding feature matrix.
[0084] That is, in the embodiment of the present disclosure, after obtaining the spatiotemporal embedding feature vector and the text embedding feature matrix, the spatiotemporal embedding feature vector can be normalized; the text embedding feature matrix can be normalized; and similarity can be calculated and determined based on the normalized spatiotemporal embedding feature vector and the normalized text embedding feature matrix. Thus, the reliability of the similarity calculation process can be effectively improved by normalizing the spatiotemporal embedding feature vector and the text embedding feature matrix.
[0085] S208: Determine the maximum value among multiple similarities.
[0086] S209: Taking the candidate behavior recognition result corresponding to the maximum value as the target behavior recognition result.
[0087] Optionally, in some embodiments, a similarity threshold may be set, and then candidate behavior recognition results with a plurality of similarities greater than the similarity threshold are all used as target behavior recognition results, and there is no limitation on this.
[0088] That is, in the embodiment of the present disclosure, after calculating and determining the similarity based on the normalized spatiotemporal embedding feature vector and the normalized text embedding feature matrix, the maximum value among multiple similarities can be determined, and the candidate behavior recognition result corresponding to the maximum value is used as the target behavior recognition result. In this way, the accuracy and reliability of the obtained target behavior recognition result can be guaranteed.
[0089] In this embodiment, the target skeleton sequence is obtained by normalizing the skeleton sequence to be identified; the target skeleton sequence is feature encoded to obtain the spatiotemporal embedding feature vector. Therefore, the standardization of the target skeleton sequence can be effectively improved by normalizing the skeleton sequence to be identified, thereby effectively improving the indication effect of the obtained spatiotemporal embedding feature vector on the user behavior characteristics. The spatiotemporal embedding feature vector is normalized; the text embedding feature matrix is normalized; and the similarity is calculated and determined based on the normalized spatiotemporal embedding feature vector and the normalized text embedding feature matrix. Therefore, the reliability of the similarity calculation process can be effectively improved by normalizing the spatiotemporal embedding feature vector and the text embedding feature matrix. By determining the maximum value among multiple similarities; the candidate behavior recognition result corresponding to the maximum value is used as the target behavior recognition result. Therefore, the accuracy and reliability of the obtained target behavior recognition result can be guaranteed.
[0090] Based on the above embodiments, the present disclosure proposes a human skeleton behavior recognition method based on multimodal data. In view of the problems existing in the existing human skeleton behavior recognition methods and the high difficulty in obtaining real abnormal behavior data, the present disclosure adopts multimodal data and comparative learning to achieve human skeleton behavior recognition tasks that meet project and task requirements by pre-training on public data sets, and training with a small amount of real scene data or zero samples.
[0091] like Figure 3 As shown, Figure 3 The following is a flow chart of a multimodal human skeleton behavior recognition method proposed in the present disclosure, wherein the specific steps are as follows:
[0092] 1_1) Normalize the input skeleton sequence;
[0093] 1_2) Generate prompt words according to the actions that need to be identified in the business scenario;
[0094] 2_1) Feature encoding of the normalized skeleton sequence into a spatiotemporal embedding feature vector;
[0095] 2_2) Encode the generated Prompt into a text embedding feature matrix;
[0096] 3_1) Normalize the spatiotemporal embedding feature vector;
[0097] 3_2) Normalize the text embedding feature matrix;
[0098] 4) Calculate the similarity between the spatiotemporal embedding feature vector and the text embedding feature matrix;
[0099] 5) Based on the similarity calculated by S4, output the behavior category with the greatest similarity.
[0100] When the behavior label to be identified is "fighting", Figure 4 As shown, Figure 4 It is an inference flow chart of the multimodal human skeleton behavior recognition method proposed in the present disclosure.
[0101] Step S1_1:
[0102] like Figure 5 As shown, Figure 5It is a joint point bitmap proposed according to the present disclosure. When the input skeleton sequence is normalized, the embodiment of the present disclosure may rotate and translate the joint point coordinates of each frame in the acquired skeleton sequence data so that the spinal joint (joint point No. 2) corresponds to the origin of the coordinate system; the line between the hip joint (joint point No. 1) and the spinal joint (joint point No. 2) is parallel to the z-axis; the line between the right shoulder joint (joint point No. 9) and the left shoulder joint (joint point No. 5) is parallel to the x-axis.
[0103] Step S1_2:
[0104] Design prompts based on the categories that need to be detected in the project or task. For example, you need to identify actions such as fighting, falling, wandering, and smoking. However, for categories that exist in the training data, you can design prompts as "This is xx action". For categories that do not exist in the training data, you need to use existing actions or simple descriptions to express them. For example: There are "walking", "fighting", and "falling" actions in the training data. Now there are project requirements for "fighting" and "wandering". You can design prompts as "This is a fighting action" and "This is the action of a person walking back and forth in an area". Send the designed prompts to the word segmenter.
[0105] Step S2_1:
[0106] The normalized skeleton sequence is sent to the skeleton sequence encoder for feature encoding. Here, any human skeleton behavior recognition network can be selected as the feature encoder. It only needs to ensure that the dimension of the feature is 1*512 when the feature is finally output. If the original network output is not 1*512, it is necessary to add 1×1 convolution or other feature dimensionality reduction (dimensionality increase) methods to the original network to transform the feature dimension.
[0107] The formula is as follows:
[0108] f st =Encoder sq (x sq ) st ∈R 1×512 1)
[0109] Step S2_2:
[0110] Send the word segmented Prompt in S1_2 to the text encoder. Here you can choose bert as the text encoder. The output text embedding matrix dimension is n*512, where n is the number of Prompt.
[0111] The formula is as follows:
[0112]
[0113] Step S3_1:
[0114] Normalize the spatiotemporal embedding feature vector output by S2_1.
[0115] The formula is as follows:
[0116]
[0117] Step S3_2:
[0118] Normalize the text embedding feature matrix output by S2_2.
[0119] The formula is as follows:
[0120]
[0121] Step S4:
[0122] Based on the normalized spatiotemporal feature vector and the normalized text embedding feature matrix, the similarity between the current skeleton sequence and all prompts is calculated.
[0123]
[0124] Step S5:
[0125] The one with the highest similarity is selected as the recognition category of the current input skeleton sequence.
[0126] The formula is as follows:
[0127] class=argmax(f similarity ) (6)
[0128] In summary of the above embodiments, the present invention proposes a new multimodal human skeleton behavior recognition method, which uses two modal data, skeleton sequence data and semantic data, for training and reasoning at the same time, to further improve the recognition accuracy of the model. The present invention is an open set human skeleton behavior recognition method, which can modify different prompt words so that the model can detect different actions, rather than being limited to the categories of data in the training samples, so that it can be applied to different projects and task scenarios. The present invention only needs to be pre-trained on a public data set, and can quickly respond to project and task requirements when little or no real scene data is required.
[0129] The multimodal human skeleton behavior recognition method proposed in the present invention uses multimodal data for model inference to improve model accuracy compared to existing methods. At the same time, this method is an open set recognition method, which can recognize categories that are not in the training data. Therefore, the demand for data samples is much smaller than that of traditional closed set recognition methods. It can respond quickly and be applicable to more new projects and business scenarios, meeting more project and business needs.
[0130] Figure 6 It is a structural schematic diagram of a multimodal human skeleton behavior recognition device proposed in an embodiment of the present disclosure.
[0131] like Figure 6 As shown, the multimodal human skeleton behavior recognition device 60 includes:
[0132] An acquisition module 601 is used to acquire a skeleton sequence to be identified of a target human skeleton and a prompt word set, wherein the prompt word set includes a plurality of candidate prompt texts, and the candidate prompt texts are associated with candidate behavior recognition results;
[0133] The first encoding module 602 is used to perform feature encoding on the skeleton sequence to be identified to obtain a spatiotemporal embedding feature vector;
[0134] The second encoding module 603 is used to perform text encoding on the candidate prompt text to obtain a text embedding feature matrix;
[0135] A first determination module 604 is used to determine the similarity between the spatiotemporal embedding feature vector and the text embedding feature matrix;
[0136] The second determination module 605 is used to determine a target behavior recognition result that matches the skeleton sequence to be recognized from a plurality of candidate behavior recognition results according to similarity.
[0137] It should be noted that the above explanation of the multimodal human skeleton behavior recognition method is also applicable to the multimodal human skeleton behavior recognition device of this embodiment, and will not be repeated here.
[0138] In this embodiment, by obtaining the skeleton sequence to be identified and the prompt word set of the target human skeleton, wherein the prompt word set includes multiple candidate prompt texts, and the candidate prompt texts are associated with the candidate behavior recognition results; feature encoding is performed on the skeleton sequence to be identified to obtain a spatiotemporal embedding feature vector; text encoding is performed on the candidate prompt text to obtain a text embedding feature matrix; the similarity between the spatiotemporal embedding feature vector and the text embedding feature matrix is determined; and the target behavior recognition result matching the skeleton sequence to be identified is determined from multiple candidate behavior recognition results according to the similarity. In this way, business needs can be quickly met with few or even zero samples, and multimodal data can further improve the accuracy of the method.
[0139] Figure 7 A block diagram of an exemplary computer device suitable for implementing embodiments of the present disclosure is shown. Figure 7 The computer device 12 shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0140] like Figure 7As shown, the computer device 12 is in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 that connects various system components (including the system memory 28 and the processing unit 16).
[0141] The bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor or a local bus using any of a variety of bus structures. For example, these architectures include but are not limited to Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus and Peripheral Component Interconnection (PCI) bus.
[0142] The computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computer device 12, including volatile and non-volatile media, removable and non-removable media.
[0143] The memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the storage system 34 may be used to read and write non-removable, non-volatile magnetic media ( Figure 7 Not shown, often called a "hard drive").
[0144] although Figure 7Not shown, a disk drive for reading and writing a removable non-volatile disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing a removable non-volatile optical disk (e.g., a compact disc read only memory (Compact Disc Read Only Memory; hereinafter referred to as: CD-ROM), a digital versatile disc read only memory (Digital Video Disc Read Only Memory; hereinafter referred to as: DVD-ROM) or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 via one or more data medium interfaces. The memory 28 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of the various embodiments of the present disclosure.
[0145] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in the memory 28, such program modules 42 including but not limited to an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment. The program modules 42 generally perform the functions and / or methods of the embodiments described in the present disclosure.
[0146] The computer device 12 may also communicate with one or more external devices 14 (e.g., keyboard, pointing device, display 24, etc.), one or more devices that enable a human body to interact with the computer device 12, and / or any device that enables the computer device 12 to communicate with one or more other computing devices (e.g., network card, modem, etc.). Such communication may be performed via an input / output (I / O) interface 22. In addition, the computer device 12 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with other modules of the computer device 12 via a bus 18. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the computer device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0147] The processing unit 16 executes various functional applications and data processing by running the programs stored in the system memory 28, such as implementing the multimodal human skeleton behavior recognition method mentioned in the above embodiment.
[0148] In order to implement the above embodiments, the present disclosure also proposes a non-temporary computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the multimodal human skeleton behavior recognition method proposed in the above embodiments of the present disclosure is implemented.
[0149] In order to implement the above embodiments, the present disclosure further proposes a computer program product. When an instruction processor in the computer program product executes, the multimodal human skeleton behavior recognition method proposed in the above embodiments of the present disclosure is executed.
[0150] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the following claims.
[0151] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
[0152] It should be noted that, in the description of the present disclosure, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of the present disclosure, unless otherwise specified, the meaning of "plurality" is two or more.
[0153] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code that includes one or more executable instructions for implementing the steps of a specific logical function or process, and the scope of the preferred embodiments of the present disclosure includes alternative implementations in which functions may not be performed in the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present disclosure belong.
[0154] It should be understood that the various parts of the present disclosure can be implemented in hardware, software, firmware or a combination thereof. In the above-mentioned embodiments, multiple steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented by any one of the following technologies known in the art or their combination: a discrete logic circuit having a logic gate circuit for implementing a logic function for a data signal, a dedicated integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0155] A person skilled in the art may understand that all or part of the steps in the method for implementing the above-mentioned embodiment may be completed by instructing related hardware through a program, and the program may be stored in a computer-readable storage medium, which, when executed, includes one or a combination of the steps of the method embodiment.
[0156] In addition, each functional unit in each embodiment of the present disclosure may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0157] The storage medium mentioned above can be a read-only memory, a magnetic disk or an optical disk, etc.
[0158] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representation of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner.
[0159] Although the embodiments of the present disclosure have been shown and described above, it is to be understood that the above embodiments are exemplary and are not to be construed as limitations of the present disclosure. A person skilled in the art may change, modify, replace and vary the above embodiments within the scope of the present disclosure.
Claims
1. A multimodal human skeleton behavior recognition method, characterized in that: include: Obtaining a skeleton sequence to be identified of a target human skeleton and a prompt word set, wherein the prompt word set includes a plurality of candidate prompt texts, and the candidate prompt texts are associated with candidate behavior recognition results; Performing feature encoding on the skeleton sequence to be identified to obtain a spatiotemporal embedding feature vector; Performing text encoding on the candidate prompt text to obtain a text embedding feature matrix; Determining the similarity between the spatiotemporal embedding feature vector and the text embedding feature matrix; A target behavior recognition result matching the skeleton sequence to be recognized is determined from the plurality of candidate behavior recognition results according to the similarity.
2. The method according to claim 1, characterized in that The candidate prompt text is generated based on the following method: Obtain business scenario demand information; Determine multiple behavior tags to be identified based on the business scenario requirement information; If the behavior tag to be identified belongs to a historical tag set, generating the corresponding candidate prompt text based on the behavior tag to be identified, wherein the historical tag set includes a plurality of historical behavior tags, and the historical behavior tags are obtained by training a skeleton behavior recognition model; If the behavior tag to be identified does not belong to the historical tag set, the behavior description text corresponding to the behavior tag to be identified is obtained, and the candidate prompt text is generated according to the behavior description text.
3. The method according to claim 1, characterized in that The step of encoding the skeleton sequence to be identified to obtain a spatiotemporal embedding feature vector includes: Normalizing the skeleton sequence to be identified to obtain a target skeleton sequence; Feature encoding is performed on the target skeleton sequence to obtain the spatiotemporal embedding feature vector.
4. The method according to claim 3, characterized in that The step of normalizing the skeleton sequence to be identified to obtain a target skeleton sequence includes: Construct a skeleton behavior recognition coordinate system; Determining a target joint point in the target human skeleton; The skeleton sequence to be identified is normalized based on the skeleton behavior identification coordinate system and the target joint points to obtain the target skeleton sequence.
5. The method according to claim 1, characterized in that Determining the similarity between the spatiotemporal embedding feature vector and the text embedding feature matrix includes: Normalizing the spatiotemporal embedding feature vector; Normalizing the text embedding feature matrix; The similarity is calculated and determined based on the normalized spatiotemporal embedding feature vector and the normalized text embedding feature matrix.
6. The method according to claim 1, characterized in that The step of determining a target behavior recognition result that matches the skeleton sequence to be recognized from the plurality of candidate behavior recognition results according to the similarity comprises: determining a maximum value among a plurality of the similarities; The candidate behavior recognition result corresponding to the maximum value is used as the target behavior recognition result.
7. A multimodal human skeleton behavior recognition device, characterized in that: include: An acquisition module, used to acquire a skeleton sequence to be identified of a target human skeleton and a prompt word set, wherein the prompt word set includes a plurality of candidate prompt texts, and the candidate prompt texts are associated with candidate behavior recognition results; A first encoding module is used to perform feature encoding on the skeleton sequence to be identified to obtain a spatiotemporal embedding feature vector; A second encoding module is used to perform text encoding on the candidate prompt text to obtain a text embedding feature matrix; A first determination module is used to determine the similarity between the spatiotemporal embedding feature vector and the text embedding feature matrix; The second determination module is used to determine a target behavior recognition result that matches the skeleton sequence to be recognized from the plurality of candidate behavior recognition results according to the similarity.
8. A computer device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that: in, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 6.
10. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 6.