Digitization of human motion data for semantic search

By employing a multi-component architecture and vector search technology, the problem of existing devices being unable to fully capture the details of physical movement and exercise has been solved, enabling efficient digital and automated searching of complex postures and expressions, and promoting the preservation and learning of traditional culture.

CN121866596APending Publication Date: 2026-04-14SONY GROUP CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-16
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing electronic devices are inadequate in capturing the complex and delicate body postures, movements, and expressions associated with physical exercise such as classical dance, yoga, or martial arts. They cannot fully capture the details of these movements, making it difficult to automate training tasks.

Method used

It adopts a multi-component architecture, including a performer separation and localization network (SL network), a multi-head spatiotemporal physical attribute recognition network (ARN), and a time tracking and integration network. By receiving image sequences, it extracts localized instances and determines multiple performance attributes, such as emotion attributes, posture attributes, and gesture attributes, generates metadata, and performs vector search in the performance database to identify unregistered or registered postures.

Benefits of technology

It enables efficient and automated digitization of complex body postures and expressions, provides detailed search capabilities, promotes the preservation and learning of traditional cultural assets, and allows for research without expert guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

An electronic device and method for digitization of a semantic search of human motion data are provided. An electronic device receives an input sequence of images depicting a physical performance associated with a physical athletic exercise, and extracts a localized instance of a performer associated with the physical performance from image frames of the sequence of image frames. The electronic device determines a plurality of performance attributes associated with the performer based on applying an attribute recognition network on the localized instance. The plurality of performance attributes includes an emotion attribute, a gesture attribute, and a gesture attribute. The electronic device searches a performance database based on the plurality of performance attributes to generate search results. The search results specify whether the localized instance represents an unregistered gesture or a registered gesture of the physical exercise.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Cross-reference to related applications / incorporation by reference

[0002] This application claims priority to U.S. Application No. 18 / 806,393, filed August 15, 2024, with the United States Patent and Trademark Office, which in turn claims priority to U.S. Provisional Patent Application Serial No. 63 / 579,057, filed August 28, 2023. The aforementioned patent applications are incorporated herein by reference in their entirety. Technical Field

[0003] Various embodiments of this disclosure relate to motion data. More specifically, various embodiments of this disclosure relate to electronic devices and methods for digitizing human motion data for semantic search. Background Technology

[0004] Advances in motion capture technology have led to the development of electronic devices capable of capturing human movement and images in digital format. These devices typically include features such as pose estimation, tracking models, and image or video processing. They enable the manipulation of various aspects of images or videos to extract pose and depth from the collected data. Currently, pose estimation, tracking models, and image or video processing are used to analyze and extract meaningful information (such as pose and expression) from digitized image or video data. However, it is important to note that these technologies typically only manage 17-32 major joints of the human body. For example, some known devices can wirelessly capture movement using various sensor systems and cameras, individually. However, when it comes to disciplines of body movement, such as classical dance forms, yoga, or martial arts, which often involve complex and subtle movements of the eyes, fingers, and facial expressions, existing electronic devices may be insufficient. These devices are limited in their ability to capture the complex and delicate body postures, movements, and expressions associated with classical Indian dance forms, yoga, or martial arts.

[0005] By comparing the described system with some aspects of this disclosure, the limitations and disadvantages of conventional and traditional methods will become apparent to those skilled in the art, as set forth more fully in the claims below and with reference to the accompanying drawings. Summary of the Invention

[0006] An electronic device and method for digitizing human motion data for semantic search are provided, which is substantially as shown in at least one figure and / or described in conjunction with at least one figure, as set forth more fully in the claims.

[0007] These and other features and advantages of this disclosure can be understood by referring to the following detailed description of the disclosure and a review of the accompanying drawings, in which the same reference numerals refer to the same parts throughout the drawings. Attached Figure Description

[0008] Figure 1 This is a block diagram illustrating an exemplary network environment for digitizing human motion data for semantic search according to embodiments of the present disclosure.

[0009] Figure 2 This illustrates an embodiment according to the present disclosure. Figure 1 A block diagram of an exemplary electronic device.

[0010] Figure 3 This is a diagram illustrating an exemplary application of a separation and localization (SL) network according to embodiments of the present disclosure.

[0011] Figure 4 This is a diagram illustrating an exemplary scenario generated from metadata according to embodiments of this disclosure.

[0012] Figure 5 This is a diagram illustrating an exemplary scenario of an emotion recognition network according to an embodiment of the present disclosure.

[0013] Figure 6 This is a diagram illustrating an exemplary scenario of a pose recognition network according to an embodiment of the present disclosure.

[0014] Figure 7 This is a diagram illustrating an exemplary scenario of a gesture recognition network according to an embodiment of the present disclosure.

[0015] Figure 8 This is a diagram illustrating an exemplary scenario of vector search according to embodiments of the present disclosure.

[0016] Figure 9 This is a diagram illustrating an exemplary scenario displayed by a user equipment according to an embodiment of the present disclosure.

[0017] Figure 10 This is a flowchart illustrating the operation of an exemplary method for digitizing human motion data for semantic search according to embodiments of the present disclosure. Detailed Implementation

[0018] The implementations described below can be found in electronic devices and methods for digitizing human motion data for semantic search. Exemplary aspects of this disclosure may provide an electronic device that receives an input image sequence depicting a physical performance associated with physical exercise (e.g., Figure 1The input image sequence 118). Next, the electronic device extracts localized instances of the performer associated with the physical performance from the image frames of the image frame sequence. Then, the electronic device determines multiple performance attributes associated with the performer based on an attribute recognition network applied to the localized instances. Furthermore, the multiple performance attributes include emotional attributes, which include location information associated with a set of facial features of the performer, posture attributes associated with the performer's body, and gesture attributes associated with each of the performer's hands. Additionally, the electronic device searches a performance database based on the multiple performance attributes to generate a first search result, wherein the first search result specifies an unregistered or registered posture representing a physical movement exercise from the localized instance.

[0019] Physical movement exercises, such as classical dance forms, martial arts, and yoga, typically involve complex, subtle, and nuanced movements of the eyes, fingers, and facial expressions. Due to the limited ability of existing electronic devices to manage only major body joints, they may not be able to fully capture such delicate combinations of expressions. Therefore, these devices may not adequately capture the complex and nuanced body postures, movements, and expressions associated with physical movement exercises, such as those in classical dance forms, martial arts, and yoga, such as Indian classical dance. The limitations of current systems make it difficult to automate the task of dance instructors because the granularity of physical movement exercises may not be sufficiently captured. For example, while body positioning and movement recorded by pose estimation, tracking models, or image processing may be crucial for automating training programs for physical movement exercises such as classical dance forms, existing recordings may not capture the required precision. Therefore, training programs based on current pose estimation, tracking models, or image processing may not be comprehensive enough to train participants according to the complex postures and movements of physical movement exercises. Failure to train participants according to these details may lead to suboptimal training results.

[0020] The electronic device disclosed herein can provide automatic, efficient, and robust digitization of human motion data for semantic search. To achieve this, the electronic device can receive a sequence of image frames associated with a performer engaging in physical movement exercises. In some embodiments, the electronic device can extract localized instances to apply an attribute recognition network to determine multiple performance attributes, such as emotional attributes, posture attributes, and gesture attributes. Gesture attributes can represent root handprints from multiple root handprints associated with the physical movement exercise. Emotional attributes can be represented by various permutations and combinations of a set of facial features. Posture attributes can be represented by keypoint detection, where the detected keypoints can correspond to all body joints of the performer. The electronic device can generate metadata based on multiple performance attributes. The electronic device can then search a performance database for the generated metadata to identify whether the localized instances represent unregistered or registered postures of the physical movement exercise (e.g., a dance form). This identification can be based on the calculation of vector similarity scores between a vector search query (based on metadata generation) and assets in the performance database.

[0021] Therefore, electronic devices can be configured to capture physical assets, such as complex and nuanced body postures, movements, and expressions. These captured physical assets can be converted into digital assets and tagged to provide search capabilities to users or participants. For example, electronic devices can be configured to detect and identify specific body parts and their configurations, and then aggregate the identified parts and configurations for a comprehensive analysis and representation of bodily performance.

[0022] In some implementations, electronic devices can employ a multi-component architecture, including performer separation and localization networks, multi-head spatiotemporal physical attribute recognition networks, and time tracking and integration networks. This architecture enables the capture of a wide range of nuances required for the digitization of physical movement training, such as traditional dance forms, including but not limited to tracking of limb coordination, gaze, eyebrow movements, lip movements, and emotions.

[0023] The publicly available system offers advantages in preserving and promoting traditional cultural assets. For example, it enables the digitization of existing content, including monocular video, with a high level of detail. The system can also facilitate the creation of searchable sub-assets, allowing students to study specific aspects of indigenous art forms without direct expert guidance.

[0024] Figure 1 This is a block diagram illustrating an exemplary network environment for digitizing human motion data for semantic search, according to embodiments of this disclosure. Reference Figure 1The diagram illustrates a network environment 100. Network environment 100 may include an electronic device 102, a server 108, a database 110, a user device 114, and a communication network 116. The electronic device 102, database 110, and user device 114 may be configured to communicate with each other via one or more communication networks (such as communication network 116). Furthermore, the electronic device 102 may be configured to store, for example, a separation and localization (SL) network 104 and an attribute recognition network (ARN) 106. The electronic device 102 may receive an input image sequence 118 for motion asset creation. Figure 1 The diagram also shows a set of static assets 112A and a set of moving assets 112B that can be stored in database 110.

[0025] Electronic device 102 may include suitable logic, circuitry, interfaces, and / or code, which may be configured to receive an input image sequence 118 depicting a physical performance associated with physical exercise (e.g., Indian dance forms such as Kathak, yoga, or martial arts). Electronic device 102 can receive image frames from the image frame sequence (e.g., ... Figure 3 Extract the localized instance of performer 118A associated with physical performance from the image frame at position 302 in the image (e.g., ...). Figure 3 (Localization instance at location 308). Electronic device 102 can determine multiple performance attributes associated with performer 118A based on the application of ARN 106 to the localization instance. Here, the multiple performance attributes include emotional attributes (e.g., Figure 4 The emotional attribute 410 includes: positional information associated with a set of facial features of the performer 118A; and posture attributes (e.g., Figure 4 The posture attribute 416), which is associated with the body of the performer 118A; and the gesture attribute (e.g., Figure 4 The gesture attribute 422 is associated with each hand of the performer 118A. The electronic device 102 can search a performance database (e.g., database 110 or performance databases 804A-804N) based on multiple performance attributes to generate a first search result (e.g., ...). Figure 8 The first search result shown is 814A-814N). The first search result can specify a localized instance representing an unregistered or registered posture of body movement exercise. Examples of electronic devices 102 may include, but are not limited to, computing devices, smartphones, cellular phones, mobile phones, gaming devices, mainframe machines, servers, computer workstations, machine learning devices (enabled or hosted with, for example, computing resources, memory resources, and network resources), wearable devices with performance attribute detection sensors, depth sensor devices, and / or consumer electronics (CE) devices.

[0026] In one embodiment, registered or unregistered postures may be referred to as mudras in physical movement exercises. As used herein, the term "mudra" in physical movement exercises can refer to symbolic hand positions used to express meaning, emotion, and rhythmic experience in performance. A mudra can essentially be a set of symbolic movements (or "movements") with a defined physical form, name, and visible shape. The meaning of a mudra is created through the movement of the hand, rather than through a traceable linguistic presence. The physical form of a mudra can be sculpted by proxies of hand movements. In performance, mudras, as the energy of hand movement, interpret meaning and experience a variety of emotions generated in the body. In Indian physical movement exercises, mudras are key dance steps, encompassing root mudras associated with gestures, postures associated with the body, and emotions expressed through the performer's facial expressions.

[0027] In one embodiment, a registered pose for physical exercise can be referred to as a static asset in static asset set 112A or a motion asset in motion asset set 112B. Furthermore, registered poses can be stored in a performance database or database 110, for example. An unregistered pose for physical exercise can be an asset that does not exist in the performance database.

[0028] For example, electronic device 102 may be able to digitize and tag assets (such as static assets 112A or such assets in a collection of motion assets 112B) to generate searchable digital assets through which a user can master physical movement exercises (such as dance forms) without the need for guidance from a trainer. For example, electronic device 102 may capture the granularity of one or more static assets 112A, such as granularity associated with fingers, eye gaze, eyebrow movement, lip movement, and emotions. For example, the granularity associated with fingers may relate to the relative position of each joint of the finger.

[0029] The SL network 104 can be configured to receive image frames from an input image sequence 118 and generate localization information associated with performer 118A present in the image frames of the input image sequence 118. The localization information can be further used to extract localized instances of performer 118A. According to an embodiment, the SL network 104 can segment the received image frames into multiple parts for identification and analyze the parts from each segment. The SL network 104 can provide pixel-wise details for each part based on the application of a clustering algorithm (such as K-means clustering). The SL network 104 can divide the image frame into multiple regions. Each region within the multiple regions can be separated from each other using multiple boundaries associated with each region within the multiple regions. (e.g., as...) Figure 3(As shown in 306). The boundaries of each region can be associated with further processing of the region for image classification or object detection. For example, the SL network 104 can perform instance segmentation or semantic segmentation to generate a label for each pixel in a plurality of pixels in an image frame. Furthermore, pixel labeling can help determine a set of pixels associated with an object type (e.g., a dancer).

[0030] In one embodiment, the SL network 104 can identify the location of each region among multiple regions of a segmented image frame. The location of each region can be identified by drawing a boundary around each region. The SL network 104 can apply a localization algorithm to predict a set of four numbers to draw boundaries around the multiple regions. Such numbers can define the coordinates (e.g., bounding box coordinates) of the localized region, defining its height and width.

[0031] For example, video (such as input image sequence 118) can be received by electronic device 102 to identify the coordinates of each object in the image frames. Electronic device 102 can apply SL network 104 to each image frame of the received video. SL network 104 can further segment each object in the image frames based on the application of clustering algorithms. Furthermore, SL network 104 can identify the coordinates of each object in the image frames based on the application of localization algorithms.

[0032] In one embodiment, the SL network 104 may be a neural network. The neural network may be a system of computational networks or artificial neurons arranged as nodes in multiple layers. Multiple layers of the neural network may include an input layer, one or more hidden layers, and an output layer. Each layer in the multiple layers may include one or more nodes (or artificial neurons, e.g., represented by circles). The outputs of all nodes in the input layer may be coupled to at least one node in the hidden layer. Similarly, the input of each hidden layer may be coupled to the output of at least one node in other layers of the neural network. The output of each hidden layer may be coupled to the input of at least one node in other layers of the neural network. Nodes in the final layer may receive input from at least one hidden layer to output a result. The number of layers and the number of nodes in each layer may be determined by the hyperparameters of the neural network. Such hyperparameters may be set on the training dataset before, during, or after training of the neural network.

[0033] Each node in a neural network can correspond to a mathematical function (e.g., a sigmoid function or a rectified linear unit) with a set of parameters that are adjustable during the network's training. This set of parameters can include, for example, weight parameters, regularization parameters, etc. Each node can use the mathematical function to compute an output based on one or more inputs from nodes in other layers of the neural network (e.g., the previous layer). All or some nodes in a neural network can correspond to the same or different mathematical functions.

[0034] In training a neural network, one or more parameters of each node can be updated based on whether the output of the final layer for a given input (from the training dataset) matches the correct result of the neural network-based loss function. This process can be repeated for the same or different inputs until the minimum of the loss function is achieved, and the training error is minimized. Several methods for training are known in the art, such as gradient descent, stochastic gradient descent, batch gradient descent, gradient boosting, metaheuristics, etc.

[0035] ARN 106 can be configured to be applied to a localized instance of performer 118A to determine multiple performance attributes associated with performer 118A. ARN 106 can be a computational network or artificial neuron system to identify specific attributes or features of performer 118A in the input image sequence 118 and learn the identified attributes. For example, ARN 106 can be a convolutional neural network (CNN), a visual transformer, or a variant thereof to extract multiple performance attributes of performer 118A. For example, the extracted multiple performance attributes may include emotion attributes, posture attributes, gesture attributes, foot attributes, etc. ARN 106 extracts features from each of the multiple performance attributes, and further, the extracted features are processed into a multi-label classification task. Here, multi-label classification helps ARN 106 map the input to a binary vector. For example, the input may include video, images, high-level meshes, etc. Alternatively, in some embodiments, the artificial neurons in ARN 106 can be implemented using a combination of hardware and software.

[0036] In one embodiment, the ARN 106 can receive a localized instance of performer 118A from the SL network 104. The localized instance of performer 118A can be determined by applying the SL network 104 to image frames received from the input image sequence 118. The details of applying the SL network 104 to the image frames are described above in the description of the SL network 104.

[0037] In one embodiment, ARN 106 is a multi-head neural network that includes multiple attribute networks, such as an emotion recognition network, a pose recognition network, and a gesture recognition network, to classify and learn multiple performance attributes. ARN 106 can be trained using supervised, semi-supervised, or unsupervised learning techniques. During training, each of the multiple performance attribute networks can learn to classify parts of a localized instance to determine multiple performance attributes, such as facial attributes, pose attributes, and gesture attributes. Each determined performance attribute can then be combined into a weighted sum of features based on weights associated with the determined performance attribute. Furthermore, the weighted sum of features (such as facial features, body poses, and gestures) can be used to train ARN 106 for each performance attribute.

[0038] Server 108 may include suitable logic, circuitry, and interfaces, and / or code, which may be configured to receive an input image sequence 118 depicting a physical performance associated with physical exercise. Server 108 may receive a localized instance of performer 118A from electronic device 102. Server 108 may receive multiple performance attributes. Server 108 may receive search queries from electronic device 102. Server 108 may search a performance database (e.g., database 110) based on the search queries received from electronic device 102. Furthermore, server 108 may obtain a first search result and provide the first search result to electronic device 102.

[0039] Server 108 can be implemented as a cloud server and can perform operations through web applications, cloud applications, HTTP requests, repository operations, file transfers, etc. Other example implementations of Server 108 may include, but are not limited to, database servers, file servers, web servers, media servers, application servers, mainframe servers, machine learning servers (enabled or hosted with, for example, computing resources, storage resources, and network resources), or cloud computing servers.

[0040] In at least one embodiment, server 108 can be implemented as multiple distributed cloud-based resources using a variety of techniques well known to those skilled in the art. Those skilled in the art will understand that the scope of this disclosure is not limited to the implementation of server 108 and electronic device 102 as two separate entities. In some embodiments, the functionality of server 108 may be wholly or at least partially incorporated into electronic device 102 without departing from the scope of this disclosure. In some embodiments, server 108 may host database 110. Alternatively, server 108 may be decoupled from database 110 and communicatively coupled to database 110.

[0041] Database 110 may include suitable logic, interfaces, and / or code, and may be configured to store a static asset set 112A and a motion asset set 112B. Database 110 may originate from data in a relational or non-relational database, or from a collection of comma-separated value (CSV) files in a conventional or big data storage. Database 110 may be stored or cached on a device, such as a server (e.g., server 108) or electronic device 102. The device storing database 110 may be configured to receive queries for static asset set 112A and motion asset set 112B from electronic device 102 or server 108. For example, a query may include a request for static assets from static asset set 112A and motion assets from motion asset set 112B. In another embodiment, a query may include a natural language query or an image-based query from user device 114. Database 110 may receive user input from user device 114 (e.g., ...). Figure 9 (User input at position 902 in the database). In response, the device of database 110 can be configured to generate a vector search query and calculate a similarity score between the vector search query and each of the multiple assets. Furthermore, the device of database 110 can be configured to provide search results, which can be generated based on the calculated similarity scores, to electronic device 102, server 108, or user device 114. In an embodiment, the search results may include the top k assets in a sequence based on the calculated similarity scores of each of the multiple assets stored in database 110.

[0042] In some embodiments, database 110 may be hosted on multiple servers located in the same or different locations. Operation of database 110 may be performed using hardware, including processors, microprocessors (e.g., for performing or controlling one or more operations), field-programmable gate arrays (FPGAs), or application-specific integrated circuits (ASICs). In other cases, database 110 may be implemented using software. In an exemplary embodiment, database 110 may be a vector database.

[0043] The static asset collection 112A may include metadata for multiple body movement exercises. Metadata may include tags, emotional attributes, posture attributes, gesture attributes, and other information associated with the image frames of the body movement exercises. For example, the metadata of an image frame might represent a tag for the Krishn pose in the Katak dance, where the emotional attributes of the Krishn pose are calm and smiling (this could be as follows). Figure 5 Encoded in the character as shown at 506A), the posture attributes of the Krishn gesture can be a slight tilt of the body, and a gesture close to the face indicates the flute (similar to...). Figure 6The pose attribute at 606A in the text), and the gesture attribute of the Krishn pose can be the Tripataka gesture with both hands (similar to...). Figure 7 (Gesture attribute at 706A in the metadata). In addition, metadata can also include the relationship between emotion attributes, posture attributes, and gesture attributes.

[0044] The motion asset set 112B may include metadata for multiple physical movement exercises. The metadata may include tags, emotional attributes, posture attributes, gesture attributes, duration, and other information associated with multiple image frames of the physical movement exercise. For example, the metadata for motion assets may include the static asset set 112A, which may represent the transition from the Krishn pose in Kathak to the Radha pose in Kathak.

[0045] User device 114 may include suitable logic, circuitry, and interfaces configured to provide user input related to physical exercise to electronic device 102. User input may include at least one of natural language queries or image-based queries. Furthermore, electronic device 102 may be configured to generate a vector search query based on a neural network trained contrastively on the user input, and to calculate a similarity score between the second vector search query and each of a plurality of assets (such as a set of static assets 112A and a set of moving assets 112B). Electronic device 102 may also be configured to generate search results comprising the top k assets among the plurality of assets for which a calculated similarity score exceeds a threshold score. Additionally, user device 114 may be configured to receive responses from electronic device 102 to at least one of natural language queries or image-based queries. User device 114 may be controlled by electronic device 102 to display the responses. Responses may include, for example, source image frames depicting a registered pose in at least one of the top k assets. Examples of user equipment 114 may include, but are not limited to, computing devices, smartphones, cellular phones, mobile phones, gaming devices, mainframe machines, servers, computer workstations, machine learning devices (enabled or hosted with, for example, computing resources, storage resources, and networking resources), wearable devices with performance attribute detection sensors, depth sensor devices, and / or consumer electronics (CE) devices.

[0046] For example, user equipment 114 may be configured to receive user input from a user and provide it to electronic device 102. User input may include natural language queries, such as queries for the Raudra Nataraja mudra associated with the Indian classical physical exercise Bharatnatyam (e.g., ...). Figure 9As shown in 902 (in the diagram). The circuitry of electronic device 102 can be configured to perform a vector search for Raudra Natraja's handprint on database 110 to calculate a similarity score (e.g., ...). Figure 9 (As shown in 904). Furthermore, based on the similarity score between a search query for Raudra Natraja's handprint and each asset in either a set of static assets 112A or a set of moving assets 112B in database 110, electronic device 102 can be configured to generate search results and display them on user device 114 (e.g., ...). Figure 9 (As shown in 906).

[0047] Communication network 116 may include a communication medium through which electronic devices 102 and server 108 can communicate with each other. Communication network 116 may be either a wired or wireless connection. Examples of communication network 116 may include, but are not limited to, the Internet, cloud networks, cellular or wireless mobile networks (such as LTE and 5G New Radio (NR)), satellite communication systems (using, for example, low Earth orbit satellites), Wi-Fi networks, personal area networks (PANs), local area networks (LANs), or metropolitan area networks (MANs). Various devices in network environment 100 may be configured to connect to communication network 116 according to various wired and wireless communication protocols. Examples of such wired and wireless communication protocols may include, but are not limited to, at least one of Transmission Control Protocol and Internet Protocol (TCP / IP), User Datagram Protocol (UDP), Hypertext Transfer Protocol (HTTP), File Transfer Protocol (FTP), ZigBee, EDGE, IEEE 802.11, Li-Fi, IEEE 802.16, IEEE 802.11s, IEEE 802.11g, multi-hop communication, wireless access point (AP), device-to-device communication, cellular communication protocols, and Bluetooth (BT) communication protocols.

[0048] In operation, electronic device 102 can receive an input image sequence 118 depicting a physical performance associated with a physical exercise. The physical exercise can be a form of classical Indian dance, yoga, or martial arts, or various dance forms such as hip-hop, samba, belly dance, ballet, tap dance, and so on. An SL network 104 can be applied to the input image sequence 118 to generate localization information associated with performers 118A. As an example, electronic device 102 can receive a Kuchipudi dance video containing multiple image frames. Electronic device 102 can apply the SL network 104 to at least one of the image frames to separate multiple performers present in the image frames. Furthermore, the SL network 104 can identify the position and coordinates of each performer present in the image frames and can predict boundaries or borders around each performer to generate localization information associated with each performer. Details regarding performer 118A are provided, for example, in... Figure 3 Further details are provided below.

[0049] Electronic device 102 can extract localized instances of performer 118A associated with physical performance from image frames of input image sequence 118. Localized instances of performer 118A can be extracted based on localization information. Here, a localized instance of a performer (such as performer 118A) can refer to specific identifiers and spatial information of a dance performer within a given frame of input image sequence 118. Such information can include the precise location, boundaries, and segmentation of the performer's image, thereby distinguishing them from the background and other performers. Each performer's localized instance can be associated with a tag. For example, tags can be associated with head, hands, feet, body, legs, eyes, nose, etc. In one embodiment, localized instances of each of a plurality of performers can also be stored on database 110. Electronic device 102 or server 108 can retrieve localized instances based on a combination of metadata and tags associated with each localized instance. For example, a localized instance of performer 118A can be identified based on head configuration to determine the performer 118A's emotion. Details about the localized instance are provided, for example, in... Figure 3 Further details are provided below.

[0050] Electronic device 102 can determine multiple performance attributes associated with performer 118A based on an attribute recognition network applied to localized instances. These performance attributes may include emotional attributes, posture attributes, gesture attributes, etc. Here, emotional attributes may include location information associated with a set of facial features of performer 118A, posture attributes associated with performer 118A's body, and gesture attributes associated with each hand of performer 118A. Details regarding these multiple performance attributes are provided, for example, in... Figure 4 Further details are provided below.

[0051] Electronic device 102 can search a performance database based on multiple performance attributes to generate a first search result. The search result specifies whether the localized instance represents an unregistered or registered pose of a body movement exercise. The search on the performance database (e.g., performance databases 804A-804N) can be a vector search based on a vector-based distance metric (such as vector dot product) between the search query and dance information stored on the performance database. In one embodiment, performance databases 804A-804N can be a static asset database that may store a static asset set 112A. In another embodiment, the performance database can be a motion asset database that may store a motion asset set 112B. Details regarding the first search result are provided in, for example... Figure 9 Further details are provided below.

[0052] Electronic device 102 provides an automated, efficient, and robust digitization process for human motion data to enable semantic search capabilities. Through this process, electronic device 102 can capture the physical assets of a performance, including complex and nuanced body postures, movements, and expressions essential to physical exercise. These captured physical assets can then be converted into digital assets. Furthermore, electronic device 102 can be equipped to tag such digital assets, thereby enhancing their searchability and accessibility for users or participants who may wish to study or reference specific aspects of physical exercise. This tagging feature can be particularly beneficial for educational and preservation purposes, as it allows for detailed study and analysis of subtle elements that define traditional physical exercise.

[0053] Figure 2 According to embodiments of this disclosure Figure 1 A block diagram of an exemplary electronic device. Combined with... Figure 1 To explain the components Figure 2 . refer to Figure 2 The diagram illustrates a block diagram 200 of an electronic device 102. The block diagram 200 of the electronic device 102 may include circuitry 202, a memory 204, an input / output (I / O) device 206, and a network interface 208. The memory 204 may store an ARN 106. The input / output (I / O) device 206 may include a display device 210. Furthermore, the ARN 106 may include an emotion recognition network 106A, a gesture recognition network 106B, and a gesture recognition network 106C.

[0054] Circuit 202 may include suitable logic, circuitry, and / or interfaces that can be configured to execute program instructions associated with different operations to be performed by electronic device 102. Circuit 202 may include one or more processing units that can be implemented as individual processors. In embodiments, one or more processing units may be implemented as an integrated processor or processor cluster that collectively performs the functions of one or more dedicated processing units. Circuit 202 may be implemented based on a variety of processor technologies known in the art. Examples of implementing circuit 202 may be x86-based processors, graphics processing units (GPUs), reduced instruction set computing (RISC) processors, application-specific integrated circuit (ASIC) processors, complex instruction set computing (CISC) processors, microcontrollers, central processing units (CPUs), and / or other control circuitry.

[0055] Memory 204 may include suitable logic, circuitry, interfaces, and / or code, which may be configured to store one or more instructions to be executed by circuitry 202. The one or more instructions stored in memory 204 may be used to perform different operations of circuitry 202 (and / or electronic device 102). Memory 204 may also be configured to store ARN 106, emotion recognition network 106A, gesture recognition network 106B, gesture recognition network 106C, and contrastive training neural network 212. Training data (e.g., static asset set 112A, motion asset set 112B) for ARN 106, emotion recognition network 106A, gesture recognition network 106B, gesture recognition network 106C, and contrastive training neural network 212 is provided by user device 114. Memory may store ARN 106, emotion recognition network 106A, gesture recognition network 106B, gesture recognition network 106C, and contrastive training neural network 212 in a file format. Memory 204 may be persistent memory, non-persistent memory, or a combination thereof. Examples of implementing memory 204 may include, but are not limited to, random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), hard disk drive (HDD), solid-state drive (SSD), CPU cache and / or security digital card (SD card).

[0056] ARN 106 can be a computational network or artificial neural network system to identify specific attributes or features of performer 118A in the input image sequence 118 and learn the identified attributes. The artificial neural network system used by ARN 106 is a CNN to extract multiple performance attributes of performer 118A. For example, the extracted multiple performance attributes may include emotional attributes, posture attributes, gesture attributes, foot attributes, etc. ARN 106 extracts features from each of the multiple performance attributes, and further, the extracted features are processed into a multi-label classification task. Here, multi-label classification helps ARN 106 map the input to a binary vector. For example, the input may include video, images, high-level meshes, etc. Alternatively, in some embodiments, the artificial neurons in ARN 106 may be implemented using a combination of hardware and software.

[0057] In this embodiment, ARN 106 is a multi-head neural network comprising an emotion recognition network 106A, a pose recognition network 106B, and a gesture recognition network 106C. Furthermore, ARN 106 can be applied to localized instances of performer 118A to determine multiple performance attributes.

[0058] The Emotion Recognition Network 106A is a network that uses machine learning (ML) algorithms and CNNs to analyze, interpret, and classify human emotions. It can analyze the location information of facial feature sets and body language based on body posture and gestures to recognize human emotions. The Emotion Recognition Network 106A can be enabled to recognize human emotions in real time based on real-time data input. Furthermore, it can use context-aware techniques to improve the accuracy of emotion detection even with noisy data input. These context-aware techniques may include CAER-Net, which analyzes facial feature sets and contextual information about joints.

[0059] In an embodiment, the emotion recognition network 106A can be applied to a localized instance of performer 118A to detect the state of each facial feature in a facial feature set. Furthermore, the emotion recognition network 106A encodes the detected state of each facial feature in the facial feature set as a character sequence. This character sequence can be included in the location information of each facial feature in the facial feature set. Additionally, the facial feature set includes eye gaze features, eyebrow features, nose features, lip features, and forehead twitching features.

[0060] The pose recognition network 106B can be a system that uses machine learning (ML) algorithms and artificial intelligence (AI) to analyze and recognize human poses. Key points of the human body in an image frame are combined with AI to estimate and recognize human poses. In another embodiment, the pose recognition network 106B can be sensor-based to collect positional information of each key point of the human body.

[0061] In an embodiment, the pose recognition network 106B can be applied to localized instances of a performer to detect key points on the body, and based on the detected key points, detect dance poses associated with body movement exercises. Here, pose attributes indicate dance poses. For example, the pose recognition network 106B can be a Kinect, a hybrid of fuzzy logic and machine learning, a convolutional neural network (CNN), etc.

[0062] The gesture recognition network 106C can be a machine learning-based system for recognizing and interpreting gestures. The gesture recognition network 106C receives input data such as localization information associated with image frames to recognize hands in the image frames. Furthermore, the gesture recognition network 106C detects the finger joints of a person in the image frames to recognize gestures. Additionally, the gesture recognition network 106C classifies the recognized gestures of a person in the image frames.

[0063] In an embodiment, the gesture recognition network 106C can be applied to a localized instance of a performer to detect the finger joints in each of the performer's hands; and gesture information is determined based on the position of the detected finger joints. Here, the gesture information indicates a symbolic hand position that expresses meaning, emotion, or rhythmic experience at a particular moment in a physical performance. For example, the gesture recognition network 106C can be a transformer-based gesture recognition engine, a fine-tuned CNN, a deep CNN, a deep Q-network, a 3D-CNN model, etc.

[0064] The contrast-trained neural network 212 can correspond to a system that can be trained to recognize similarities between data points. For example, data points can be metadata and data stored in database 110. Similarity can be identified using a contrastive loss function that generates similar embeddings for similar inputs and dissimilar embeddings for dissimilar inputs. For example, a vector similarity score can be calculated between an input vector search query and each of multiple assets in a performance database (e.g., performance databases 804A-804N). The input vector search query can be generated based on the application of metadata to the contrast-trained neural network 212. Here, metadata can be generated for image frames based on multiple performance attributes.

[0065] In an embodiment, the contrast-trained neural network 212 can be applied to metadata of image frames. Here, the metadata can be generated based on multiple performance attributes. The contrast-trained neural network 212 can be applied to the metadata to generate a first vector search query (e.g., at least one of first vector search queries 812A-812N associated with frames 0 to N). The generated first vector search query can be input to a performance database to compute a vector similarity score between the input first vector search query and each of multiple assets in the performance database.

[0066] In an embodiment, the contrast-trained neural network 212 allows the electronic device 102 to uniquely capture multiple performance attributes of highly expressive physical movement exercises (such as Kuchipudi, Kathakali, Bharatnatyam, etc.).

[0067] In an embodiment, during contrastive training of a neural network, one or more parameters of each node of the neural network can be updated based on whether the output of the final layer for a given input (from the training dataset) matches the correct result according to the loss function of the neural network. This process can be repeated for the same or different inputs until the minimum value of the loss function can be achieved, and the training error can be minimized. Several methods for training are known in the art, such as gradient descent, stochastic gradient descent, batch gradient descent, gradient boosting, metaheuristics, etc.

[0068] The contrastive training neural network 212 may include electronic data, which may be implemented as a software component of an application executable, for example, on electronic device 102. The contrastive training neural network 212 may rely on libraries, external scripts, or other logic / instructions for execution by a processing device (such as circuitry 202). The contrastive training neural network 212 may include code and routines configured to enable the computing device (such as circuitry 202) to perform one or more operations for uniquely capturing multiple performance attributes of highly expressive physical movement exercises (such as Kuchipudi, Kathakali, Bharatnatyam, etc.). Alternatively or additionally, the contrastive training neural network 212 may be implemented using hardware including a processor, a microprocessor (e.g., for performing or controlling the execution of one or more operations), a field-programmable gate array (FPGA), or an application-specific integrated circuit (ASIC). Alternatively, in some embodiments, the neural network may be implemented using a combination of hardware and software.

[0069] I / O device 206 may include suitable logic, circuitry, interfaces, and / or code that can be configured to receive input and provide output based on the received input. For example, I / O device 206 may receive a sequence of input images depicting a physical performance. I / O device 206 may also be configured to display or render multiple performance attributes associated with a performer based on the application of an attribute recognition network to localized instances (such as handprints). I / O device 206 may include display device 210. Examples of I / O device 206 may include, but are not limited to, a display (e.g., a touchscreen), a keyboard, a mouse, a joystick, a microphone, or a speaker. Examples of I / O device 206 may also include Braille I / O devices, such as Braille keyboards and Braille readers.

[0070] Network interface 208 may include suitable logic, circuitry, interfaces, and / or code that can be configured to facilitate communication between electronic device 102 and server 108 via communication network 116. Network interface 208 can be implemented using various known technologies to support wired or wireless communication between electronic device 102 and communication network 116. Network interface 208 may include, but is not limited to, antennas, radio frequency (RF) transceivers, one or more amplifiers, tuners, one or more oscillators, digital signal processors, codec chipsets, subscriber identity module (SIM) cards, or local buffer circuitry.

[0071] Network interface 208 can be configured to communicate wirelessly with networks such as the Internet, intranets, wireless networks, cellular telephone networks, wireless local area networks (LANs), or metropolitan area networks (MANs). Wireless communication can be configured to use one or more of multiple communication standards, protocols, and technologies, such as Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), Wideband Code Division Multiple Access (W-CDMA), Long Term Evolution (LTE), Fifth Generation (5G) New Radio (NR), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi (such as IEEE 802.11a, IEEE 802.11b, IEEE 802.11g, or IEEE 802.11n), Voice over Internet Protocol (VoIP), Li-Fi, Wi-MAX, email, instant messaging, and Short Message Service (SMS) protocols.

[0072] Display device 210 may include suitable logic, circuitry, and interfaces that can be configured to display or render search results in response to vector search queries. Display device 210 may be a touchscreen that allows a user to provide user input via the display device 210. The touchscreen may be at least one of a resistive touchscreen, a capacitive touchscreen, or a thermal touchscreen. Display device 210 may be implemented using a variety of known technologies, such as, but not limited to, liquid crystal display (LCD), light-emitting diode (LED), plasma display, or organic LED (OLED) display technologies, or at least one of other display devices. According to embodiments, display device 210 may refer to the display screen of a head-mounted device (HMD), smart glasses device, see-through display, projection-based display, electrochromic display, or transparent display. Figure 4 The neural network 212, trained in contrast, is used to generate the first vector query, for example in... Figure 8 In the middle, various operations of circuit 202 are used to implement a multi-head neural network for the recognition of multiple attributes.

[0073] Figure 3 An exemplary application of a separation and localization (SL) network according to embodiments of the present disclosure is illustrated. Combined with information from... Figure 1 and Figure 2 To explain the elements Figure 3 . refer to Figure 3 Figure 300 illustrates an exemplary scenario depicting the process of SL network 304. In this case, SL network 304 is shown being applied to an image frame to separate at least one performer (such as performer 118A) from the image frame and generate localization information for producing localized instances.

[0074] The input image sequence 118 can be received by circuit 202. Circuit 202 can select at least one image frame (such as image frame 302) from the input image sequence 118. Image frame 302 depicts a physical performance associated with physical movement exercises. In this case, image frame 302 may include multiple performers associated with the physical performance. Examples of physical movement exercises associated with physical performance may include, but are not limited to, forms of Indian classical dance such as Bharatanatyam, Kuchipudi, Kathakali, Kathak, Odissi, Sattriya, Manipuri, and Mohiniyattam; forms of yoga such as Kundalini Yoga, Vinyasa Yoga, Ashtanga Yoga, Bikram Yoga, and surya namaskar; and martial arts such as kung fu, judo, karate, wing chun, kendo, and tai chi.

[0075] In some respects, the multiple performers in image frame 302 may be in different poses relative to each other. For example, in image frame 302, all the performers shown may be in different poses, each depicting a different aspect of a physical exercise. In other respects, all the performers in the selected image frames may be in the same pose (not shown).

[0076] Circuit 202 can be configured to apply SL network 304 to a selected image frame 302. SL network 304 can analyze image frame 302 pixel by pixel and identify the boundaries of each performer among multiple performers, as shown in localization example 306. Furthermore, SL network 304 can be configured to isolate at least one performer for further analysis.

[0077] The SL network 304 can be further configured to generate localization information. Based on this localization information, multiple localization instances 308 can be extracted. For example, the multiple localization instances 308 may include localization instances that identify and label the boundaries of the performer's head, localization instances that identify and label the boundaries of the performer's pose, and localization instances that identify and label the boundaries of the performer's hands (as shown in 308). These localization instances 308 can be identified and labeled with boundaries for further analysis of the performer in the selected image frame 302. Localization instances around the head can be processed for emotion recognition, localization instances around the pose can be processed for pose recognition, and localization instances around the hands can be processed for gesture recognition.

[0078] For example, the SL network 304 can separate each performer from multiple performers in the selected image frame 302 and further provide localization information for each performer. This localization information may include bounding boxes as shown at 308.

[0079] Figure 4 This is a diagram illustrating an exemplary application of metadata generation based on embodiments of this disclosure. Combined with... Figure 1 , Figure 2 and Figure 3 To explain the components Figure 4 . refer to Figure 4 Figure 400 illustrates an exemplary scenario for generating metadata. In this scenario, ARN 106 is shown as a localized instance applied to each performer in an image frame to generate metadata associated with each performer (such as performer 118A).

[0080] In Figure 400, ARN 106 can be applied to the extracted localized instances of the performer to determine multiple performance attributes. To determine each of the multiple performance attributes, ARN 106 can be used, which includes an emotion recognition network 404, a pose recognition network 412, and a gesture recognition network 418.

[0081] A localized instance 402 of the performer can be received to generate metadata associated with the performer 118A, and this metadata can be associated with localized instance 308. Furthermore, based on the application of SL network 304, the localized instance of performer 402 can be identified and localized (e.g., within a bounding box). For example, the localized instance of performer 402 can be associated with at least one localized instance (such as...). Figure 3 The bounding box of the localized instance in 308 is associated with it.

[0082] An emotion recognition network 404 can be applied to one of the localized instances of performer 402 for feature detection 406. Feature detection 406 may include the detection of facial feature sets, gaze detection 408A, eyebrow detection 408B, nose detection 408C, lip detection 408D, forehead detection 408E, etc. Furthermore, each feature from feature detection 406 can be combined to correspond to an emotion attribute 410. For example, the emotion attribute 410 can be determined by tracking the performer's gaze, eyebrow movements, lip movements, and overall facial expressions.

[0083] The pose recognition network 412 can be applied to one of the localized instances of the performer 402 for keypoint detection 414. Keypoints can include points on various parts of the body. Based on keypoint detection 414, pose attributes 416 can be determined.

[0084] The gesture recognition network 418 can be applied to one of the localized instances of the performer 402 for knuckle detection 420 of each finger associated with both hands. Gesture attributes 422 of each hand can be generated based on the detected knuckles.

[0085] ARN 106 can generate metadata (such as metadata generation 424) for image frames based on multiple determined performance attributes (including emotion attribute 410, posture attribute 416, and gesture attribute 422). For example, metadata can be generated based on various combinations and permutations of multiple performance attributes. The generated metadata can uniquely identify a specific posture, divine representation, or situation depicted by a localized instance in an image frame associated with physical movement exercise.

[0086] Figure 5 This is a diagram illustrating an exemplary application of an emotion recognition network according to embodiments of this disclosure. Combined with... Figure 1 , Figure 2 , Figure 3 and Figure 4 To explain the components Figure 5 . refer to Figure 5 Figure 500 illustrates an exemplary scenario for generating emotional attributes. In this exemplary scenario, ARN 106 is shown as a face 506 applied to a localized instance of a performer (such as performer 118A) in an image frame to generate emotional attributes.

[0087] In Figure 500, the emotion recognition network 504 can receive a localized instance 502 of the performer associated with the head of performer 118A. The emotion recognition network 504 can use machine learning (ML) algorithms and convolutional neural networks (CNNs) to analyze, interpret, and classify the set of facial expressions of performer 118A.

[0088] In some aspects, the emotion recognition network 504 can be applied to a localized instance of performer 118A to detect the state of each facial feature in the facial feature set. Furthermore, the emotion recognition network 504 can use circuitry 202 to encode the state of each detected facial feature into a character sequence (as shown in Table 506A). This character sequence can be included in the location information of each facial feature. The facial feature set can include, for example, eye gaze features, eyebrow features, nose features, lip features, and forehead features. In some embodiments, the facial feature set can be combined to determine emotional attributes. Different permutations and combinations of the facial feature set can represent multiple emotional attributes. For example, eye gaze can be encoded as EGRWOLWO, as shown in Table 506A, which is provided in description 506B as “EmotionHead-Gaze-RightWideOpen-LeftWideOpen”. Similarly, eyebrows can be encoded as EERSSLS, as shown in Table 506A. Similar to eye gaze encoding and eyebrow encoding, the emotion recognition network 504 can encode other facial features in the facial feature set. Furthermore, each facial feature in the facial feature set can be encoded as a character sequence to determine the emotional attribute 508 associated with the performer 118A. For example, as Figure 5 As shown in 506, based on Table 506A and Description 506B, the emotional attribute 508 of performer 118A can be determined to be fear.

[0089] Figure 6 This is a diagram illustrating an exemplary application of a pose recognition network according to embodiments of the present disclosure. Figure 6 With from Figure 1 , Figure 2 , Figure 3 , Figure 4 and Figure 5 The components are explained in combination. (See reference.) Figure 6 Figure 600 illustrates an exemplary scenario depicting the generation of pose attributes. In this exemplary scenario, ARN 106 is shown as a localization instance 602 applied to a performer (e.g., performer 118A) in an image frame to generate pose attributes.

[0090] In Figure 600, the pose recognition network 604 can receive a localized instance 602 of the performer associated with the pose of the performer 118A's body. The pose recognition network 604 can use machine learning (ML) algorithms or artificial intelligence (AI) techniques to analyze and recognize human poses. In some aspects, key points of the performer in the image frame (such as...) Figure 6 As shown in 606A, it can be combined with AI technology to estimate and identify human posture.

[0091] Each keypoint of the human body can be represented relative to other keypoints. For example, keypoints can be detected based on a localized instance 602 of performer 118A (as shown in 606) and the relationships between keypoints (as shown in 606A). A pose recognition network 604 can generate pose attributes 608 based on the detected keypoints and the relationships between them. In this context, pose attributes can indicate a dance pose or a standing posture. For example, the detected keypoints can include identifiers of various body joints of performer 118A that collectively define the overall pose.

[0092] Figure 7 This is a diagram illustrating an exemplary application of a gesture recognition network according to embodiments of the present disclosure. Figure 7 With from Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 and Figure 6 The components are explained in combination. (See reference.) Figure 7 Figure 700 illustrates an exemplary scenario depicting the generation of gesture attributes 708. In this exemplary scenario, ARN 106 is shown as a localization instance 702 applied to a performer (e.g., performer 118A) in an image frame to generate gesture attributes.

[0093] The gesture recognition network 704 can receive localized instances of performer 702 associated with the performer's hands. The gesture recognition network 704 can use machine learning (ML) or artificial intelligence (AI) based algorithms to analyze and recognize the performer's gestures, as shown in 706. In some aspects, the performer's knuckles (as shown in 706A) can be used to recognize gestures. Furthermore, the gesture recognition network 704 can classify the recognized gestures of the performer in image frames.

[0094] In some embodiments, each of the performer's finger joints may be represented relative to other finger joints. For example, the finger joints of each hand may be detected based on the application of the gesture recognition network 704 to a localized instance of the performer 702. Furthermore, the gesture recognition network 704 may determine gesture information based on the position of the detected finger joints. In this context, gesture information may indicate symbolic hand positions that express meaning, emotion, or rhythmic experience at a particular moment in a physical performance.

[0095] Gesture attribute 708 can represent the root mudra among multiple root mudras associated with physical exercise. For example, gesture attribute 708 can include various mudras such as Tripataka, Ardhachandra, Mushti, Pataka, Mayura, Katakamukha, Padmakosha, Suchi, Chakra, etc.

[0096] Figure 8 This is a diagram illustrating an exemplary application of vector search according to embodiments of the present disclosure. Figure 8 With from Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 , Figure 6 and Figure 7 The components are explained in combination. (See reference.) Figure 8 Figure 800 illustrates an exemplary scenario depicting vector search for stationary and moving assets.

[0097] In some aspects, metadata (such as metadata 802A-802N) can be generated for image frames (such as frames 0 to N) based on multiple performance attributes. After generation, the metadata can be tagged. The metadata associated with each image frame can include information and details of multiple performance attributes of that frame. A contrast-trained neural network 212 can be applied to the metadata 802A-802N to generate first vector search queries 812A-812N. These first vector search queries 812A-812N can be used as input to performance databases 804A-804N. The performance databases 804A-804N can include multiple assets, such as a set of static assets 112A. These assets can include embedding information, metadata 802A-802N, and labels associated with the metadata 802A-802N. The embedding information can encode registered poses for body movement exercises. The metadata 802A-802N can be associated with each registered pose and the source image frame depicting that pose.

[0098] In some embodiments, the registered gesture may represent a handprint among multiple handprints in a physical exercise. The handprint may be a key dance step, which includes a root handprint associated with gesture attribute 422, a body-related posture based on posture attribute 416, and an emotion associated with emotion attribute 410.

[0099] Circuit 202 can be configured to calculate a vector similarity score between the generated first vector search queries 812A-812N and each asset in the performance database 804A-804N. Circuit 202 can then generate first search results 814A-814N based on vector similarity scores below a threshold. These first search results 814A-814N can indicate that metadata (such as metadata 802A-802N) represents an unregistered pose in a body movement exercise. Circuit 202 can be configured to check whether a static asset (e.g., static asset 808A-808N) exists in the performance database 804A-804N. Here, when a static asset exists in the performance database 804A-804N (e.g., static asset 808A-808N), the metadata is represented as a registered pose. Furthermore, when a static asset does not exist in the performance database 804A-804N, the metadata is represented as an unregistered pose. If the metadata indicates an unregistered pose, circuit 202 can update performance databases 804A-804N to include that metadata. The updated metadata can be referred to as a newly registered pose for body movement exercises (in...). Figure 8 (As shown in 810A-810N). Therefore, a newly registered pose can be used as a tag assigned to a previously unregistered pose.

[0100] For example, when metadata 802A-802N is identified as a unique combination of emotional attributes, posture attributes, and gesture attributes that have not been previously registered in performance database 804A-804N, circuit 202 can update performance database 804A-804N to register metadata 802A-802N.

[0101] In some respects, it is possible to perform for each performer across image frames (e.g.) Figure 3 The metadata of performers at position 302 in the image is combined and integrated to generate motion metadata (such as...). Figure 8 (As shown in 816). The motion metadata can be generated based on input from each image frame in the sequence (i.e., multiple newly registered poses of body movement exercises in performance databases 804A-804N). Circuit 202 can generate a third vector search query 824 by applying a contrastive training neural network 212 to the generated motion metadata. The third vector search query 824 can be used as input to performance database 818 to compute a vector similarity score between the third vector search query 824 and each asset in performance database 818. Performance database 818 may include multiple motion assets, such as motion asset set 112B. These motion assets may include embedding information, metadata, and associated labels. The embedding information may encode the registered motions of the body movement exercises. The metadata may be associated with each registered motion and a set of source image frames depicting that motion.

[0102] In one embodiment, performance databases 804A-804N may store static asset set 112A, and performance database 818 may store motion asset set 112B. Alternatively, in another embodiment, performance databases 804A-804N may store motion asset set 112B, and performance database 818 may store static asset set 112A.

[0103] Circuit 202 can calculate a vector similarity score between the generated third vector search query 824 and each motion asset in the performance database 818. It can then generate a third search result 826 based on vector similarity scores below a threshold. This third search result 826 can indicate that the motion metadata (as shown in 816) represents an unregistered motion for a physical exercise. Circuit 202 can be configured to check if a motion asset exists in the performance database 818 (e.g., motion asset 820 exists). Here, when a motion asset exists in the performance database 818 (e.g., motion asset 820 exists), the metadata is represented as a registered motion. Furthermore, when a static asset does not exist in the performance database 818, the metadata is represented as an unregistered motion. If the motion metadata represents an unregistered motion, circuit 202 can update the performance database 818 to include that motion metadata as a newly registered motion for a physical exercise (as shown in 822). Therefore, the newly registered motion can be used as a tag assigned to a previously unregistered motion.

[0104] For example, circuit 202 of electronic device 102 can update performance databases 804A-804N or 818 to store new assets, such as newly registered poses or newly registered movements. These databases can be updated when unregistered poses or movements are detected based on first search results 814A-814N or third search results 826, respectively.

[0105] In one exemplary embodiment, a set of non-fungible tokens (NFTs) can be generated for each new asset stored in performance databases 804A-804N. Alternatively, NFTs can be generated for each new motion asset stored in performance database 818, such as... Figure 8As shown in 824, NFTs can be tokenized using blockchain technology. The generation of NFTs can leverage blockchain technology, which ensures the authenticity and uniqueness of NFTs. Each NFT in the NFT set can be associated with metadata (such as metadata 802A-802N or metadata 816). Metadata can define a new asset or new motion asset and include information about physical performance, such as emotional attributes 410, posture attributes 416, or gesture attributes 422. The metadata associated with each NFT in the NFT set can enhance the value of each NFT and provide additional context for each NFT. Furthermore, NFTs can be allocated to users or participants who may create new assets or new motion assets. The allocation of NFTs can be performed through a secure authentication process that verifies the ownership of users or participants.

[0106] For example, a marketplace can be established for users to participate in and showcase their generated NFTs, which can then be exchanged for other applications. This marketplace could allow users to buy, sell, or exchange NFTs related to new assets or sports assets associated with physical performance, such as Bharatnatyam pose assets, martial arts pose assets, yoga pose assets, etc. Furthermore, smart contracts can be implemented to manage transactions and ownership transfers within the NFT marketplace. Smart contracts can ensure ownership and also guarantee that NFT transactions are executed securely and transparently. Additionally, a royalty mechanism can be integrated into the NFT marketplace, allowing users or performers to receive royalties for NFTs sold or traded. Royalties can incentivize users or performers to create and share new assets or sports assets, and also allow them to benefit from them.

[0107] Figure 9 This is a diagram illustrating an example scenario displayed by a user equipment according to an embodiment of the present disclosure. Figure 9 With from Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 and Figure 8 The components are explained in combination. (See reference.) Figure 9 Figure 900 shows an example scenario depicting the display of user device 114 for a search query.

[0108] Circuit 202 can be configured to receive user input via a search bar 902 associated with physical exercise. The user input may include at least one of a natural language query or an image-based query. Circuit 202 may generate a vector search query based on a contrast-trained neural network 212 applied to the user input. Circuit 202 may calculate a similarity score between the vector search query and each of a plurality of assets (such as static assets from a set 112A of static assets stored in a performance database (such as performance databases 804A-804N or performance database 818) or motion assets from a set 112B of motion assets stored in a performance database (such as performance databases 804A-804N or performance database 818). Circuit 202 may generate search results for the top k assets whose calculated similarity scores are higher than a threshold score. Circuit 202 can be configured to control user device 114 to display a response based on a second search result. The response may include a source image frame depicting a registered pose in at least one of the top k assets.

[0109] In some aspects, the display of user device 114 may include a search bar 902, UI element 904, preview UI element 906, a buy button 908, a generate new button 910, a sell button 912, and navigation arrows. The navigation arrows allow the user of user device 114 to navigate between multiple results, such as the top k assets. User device 114 may receive user input on the search bar 902, such as a query for the dance pose “Raudra Nataraja”. Circuit 202 may generate a second search result, as shown in UI element 904, including the top k assets from among the multiple assets. The display of user device 114 may include information such as a similarity score of 0.89, the tag “Nataraja pose”, the emotion attribute “Rudra roop (angry or shocked)”, the root handprint (gesture attribute) of “Pataka handprint on the left hand and Dola on the right hand”, and the pose attribute of “Abhinaya”. Preview UI element 906 may present one of the top k assets at a time, and the navigation arrows allow the user to browse multiple top k assets.

[0110] In some embodiments, user device 114 may be configured to display a second search result based on a sequence of similarity scores between a second vector search query and each asset in database 110.

[0111] In one embodiment, a purchase button 908 on the display of user device 114 allows a user to preview NFTs associated with assets in database 110 and purchase user input on search bar 902. User input on search bar 902 can be associated with NFTs associated with assets in database 110. A generate new button 910 on the display of user device 114 allows a user to generate new NFTs that may include new assets. New assets can be stored in database 110 after generation. A sell button 912 on the display of user device 114 allows a user to interact with NFTs associated with assets in database 110 and sell newly generated NFTs (new NFTs can be generated using the generate new button 910 on the display of user device 114). User input on search bar 902 can also be associated with newly generated NFTs and can be stored as assets in database 110.

[0112] In one embodiment, the display of user device 114 may be associated with a user-friendly interface that allows the user to have seamless navigation and intuitive features for managing NFTs (using the Generate New button 910 to create a new NFT or...). Figure 8 (824 NFTs were generated in the text). Alternatively, by associating NFTs with assets created by users using metadata associated with physical performance (such as Bharatnatyam pose images, yoga poses, martial arts stances, etc.), users or performers can establish verifiable ownership, showcase their talents, and potentially monetize digital assets in a secure and transparent manner.

[0113] Figure 10 This is a flowchart illustrating the operation of an example method for digitizing machine learning-based motion data according to embodiments of the present disclosure. Figure 10 With from Figure 1 , Figure 2 , Figure 3 , Figure 4 , Figure 5 , Figure 6 , Figure 7 , Figure 8 and Figure 9 The components are described in combination. (See reference.) Figure 10 The flowchart 1000 is shown. Flowchart 1000 may include operations from 1002 to 1010 and can be... Figure 1 Electronic device 102 or Figure 2 The circuit 202 is implemented. Flowchart 1000 can start from 1002 and proceed to 1004.

[0114] At 1004, an input image sequence 118 can be received. Circuit 202 can be configured to receive the input image sequence 118 depicting a physical performance associated with physical exercise. Details regarding the input image sequence 118 are provided, for example, in... Figure 1 Further description is provided at (location 118).

[0115] At position 1006, a localized instance 306 of the performer can be extracted. Circuit 202 can be configured to extract a localized instance of the performer 118A associated with the physical performance from image frames in an image frame sequence. Details regarding the extraction of localized instances are provided, for example, in... Figure 3 (exist Figure 3 Further description is provided in the 306 locations mentioned above.

[0116] At point 1008, multiple performance attributes (such as, Figure 4 (At locations 410, 416, and 422). Circuit 202 can be configured to determine multiple performance attributes associated with performer 118A based on the application of ARN 106 to the localized instance. At 1008A, the multiple performance attributes may include emotional attributes, posture attributes, and gesture attributes. Details regarding the performance attributes are as follows: Figure 4 Further details are provided below.

[0117] You can search at 1010. Figure 8 The performance database is located at 804A-804N or 818. Circuit 202 can be configured to search the performance database based on multiple performance attributes to generate search results, where the search results specify localized instances representing unregistered or registered poses of body movement exercises. Details about the performance database are provided, for example, in... Figure 8 (Further described in 804A-804N or 818). Control can be passed to the end.

[0118] Although flowchart 1000 is shown as discrete operations, such as 1004, 1006, 1008, 1008A, and 1010, this disclosure is not limited thereto. Therefore, in some embodiments, such discrete operations may be further divided into additional operations, combined into fewer operations, or eliminated, depending on the implementation, without departing from the essence of the disclosed embodiments.

[0119] Various embodiments of this disclosure may provide non-transitory computer-readable media and / or storage media having computer-executable instructions stored thereon, which are executable by a machine and / or a computer, to operate electronic devices (e.g., Figure 1 The electronic device 102). Such instructions can cause the electronic device 102 to perform operations, which may include receiving a sequence of input images depicting a physical performance associated with physical exercise (e.g., Figure 1The input image sequence 118). The operation may further include extracting localized instances of the performer associated with the physical performance from image frames in the image frame sequence. The operation may further include determining multiple performance attributes associated with the performer 118A based on applying an attribute recognition network to the localized instances. The operation may further include multiple performance attributes, including: emotional attributes, which include positional information associated with a set of facial features of the performer 118A; posture attributes associated with the body of the performer 118A; and gesture attributes associated with each hand of the performer 118A. The operation may further include searching a performance database based on the multiple performance attributes to generate search results, wherein the search results specify unregistered or registered postures representing physical movement exercises by the localized instances.

[0120] Exemplary aspects of this disclosure may provide an electronic device (such as...) Figure 1 An electronic device 102 includes circuitry (such as circuitry 202). Circuitry 202 can be configured to receive a sequence of input images depicting a physical performance associated with physical exercise (e.g., Figure 1 The input image sequence 118). Circuit 202 can be configured to extract localized instances of the performer associated with physical performance from image frames in the image frame sequence. Circuit 202 can be configured to determine multiple performance attributes associated with performer 118A based on applying an attribute recognition network to the localized instances. The multiple performance attributes include emotional attributes, which include positional information associated with a set of facial features of performer 118A, posture attributes associated with the body of performer 118A, and gesture attributes associated with each hand of performer 118A. Circuit 202 can be configured to search a performance database based on the multiple performance attributes to generate a first search result, wherein the search result specifies an unregistered or registered posture representing a physical movement exercise by the localized instance.

[0121] In one embodiment, physical exercise is Indian physical exercise.

[0122] In one embodiment, the circuit is also configured to apply a separation and localization (SL) network on the image frame to generate localization information associated with performer 118A, wherein localized instances of performer 118A are extracted based on the localization information.

[0123] In one embodiment, the attribute recognition network is a multi-head neural network, which includes an emotion recognition network, a pose recognition network, and a gesture recognition network.

[0124] In one embodiment, the attribute recognition network includes an emotion recognition network, and the circuitry is further configured to apply the emotion recognition network to a localized instance of performer 118A to detect the state of each facial feature in a set of facial features and to encode the detected state of each facial feature in the set of facial features into a character sequence, wherein the location information includes the character sequence for each facial feature in the set of facial features.

[0125] In one embodiment, a set of facial features includes eye gaze features, eyebrow features, nose features, lip features, and forehead twitching features.

[0126] In one embodiment, the attribute recognition network includes a pose recognition network, and the circuitry is further configured to apply the pose recognition network to a localized instance of the performer 118A to detect key points on the body and to detect dance poses associated with body movement exercises based on the detected key points, wherein the pose attributes indicate the dance poses.

[0127] In one embodiment, the attribute recognition network includes a gesture recognition network, and the circuitry is further configured to apply the gesture recognition network to a localized instance of performer 118A to detect the knuckles in each hand of performer 118A and determine gesture information based on the position of the detected knuckles, wherein the gesture information indicates a symbolic hand position that expresses meaning, emotion, or rhythmic experience in a physical performance at a given moment.

[0128] In one embodiment, the gesture attribute represents the root mudra among multiple root mudras in a dance form, yoga, or martial arts.

[0129] In one embodiment, the circuit is further configured to generate metadata for an image frame based on multiple performance attributes, generate a first vector search query based on a contrast-trained neural network applied to the generated metadata, and input the first vector search query into a performance database.

[0130] In one embodiment, the performance database includes multiple assets, each asset including: embedded information that encodes registered poses of a body movement exercise; metadata associated with each of the registered poses and source image frames depicting the registered poses; and tags associated with the metadata.

[0131] In one embodiment, the registered gesture represents a handprint among multiple handprints of a dance form, and is a key dance step consisting of a root handprint associated with gesture attributes, a posture associated with the body, and an emotion associated with emotional attributes.

[0132] In one embodiment, the circuitry is further configured to calculate a vector similarity score between the input first vector search query and each of a plurality of assets in the performance database, and to generate a first search result based on the vector similarity scores below a threshold. The first search result indicates a localized instance representing an unregistered pose of a body movement exercise, and the performance database is updated based on metadata from the image frame to include localized instances as newly registered poses of body movement exercises.

[0133] In one embodiment, the circuitry is further configured to receive user input associated with physical exercise from a user device, wherein the user input includes at least one of a natural language query or an image-based query. The circuitry is also configured to generate a second vector search query based on a neural network trained contrastively on the user input. The circuitry is further configured to compute a similarity score between the second vector search query and each of a plurality of assets. The circuitry is also configured to generate a second search result comprising the top k assets among the plurality of assets, for which a computed similarity score is above a threshold score, and to control the user device to display a response based on the second search result, wherein the response includes source image frames depicting a registered pose in at least one of the top k assets.

[0134] This disclosure can also be positioned in a computer program product that contains all the features enabling the implementation of the methods described herein and is capable of executing those methods when loaded into a computer system. In the present context, a computer program means any expression, in any language, code, or notation, representing a set of instructions intended to cause a system with information processing capabilities to perform a particular function directly or after one or both of the following: a) being translated into another language, code, or notation; or b) being reproduced in a different form of material.

[0135] Although this disclosure has been described with reference to certain embodiments, those skilled in the art will understand that various changes can be made and equivalents can be substituted without departing from the scope of this disclosure. Furthermore, many modifications can be made to adapt particular situations or materials to the teachings of this disclosure without departing from its scope. Therefore, it is intended that this disclosure be limited to the disclosed embodiments, but rather that it encompass all embodiments falling within the scope of the appended claims.

Claims

1. An electronic device, comprising: The circuit is configured as follows: Receive an input image sequence depicting physical performance associated with physical exercise; Extract localized instances of performers associated with physical performance from image frames in an image frame sequence; The application of an attribute recognition network to localized instances determines multiple performance attributes associated with the performer, wherein the multiple performance attributes include: Emotional attributes, including locational information associated with the performer's facial feature set, Postural attributes associated with the performer's body, and The gesture attributes associated with each of the performer's hands; and The performance database is searched based on the aforementioned multiple performance attributes to generate a first search result. The first search result specifies a localized instance representing an unregistered or registered pose for physical exercise.

2. The electronic device of claim 1, wherein the physical exercise corresponds to an Indian dance form, martial arts, or yoga.

3. The electronic device of claim 1, wherein the circuitry is further configured to apply a separation and localization (SL) network to the image frames to generate localized information associated with the performer. Among them, localized instances of performers are extracted based on localized information.

4. The electronic device according to claim 1, wherein the attribute recognition network is a multi-head neural network, which includes an emotion recognition network, a posture recognition network, and a gesture recognition network.

5. The electronic device of claim 1, wherein the attribute recognition network includes an emotion recognition network, and wherein the circuitry is further configured to: An emotion recognition network is applied to localized instances of the performer to detect the state of each facial feature in the facial feature set; and The detected state of each facial feature in the facial feature set is encoded into a character sequence. The location information includes the character sequence of each facial feature in the facial feature set.

6. The electronic device of claim 4, wherein the facial feature set includes eye gaze features, eyebrow features, nose features, lip features, and forehead twitching features.

7. The electronic device of claim 1, wherein the attribute recognition network includes a pose recognition network, and wherein the circuit is further configured to: Applying pose recognition networks to localized instances of performers to detect key points on their bodies; and Dance postures associated with physical movement exercises are detected based on the detected key points, where posture attributes indicate dance postures.

8. The electronic device of claim 1, wherein the attribute recognition network includes a gesture recognition network, and wherein the circuitry is further configured to: A gesture recognition network is applied to localized instances of the performer to detect the knuckles of each hand; and Gesture information is determined based on the detected position of the knuckles, where the gesture information indicates the symbolic hand position that expresses meaning, emotion, or rhythmic experience in a physical performance at a given moment.

9. The electronic device of claim 1, wherein the gesture attribute represents the root handprint among a plurality of root handprints of a physical exercise.

10. The electronic device of claim 1, wherein the circuit is further configured to: Metadata is generated for the image frame based on the multiple performance attributes; A contrast-trained neural network is applied to the generated metadata to generate a first vector search query; as well as Input the first vector search query into the performance database.

11. The electronic device of claim 10, wherein the performance database comprises a plurality of assets, each asset comprising: Embedded information encoding registered postures during physical exercise. Metadata associated with each of the registered poses and the source image frames depicting the registered poses, and Tags associated with metadata.

12. The electronic device of claim 11, wherein the circuitry is further configured to generate a set of non-fungible tokens (NFTs) for each of a plurality of assets stored in a performance database.

13. The electronic device of claim 11, wherein the registered posture represents a handprint among a plurality of handprints of physical exercise, and is a key dance step consisting of a root handprint associated with a gesture attribute, a posture associated with a body, and an emotion associated with an emotional attribute.

14. The electronic device of claim 11, wherein the circuit is further configured to: Calculate the vector similarity score between the input first vector search query and each of the plurality of assets in the performance database; and The first search result is generated based on the vector similarity score, which is below a threshold. The first search result indicates a localized instance representing an unregistered posture during body movement exercise; and The performance database is updated based on the metadata of the image frames to include the localized instances as new registered poses for body movement exercises.

15. The electronic device of claim 11, wherein the circuit is further configured to: Receive user input related to physical exercise from the user device. The user input includes at least one of natural language queries or image-based queries; The second vector search query is generated by applying a neural network trained on contrast with user input. Calculate the similarity score between the second vector search query and each of the plurality of assets; Generate a second search result, which includes the top k assets among the plurality of assets whose calculated similarity scores are higher than a threshold score; as well as The user device is controlled to display a response based on a second search result, wherein the response includes a source image frame depicting a registered pose in at least one of the top k assets.

16. A method comprising: In electronic devices: Receive an input image sequence depicting physical performance associated with physical exercise; Extract localized instances of performers associated with physical performance from image frames in an image frame sequence; The application of an attribute recognition network to localized instances determines multiple performance attributes associated with the performer, wherein the multiple performance attributes include: Emotional attributes, including locational information associated with the performer's facial feature set, Postural attributes associated with the performer's body, and The gesture attributes associated with each of the performer's hands; and The performance database is searched based on the aforementioned multiple performance attributes to generate a first search result. The first search result specifies a localized instance representing an unregistered or registered pose for physical exercise.

17. The method of claim 16, wherein the physical exercise corresponds to an Indian dance form, martial arts, or yoga.

18. The method of claim 16, wherein the electronic device is further configured to apply a separation and localization (SL) network to the image frames to generate localization information associated with the performer. Among them, localized instances of performers are extracted based on localized information.

19. The method of claim 16, wherein the gesture attribute represents the root handprint among a plurality of root handprints of a physical exercise.

20. A non-transitory computer-readable medium storing computer-executable instructions thereon that, when executed by a first electronic device, cause the first electronic device to perform operations, said operations including: Receive an input image sequence depicting physical performance associated with physical exercise; Extract localized instances of performers associated with physical performance from image frames in an image frame sequence; The application of an attribute recognition network to localized instances determines multiple performance attributes associated with the performer, wherein the multiple performance attributes include: Emotional attributes, including locational information associated with the performer's facial feature set, Postural attributes associated with the performer's body, and The gesture attributes associated with each of the performer's hands; and The performance database is searched based on the aforementioned multiple performance attributes to generate a first search result. The first search result specifies a localized instance representing an unregistered or registered pose for physical exercise.