System and method for analyzing, documenting, authenticating, archiving, and crediting dances
A machine learning-based method for pose estimation and blockchain authentication addresses the challenge of crediting and archiving dance creators, improving accuracy and reliability in dance video analysis and attribution.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2026-03-26
AI Technical Summary
Current systems lack an automated and intelligent method for tracking and crediting dance creators, particularly on social media platforms, and do not provide structured ways to analyze and archive dance works for attribution, royalty payment, and historical context.
A computer-implemented method using machine learning models for pose estimation, combined with contextual data, to create reference videos, compare dance videos, and utilize blockchain for authentication and crediting, with a scalable similarity scoring system.
Enhances accuracy and reliability in detecting and comparing dance poses, facilitates attribution and royalty payments, and provides a platform for archiving and crediting dance creators, promoting fair compensation and historical preservation.
Smart Images

Figure CA2025051244_26032026_PF_FP_ABST
Abstract
Description
SYSTEM AND METHOD FOR ANALYZING, DOCUMENTING, AUTHENTICATING,ARCHIVING, AND CREDITING DANCESFIELD
[0001] The present disclosure relates to methods and systems for detecting the position and orientation of a person's body parts in an image or video. More particularly, it relates to a pose estimation model for recognizing, authenticating, crediting and archiving original dances.BACKGROUND
[0002] After years of appropriation, there is a growing awareness and movement to better compensate and credit dancers. For example, in 2019, a new dance creation, known as the Renegade, went viral on social media platforms including TikTok® and Instagram®. In a few weeks, the dance attracted the attention of TikTok dancers with many followers, and celebrities came up with their own renditions of the dance. While the popularity of the Renegade dance seemed positive on the surface, a huge challenge was that no one was acknowledging, crediting, and referencing the original creator. In fact, the dance helped catapult a contingent TikTokers including Charli D’Amelio and Addison Rae to stardom. These TikTok dancers appeared on late-night television show, broke video viewership records, and profited from reality series deals. (Rosenblatt, 2021; Wicker, 2020). Rosenblatt (2021).
[0003] A significant step towards crediting and compensating dance creators relies on the ability to establish an easy-to-implement copyright framework. While such a framework is well-established for protecting copyright or intellectual property (IP) rights for songs, books, and films, the same cannot be said for dance / choreography because of the complexity in filing for copyright protection and determining what qualifies as copyrightable dance. Some creators have begun advocating for “choreographers to be recognized and receive appropriate credit and protection” (Robert, 2020) and some organizations have begun developing several online resources for archiving (e.g., Dance Collection Danse Archives (https: / / www.dcd.ca / archives.html#), sharing dance routines (e.g., Video Can (https: / / videocan.ca / ) and creating professional standards for dancers (https: / / cadaontario.wildapricot.org / psd 11 copyright#howcopyright). Most of theseinitiatives and resources focus on professional choreographers or specific genres (e.g., K-POP, Kim et al 2016 US 2016 / 0110453 Al) who, arguably, are a small subset of the massively exploding dance industry in the internet era. A more inclusive approach is required to capture creators of all ages and cultures who are uploading their creative productions on social media platforms such as TikTok.
[0004] Another limitation of the ongoing initiatives is that there is no automated or intelligent system for tracking the copying and reuse of original dance work (especially on social media) to facilitate attribution and royalty payment. TikTok, for example, allows the platform users to acknowledge creators and videos that have inspired the users’ dance or work by tagging and mentioning them when describing a video or posting a comment (https : / / support.tiktok. com / en / using-tiktok / creating- videos / credit-a-video).
[0005] Another challenge is that the current approaches to addressing exploitation are narrowly framed as they focus on the immediate content creators. The issue goes far beyond crediting and attribution but also includes tracing history (such as genealogy), connecting dances to context and archiving of dance work for research, education, and posterity. Related to this gap is that existing platforms do not have structured and systematic ways of collecting contextual information to allow building robust models for analyzing traceability and context (e.g., motivation and mood of dancers).
[0006] Pose estimation has numerous applications, including sports analysis, human-computer interaction, and dance video analysis
[0012] , Several models have been developed for pose estimation, each with its own strengths and weaknesses, the more commonly models used include but not limited to: OpenPose, developed by researchers at Carnegie Mellon University, is a real-time multi-person 2D pose estimation model that uses part affinity fields to detect keypoints and construct skeletal representations. It is known for its accuracy and ability to handle multiple people in a single frame [1], PoseNet, developed by Google, is a lightweight, efficient model designed for real-time pose estimation on mobile and web platforms. PoseNet uses convolutional neural networks to detect keypoints and is optimized for performance on a wide range of devices [2],SUMMARY
[0007] In one of its aspects, a computer-implemented method for creating a reference video associated with at least one first action by processing circuitry and a memory device having instructions executable by the processing circuitry to carry out the steps of: receiving a plurality of first sequential images associated with the at least one first action, preprocessing the plurality of first sequential images; extracting frames from the plurality of sequential images to generate feature vectors associated with at least one pose; with at least one trained machine learning model, detecting first keypoints pertaining to the at least one first pose and generating a first pose data set associated with the at least one first pose; filtering first keypoints based on confidence scores to generate a refined first pose data set; and on the memory device, storing the refined first pose data set, the plurality of first sequential images, extracted frames, and audio information descriptive of the at least one first action, and metadata.
[0008] In another aspect, a computer system comprising a hardware processor and a memory device on which instructions are encoded to cause the hardware processor to perform the operations of: receiving a plurality of first sequential images associated with the at least one first action, preprocessing the plurality of first sequential images; extracting frames from the plurality of first sequential images to generate feature vectors associated with at least one pose; with at least one trained machine learning model, detecting first keypoints pertaining to the at least one first pose and generating a first pose data set associated with the at least one first pose; filtering first keypoints based on confidence scores to generate a refined first pose data set; and on a computer readable medium, storing the refined first pose data set, plurality of first sequential images, extracted frames, and audio information.
[0009] In another aspect, a computer readable medium storing instructions executable by a processor to carry out the operations comprising: receiving a plurality of first sequential images associated with the at least one first action, preprocessing the plurality of first sequential images; extracting frames from the plurality of sequential images to generate feature vectors associated with at least one pose; with at least one trained machine learning model, detecting first keypoints pertaining to the at least one first pose and generating a first pose data set associated with the at least one first pose; filtering first keypoints based on confidence scores to generate a refined first pose data set; and on the computer readable medium, storing the refined first pose data, plurality of first sequential images, extracted frames, and audio information descriptive of the at least one first action, and metadata, as a reference video.
[0010] In another aspect, a computer-implemented method for comparing a reference video associated with at least one reference action to a target video associated with at least one target action, the method comprising the steps of: receiving a plurality of first sequential images associated with the at least one reference action, preprocessing the plurality of first sequential images; receiving a plurality of second sequential images associated with the at least one target action, preprocessing the plurality of second sequential images; and using at least another machine learning model to determine a similarity score between the reference video and the target video based on at least one of a benchmark threshold for standardizing similarity classification, a weighted distance, and a difference.
[0011] In another aspect, a computer-implemented method for assigning a unique identifier to a dance video, the method comprising the steps of: extracting the dance video data from a data repository comprising a database record associated with the dance video; preparing blockchain data for blockchain authentication;creating metadata from the blockchain data for creating a non-fungible token (NFT) using a cover image of the dance video as an NFT image and storing the NFT on a distributed file system; minting the NFT on the blockchain using a blockchain network; generating a blockchain hash associated with the dance video stored on the blockchain; and updating the database record of the dance video using the blockchain hash in the data repository.
[0012] In another aspect, the present disclosure comprises a scalable similarity scoring system “Idometrics” that runs on multiple devices, promoting both dance content sharing and systematically addressing exploitation of creators.
[0013] In another aspect the methods and systems described herein automatically generate dance impact and metrics like bibliometrics used for citation analysis in research publications (Wood and Wilson, 2001, https: / / link.springer.eom / article / 10.1023 / A:1017919924342).
[0014] In another aspect, the methods and systems described herein leverage the growth in hardware power and data that now allows modem machine learning models to not rely on a single source of information, such as text, audio, or images. Much like humans, correlating information from multiple sources can enhance the predictive capacity of these models.
[0015] In another aspect, the methods and systems described herein utilize a dual model combining vision and text. This model comprises a context learning framework which leams to detect dance steps and poses, while the text model provides additional descriptive information, serving as contextual data. This approach to multimodal models has motivated research in In-Context Learning ([2301,00234] A Survey on Incontext Learning (arxiv.org)) ([2403,00231] Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models). With advancements in large language models, there is no need to train the text model from scratch; instead, models such as GPT-3.5, GPT-4, and Llama3 may be leveraged as text encoders for the contextual component. To standardize the text information, the platform collects questionnaires covering aspects such as mood, motivation, location,cultural background, description and genre of the dances. The dual model of vision and text allows for creating a closer mapping between similar dances, not only through visual information, but also with the aid of contextual data. Aspects of the methods and systems described herein achieve this objective using unsupervised algorithms. Furthermore, aspects of the methods and systems described herein use text autocompletion, meaning that given a dance video, the model could provide information about the possible genre or cultural background of the dance and vice versa.
[0016] One challenge for the vision model is to accurately detect dance poses, a process known as pose estimation ([2308,13872] Vision-Based Human Pose Estimation via Deep Learning: A Survey (arxiv.org)). One of the main hurdles is effectively capturing the time sequence of the dancers' poses. The methods and systems described herein advance accuracy of pose estimation, pose comparison, and similarity analysis of dance videos, and provides an innovative platform that supports various dance genres through contextual learning and dance genealogy. The methods and systems described herein leverage advanced techniques to create a scalable solution that enhances the accuracy, reliability, and practical applicability of the dual model.
[0017] BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 shows an operating environment comprising a computerized platform for a method for analyzing, documenting, authenticating, archiving, and crediting a dance choreography;
[0019] Figure 2 shows a flow chart with example steps for creating a reference video;
[0020] Figure 3 a shows a flow chart with example steps for creating a target video by a user;
[0021] Figure 3b shows an example workflow for the context learning model with input and output parameters;
[0022] Figure 4 shows 17 keypoints detected by the PoseNet model for human pose estimation;
[0023] Figure 5 shows a similarity analysis process flow, including the steps of spatial alignment and normalization, weighted distance calculation, and the final similarity scoring using an Ido constant;
[0024] Figure 6 shows a similarity score generated after completing the similarity analysis in Figure 5;
[0025] Figure 7 shows a microservice architecture of the platform which supports the storage and analysis of dance videos on the platform;
[0026] Figure 8 shows a home page with dance videos uploaded to a desktop version of an application, wherein each video is accompanied by its corresponding similarity scores and other metadata, allowing users to explore and interact with different dance performances;
[0027] Figure 9 shows a donation feature of the application running on a mobile device, where users can support their favorite dance creators through monetary donations;
[0028] Figure 10 shows a user interaction interface of the Idometrics platform on a tablet, where users can connect and relate with each other through comments, likes, and shares, thereby fostering a community environment and encourages engagement among users;
[0029] Figure 11 shows a private messaging feature of the application on a tablet, allowing users to communicate directly with each other;
[0030] Figure 12a shows reference poses and Figure 12b shows target poses, which are similar and in the same direction;
[0031] Figure 13a shows reference poses and Figure 13b shows target poses, which are similar and in the opposite direction;
[0032] Figure 14a shows reference poses and Figure 14b shows target poses, which are dissimilar; and
[0033] Figure 15 shows a flowchart with example steps for registering a created dance on the platform on the blockchain.DETAILED DESCRIPTION
[0034] The following detailed description refers to the accompanying drawings. Wherever possible, the same reference numbers are used in the drawings and thefollowing description to refer to the same or similar elements. While embodiments of the disclosure may be described, modifications, adaptations, and other implementations are possible. For example, substitutions, additions, or modifications may be made to the elements illustrated in the drawings, and the methods described herein may be modified by substituting, reordering, or adding stages to the disclosed methods. Accordingly, the following detailed description does not limit the disclosure. Instead, the proper scope of the disclosure is defined by the appended claims.
[0035] Moreover, it should be appreciated that the particular implementations shown and described herein are illustrative of the disclosure and are not intended to otherwise limit the scope of the disclosure in any way. Indeed, for the sake of brevity, certain sub-components of the individual operating components, and other functional aspects of the systems may not be described in detail herein. Furthermore, the connecting lines shown in the various figures contained herein are intended to represent example functional relationships and / or physical couplings between the various elements. It should be noted that many alternative or additional functional relationships or physical connections may be present in a practical system.
[0036] Figure 1 shows a system 10, or a computerized platform, for method for analyzing, documenting, authenticating, archiving, and crediting a dance choreography. The system 10 comprises machine or computing system 12 having processing circuitry 14 such as one or more processors, at least one memory device such as memory 16, input / output (I / O) module 18, which are in communication with each other via centralized circuit system 19. An image capture device 20 is coupled to machine 12 for recording the dance choreography comprising a series of movements as data 21. Although computing system 12 is depicted to include only one processor 14, computing system 12 may include a number of processors therein. In an embodiment, memory 16 is capable of storing machine executable instructions 22, and data 21, including data models and process models. Database 26 is coupled to computing system 12 and stores pre-processed data, model output data, and so forth. Further, the processor 14 is capable of executing the instructions in memory 16 to implement aspects of processes described herein. For example, processor 14 may be embodied as an executor of software instructions, wherein the software instructionsmay specifically configure processor 14 to perform algorithms and / or operations described herein when the software instructions are executed. Alternatively, processor 14 may be configured to execute hard-coded functionality. Computerized system 10 may be software (e.g., code segments compiled into machine code), hardware, embedded firmware, or a combination of software and hardware, according to various embodiments.
[0037] Examples of the I / O module 18 include, but are not limited to, an input interface and / or an output interface. Some examples of the input interface may include, but are not limited to, a keyboard, a mouse, a joystick, a keypad, a touch screen. Some examples of the output interface may include, but are not limited to, a microphone, a speaker, a ringer, a graphical user interface 28. In an example embodiment, processor 14 may include I / O circuitry configured to control at least some functions of one or more elements of I / O module 18, such as, for example, a speaker, a microphone, a display 28, and / or the like. Processor 14 and / or the I / O circuitry may be configured to control one or more functions of the one or more elements of I / O module 18 through computer program instructions, for example, software and / or firmware, stored on a memory, for example, the memory 16, and / or the like, accessible to the processor 14.
[0038] A communication interface associated with the I / O module 18 enables computing system 12 to communicate with other entities over various types of wired, wireless or combinations of wired and wireless networks 30, such as for example, the Internet. The communication interface facilitates communication between the computing system 12 and I / O peripherals. In at least one example embodiment, the communication interface includes a transceiver circuitry configured to enable transmission and reception of data signals over the various types of communication networks. In some embodiments, the communication interface may include appropriate data compression and encoding mechanisms for securely transmitting and receiving data over the communication networks.
[0039] Centralized circuit system 19 may be various devices configured to, among other things, provide or enable communication between the components (14-18) of computing system 12. In certain embodiments, centralized circuit system 19 may bea central printed circuit board (PCB) such as a motherboard, a main board, a system board, or a logic board. Centralized circuit system 19 may also, or alternatively, include other printed circuit assemblies (PCAs), communication channel media or bus.
[0040] A plurality of user computing devices 32 and data sources 34 are coupled to computing system 12 with a communication network 30. User computing devices 32 can therefore access computing system 12, and system 10 is operable to register and authenticate users (using a login, unique identifier, and password for example) prior to providing access to applications, a local network, network resources, other networks and network security devices.
[0041] In an example, the processor 14 can execute instructions in memory 16 to configure a data pre-processing module 40, a matching module 42, and a similarity engine / module 44. In more detail, the matching module 42, and the similarity module 44 comprise a suite of algorithms which receive pre-processed data derived from a plurality of raw data sources 26. The processor 14 is configured by the machine executable instructions to initiate and acquire raw data for input into the data preprocessing module 40; and the processor 14 implements the machine executable instructions to configure the matching module 40, and the similarity module 44 using data models to provide a similarity score between a target video and reference video.
[0042] Looking at Figure 2, there is a shown a flow chart 100 with example steps for creating a reference video. In one example, in step 102, a first user, or dance challenge creator, sets out to create a dance challenge comprising a series of movements to be executed. The dance may be performed to music, as such the movement patterns may be associated with appropriate beat counts for timing.
[0043] In step 104, the user records a video of the series of movements using an image capture device 20, 50 such as a camera or smartphone, or user computing device 32.
[0044] In step 105, the user records accompanying audio / story telling and using an audio capture device 20b, such as a camera or smartphone, to generate storytelling data. To aid in the storytelling, the user may be presented with a form with questions for the user to answer, selections e.g. drop-down menu or radio buttons, or promptsdirected at the user geared to spur creativity. In one example, the questions may include: "Have you posted this video anywhere else? What is the title of your dance? Does the title have any meaning? Do any of the dance moves have any meaning? What aspect of this dance did you create? Which cultural heritage inspired this dance? Select a dance genre." The user then uploads the video to machine 12.
[0045] In step 106, the processing circuity 14 executes a set of instructions in memory 16 to format the videos, process the videos in the background to extract image frames (step 108), and generate pose data.
[0046] In step 110, the processing circuity 14 executes a set of instructions in memory 16 to process and format the image frames.
[0047] In step 112, the processed video and images are uploaded to a data repository 26, such as Cloudinary CDN™, an image and video API platform, from Cloudinary, Inc., San Jose, CA, U.S.A.
[0048] In step 114, data repository 26 returns a JSON object containing image and video URLs, and uses these URLs to create image objects for pose estimation.
[0049] In step 116, the processing circuity 14 executes a set of instructions in memory 16 to perform pose estimation and confidence screening to filter unreliable keypoints (step 118).
[0050] Next, in step 120, the processing circuity 14 executes a set of instructions in memory 16 to generate the refined pose data and landmark data.
[0051] Next, in step 122, the processing circuity 14 executes a set of instructions in memory 16 to combine and store the refined pose data pose data and landmark data with user-provided storytelling information in the database, step 124. The reference video is then posted on the Internet or social media for other users, such as, the general population or dance challenge participants to view and perform. In addition, the posting the reference video also creates a record of its creation by the user for proper accreditation, and may include a date-stamp to mark the date and time of creation. Such information may be included in metadata.
[0052] Looking at Figure 3 a, there is a shown a flow chart 200 with example steps for creating a target video by a user, or dance challenge participant. In one example, in step 202, the dance challenge participant creates a dance challengecomprising a series of movements to be executed. The dance may be performed to music, as such the movement patterns may be associated with appropriate beat counts for timing.
[0053] In step 204, the user views the reference video from Figure 2, records a video associated with the reference video series of movements using an image capture device 20, 50, such as a camera or smartphone, or user computing device 32, and uploads the video to machine 12.
[0054] In step 206, the processing circuity 14 executes a set of instructions in memory 16 to format the videos, process the videos in the background to extract image frames (step 208), and generate pose data.
[0055] In step 210, the processing circuity 14 executes a set of instructions in memory 16 to process and format the image frames.
[0056] In step 212, the processed video and images are uploaded to a data repository 26, such as Cloudinary CDN, an image and video API platform, from Cloudinary, Inc., San Jose, CA, U.S.A.
[0057] In step 214, data repository 26 returns a JSON object containing image and video URLs, and uses these URLs to create image objects for pose estimation.
[0058] In step 216, the processing circuity 14 executes a set of instructions in memory 16 to perform pose estimation and confidence screening to filter unreliable keypoints (step 218).
[0059] Next, in step 220, the processing circuity 14 executes a set of instructions in memory 16 to generate the refined pose data and landmark data.
[0060] Next, in step 222, the processing circuity 14 executes a set of instructions in memory 16 to combine and store the refined pose data pose data and landmark data with user-provided storytelling information in the database, step 224.
[0061] As such, unlike the reference video process of Figure 2, the example steps of Figure 3a do not include storytelling form completion.
[0062] Pose Estimation
[0063] In more detail, aspect of the methods and systems described herein improve and use pose estimation models for detecting the position and orientation of a person's body parts in an image or video. For example, in step 116 of Figure 2 andstep 216 of Figure 3a, the processing circuity 14 executes a set of instructions, such as, pose estimation models, to identify keypoints on the human body, such as the joints, and using these key points to construct a skeletal representation.
[0064] Regardless of the pose estimation model, the methods and systems described herein prioritize models based on the following factors:
[0065] Efficiency: quick and responsive pose estimation for real-time performance. This efficiency is critical for dance video analysis, where multiple frames need to be processed in a short time [2],
[0066] Lightweight: designed to run efficiently on both mobile and web platforms, providing flexibility and accessibility for users on different devices as our platform aims to be accessible to a broad audience [2],
[0067] Accuracy: in detecting keypoints, ensuring reliable pose estimation for various applications, including dance analysis
[0013] ,
[0068] Ease of Integration: architecture and application program interface (API) that are straightforward to integrate into various applications, allowing for seamless implementation within our platform
[0014] ,
[0069] License flexibility: to allow commercial use and modification of pose estimation model.
[0070] The methods and systems described herein improve and utilize the publicly available PoseNet, as an example of a post estimation tool, as it meets the above criteria (i.e., efficiency, lightweight, accuracy, integration ease and license flexibility). As the platform’s users grow, the database will be large enough to retrain the model or build a new pose estimation method. To extract key points from dance videos, PoseNet can detect 17 keypoints on the human body, such as the nose, eyes, ears, shoulders, elbows, wrists, hips, knees, and ankles as shown in Figure 4. The methods and systems described herein take the following steps to apply pose estimation (e.g., PoseNet) creating, extracting, and storing dancers’ poses:
[0071] In more detail, in example steps 106, 108, 206, and 208, the processing circuitry 14 executes a set of instructions in memory 16 to format the videos, process the videos in the background to extract image frames and generate pose data. As such, in one example, dance videos are divided into frames at intervals of 500 milliseconds,resulting in approximately 60 images for a 30-second video. This approach to video frame extraction significantly improves upon the methods of Cao et al. and Wei et al. by addressing the temporal consistency and sequencing challenges inherent in pose estimation. While Cao et al.'s OpenPose and Wei et al.'s Convolutional Pose Machines effectively detect poses within individual frames, they struggle with maintaining consistency across a video sequence due to independent frame processing. The methods and systems described herein address this gap by extracting frames at regular 500-millisecond intervals, which standardizes the frame rate, ensuring smoother and more reliable pose tracking over time. This consistent interval maintains continuity and reduces jittering as well as enhances the accuracy of sequencing poses throughout the dance video, offering a robust solution for analyzing and matching dance movements [1][2],
[0072] Keypoint Detection and Confidence Screening: PoseNet estimates keypoints for each frame and assigns confidence scores with debatable reliability for some poses. The methods and systems described herein derive and introduce a confidence screening threshold of 0.4 to filter out unreliable keypoints, ensuring higher accuracy and reliability in the pose estimation process. This introduced threshold, which ensures the reliability of detected keypoints, addresses reliability issues identified in related previous works [1], [2],
[0073] Data Storage: The processed data, including images, pose landmarks, and confidence scores, are stored in a cloud-based database. This integration with cloud services like Cloudinary CDN facilitates easy video uploading, processing, and storage, enhancing scalability and ease of use over traditional storage methods.
[0074] The platform’s database and repository support contextual learning modeling and dance genealogy, allowing for the classification and comparison of dance moves across different genres, styles, location, and time [6], The objective of the repository is to link to the database digital narratives by dances on the processes behind developing their dances. The methods and systems described herein capture digital stories through short questionnaires on the platform for two purposes:
[0075] 1) educating audiences as well, enhancing the educational and historical value of the platform, providing users with deeper insights into dance movements and their origins, leading to cultural and labour appreciations [6]; and
[0076] 2) adding a layer of contextual understanding and historical relevance to dance analysis, which is not typically covered in related works and providing an additional layer to authenticate creative work.
[0077] With the rich database and repository, the the methods and systems described herein use exploratory statistical methods to reveal trends in storytelling data and predict the motivation and inspiration for a dance. The analysis can gainfully answer research questions including: What are the motivations behind various dances routines? How does motivation for the dance type vary with socio-economic, demographic factors (e.g., age, gender) and geographical locations? What story is a dance video telling (i.e., storydancing)? In one aspect, the methods and systems described herein use unsupervised learning approaches (e.g., hierarchical clustering, k-nearest neighbours and network analysis) to group dances into trees, linking dances to original sources. This feature will be useful when a creator posts a dance challenge and people respond to it.
[0078] The context learning model uses parameters images as input (e.g. the pose estimation data, JSON files) and outputs text (e.g. history, motivation and description). Figure 3b shows an example workflow 300 for the context learning model with input and output parameters. In one example, in steps 302, 304, a dance video and accompanying story telling data and description of the dance video is retrieved from a data store. The storytelling data may include answer to questions such as: “ Have you posted this video anywhere else? What is the title of your dance? Does the title have any meaning? Do any of the dance moves have any meaning? What aspect of this dance did you create? Which cultural heritage inspired this dance? Select a dance genre. ” The video may include a video clip, images and pose features with landmarks, as described before.
[0079] In step 306, the dance video and accompanying story telling data and description of the dance video are input into a machine learning engine whichcomprises a set of instructions in memory 16 executable by the processing circuity 14 to generate a new dance video (step 308) and anew story (step 310), as outputs.
[0080] The methods and systems described herein provide a platform that integrates seamlessly with cloud services such as Cloudinary CDN, Stripe payment and Polygon blockchain network, ensuring scalability and ease of use. This practical aspect is a significant enhancement over more complex or less user-friendly implementations such as platforms that require on-premises hardware and manual updates. For example, traditional video processing platforms like Adobe Media Server and Red5 Server often involve complex configurations and do not offer the same level of scalability and convenience as cloud-based solutions
[0015] ,
[0016] , These systems typically require extensive technical knowledge for setup and maintenance, making them less accessible to a wider audience. In contrast, the methods and systems described herein comprise a platform with cloud integration that allows for automatic scaling, easier updates, and maintenance-free operation. For instance, cloud-based solutions such as AWS Media Services and Google Cloud Video Intelligence automatically manage server resources based on demand, eliminating the need for manual intervention
[0017] ,
[0018] , This significantly reduces operational overhead and ensures that the platform can handle varying loads efficiently. Additionally, the use of cloud services provides robust data security, backup, and disaster recovery options, enhancing the platform's reliability and user experience
[0019] ,
[0081] Aspects of the methods described herein are shown in a micro-service based full stack application, such as the Idometrics™ application developed by the Applicant, as shown in Figure 7. The application comprises data storage capabilities to store user information, including dance videos and similarity scores, and provides creators with copyright protection over their dances through blockchain authentication for security and transparency. The copyright and unique identifier in the form of a non-fungible token (NFT) and similarity report generated on the platform can be used as proof of originality, ownership and authenticity.
[0082] Additionally, the application, or platform, offers users the opportunity to be entertained by creators' dances, as shown in Figure 8. The users may also participate in dance challenges geared towards similarity analysis, and supportcreators via donations and other ways for monetization. For example, Figure 9 shows a process flow 400 with screenshots pertaining to a donation feature of the application running on a mobile device, where users can support their favorite dance creators through monetary donations by following steps 402-412. In one example, this feature is supported with a one-time password (OTP) authentication to ensure secure and transparent transactions.
[0083] Users can also interact with each other and develop friendships on the platform, as shown in Figure 10. The user may also interact directly with creators through private messages, as shown in Figure 11. The Idometrics platform will be available for all devices, accessible on the web and mobile. This comprehensive application showcases the effectiveness, user engagement, and practical utility of the technology, which may not be as explicitly demonstrated in related works like Move Mirror by Google Al [7],
[0084] Pose Comparison
[0085] To ensure consistent pose comparison, the methods and systems described herein performs the preprocessing steps. In one example, the data pre-processing module 40 comprises the steps of:
[0086] (a) resizing by cropping and scaling each image based on the bounding box coordinates of the person;
[0087] (b) normalizing keypoint coordinates as L2 normalized vector arrays, enabling consistent pose representation across different image sizes.
[0088] This preprocessing approach, which involves L2 normalization of keypoint coordinates, improves Toshev and Szegedy's [3] method by ensuring consistent pose representation across different image sizes and positions. While Toshev and Szegedy's DeepPose method effectively regresses joint coordinates using a deep neural network, it does not explicitly normalize these coordinates, leading to potential variations in pose estimation due to changes in scale, rotation, and translation of the input images. In contrast, the methods and systems described herein normalize keypoint coordinates as L2 normalized vector arrays, which scale the coordinates to have a unit norm, making the poses comparable across different frames and conditions. This process mitigates inconsistencies and enhances robustness tovariations in image size and position, resulting in more reliable pose matching and comparison. The normalization steps, including resizing, scaling, and vector normalization, ensure that the pose representations are standardized, thereby improving the overall accuracy and temporal consistency of pose estimation in dynamic scenarios.
[0089] The methods and systems described herein apply weighted distance metric to account for variability in confidence levels of detected keypoints. The weighted distance formula for pose matching is designed to improve the accuracy of comparing two sets of pose keypoints. It accounts for the confidence scores of each keypoint detected by the pose estimation model. The formula used in this approach is:
[0090] n D(F, G') =1* y FckI \Fxyk— Gxyk11 ... (1) k = lck
[0091] where: DFG : Weighted distance between the reference pose F and the target pose G; Fck : Confidence score of the k-th key point in the reference pose; Fxyk : Coordinates of the k-th keypoint in the reference pose; Gxyk: Coordinates of the k- th keypoint in the target pose; n: Total number of keypoints of the reference pose F.
[0092] The formula (1) above gives higher importance to keypoints with higher confidence scores, ensuring that reliable keypoints have a greater influence on the pose matching result. Keypoints with lower confidence scores are given less weight, which minimizes the impact of potentially inaccurate keypoints or missing keypoints. In addition, by focusing on confident keypoints, the method becomes more robust to noise and inaccuracies in pose detection. This approach minimizes the impact of low- confidence keypoints on the overall similarity score. Looking at Figure 5, aspects of the methods and systems described herein refine the formula for pose matching to handle cases where some keypoints are not visible (resulting in NaN values). The refined formula ensures accurate similarity scoring even with missing keypoints:
[0093] The refined formula assigns zeros to NaN (non-available number values), ensuring more reliable similarity scores relative to previous methods (Jang et al. 2015 US20160110453A1) [5], To further refine the weighted distance metric and addresspractical situations where the summation could result in NaN values, the methods and systems described herein introduce the following set of equations:
[0094]
[0095]
[0096]
[0097]
[0098]
[0099] n11. . (7)
[0100] If Pi or Qi results in NaN, assign zero to these values to avoid computational errors.
[0101] if P1= NaN > P1= 0 ... (8)
[0102] if Q. = NaN > Q1 =0 ... (9)
[0103] if Pn= NaN > Pn= 0 ... (10)
[0104]
[0105] where: P is a first summation for the k-th comparison; Q is a second summation for the k-th comparison; and n is a total number of comparisons.
[0106] The methods and systems described herein comprises improvements over Jang’s method by assigning Zero to NaN Values. The improvement comes fromhandling cases where some keypoints may not be detected in either the reference or the target pose, resulting in NaN (Not a Number) values during the calculation. This is addressed by assigning zero to these NaN values.
[0107] In existing methods, if NaN values are present, the summation and multiplication in the formula would result in NaN, leading to an invalid or undefined similarity score. However, in the methods and systems described herein, by assigning zero to NaN values, the calculation can proceed without interruption, ensuring that all keypoints, whether detected or missing, contribute to the final similarity score in a meaningful way, and allows for stability in calculations.
[0108] As such, these improvements lead to increased reliability, as the adjustment ensures that the similarity score is always computable, even if some keypoints are missing. This is desirable in real-world applications where not all keypoints may be reliably detected. These improvements also provide a consistent way to handle missing data, which can be especially important when comparing poses in sequences of frames in videos.
[0109] Furthermore, Jang discloses a pose matching technique does not explicitly handle NaN values. This could lead to inconsistencies or failures in the pose matching process when keypoints are missing or detected with very low confidence. The methods and systems described herein comprises further improvements over Jang’s method by explicitly managing NaN values and assigning zeros, and therefore offers a more robust and reliable pose matching process, which increases the overall effectiveness of the system in real-world scenarios where data imperfections are common.
[0110] The methods and systems described herein comprises a benchmark threshold for standardizing similarity classification (“Ido constant” (z). It determines the similarity score (5) based on the weighted distance (w), Ido constant (z), and a difference (d). It applies a fuzzy state to describe the similarity limits:
[0111] i = 0.017 ... (12)
[0112] Limit: In this range, the reference and target poses are similar and in the same direction, as shown in Figures 12a and 12b, respectively. It can be expressed as:
[0113] 0 < w < i ... (13)
[0114] where the most identical is at w equal 0.
[0115] Upper limit: In this range, the reference and target poses are similar but in opposite directions, as shown in Figures 13a and 13b, respectively. It can be expressed as:
[0116] i < w < 0.1 ... (14)
[0117] Over limit: In this range, the reference and target poses are dissimilar, as shown in Figures 14a and 14b, respectively. It can be expressed as:
[0118] 0.1 < w > 0.1 ... (15)
[0119] The Ido’s constant is an anchor point in an interpolation function, ensuring that numerical thresholds correspond to human perception of pose similarity. In one example, an Ido constant of 0.017 maps to an 85% similarity score, forming the lower bound of the similarity region.
[0120] The similarity scores are determined based on the weighted distance w and interpolate across predefined anchor points:
[0121] A ={ (0.0, 100), (0.017, 85), (0.018, 70), (0.020, 50), (1.0, 0) } ... (16)
[0122] For a given distance w, the similarity score is obtained by linear interpolation:
[0123] s(w) = yt+ ((yf+1- yt / (xf+1- x ) * (w - x<) ... (17)
[0124] where (x^yj, (xi+1,yi+1) 6 A and xt< w < xi+1.
[0125] The total similarity score St for a sequence of n pose comparisons can then be expressed as an average of the total percentage similarity across all frames.
[0126]
[0127] Where A: ordered set of anchor points mapping distance to similarity; x: weighted distance axis. distance; y: similarity percentage axis; w Weighted distance;i: Ido constant, set to 0.017; s: Similarity score; SnSimilarity score for the 77-th comparison; St Total similarity score; and n: Total number of comparisons.
[0128] To simplify interpretation, the similarity outcome is grouped into two categories using the interpolation framework and grading thresholds: (1) Similar: for results within the classified limit and (2) Dissimilar: for every other case outside the limit, which includes both under limit and over limit.
[0129]
[0130] This classification method provides a robust benchmark for similarity analysis. By introducing the Ido constant as a benchmark threshold for pose similarity analysis, the methods and systems described herein address several limitations found in earlier methods (e.g., Seo (2021) and Tobesoft), which lacked consistent criteria for similarity evaluation. The methods and systems described herein improve similarity analysis by adhering to these concepts, among others:
[0131] (a) consistent threshold: previous methods (e.g., Seo (2021) and Tobesoft) often used inconsistent or arbitrary thresholds, leading to unreliable results. The Ido constant (0.017) provides a fixed, standardized threshold that ensures robust and repeatable similarity analysis across different datasets and scenarios;
[0132] (b) robust benchmark: using the Ido constant allows for a clear and objective distinction between similar and dissimilar poses, enhancing the accuracy and reliability of the similarity scores compared to previous methods with less defined criteria;
[0133] (c) improved similarity calculations: the methods and systems described herein calculate the total similarity score as an average of the total percentage similarity, making the results easier to interpret and compare. This approach overcomes the ambiguity and lack of standardization in earlier methods.
[0134] When compared to current industry practices, the methods and systems described herein with the Ido constant offer unique advantages, such providing a consistent benchmark, making the similarity analysis more reliable over the existing methods, such as OpenPose (Cao et al.) which lacks a standardized threshold for pose similarity analysis. The fixed benchmark of the Ido constant ensures consistentsimilarity classification, improving reliability across applications, unlike PoseNet (Google) which does not employ a fixed threshold like the Ido constant, leading to potential inconsistencies. The Ido constant also offers a standardized measure for similarity, enhancing the objectivity and consistency of pose evaluations, unlike DeepPose (Toshev and Szegedy), which does not incorporate a fixed threshold for pose similarity, resulting in variability in assessments.
[0135] Using the similarity calculations, the total similarity score for the full analysis in Table 1 is 91.25%.
[0136] To further improve usability, we introduce a grading system in Table 2 that categorizes similarity percentages into qualitative levels. This enables both technical and non-technical users to interpret results intuitively:
[0137] Table 1: Similarity result showing, target and reference keypoints, and similarity scores.
[0138]
[0139] Table 2: Similarity grading system
[0140] Figure 15 shows a flowchart 500 with example steps for registering a created dance on the system 10, such as the Idometrics platform, on the blockchain, thereby providing a digital copyright of the creator’s dance that is immutable. The example steps comprise extracting the dance data from the database e.g. Idometrics database (step 502) and preparing the blockchain data for blockchain authentication (step 504). The blockchain data may comprise information pertaining to the (i) creator, such as name, user ID; (ii) a URL to the dance video asset in a Cloudinary CDN, or equivalent; (iii) a cover image of the created video; (iv) a title of the dance video and (v) a unique identifier of the created video. This information may be stored in the database, such as the Idometrics database. In the next step metadata is created from the blockchain data to create an NFT using the cover image as the NFT image (step 506). The NFT may be stored on a decentralized data store, such as Inter- Planetary File System (IPFS) using an IPFS URL Each file or dance video has a unique hash that can be compared to a fingerprint, and using its hashvalue to address a file minimizes tampering. Next, the NFT is minted on the blockchain using blockchain network e.g. Polygon platform, generating the blockchain hash (a unique ID used to identify and retrieve an item stored on the blockchain) (step 508). As such, the dance video is associated with a smart contract which includes ownership details of the contract by the original copyright holders. In step 510, the database record of the dance using the blockchain hash e.g. blockchain ID, is stored in the database e.g. Idometrics database.
[0141] In one example, different machine learning classifiers or algorithms are used for building the models, such as, supervised learning algorithms, unsupervised learning algorithms and reinforcement learning algorithms. Examples of supervised learning algorithm systems include support vector machine, decision tree, linearregression, logistic regression, naive Bayes, ^-nearest neighbor, random forest, AdaBoost, XGBoost, and neural network methods. Examples of unsupervised learning algorithm systems include K-means, mean shift, affinity propagation, hierarchical clustering, DBSCAN (density-based spatial clustering of applications with noise), Gaussian mixture modeling, Markov random fields, ISODATA (iterative self-organizing data), and fuzzy C-means systems. Examples of reinforcement learning algorithm systems include Maja and Teaching-Box systems. Generally, training the predictive models involves optimizing the parameters of a predictive system to minimize the loss function. In addition to the training step, the models also undergo validation using test datasets. In one example, the models are trained using supervised learning with tagged input-output pairs, learning to recognize patterns and make informed recommendations. Weights are applied to variables, and the models generate recommendations and insights for matching.
[0142] In one implementation, processor 14 may be embodied as a multi-core processor, a single core processor, or a combination of one or more multi-core processors and one or more single core processors. For example, processor 14 may be embodied as one or more of various processing devices, such as a coprocessor, a microprocessor, a controller, a digital signal processor (DSP), a processing circuitry with or without an accompanying DSP, or various other processing devices including integrated circuits such as, for example, an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), a microcontroller unit (MCU), a hardware accelerator, a special-purpose computer chip, Application-Specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), Programmable Logic Controllers (PLC), Graphics Processing Units (GPUs), and the like. For example, some or all of the device functionality or method sequences may be performed by one or more hardware logic components.
[0143] Memory 16 may be embodied as one or more volatile memory devices, one or more non-volatile memory devices, and / or a combination of one or more volatile memory devices and non-volatile memory devices. The term “machine readable medium” can include any medium that is capable of storing, encoding, orcarrying instructions for execution by the machine 12 and that cause the machine 12 to perform any one or more of the techniques of the present disclosure, or that is capable of storing, encoding or carrying data structures used by or associated with such instructions. Nonlimiting machine-readable medium examples can include solid- state memories, and optical and magnetic media. Specific examples of machine- readable media can include: non-volatile memory, such as semiconductor memory devices (e.g., Electrically Programmable Read-Only Memory (EPROM), Electrically Erasable Programmable Read-Only Memory (EEPROM)) and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; Random Access Memory (RAM); Solid State Drives (SSD); and CD-ROM and DVD-ROM disks. In some examples, machine readable media can include non- transitory machine-readable media. In some examples, machine readable media can include machine readable media that is not a transitory propagating signal.
[0144] The communication interface enables computing system 12 to communicate with other entities over various types of wired, wireless or combinations of wired and wireless networks, such as for example, the Internet. In at least one example embodiment, the communication interface includes a transceiver circuitry configured to enable transmission and reception of data signals over the various types of communication networks. In some embodiments, communication interface may include appropriate data compression and encoding mechanisms for securely transmitting and receiving data over the communication networks. Communication interface facilitates communication between computing system 12 and I / O peripherals.
[0145] It is noted that various example embodiments as described herein may be implemented in a wide variety of devices, network configurations and applications.
[0146] Those of skill in the art will appreciate that other embodiments of the disclosure may be practiced in network computing environments with many types of computer system configurations, including personal computers (PCs), industrial PCs, desktop PCs), hand-held devices, multi-processor systems, microprocessor-based or programmable consumer electronics, network PCs, server computers, minicomputers, mainframe computers, and the like. Accordingly, system 10 may be coupled to theseexternal devices via the communication, such that system 10 is controllable remotely. Embodiments may also be practiced in distributed computing environments where tasks are performed by local and remote processing devices that are linked (either by hardwired links, wireless links, or by a combination thereof) through a communications network. In a distributed computing environment, program modules may be located in both local and remote memory storage devices.
[0147] In another implementation, system 10 follows a cloud computing model, by providing an on-demand network access to a shared pool of configurable computing resources (e.g., servers, storage, applications, and / or services) that can be rapidly provisioned and released with minimal or nor resource management effort, including interaction with a service provider, by a user (operator of a thin client).
[0148] Further, while only a single machine is illustrated, the term “machine” shall also be taken to include any collection of machines that individually or jointly execute a set (or multiple sets) of instructions to perform any one or more of the methodologies discussed herein, such as cloud computing, software as a service (SaaS), other computer cluster configurations.
[0149] Examples, as described herein, can include, or can operate on, logic or a number of components, modules, or mechanisms (all referred to hereinafter as “modules”). Modules are tangible entities (e.g., hardware) capable of performing specified operations and is configured or arranged in a certain manner. In an example, circuits are arranged (e.g., internally or with respect to external entities such as other circuits) in a specified manner as a module. In an example, the whole or part of one or more computer systems (e.g., a standalone, client or server computer system) or one or more hardware processors are configured by firmware or software (e.g., instructions, an application portion, or an application) as a module that operates to perform specified operations. In an example, the software can reside on a non- transitory computer readable storage medium or other machine-readable medium. In an example, the software, when executed by the underlying hardware of the module, causes the hardware to perform the specified operations.
[0150] Accordingly, the term “module” is understood to encompass a tangible entity, be that an entity that is physically constructed, specifically configured (e.g.,hardwired), or temporarily (e.g., transitorily) configured (e.g., programmed) to operate in a specified manner or to perform part or all of any operation described herein. Considering examples in which modules are temporarily configured, each of the modules need not be instantiated at any one moment in time. For example, where the modules comprise a general-purpose hardware processor configured using software, the general-purpose hardware processor is configured as respective different modules at different times. Software can accordingly configure a hardware processor, for example, to constitute a particular module at one instance of time and to constitute a different module at a different instance of time.
[0151] Examples, as described herein, can include, or can operate on, logic or a number of components, modules, or mechanisms. Modules are tangible entities (e.g., hardware) capable of performing specified operations and are configured or arranged in a certain manner. In an example, circuits are arranged (e.g., internally or with respect to external entities such as other circuits) in a specified manner as a module. In an example, the whole or part of one or more computer systems (e.g., a standalone, client, or server computer system) or one or more hardware processors are configured by firmware or software (e.g., instructions, an application portion, or an application) as a module that operates to perform specified operations. In an example, the software can reside on a machine-readable medium. In an example, the software, when executed by the underlying hardware of the module, causes the hardware to perform the specified operations.
[0152] Implementations of the subject matter described in this specification can be implemented in a computing system that includes a back-end component, e.g., as a data server, or that includes a middleware component, e.g., an application server, or that includes a front-end component, e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the subject matter described in this specification, or any combination of one or more such back-end, middleware, or front-end components. The components of the system 10 can be interconnected by any form or medium of wireline and / or wireless digital data communication, e.g., a communications network 30. Examples of communication networks include a local area network (LAN), a radio access network(RAN), a metropolitan area network (MAN), a wide area network (WAN), Worldwide Interoperability for Microwave Access (WIMAX), a wireless local area network (WLAN) using, for example, 802.11 a / b / g / n and / or 802.20, all or a portion of the Internet, and / or any other communication system or systems at one or more locations, and free-space optical networks. The network may communicate with, for example, Internet Protocol (IP) packets, Frame Relay frames, Asynchronous Transfer Mode (ATM) cells, voice, video, data, and / or other suitable information between network addresses.
[0153] There may be any number of computers associated with, or external to, the system 10 and communicating over network 30. Further, the terms “client,” “user,” and other appropriate terminology may be used interchangeably, as appropriate, without departing from the scope of this disclosure.
[0154] In another implementation, system 10 follows a cloud computing model, by providing an on-demand network access to a shared pool of configurable computing resources (e.g., servers, storage, applications, and / or services) that can be rapidly provisioned and released with minimal or nor resource management effort, including interaction with a service provider, by a user (operator of a thin client).
[0155] The flowchart and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagrams may represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logical function(s). In some alternative implementations, the functions noted in the block may occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flowchart illustration, and combinations of blocks in the block diagrams and / or flowchart illustration, can be implemented by special purpose hardware-based systems that perform the specified functions or acts or carry out combinations of special purpose hard-ware and computer instructions.
[0156] Benefits, other advantages, and solutions to problems have been described above with regard to specific embodiments. However, the benefits, advantages, solutions to problems, and any element(s) that may cause any benefit, advantage, or solution to occur or become more pronounced are not to be construed as critical, required, or essential features or elements of any or all the claims. As used herein, the terms "comprises," "comprising," or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but may include other elements not expressly listed or inherent to such process, method, article, or apparatus. Further, no element described herein is required for the practice of the disclosure unless expressly described as "essential" or "critical."
[0157] The preceding detailed description of example embodiments of the disclosure makes reference to the accompanying drawings, which show the example embodiment by way of illustration. While these example embodiments are described in sufficient detail to enable those skilled in the art to practice the disclosure, it should be understood that other embodiments may be realized, and that logical and mechanical changes may be made without departing from the spirit and scope of the disclosure. For example, the steps recited in any of the method or process claims may be executed in any order and are not limited to the order presented. Thus, the preceding detailed description is presented for purposes of illustration only and not of limitation, and the scope of the disclosure is defined by the preceding description, and with respect to the attached claims.
[0158] REFERENCES[1] Z. Cao, T. Simon, S. E. Wei, and Y. Sheikh, "OpenPose: Realtime Multi-Person 2D Pose Estimation using Part Affinity Fields," in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, no. 1, pp. 172-186, Jan. 2021.[2] S. E. Wei, V. Ramakrishna, T. Kanade, and Y. Sheikh, "Convolutional Pose Machines," in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016, pp. 4724-4732.[3] A. Toshev and C. Szegedy, "DeepPose: Human Pose Estimation via Deep Neural Networks," in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2014, pp. 1653-1660.[4] M. Zhang, Y. Xiong, and G. Li, "Dance Recognition Using Enhanced Keypoint Matching and Pose Tracking," in Journal of Visual Communication and Image Representation, vol. 58, pp. 486-496, Jan. 2019.[5] M. S. Jang et al., "System and Method for Searching Choreography Database Based on Motion Inquiry," U.S. Patent 9,349,847, Mar. 24, 2015.[6] T. B. Seo, "Al-based PoseEstimation Models to Evaluate Dance Cover Similarity," Brunch, 2021. [Online], Available: https : / / brunch, co. kr / @tobesoft-ai / 18[7] "Move Mirror," Google Al Experiments, 2018. [Online]. Available: https:ZZmedium.comZtensorflowZmove-mirror-an-ai-experiment-with-pose- estimation-in-the-browser-using-tensorflow-is-2f7b769f9b23[8] Patent US20180131424A1: "System and Method for Real-Time Human Pose Estimation," U.S. Patent Application Publication, May 10, 2018.[9] Patent US10628501B2: "Method and Apparatus for Human Pose Detection," U.S. Patent, Apr. 21, 2020.
[0010] Patent W02020204367A1: "Human Pose Comparison System and Method," World Intellectual Property Organization, Oct. 8, 2020.
[0011] TensorFlow Lite Examples: Pose Estimation, [Online], Available: https:ZZwww.tensorflow.orgZliteZexamplesZpose estimationZoverview
[0012] G. Papandreou, T. Zhu, N. Kanazawa, A. Toshev, J. Tompson, C. Bregler, and K. Murphy, "Towards accurate multi-person pose estimation in the wild," in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 4903- 4911.
[0013] D. Mehta, S. Sridhar, O. Sotnychenko, H. Rhodin, M. Shafiei, W. Xu, D. Casas, and C. Theobalt, "VNect: Real-time 3D human pose estimation with a single RGB camera," in ACM Transactions on Graphics (TOG), vol. 36, no. 4, Article 44, 2017.
[0014] Google Al Blog, "On-Device, Real-Time Body Pose Tracking with MediaPipe BlazePose," Google, 2020. [Online], Available: https: / / ai.googleblog.com / 2020 / 08 / on-device-real-time-body-pose-tracking.html
[0015] "Adobe Media Server," Adobe, 2023. [Online], Available: https: / / www.adobe.com / products / adobemedia-server.html
[0016] "Red5 Server," Red5, 2023. [Online], Available: https: / / www.red5pro.com /
[0017] "AWS Media Services," Amazon Web Services, 2023. [Online], Available: https: / / aws.amazon.com / media-services /
[0018] "Google Cloud Video Intelligence," Google Cloud, 2023. [Online], Available: https: / / cloud.google.com / video-intelligence
[0019] "Data Security in Cloud Platforms," Security Today, 2023. [Online], Available: https : / / securitytoday. example, com / data-security-cloud
[0020] D. Papandreou, T. Zhu, N. Kanazawa, A. Toshev, J. Tompson, C. Bregler, and K. Murphy, "PoseNet: A Convolutional Network for Real-Time 6-DOF Camera Relocalization," in IEEE International Conference on Computer Vision (ICCV), 2015.
[0021] Caba Heilbron, F., Escorcia, V., Ghanem, B., & Carlos Niebles, J. (2015). Activitynet: A large-scale video benchmark for human activity understanding. Proceedings of the IEEE conference on computer vision and pattern recognition,USA,961-970. Doi: 10.1109 / CVPR.2015.7298698.
[0022] Government of Canada. (2022). 2022 Exploration Competition. Government ofCanada.https: / / www.sshrc-crsh.gc.ca / funding-financement / nfrf fnfr / exploration / 2022 / competition-concours-eng.aspx
[0023] Hendry, D., Chai, K., Campbell, A., Hopper, L., O’Sullivan, P., & Straker, L.(2020).Development of a human activity recognition system for ballet tasks. Sports medicineopen, 6(1), 1-10.
[0024] Johnson, A. (2021). Copyrighting TikTok Dances: Choreography in the Internet Age. Washington Law Review, 96(3), 1225-1274.
[0025] Joshi, M., & Chakrabarty, S. (2021). An extensive review of computational dance automation techniques and applications. Proceedings of the Royal Society A,477(2251),20210071.
[0026] Kay, W., Carreira, J., Simonyan, K., Zhang, B., Hillier, C., Vijayanarasimhan, s.,Viola, F., Green, T., Back, T., Natsev, P., Suleyman, M. & Zisserman, A. (2017). The kinetics human action video dataset. https: / / doi.org / 10.48550 / arXiv.1705.06950
[0027] Robert, Y. (2020). JaQuel Knight is paving the way for the future of copyrighting dance. Forbes Magazine. JaQuel Knight Is Paving The Wav For The Future OfCopyrighting Dance (forbes.com)
[0028] Rosenblatt, K. (2021). The Renegade dance made Jalaiah Harmon a star. ‘I Am: Jalaiah’ explores her life since. NBC News, https: / / www.nbcnews.com / pop- culture / pop-culture-news / renegade-dance-made-ialaiah-harmon-star-i-am- ialaiah-explores-n!281412
[0029] Samuel, R. E. (2021). ‘Give Black people credit’: Black TikTok stars strike, demand credit for their work. Los Angeles Times, https : / / www. latimes . com / entertainment- arts / story / 2021-07-02 / give-black-people-credit-black-tiktok-creator-are-on- strike-and-demand-change
[0030] The Canadian Alliance of Dance Artists. (2022). Professional Standards forDance.Canadian Alliance of Dance Artists. https: / / cadaontario.wildapricot.org / psd 11 copyright
[0031] Simpson, T. T., Wiesner, S. L., & Bennett, B. C. (2014). Dance recognition system using lower body movement. Journal of Applied Biomechanics, 30(1), 147-153.
[0032] Soomro, K., Zamir, A. R., & Shah, M. (2012). UCF101: A dataset of 101 human actions classes from videos in the wild. https: / / www.crcv.ucf.edu / papers / UCF101 CRCV-TR-12-01.pdf
[0033] Vargas, S. (2021). What Copyright Protections Do Choreographers Have OverTheirWork? Dance Magazine, https : / / www. dancemagazine, com / choreography-copyright /
[0034] Wicker, J. (2020). Renegade Creator Jalaiah Harmon on Reclaiming the Viral Dance. Teen Vogue, https: / / www.teenvogue.com / story / ialaiah-harmon-renegade- creator-viral-dance
Claims
CLAIMS:
1. A method for creating a reference video associated with at least one first action by processing circuitry and a memory device having instructions executable by the processing circuitry to carry out the steps of: receiving a plurality of first sequential images associated with the at least one first action, preprocessing the plurality of first sequential images; extracting frames from the plurality of sequential images to generate feature vectors associated with at least one pose; with at least one trained machine learning model, detecting first keypoints pertaining to the at least one first pose and generating a first pose data set associated with the at least one first pose; filtering first keypoints based on confidence scores to generate a refined first pose data set; and on the memory device, storing the refined first pose data set, the plurality of first sequential images, extracted frames, and audio information descriptive of the at least one first action, and metadata.
2. The method of claim 1, wherein the at least one first pose is associated with a dance choreography.
3. The method of claim 2, wherein the first keypoints correspond to anatomical body parts.
4. The method of claim 3, wherein the anatomical body parts comprise at least one of a nose, eye, ear, shoulder, elbow, wrist, hip, knee, and ankle.
5. The method of claim 3, wherein the first key points are linked to form a first skeletal structure.
6. The method of claim 5, the processing circuity executes a set of instructions in memory to filter unreliable keypoints based on the confidence scores thereby increasing accuracy and reliability of pose estimation.
7. The method of claim 6, wherein a target video is compared to the reference video using at least another machine learning model to determine a similarity score between the reference video and the target video.
8. The method of claim 6, wherein a target video is compared to other videos stored in a data repository using at least another machine learning model to determine a similarity score between the target video and the other videos.
9. The method of claim 6, wherein the processing circuity executes a set of instructions to combine and store the refined first pose data pose data set and landmark data with user-provided storytelling information.
10. The method of claim 9, wherein the metadata comprises at least one of details of a creator of the reference video, date stamp and time stamp.
11. The method of claim 1 , wherein the step of extracting frames from the plurality of sequential images comprises further steps of at least formatting the videos by extracting frames at regular intervals to standardize the frame rate, thereby enhancing accuracy of pose tracking throughout the video.
12. The method of claim 1, wherein preprocessing the plurality of first sequential images comprises preprocessing steps of:(a) resizing by cropping and scaling each image based on bounding box coordinates of a person;(b) normalizing keypoint coordinates as L2 normalized vector arrays, thereby enabling consistent pose representation across different image sizes and making the poses comparable across different frames and conditions.
13. The method of claim 12, wherein a weighted distance metric is applied to account for variability in confidence levels of the detected keypoints and improves the accuracy of comparing two sets of pose key points.
14. The method of any one of claims 7 to 10, wherein the target video is generated with the steps of: receiving a plurality of second sequential images associated with at least one second action, preprocessing the plurality of second sequential images; extracting frames from the plurality of second sequential images to generate feature vectors associated with at least one second pose; with at the least one trained machine learning model, detecting second keypoints pertaining to the at least one second pose and generating a second pose data set associated with the at least one second pose; filtering the second keypoints based on confidence scores to generate a refined second pose data set; and on the computer readable medium, storing the refined second pose data, plurality of second sequential images, extracted frames, and metadata.
15. The method of claim 14, wherein the at least another machine learning model determines the similarity score based on at least performing spatial alignment and normalization of the first keypoints and the second keypoints, performing weighted distance calculation of first keypoints and the second keypoints, and a similarity scoring constant, thereby providing a consistent threshold for comparing the reference video and the target video.
16. The method of claim 15, wherein when keypoints are not detected in either the reference or the target pose, assigning to NaN (Not a Number) values or missing data thereby allowing a calculation to proceed without interruption, such that all keypoints, whether detected or missing, contribute to a final similarity score.
17. The method of claim 16, comprising a benchmark threshold for standardizing similarity classification, such that the similarity score is based on a weighted distance (w), the benchmark threshold (z) and a difference (d).
18. The method of claim 7, wherein the least one trained machine learning model analyzes the target video to provide information about a possible genre or cultural background of the dance choreography.
19. The method of claim 7, wherein the least one trained machine learning model analyzes the target video to provide information pertaining to a motivation and a mood of a dancer.
20. A computer system comprising a hardware processor and a memory device on which instructions are encoded to cause the hardware processor to perform the operations of: receiving a plurality of first sequential images associated with the at least one first action, preprocessing the plurality of first sequential images; extracting frames from the plurality of first sequential images to generate feature vectors associated with at least one pose; with at least one trained machine learning model, detecting first keypoints pertaining the at least one first pose and generating a first pose data set associated with the at least one first pose; filtering the first keypoints based on confidence scores to generate a refined first pose data set; and on a computer readable medium, storing the refined first pose data set, plurality of first sequential images, extracted frames, and audio information descriptive of the at least one first action, and metadata.
21. The computer system of claim 20, wherein the hardware processor executes a set of instructions to preprocess the plurality of first sequential images by croppingand scaling each image based on a bounding box coordinates of a person; normalizing key point coordinates as L2 normalized vector arrays.
22. The computer system of claim 20, wherein the hardware processor executes a set of instructions to generate refined first pose data set and landmark data, and combine the refined pose data pose data and the landmark data with user-provided storytelling information to form a target video.
23. The computer system of claim 20, wherein the hardware processor executes a set of instructions to associate the target video with at least one of a date-stamp, a time stamp, user information for proper accreditation.
24. The computer system of claim 20, wherein the hardware processor executes a set of instructions to associate the target video with a unique identifier.
25. The computer system of claim 24, wherein the unique identifier is a non- fungible token that is recorded on a blockchain to certify ownership and authenticity.
26. The computer system of claim 20, wherein when the first keypoints are not detected in the at least one first pose, the hardware processor executes a set of instructions to assign zeros to NaN (Not a Number) values.
27. The computer system of claim 22, wherein the target video is compared to a reference video based on a benchmark threshold for pose similarity analysis.
28. The computer system of claim 27, wherein the benchmark threshold provides a standardized measure for similarity, thereby enhancing objectivity and consistency of pose evaluations.
29. A computer-implemented method for comparing a reference video associated with at least one reference action to a target video associated with at least one target action, the method comprising the steps of: receiving a plurality of first sequential images associated with the at least one reference action, preprocessing the plurality of first sequential images; receiving a plurality of second sequential images associated with the at least one target action, preprocessing the plurality of second sequential images; and using at least another machine learning model to determine a similarity score between the reference video and the target video based on at least one of a benchmark threshold for standardizing similarity classification, a weighted distance, and a difference.
30. The computer-implemented method of claim 29, wherein a total similarity score (St ) for a sequence of n pose comparisons can then be expressed as an average of the total percentage similarity across all frames of the plurality of first sequential images and the the plurality of second sequential images.
31. The computer-implemented method of claim 30, wherein the total similarity score (St ) for a sequence of n pose comparisons is defined as:
32. A computer-implemented method for assigning a unique identifier to a dance video, the method comprising the steps of: extracting the dance video data from a data repository comprising a database record associated with the dance video; preparing blockchain data for blockchain authentication; creating metadata from the blockchain data for creating a non-fungible token (NFT) using a cover image of the dance video as an NFT image and storing the NFT on a distributed file system; minting the NFT on the blockchain using a blockchain network;generating a blockchain hash associated with the dance video stored on the blockchain; and updating the database record of the dance video using the blockchain hash in the data repository.
33. The computer-implemented method of claim 32, wherein the blockchain hash enables appropriate attribution and compensation to an owner of the dance video.
Citation Information
Patent Citations
Method for comparing similarity of motion postures in video
CN110458235A
Audio analysis learning with video data
US10573313B2
Lighttrack: system and method for online top-down human pose tracking
US11288835B2
Collaborative video non-fungible tokens and uses thereof
US11367060B1