Interaction method, electronic device, storage medium, and computer program product

By displaying the first multimedia content and user request information, and using image recognition and deep learning models to generate the second multimedia content and response information, the problem of motion feedback for users when learning dance or yoga on their own is solved, improving learning efficiency and the immediacy of guidance.

WO2025260276A1PCT designated stage Publication Date: 2025-12-26BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/100058
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-06-19
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

When users learn dance or yoga on their own, it is difficult to obtain efficient feedback and guidance on movements, especially when the movements are not standard and it is difficult for them to detect and correct them on their own. Online teaching also cannot provide effective guidance anytime and anywhere.

Method used

By displaying first multimedia content and user request information, and using image recognition and deep learning models to process actions, second multimedia content and response information are generated to provide action evaluation and choreography suggestions.

Benefits of technology

It improves users' learning efficiency in dance and yoga movements, achieves accurate movement analysis and efficient feedback, and meets users' needs for instant guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024100058_26122025_PF_FP_ABST
    Figure CN2024100058_26122025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to the technical field of computers, and relates to an interaction method, an electronic device, a storage medium, and a computer program product. The interaction method comprises: displaying first multimedia content from a user and request information of the user based on the first multimedia content, the first multimedia content comprising one or more actions; on the basis of the request information, processing the one or more actions in the first multimedia content to determine second multimedia content and response information for the request information; and displaying the second multimedia content and the response information.
Need to check novelty before this filing date? Find Prior Art

Description

Interaction method, electronic device, storage medium and computer program product TECHNICAL FIELD

[0001] The present disclosure relates to the field of computer technology, and in particular, to an interaction method, an electronic device, a storage medium and a computer program product. BACKGROUND

[0002] With the development of Internet technology, users can learn and practice dance, yoga and other sports by using electronic devices such as mobile phones and computers. For example, a user can imitate the actions in a sports video of another person. Alternatively, the user can also participate in some online courses and be guided by a teacher in real time through video connection.

[0003] SUMMARY

[0004] This summary is provided to introduce a selection of concepts, which are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used in limiting the scope of the claimed subject matter.

[0005] According to some embodiments of the present disclosure, an interaction method is provided, comprising: displaying first multimedia content from a user and request information of the user based on the first multimedia content, the first multimedia content comprising one or more actions; processing the one or more actions in the first multimedia content based on the request information to determine second multimedia content and response information of the request information; and displaying the second multimedia content and the response information.

[0006] According to some embodiments of the present disclosure, an electronic device is provided, comprising: a memory; and a processor coupled to the memory, the processor being configured to execute an interaction method of any of the embodiments described in the present disclosure based on instructions stored in the memory.

[0007] According to some embodiments of the present disclosure, a computer-readable storage medium is provided, having stored thereon a computer program, which, when executed by a processor, performs an interaction method of any of the embodiments described in the present disclosure.

[0008] According to some embodiments of the present disclosure, a computer program is provided, comprising: instructions which, when executed by a processor, cause the processor to perform an interaction method of any of the embodiments described in the present disclosure.

[0009] Other features, aspects, and advantages of the present disclosure will become apparent from the following detailed description of the exemplary embodiments with reference to the following accompanying drawings. BRIEF DESCRIPTION OF DRAWINGS

[0010] Preferred embodiments of the present disclosure are described herein below with reference to the accompanying drawings. The accompanying drawings are used to provide further understanding of the present disclosure, and together with the specific description below, form a part of the description of the present disclosure, and are used to explain the present disclosure. It should be understood that the accompanying drawings described below only relate to some embodiments of the present disclosure, and do not constitute a limitation on the present disclosure. In the drawings:

[0011] FIG. 1 shows a flow diagram of an interaction method according to some embodiments of the present disclosure.

[0012] FIG. 2 shows a flow diagram of a method of second multimedia content and response information according to some embodiments of the present disclosure.

[0013] FIG. 3 shows a flow diagram of a method of second multimedia content and response information according to some other embodiments of the present disclosure.

[0014] FIG. 4 shows a flow diagram of a method of second multimedia content and response information according to yet some other embodiments of the present disclosure.

[0015] FIG. 5 shows a structural diagram of an interaction device according to some embodiments of the present disclosure.

[0016] FIG. 6 shows a structural diagram of an electronic device according to some embodiments of the present disclosure.

[0017] FIG. 7 shows a structural diagram of a computer system according to some embodiments of the present disclosure.

[0018] It should be understood that the size of each part shown in the drawings is not necessarily drawn according to the actual proportion relationship. The same or similar reference numerals are used in the drawings to represent the same or similar parts. Therefore, once a part is defined in one drawing, it can not be further discussed in subsequent drawings. DETAILED DESCRIPTION

[0019] The technical solutions in the embodiments of the present disclosure will be described clearly and completely below with reference to the drawings in the embodiments of the present disclosure. However, it is obvious that the described embodiments are only some of the embodiments of the present disclosure, but not all the embodiments. The description of the embodiments below is actually only illustrative, and should not be considered as any limitation on the present disclosure and its application or use. It should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments described herein.

[0020] It should be understood that various steps in the method implementations of the present disclosure can be performed in different order and / or in parallel. Additionally, the method implementations can include additional steps and / or omit performing the steps shown. The scope of the present disclosure is not limited in this regard. The relative arrangement of components and steps, numerical expressions, and numerical values set forth in these examples should be construed only as examples, not limiting the scope of the present disclosure, unless otherwise specifically stated.

[0021] The term "include," and derivations thereof, means the term "comprise" or "contain" and variations as an open-ended term such that when the phrase "includes (comprises, contains)" is used, it means at least the stated elements, but not excluding others. Further, the term "comprise" and variations thereof as used in the present disclosure means the term "comprise" or "contain" and variations as an open-ended term such that when the phrase "comprises (contains)" is used, it means at least the stated elements, but not excluding others. Thus, include and comprise are synonymous. The term "based on" means "based, at least in part, on."

[0022] Reference throughout this specification to "an embodiment", "some embodiments" or "embodiments" means that a particular feature, structure or characteristic described in connection with the embodiment is included in at least one embodiment of the application. The appearances of the phrases "in one embodiment" or "in some embodiments" or "in embodiments" in various places in the specification are not necessarily all referring to the same embodiment, although they can. Furthermore, the terms "a" or "an", as used herein, mean "one or more" when used in the context of expressing an quantity of objects unless otherwise indicated.

[0023] It should be noted that the terms "first", "second", and so on used in the present disclosure are merely used to distinguish different apparatuses, modules or units, and do not imply the order or the mutual dependency of the functions performed by these apparatuses, modules or units. Unless otherwise specified, the terms "first", "second", and so on are not intended to imply a given order or any other way of given order.

[0024] It should be noted that the terms "one", "multiple", and the like used in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise explicitly specified in the context, "one" or "multiple" should be understood as "one or more".

[0025] The names of the messages or information exchanged between the plurality of apparatuses in the embodiments of the present disclosure are merely used for illustrative purposes, and are not intended to limit the scope of the messages or information.

[0026] Embodiments of the present disclosure will be described in detail below with reference to the drawings, but the present disclosure is not limited to these specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes can not be described in some embodiments. In addition, in one or more embodiments, specific features, structures, or characteristics can be combined by any suitable means apparent from the present disclosure to those of ordinary skill in the art.

[0027] It should be understood that the present disclosure does not limit how to obtain the image to be applied / processed. In one embodiment of the present disclosure, the image can be obtained from a storage device, such as an internal memory or an external storage device, and in another embodiment of the present disclosure, a photographing component can be mobilized to take a picture. It should be noted that the obtained image can be an image collected or a frame image of a video collected, and is not particularly limited thereto.

[0028] In the context of the present disclosure, the image can refer to any of a variety of images, such as a color image, a grayscale image, etc. It should be noted that the type of image is not specifically limited in the context of the present specification. In addition, the image can be any appropriate image, such as an original image obtained by a camera, or an image that has been subjected to a specific process, such as preliminary filtering, de-aliasing, color adjustment, contrast adjustment, normalization, etc. It should be noted that the pre-processing operation can also include other types of pre-processing operations known in the art, which will not be described in detail here.

[0029] When a user learns dance, yoga, or the like from a video in a network, there can be a situation where the motion is not standard, but the user is difficult to self-recognize and correct, and the motion that a beginner can learn is also limited. Online teaching is difficult to initiate at any time and any place, and users are also difficult to obtain guidance at any time and any place, such as correction of non-standard motion, or inspiration for adapting or expanding motion. Therefore, in the related art, a user is difficult to efficiently obtain feedback on multimedia content containing motion.

[0030] To solve the above technical problem, the present disclosure provides an interaction method, an electronic device, a storage medium, and a computer program product. Embodiments of the interaction method of the present disclosure are described below with reference to FIG. 1.

[0031] FIG. 1 shows a flowchart of an interaction method according to some embodiments of the present disclosure. As shown in FIG. 1, the interaction method of this embodiment includes steps S102 to S106.

[0032] In step S102, first multimedia content from a user and request information of the user based on the first multimedia content are displayed, and the first multimedia content includes one or more motions.

[0033] The first multimedia content is image-based multimedia content, for example, one or more images, or one or more videos. The first multimedia content can carry various actions, such as yoga actions, dance actions, fitness actions, martial arts actions, sports actions, and the like. These actions can be demonstrated by the user in step S102, or can be demonstrated by other objects. That is, the first multimedia content can be an image or a video of one or more actions performed by the user in step S102, or can be an image or a video of one or more actions performed by other objects (such as other objects, or virtual objects). The first multimedia content can be locally photographed or uploaded by the user from the user device, or stored in the cloud or server by the user and specified by the user.

[0034] The content of the request information is associated with the first multimedia content, and can be text or voice. The user can input the request information through a text input control, a voice input control, or a selection control. The request information is, for example, an evaluation request, indicating that the first multimedia content is evaluated as a whole or in part. For another example, a generation request indicates that new multimedia content is generated based on the first multimedia content. The generated multimedia content can be a continuation or an extension of the first multimedia content. According to needs, the request information can also include other contents, which are not described herein.

[0035] The actions in the first multimedia content can be determined according to the poses of the objects in the first multimedia content. In some embodiments, the objects in the first multimedia content are identified; and one or more actions of the objects in the first multimedia content are determined according to the poses of the objects. For example, the objects and the poses of the objects in the first multimedia content can be determined by image recognition of the images in the first multimedia content, for example, using target detection and tracking algorithms, convolutional neural networks (CNN), and the like to recognize the poses, and then classifying based on the poses. In addition, when the first multimedia content is a video, or includes multiple images, the first multimedia content can be regarded as sequential data, and a recurrent neural network (RNN) or the like model for processing sequential data can be used to obtain the determination result of the actions.

[0036] In step S104, one or more actions in the first multimedia content are processed based on the request information to determine the second multimedia content and the response information of the request information.

[0037] Based on the request information, a processing manner of one or more actions in the first multimedia content can be determined. In some embodiments, which action in the one or more actions in the first multimedia content is processed can also be determined based on the request information. The processing manner can be to make an evaluation; or, it can also be to make a prediction, such as predicting other actions to generate new multimedia content.

[0038] The second multimedia content is a video or one or more images. The second multimedia content is related to the first multimedia content, and also includes one or more actions. For example, it includes the same actions as the first multimedia content, or related actions. The type of the second multimedia content can be the same as the type of the first multimedia content. For example, in the case where the first multimedia content is a video, the second multimedia content is a video; in the case where the first multimedia content is an image, the second multimedia content is one or more images. Of course, the type of the second multimedia content can also be different from the type of the first multimedia content, as needed.

[0039] The response information is used to interpret or describe at least one of the first multimedia content, the second multimedia content, and is capable of responding to the request information. In other words, the request information and the response information are based on the same subject. For example, the request information indicates to make an evaluation on the actions in the first multimedia content, and the response information can include the evaluation content in combination with at least one of the first multimedia content, the second multimedia content; for another example, the request information indicates to generate the second multimedia content based on the actions in the first multimedia content, and the response information can include an interpretation of the generated second multimedia content. The response information can be text or voice.

[0040] In some embodiments, in the case where the request information is voice, the text in the voice can be obtained through voice recognition first, and then the text is processed. Of course, in the case where the model for processing the request information supports voice input, the conversion from voice to text can also not be performed in advance.

[0041] In determining the second multimedia content, a generative model can be used. For example, based on the request information, one or more generative models are used to process one or more actions in the first multimedia content to determine the second multimedia content and the response information of the request information. The generative model is used to output target content based on input information. The input information includes the basis for processing in the generation process of the generative model, such as the request information (or the text corresponding to the request information), the requirement for the output content, and the like. The generative model used in the present application includes, for example, a model for generating text (or voice) and image (or video) based on text (or voice) and image (or video), that is, a generative model that supports multi-modal input and multi-modal output. Of course, a generative model with multi-modal input and single-modal output, or multiple types of models can also be used. For example, first, the request information (or the text corresponding to the request information) and the first multimedia content are processed by using a first generative model to generate the second multimedia content; and then the second multimedia content is processed by using a second generative model to generate the response information. In some embodiments, the second multimedia content can be generated by using generative adversarial networks (GAN) and variation autoencoder (VAE), policy gradient algorithm, Q-learning algorithm, and the like.

[0042] In some embodiments, a deep learning model is used to construct an action prediction model to determine the action included in the second multimedia content. The model includes an input layer, a hidden layer, and an output layer. The input layer is used to receive the feature vector of the music (optional) and the action, which can be the action in the first multimedia content or the determined action in the second multimedia content to be generated; the hidden layer is used to learn the characteristics and rules of the music (optional) and the action; and the output layer is used to generate the prediction result of the action. During the model training process, a large amount of music and action data can be used as training data, and a back propagation algorithm can be used to optimize the parameters of the model. By continuously adjusting the parameters of the model, the model can better learn the characteristics and rules of the music and the action, thereby improving the accuracy of action prediction. Then, the accuracy, recall rate, F1 value and the like can be used to evaluate the performance of the model. Through the evaluation of the model, the advantages and disadvantages of the model can be understood, and the model can be further optimized and improved.

[0043] In step S106, the second multimedia content and the response information are displayed.

[0044] Steps S102 to S106 can be all executed by the user device, or steps S102 and S106 can be executed at the user device, and step S104 is executed on another device different from the user device and sends the execution result to the user device. The user device can be a tablet computer, a mobile phone (such as a folding screen mobile phone, a large-screen mobile phone, etc.), a wearable device, a vehicle-mounted device, an augmented reality (AR) / virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a smart television, a smart screen, a high-definition television, a 4K television, a smart projector, and the like Internet of Things (IOT) device, and the specific type of the user device is not limited in the present disclosure.

[0045] In the above embodiments, by processing the multi-modal input from the user, the first multimedia content is parsed based on the request information, and the multi-modal output is generated and displayed to the user, so as to respond to the request of the user in the form of the second multimedia content and the response information. Therefore, the embodiments of the present disclosure can automatically parse the multimedia content including actions, and generate the corresponding multimedia content and response information, so as to facilitate the user to more efficiently obtain the parsing result of the action in the first multimedia content, and improve the learning efficiency of the user on the action such as yoga and dance.

[0046] In some embodiments, the processing manner can be determined based on the topic of the request information. The embodiments of the present disclosure for determining the second multimedia content and the response information are described below with reference to FIG. 2.

[0047] FIG. 2 shows a flow diagram of a method for generating the second multimedia content and the response information according to some embodiments of the present disclosure. As shown in FIG. 2, the method of this embodiment includes steps S202 to S204.

[0048] In step S202, the request information is subjected to intent recognition to determine the topic of the request information.

[0049] The topic of the request information reflects the key content in the request information, for example, the topic can be the category to which the request information belongs, the keyword in the request information, the abstraction of the content of the request information, etc. The topic of the request information can be determined from several preset topics, or can be freely generated.

[0050] In some embodiments, the request information can be processed by a topic analysis model to obtain a topic thereof; or processed by a classification model to obtain a category to which the request information belongs as the topic. In some embodiments, a generative model can also be used to determine the topic of the request information. For example, the request information and a processing instruction are output to a generative model for determining the topic, the processing instruction is used to instruct the generative model to determine the topic of the request information, and the result returned by the generative model is the topic.

[0051] In step S204, one or more actions in the first multimedia content are processed based on the topic of the request information to determine the second multimedia content and the response information of the request information.

[0052] Different topics can correspond to the same processing manner, but affect the processing basis in the process of generating the second multimedia content and the response information. For example, in the process of using one or more generative models for generating the second multimedia content and the response information, the topic (or the processing instruction determined based on the topic) and the one or more actions are input to the generative model, so that different topics affect the generation result.

[0053] Different topics can also correspond to different processing manners. For example, for several known topics, a generative model corresponding to each topic can be respectively set. In generating the second multimedia content and the response information, the generative model corresponding to the topic of the request information is selected to process the one or more actions.

[0054] Through the above embodiments, the one or more actions can be processed according to the topic of the request information, so that the generated second multimedia content and the response information can accurately respond to the user, and the information acquisition efficiency and the learning efficiency are improved.

[0055] The disclosure is exemplarily described in more detail in the two scenarios of "evaluation" and "arranging actions". Of course, it should be clear to those skilled in the art that the disclosure is also applicable to other application scenarios, which will not be described here.

[0056] FIG. 3 shows a flow diagram of a method of generating the second multimedia content and the response information according to some other embodiments of the disclosure. In this embodiment, the request information represents that the user wants to evaluate the actions in the first multimedia content. For example, the topic of the request information is evaluation. As shown in FIG. 3, the method of this embodiment includes steps S302 to S306.

[0057] In step S302, a target action in the one or more actions of the first multimedia content is determined.

[0058] In some embodiments, the target action can not be specified by the user in the request information. The target action can be an action that has a problem. For example, the request information is "help me check if there is a problem in this dance", which does not include an indication of a certain action, and needs to be processed by the machine to identify the action that has a problem.

[0059] The determination of whether the target action needs to be improved can be whether the target action itself is standard, or whether the target action matches the music in the first multimedia content.

[0060] In some embodiments, the matching result of one or more actions in the first multimedia content and the sound in the first multimedia content is determined; the action that does not match the sound is determined as the target action.

[0061] For example, the rhythm of the action and the rhythm of the sound (such as the background music of the video) in the first multimedia content can be referred to for matching. The action that produces inconsistent rhythm is the action that does not match the sound. The rhythm of the action can be determined according to the changes between different actions, the amplitude of the action, etc., for example, by processing the video or image recognition model to obtain the generation time of each action (or sub-action) in one or more actions and sub-actions. The rhythm of the sound can be determined according to the beat or accent of the sound, for example, by processing the audio signal by a sound processing model.

[0062] For another example, the target action can be determined by referring to whether the emotion or semantics of the action matches the emotion or semantics of the sound. If the action is hot and vigorous, but the background music when the action is performed is slow, the action can also be determined as the target action. When analyzing the emotion or semantics of the action or the sound, the action or the sound can be first converted into a feature vector, and then a semantic analysis model is used to process the feature vector to obtain a semantic analysis result.

[0063] In the case where the action does not match the sound, even if the action itself can be standard, it can also be determined as the target action that has a problem.

[0064] In some embodiments, reference multimedia content corresponding to one or more actions is determined; the reference multimedia content is compared with first multimedia content to determine the target action with a problem. The reference multimedia content may be multimedia content including the correct action. It may be read from a database, searched through a search engine, or generated based on stored or searched descriptions of the correct action. During the comparison, features of the action in the first multimedia content (e.g., force, angle, amplitude, etc.) and features of the action in the reference multimedia content can be extracted separately, and the presence of a problem with the action can be determined based on the comparison results of the features. Of course, if the features of the correct action are pre-stored, the features of the action in the first multimedia content can be directly compared with them without the aid of the reference multimedia content.

[0065] The target action can also be an action specified in the request information. In some embodiments, the action that matches the action indicated in the request information among one or more actions is determined as the target action. For example, if the request information is "Help me check if my cartwheel is standard" or "Is there a problem with the second action?", then the target actions are "cartwheel" and "the second action" explicitly indicated in the request information, respectively. The machine can then determine the position of the target action in the second multimedia content based on this indication.

[0066] In some embodiments, a dance movement evaluation model can be used to identify target movements or determine problems with those movements. The dance movement evaluation model may include an input layer, a hidden layer, and an output layer. The input layer receives feature vectors of the movement, such as the type, speed, and intensity of the movement. The hidden layer learns the characteristics and patterns of the movement, such as its fluidity, coordination, and rhythm. The output layer evaluates the standardization of the movement and provides guidance, such as the correctness, aesthetics, and suggestions for improvement of the dance movement.

[0067] In step S304, the second multimedia content is determined based on the target action.

[0068] The second multimedia content includes the target action. The second multimedia content may include the part of the first multimedia content involved in the target action, such as the relevant image or video segment; it may also include multimedia content after the action has been corrected; or it may include both of the above.

[0069] For the case where the action and the sound do not match, in some embodiments, the second multimedia content is generated by adjusting the target action in the first multimedia content so that the target action matches the sound. That is, the second multimedia content is the result of adjusting the first multimedia content so that the user can more easily compare and learn, and the efficiency of information acquisition is improved. For example, the first multimedia content is a video of object A practicing dancing in a classroom, and the second multimedia content can still be the video of object A practicing dancing in the classroom, except that some non-standard actions have been corrected. Of course, the second multimedia content can or can not include the part corresponding to the non-target action in the first multimedia content, i.e., the part where the action is not a problem.

[0070] In some embodiments, the second multimedia content including the target action can also be generated or searched. For example, according to the type, name or feature of the target action, the features of the correct action corresponding to the target action can be determined, and the second multimedia content including the correct action is generated using these features. For another example, the second multimedia content including the correct action can be searched using the type, name or feature of the target action.

[0071] In step S306, at least one of the evaluation information of the target action in the first multimedia content and the explanation information of the second multimedia content is generated as the response information.

[0072] The evaluation information is used to describe at least one of the problem of the target action, the cause of the problem, and the method of correction, and the explanation information is used to describe at least one of the key points of the correct action corresponding to the target action and the method of correction of the target action. For example, the evaluation information or the explanation information can be generated by the difference information between the target action in the first multimedia content and the correct action corresponding to the target action in the second multimedia content.

[0073] According to the need, the evaluation information and the explanation information can also appear in the second multimedia content in the form of sound, subtitles, indicators, etc.

[0074] Through the above embodiments, the target action in the first multimedia content can be accurately located and analyzed based on the action in the first multimedia content, and the information acquisition efficiency and the learning efficiency of the user are improved.

[0075] FIG. 4 shows a flowchart of a method of the second multimedia content and the response information according to yet some embodiments of the present disclosure. In this embodiment, the request information indicates that the user wants to evaluate the action in the first multimedia content. For example, the subject of the request information is evaluation. As shown in FIG. 4, the method of this embodiment includes steps S402 to S406.

[0076] In step S402, in response to the subject of the request information being choreography action, one or more target actions are predicted based on the one or more actions.

[0077] The choreography action refers to generating a combination of one or more target actions, which can or can not include part or all of the one or more actions in the first multimedia content. In the process of prediction, in addition to referring to the one or more actions, the content in the request information and at least one of the first multimedia content can also be referred to, for example, a description of the choreography action in the request information, background music desired to be used in the choreographed target action, music in the first multimedia content, and the like.

[0078] In some embodiments, one or more target actions are predicted after the one or more actions. Thus, the first multimedia content can be continued. For example, the user can give the beginning of a dance or a series of actions, and the machine can perform subsequent action choreography.

[0079] In some embodiments, one or more target actions including one or more actions are generated. Thus, the first multimedia content can be extended. For example, the user can give a key action in a dance or a series of actions, and the machine can extend it.

[0080] In predicting the target action, a prediction model such as a deep learning model or the like can be used. For example, input data can be generated according to the one or more actions in the first multimedia content; the input data is input into the prediction model; and one or more target actions output by the prediction model are obtained. The prediction model can include, for example, an input layer, a hidden layer, and an output layer. The input layer is used to receive a feature vector of the one or more actions, the hidden layer is used to learn the features and rules of music and dance actions, and the output layer is used to generate a prediction result of the dance action.

[0081] The generated one or more target actions can refer to generating the name, description information, specific features (such as the position, angle, amplitude, duration, and the like of the action) or vector-represented information of the target action, or generating a video or image of each action in the one or more target actions.

[0082] In step S404, the second multimedia content is generated based on the target action. For example, a generative model capable of generating multimedia content can be used to process the target action to obtain the generated second multimedia content.

[0083] For example, a model capable of generating multimedia content can be used to process the name, description information, or image of the target action to generate a video or image corresponding to the target action.

[0084] In step S406, the explanation information of the action in the second multimedia content is generated as the response information.

[0085] The explanation information can be generated in units of target actions, that is, the explanation information is generated for each of part or all of the target actions, or can be generated based on the whole of the second multimedia content, or both. For example, description information of the target action can be generated or searched as the explanation information, or description information of the second multimedia content can be generated based on the description information of the target action as the explanation information.

[0086] Through the above embodiments, new actions can be arranged based on the first multimedia content to generate new second multimedia content. Therefore, the user can efficiently realize action arrangement such as choreography, and improve the efficiency of information acquisition and action learning.

[0087] In some embodiments, the subject of the request information can also be "recommendation". For example, the user can provide one or more personal photos as the first multimedia content, and the request information indicates that appropriate dances or exercises are recommended according to the photos, and can also include some specific descriptions or requirements. Then, the second multimedia content can be a second multimedia content recommended for the user, which is suitable for the physical condition of the object in the photo, and the response information can be the reason for recommending the second multimedia content, and so on.

[0088] Some application scenarios of the embodiments of the present disclosure are described below.

[0089] In some embodiments, the user can send the request information and the first multimedia content through the dialogue interface with the robot. After determining the second multimedia content and the response information, the robot can send the second multimedia content and the response information in the dialogue interface. The robot can generate content corresponding to the dialogue sent by other subjects (such as other users or robots participating in the dialogue) in the dialogue scene. It can be implemented in software, hardware, or a combination of software and hardware. The robot can also be called a digital person, a virtual agent of a machine learning model. The robot can be implemented based on a machine learning model, such as a large language model (Large Language Model, LLM) or a foundation model (Foundation Model). The machine learning model can be a generative model, which is used to output target content based on input information.

[0090] In some embodiments, the user can also input the request information through a form in the application interface, and upload or specify the first multimedia content. Then, the second multimedia content and the response information can be displayed to the user in the same interface or a new interface.

[0091] Embodiments of the present disclosure can be applied to dialogue applications, sports applications, dance applications, etc. of robots. Of course, other applications can also be applied as needed, which will not be described here.

[0092] The above is an embodiment of the interactive method of the present disclosure. Next, the device and apparatus for performing the interactive method of the present disclosure are introduced.

[0093] FIG. 5 shows a structural schematic diagram of an interactive device according to some embodiments of the present disclosure. As shown in FIG. 5, the interactive device 5 of this embodiment includes a first display module 501 configured to display first multimedia content from a user and request information of the user based on the first multimedia content, the first multimedia content including one or more actions; a processing module 502 configured to process the one or more actions in the first multimedia content based on the request information to determine second multimedia content and response information of the request information; and a second display module 503 configured to display the second multimedia content and the response information.

[0094] In some embodiments, the processing module 502 is further configured to perform intent recognition on the request information to determine a subject of the request information; and process the one or more actions in the first multimedia content based on the subject of the request information to determine the second multimedia content and the response information of the request information.

[0095] In some embodiments, the processing module 502 is further configured to determine a target action in the one or more actions of the first multimedia content in response to the subject of the request information being an evaluation, wherein the target action is an action with a problem or an action indicated by the request information; determine the second multimedia content based on the target action; and generate at least one of evaluation information of the target action in the first multimedia content or explanation information of the second multimedia content as the response information.

[0096] In some embodiments, the processing module 502 is further configured to determine a matching result of the one or more actions in the first multimedia content with a sound in the first multimedia content; and determine an action that does not match the sound as the target action.

[0097] In some embodiments, the processing module 502 is further configured to generate the second multimedia content by adjusting the target action in the first multimedia content so that the target action matches the sound.

[0098] In some embodiments, the processing module 502 is further configured to determine reference multimedia content corresponding to the one or more actions; and compare the reference multimedia content with the first multimedia content to determine the target action with a problem.

[0099] In some embodiments, the processing module 502 is further configured to determine, from the one or more actions, an action matching an action indicated by the request information as a target action.

[0100] In some embodiments, the processing module 502 is further configured to generate or search for the second multimedia content including the target action.

[0101] In some embodiments, the processing module 502 is further configured to, in response to a subject of the request information being an arrangement action, predict one or more target actions based on the one or more actions; generate the second multimedia content based on the target action; and generate explanation information of the action in the second multimedia content as the response information.

[0102] In some embodiments, the processing module 502 is further configured to predict one or more target actions after the one or more actions; or generate one or more target actions including the one or more actions.

[0103] In some embodiments, the interaction device 5 further includes a determination module 504 configured to identify an object in the first multimedia content; and determine one or more actions of the object in the first multimedia content according to a pose of the object.

[0104] In some embodiments, the request information is text or voice.

[0105] In some embodiments, the first multimedia content is a video or one or more images; and the second multimedia content is a video or one or more images.

[0106] It should be noted that each unit described above is only a logical module according to the specific function implemented by it, and is not used to limit the specific implementation manner, for example, it can be implemented in software, hardware or a combination of software and hardware. In actual implementation, each unit described above can be implemented as an independent physical entity, or can be implemented by a single entity (for example, a processor (CPU or DSP, etc.), an integrated circuit, etc.). In addition, each unit described above is indicated by a dashed line in the drawings, indicating that these units can not actually exist, and the operations / functions implemented by them can be implemented by the processing circuit itself.

[0107] Further, although not shown, the device can also include a memory that can store various information generated by the device, the various units included in the device in operation, programs and data for operation, data to be transmitted by the communication unit, and the like. The memory can be a volatile memory and / or a non-volatile memory. For example, the memory can include, but is not limited to, a random access memory (RAM), a dynamic random access memory (DRAM), a static random access memory (SRAM), a read only memory (ROM), a flash memory. Of course, the memory can also be located outside the device. Alternatively, although not shown, the device can also include a communication unit that can be used to communicate with other devices. In one example, the communication unit can be implemented in an appropriate manner known in the art, for example, including a communication component such as an antenna array and / or a radio frequency link, various types of interfaces, communication units, and the like. Here will not be described in detail. In addition, the device can also include other components not shown, such as radio frequency links, baseband processing units, network interfaces, processors, controllers, and the like. Here will not be described in detail.

[0108] Some embodiments of the present disclosure also provide an electronic device. FIG. 6 shows a structural schematic diagram of an electronic device according to some embodiments of the present disclosure. For example, in some embodiments, the electronic device 6 can be various types of devices, for example, can include but is not limited to various types of mobile terminals such as mobile phones, notebook computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablets), PMPs (portable multimedia players), vehicle terminals (such as vehicle navigation terminals), and the like, as well as fixed terminals such as digital TVs, desktop computers, and the like. For example, the electronic device 6 can include a display panel for displaying data and / or execution results utilized in the scheme according to the present disclosure. For example, the display panel can be various shapes, such as a rectangular panel, an oval panel, or a polygonal panel, and the like. In addition, the display panel can not only be a flat panel, but also a curved panel, or even a spherical panel.

[0109] As shown in FIG. 6, the electronic device 6 of this embodiment includes a memory 61 and a processor 62 coupled to the memory 61. It should be noted that the components of the electronic device 6 shown in FIG. 6 are only exemplary and are not limiting, and the electronic device 6 can also have other components according to actual application needs. The processor 62 can control other components in the electronic device 6 to perform the desired functions.

[0110] In some embodiments, the memory 61 is configured to store one or more computer readable instructions. When the processor 62 executes the computer readable instructions, the computer readable instructions are executed by the processor 62 to implement the method according to any of the above embodiments. For specific implementation of each step of the method and related explanations, please refer to the above embodiments, and repeated parts will not be described here.

[0111] For example, the processor 62 and the memory 61 can communicate with each other directly or indirectly. For example, the processor 62 and the memory 61 can communicate through a network. The network can include a wireless network, a wired network, and / or any combination of a wireless network and a wired network. The processor 62 and the memory 61 can also communicate with each other through a system bus, and the present disclosure does not limit this.

[0112] For example, the processor 62 can be embodied as various appropriate processors, processing devices, etc., such as a central processing unit (CPU), a graphics processing unit (GPU), a network processing unit (NP), etc.; and can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic, a discrete hardware component. The central processing unit (CPU) can be an X86 or ARM architecture, etc. For example, the memory 61 can include any combination of various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The memory 61 may, for example, include a system memory, which stores, for example, an operating system, application programs, a boot loader, a database, and other programs, etc. Various application programs and various data, etc. can also be stored in the storage medium.

[0113] In addition, according to some embodiments of the present disclosure, various operations / processes according to the present disclosure, when implemented by software and / or firmware, can install programs constituting the software from a storage medium or a network to a computer system having a dedicated hardware structure, such as the computer system 70 shown in FIG. 7, which, when various programs are installed, can perform various functions, including functions such as those described above, etc. FIG. 7 shows a structural schematic diagram of a computer system according to some embodiments of the present disclosure.

[0114] In FIG. 7, a central processing unit (CPU) 701 executes various processes in accordance with a program stored in a read only memory (ROM) 702 or a program loaded from a storage section 707 to a random access memory (RAM) 703. In the RAM 703, data required when the CPU 701 executes various processes and the like is also stored as necessary. The central processing unit is merely exemplary, and can also be other types of processors, such as the various processors described above. The ROM 702, the RAM 703, and the storage section 707 can be various forms of computer readable storage media, as described below. Note that, although the ROM 702, the RAM 703, and the storage 707 are shown separately in FIG. 7, one or more of them can be combined or located in the same or different memory or storage modules.

[0115] The CPU 701, the ROM 702, and the RAM 703 are connected to each other via a bus 704. An input / output interface 705 is also connected to the bus 704.

[0116] The following components are connected to the input / output interface 705: an input section 706 including a touch panel, a touch pad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, and the like; an output section 707 including a display such as a cathode ray tube (CRT), a liquid crystal display (LCD), a speaker, a vibrator, and the like; a storage section 707 including a hard disk, a magnetic tape, and the like; and a communication section 709 including a network interface card such as a LAN card, a modem, and the like. The communication section 709 allows communication processing to be performed via a network such as the Internet. It is easily understood that, although the various devices or modules in the computer system 70 are shown in FIG. 7 as communicating via the bus 704, they can also communicate through a network or other means, where the network can include a wireless network, a wired network, and / or any combination of a wireless network and a wired network.

[0117] A drive 710 is also connected to the input / output interface 705 as necessary. A removable medium 711 such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, and the like is attached to the drive 710 as necessary, so that a computer program read therefrom is installed in the storage section 707 as necessary.

[0118] In the case where the above series of processes are implemented by software, the program constituting the software can be installed from a network such as the Internet or a storage medium such as the removable medium 711.

[0119] According to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a computer readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network by the communication device 709, or installed from the storage device 707, or installed from the ROM 702. When the computer program is executed by the CPU 701, the above-described functions defined in the methods of the embodiments of the present disclosure are executed.

[0120] It should be noted that, in the context of the present disclosure, a computer readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. The computer readable medium can be a computer readable signal medium or a computer readable storage medium or any combination of the two. The computer readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples of the computer readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer readable storage medium can be any tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer readable signal medium can include a data signal carried in a baseband or as part of a carrier wave, in which the computer readable program code is carried. Such a propagated data signal can take a variety of forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer readable signal medium can also be any computer readable medium that is not a computer readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF (radio frequency), or the like, or any suitable combination of the above.

[0121] The above computer readable medium can be included in the above electronic device; or can exist separately without being assembled into the electronic device.

[0122] In some embodiments, a computer program including instructions which, when executed by a processor, causes the processor to carry out the method of any of the above embodiments is also provided. For example, the instructions can be embodied in a computer program code.

[0123] In an embodiment of the disclosure, the computer program code for carrying out operations of the disclosure can be written in one or more programming languages or combinations thereof including an object oriented programming language such as Java, Smalltalk, C++ or the like and conventional procedural programming languages such as "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0124] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a procedure, or a part of a code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur in a different order than that noted in the figures. For example, two blocks depicted contiguously can actually be substantially parallel, or can occur in the reverse order, depending on the functionality involved. It should also be noted that each block in the block diagrams and / or flow diagrams, and combinations of blocks in the block diagrams and / or flow diagrams, can be implemented by dedicated hardware-based systems that perform the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0125] The modules, components or units described in the embodiments of the disclosure can be implemented by software or by hardware. In some cases, the name of the module, component or unit does not constitute a limitation on the module, component or unit itself.

[0126] The functionality described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, non- transitory machine-readable media can include RAM, ROM, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory, compact disk read-only memory (CD-ROM), digital versatile disk (DVD), Blu-ray, or another non-transitory medium suitable for storing non-transitory program code, wherein the above aforementioned media can be read and written by, for example, a machine- readable media reader and a machine-readable media recording device, respectively. In this document, the terms "non-transitory machine- readable storage medium" and "non-transitory machine-readable storage media" mean the above-described media except for a transitory signal per se.

[0127] The above description is merely exemplary of some embodiments of the disclosure and of the principles thereof. It is understood that those skilled in the art will be able to devise various embodiments of the disclosure without departing from the concepts disclosed in the present disclosure. Accordingly, the disclosure is not limited to the specific embodiments described herein, but only by the claims that follow, although various embodiments can be practiced within the scope of the claims.

[0128] In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the application can be practiced without these specific details. In other instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure an understanding of this description.

[0129] In addition, while operations are depicted in a particular, chronological sequence, this should not be understood as requiring such order unless specifically specified. Certain activities can be performed in different order or concurrently with each other. Additionally, certain activities can be performed in a different order or omitted entirely without departing from the scope of the disclosure. Similarly, while several specific implementations have been discussed, this should not be understood as excluding other implementations falling within the scope of the claims. Certain features that are, for brevity, described in the context of separate embodiments can also be implemented in a combination of embodiments. Conversely, various features that are, for brevity, described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.

[0130] While certain features of the disclosure have been illustrated and described, many modifications, substitutions, changes, and equivalents will now occur to those skilled in the art. It is, therefore, to be understood that the appended claims are intended to cover all such modifications and changes as fall within the true scope of the disclosure. The disclosure is not limited to the illustrative embodiments set forth herein but embraces all within the scope of the claims.

Claims

1. An interaction method comprising: displaying first multimedia content from a user and request information of the user based on the first multimedia content, the first multimedia content comprising one or more actions; processing the one or more actions in the first multimedia content based on the request information to determine second multimedia content and response information of the request information; displaying the second multimedia content and the response information.

2. The interaction method of claim 1, wherein, The processing the one or more actions in the first multimedia content based on the request information to determine second multimedia content and response information of the request information comprises: performing intent recognition on the request information to determine a topic of the request information; processing the one or more actions in the first multimedia content based on the topic of the request information to determine second multimedia content and response information of the request information.

3. The interaction method of claim 2, wherein, The processing the one or more actions in the first multimedia content based on the topic of the request information to determine second multimedia content and response information of the request information comprises: in response to the topic of the request information being an evaluation, determining a target action of the one or more actions in the first multimedia content, wherein the target action is an action having a problem or an action indicated by the request information; determining the second multimedia content based on the target action; generating at least one of evaluation information of the target action in the first multimedia content or explanation information of the second multimedia content as the response information.

4. The interaction method of claim 3, wherein, The determining a target action of the one or more actions in the first multimedia content in response to the topic of the request information being an evaluation comprises: determining a matching result of the one or more actions in the first multimedia content and a sound in the first multimedia content; determining an action not matching the sound as the target action.

5. The interaction method of claim 4, wherein, The determining the second multimedia content based on the target action comprises: generating the second multimedia content by adjusting the target action in the first multimedia content such that the target action matches the sound.

6. The interaction method of claim 3, wherein, The determining a target action of the one or more actions in the first multimedia content in response to the topic of the request information being an evaluation comprises: determining reference multimedia content corresponding to the one or more actions; comparing the reference multimedia content with the first multimedia content to determine a target action having a problem.

7. The interaction method of claim 3, wherein, The determining a target action of the one or more actions in the first multimedia content in response to the topic of the request information being an evaluation comprises: determining an action of the one or more actions matching an action indicated by the request information as the target action.

8. The interaction method according to claim 6 or 7, wherein, The determining the second multimedia content based on the target action comprises: generating or searching the second multimedia content comprising the target action.

9. The interaction method of claim 2, wherein, The subject based on the request information processes the one or more actions in the first multimedia content to determine a second multimedia content and a response information of the request information includes: In response to the subject of the request information being a choreographed action, predicting one or more target actions based on the one or more actions; Generating a second multimedia content based on the target actions; Generating a commentary information of actions in the second multimedia content as the response information.

10. The interaction method of claim 9, wherein, The predicting one or more target actions based on the one or more actions includes: Predicting one or more target actions after the one or more actions; or Generating one or more target actions including the one or more actions.

11. The interaction method according to any one of claims 1 to 10, further comprising: Identifying an object in the first multimedia content; Determining one or more actions of the object in the first multimedia content according to a pose of the object.

12. The interaction method of any one of claims 1 to 11, wherein, The request information is text or voice.

13. The interaction method according to any one of claims 1 to 12, wherein: The first multimedia content is a video or one or more images; The second multimedia content is a video or one or more images.

14. An electronic device comprising: a memory; and a processor coupled to the memory, the processor configured to perform the interaction method according to any one of claims 1 to 13 based on instructions stored in the memory.

15. A computer readable storage medium having stored thereon a computer program which, when executed by a processor, implements the interaction method according to any one of claims 1 to 13.

16. A computer program product which, when executed on a computer, causes the computer to implement the interaction method according to any one of claims 1 to 13.

17. A computer program comprising: instructions which, when executed by a processor, cause the processor to perform the interaction method according to any one of claims 1 to 13. ​

Citation Information

Patent Citations

  • Yoga motion guidance system and method based on computer vision

    CN111652078A

  • Interaction method and device based on video decomposition processing, equipment and storage medium

    CN116980717A

  • Simulation role-based information interaction method and device and storage medium

    CN117560340A

  • KR20230026228A