An information output method and apparatus

By dividing the feedback text into multiple parts and outputting the next part of audio and video data according to the user's response matching, the problem of inconvenient user response text indication is solved and the convenience of response is improved.

CN119336164BActive Publication Date: 2025-10-28LENOVO (BEIJING) LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411389047.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2025-10-28
Estimated Expiration
2044-09-30

AI Technical Summary

Technical Problem

When responding to text instructions, users need to spend time finding the corresponding instructions from the text, resulting in poor convenience.

Method used

The feedback text is divided into multiple parts, each part corresponds to audio and video data. After the user responds to the audio and video data, the device matches and outputs the next part of the audio and video data according to the feature information until a match is successful.

Benefits of technology

It improves the ease of user response to text instructions and reduces the time spent searching for the next instruction after each response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119336164B_ABST
    Figure CN119336164B_ABST
Patent Text Reader

Abstract

This application discloses an information output method and apparatus. The method includes: obtaining a first feedback text based on a user's query information, the first feedback text including at least a first part and a second part, each part of the first feedback text corresponding to audio-visual data, the audio-visual data including at least one of audio data, video data and image data; outputting the first audio-visual data corresponding to the first part so that the user can respond to the first audio-visual data; obtaining feature information of the user's response to the first audio-visual data, and if the user's feature information matches the first part, outputting the second audio-visual data corresponding to the second part of the first feedback text.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information output technology, and in particular to an information output method and apparatus. Background Technology

[0002] Currently, some electronic devices can obtain specific text based on user needs, allowing users to browse the obtained text and respond accordingly. For example, users can perform corresponding physical actions based on instructions given in the text.

[0003] The problem with this approach is that the text obtained based on the user's needs usually contains a lot of information, and each time the user responds, they need to spend some time finding the corresponding instructions from the text, which is not very convenient. Summary of the Invention

[0004] Therefore, this application discloses the following technical solution:

[0005] The first aspect of this application provides an information output method, including:

[0006] The first feedback text is obtained based on the user's query information. The first feedback text includes at least a first part and a second part. Each part of the first feedback text corresponds to audio-visual data, and the audio-visual data includes at least one of audio data, video data, and image data.

[0007] Output the first audio-visual data corresponding to the first part so that the user can respond to the first audio-visual data;

[0008] Obtain the feature information of the user's response to the first audio-visual data. If the user's feature information matches the first part, output the second audio-visual data corresponding to the second part of the first feedback text.

[0009] Optionally, obtaining the first feedback text based on the user's query information includes:

[0010] The first feedback text is obtained based on the user's query information and the user's interaction device information, wherein the interaction device information instructs the user to operate the control device.

[0011] Optionally, obtaining the first feedback text based on the user's query information and the user's interaction device information includes at least one of the following:

[0012] When the interactive device information instructs the user to use the wearable device, a first feedback text related to the action of the target body part is obtained based on the user's query information, wherein the target body part includes the body part of the user wearing the wearable device.

[0013] When the interactive device information instructs the user to use the controller device, a first feedback text related to the target application is obtained based on the user's query information. The target application includes an application that the user operates through the controller device.

[0014] Optional, also includes:

[0015] Identify the characteristic text contained in the first feedback text;

[0016] Identify the semantic relationships between the feature texts contained in the first feedback text;

[0017] Based on the semantic correlation between the feature texts, the first feedback text is divided into at least a first part and a second part, each part including at least one feature text with semantic correlation.

[0018] Optional, also includes:

[0019] If the user's feature information does not match the first part, a prompt message is output based on the first part to prompt the user to respond to the first audio-visual data in a way that matches the first part.

[0020] Alternatively, if the user's feature information does not match the first part, a third part that matches the feature information can be determined from among the multiple parts contained in the first feedback text.

[0021] Output the audio and video data corresponding to the third part.

[0022] Optional, also includes:

[0023] Among the multiple parts contained in the first feedback text, a third part that matches the feature information is determined. Based on the position of the third part in the first feedback text, the audio-visual data corresponding to the text part located after the position in the first feedback text is output.

[0024] Optional, also includes:

[0025] If multiple parts of the first feedback text do not match the feature information, a second feedback text is obtained.

[0026] The fourth part that matches the feature information is identified in the second feedback text;

[0027] Output the audio and video data corresponding to the fourth part.

[0028] Optionally, obtaining the second feedback text includes at least one of the following:

[0029] The second feedback text is obtained based on the query information;

[0030] The second feedback text is obtained based on the aforementioned feature information.

[0031] Optionally, each part of the first feedback text corresponds to audio-visual data, including at least one of the following:

[0032] Generate corresponding audio-visual data based on each part of the first feedback text;

[0033] The corresponding audio and video data is retrieved from the database based on each part of the first feedback text.

[0034] A second aspect of this application provides an information output device, comprising:

[0035] The obtaining unit is configured to obtain a first feedback text based on the user's query information. The first feedback text includes at least a first part and a second part. Each part of the first feedback text corresponds to audio-visual data, and the audio-visual data includes at least one of audio data, video data, and image data.

[0036] An interactive unit is used to output the first audio-visual data corresponding to the first part, so that the user can respond to the first audio-visual data;

[0037] The interaction unit is used to acquire the feature information of the user's response to the first audio-visual data. If the user's feature information matches the first part, the second audio-visual data corresponding to the second part of the first feedback text is output. Attached Figure Description

[0038] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0039] Figure 1 This is a flowchart of an information output method provided in an embodiment of this application;

[0040] Figure 2 This is a flowchart of another information output method provided in an embodiment of this application;

[0041] Figure 3 This is a flowchart of a method for segmenting first feedback text provided in an embodiment of this application;

[0042] Figure 4 This is a schematic diagram of an information output device provided in an embodiment of this application. Detailed Implementation

[0043] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0044] This application provides an information output method. Please refer to [link to relevant documentation]. Figure 1 Here is a flowchart of the method, which may include the following steps.

[0045] S101, obtain the first feedback text based on the user's query information. The first feedback text includes at least a first part and a second part. Each part of the first feedback text corresponds to audio-visual data, which includes at least one of audio data, video data, and image data.

[0046] In step S101, the device used to execute this embodiment can obtain query information in various ways. For example, the device can obtain query information in text form entered by the user through a keyboard or touch screen, or it can collect the user's voice and obtain query information in the voice through voice recognition.

[0047] The query information can represent the type of text that the user wants to obtain. For example, the query information may include "to do boxing training", "to practice singing a certain song", "the gameplay guide for XX game", "please provide a yoga practice method. I am 25 years old and have been practicing yoga for 2 years. Please provide a content on practicing yoga for 30 minutes in the living room", etc.

[0048] After obtaining the query information, the device can search the Internet based on the query information to obtain the first feedback text that matches the query information; or, the device can input the query information into a data processing model with text generation capabilities (such as a large language model) and use the data processing model to generate the first feedback text that matches the query information.

[0049] As an example, suppose the query is "to train in boxing". The first feedback text obtained based on this query could be "Start training with the basic movements of boxing. The following training begins: probing, probing, jab, side squat...".

[0050] In some optional embodiments, when the device obtains the first feedback text, it may further combine the current user's user attribute information and historical behavior information to obtain the first feedback text.

[0051] The device can obtain user attribute information and historical behavior information based on user input. For example, the user's voice collected by the device could be: "Please provide a yoga practice method. I am 25 years old and have been practicing yoga for 2 years. Please provide a method for practicing yoga for 30 minutes in the living room." Through this voice, the device can recognize the query information "yoga practice method", the user attribute information "25 years old" and "practicing yoga for 30 minutes in the living room", and the historical behavior information "practicing yoga for 2 years".

[0052] The device can also identify the user and then obtain pre-stored user attribute information and historical behavior information based on that user's identity. For example, before obtaining the query information, the device can determine the identity information of the user currently entering the query information through identity recognition technologies such as password recognition and fingerprint recognition. Alternatively, when collecting the user's voice, the device can determine the identity information of the user who made the voice through voiceprint recognition.

[0053] With access to user attribute information and historical behavior information, the device can combine these information to search the internet and obtain a matching first feedback text. Alternatively, it can input this information into a data processing model and use the model to generate the corresponding first feedback text.

[0054] Based on the aforementioned example, when the device receives the voice message "Please provide a yoga practice method. I am 25 years old and have been practicing yoga for 2 years. Please provide a 30-minute yoga practice in the living room," the device can obtain the following initial feedback text based on the information contained in the voice message:

[0055] "Okay, let's start with Mountain Pose. Stand with your feet hip-width apart, lift your sternum while relaxing your lower ribs, externally rotate your arms so that your palms face forward and let your arms hang down at your sides. Begin in Mountain Pose, inhale, straighten your arms, exhale, fold forward, starting from the hips. Place your hands beside your feet or on the floor..."

[0056] The initial feedback text can be either text that has already been divided into multiple parts, or it can be continuous text. In the latter case, the device can divide the initial feedback text into multiple parts after receiving it.

[0057] The aforementioned multiple parts may include two or more parts.

[0058] After receiving multiple parts of the first feedback text, the device can acquire the audio-visual data corresponding to each part in at least one of the following ways.

[0059] Firstly, the device can generate corresponding audio-visual data based on each part of the first feedback text. In this acquisition method, the device can call a locally deployed data processing model capable of generating audio-visual data to process any part of the first feedback text, thereby generating the corresponding audio-visual data. Alternatively, the device can send any part of the first feedback text to the server, allowing the server to use the data processing model capable of generating audio-visual data to generate the corresponding audio-visual data.

[0060] For example, in the new first feedback text, one part could be "externalize your arm so that your palm faces forward." The device can process this part of the text using a data processing model with text-to-speech capabilities to obtain the corresponding speech.

[0061] Secondly, the device can search for the corresponding audio-visual data in the database based on each part of the first feedback text. In this acquisition method, the device or the server connected to the device can pre-establish a database storing a large amount of audio-visual data, where each piece of audio-visual data can correspond to one or more keywords. For any part of the first feedback text, the device can match that part with the keywords of multiple pieces of audio-visual data in the database, and determine the audio-visual data that matches successfully as the audio-visual data corresponding to that part.

[0062] The first part in S101 may refer to any part contained in the first feedback text, and the second part may refer to the latter part of the first part in the first feedback text.

[0063] S102, output the first audio-visual data corresponding to the first part so that the user can respond to the first audio-visual data.

[0064] Each part of the first feedback text may include feature text that instructs the user to make a corresponding response.

[0065] As examples, the first feedback text can be text used to instruct the user to perform a specific physical action, and the feature text can be a verb contained in the first feedback text. For example, one part of the first feedback text can be "straighten arm", and the feature text can be "straighten".

[0066] The first feedback text can be text used to instruct the user to control a target object. The target object can be an object displayed on the screen by a specific application, such as a virtual character model or a virtual vehicle model displayed on the screen. The feature text can be instructions contained in the first feedback text for controlling the target object. As an example, one part of the first feedback text can be "jump to the platform to the left front", where the feature text can be "jump".

[0067] The audio-visual data corresponding to each part can include content corresponding to the feature text of that part. For example, the speech data generated based on the text of the first example can include the speech of "straightening up", and the video data obtained based on the text of the second example can include several video frames representing the target object jumping.

[0068] Therefore, the device can instruct the user to make a response corresponding to the feature text based on the first audio-visual data corresponding to the first part. For example, the audio-visual data of the first example can instruct the user to straighten their arm, and the audio-visual data of the second example can instruct the user to control the target object to jump to the left and forward.

[0069] S103, obtain the feature information of the user's response to the first audio-visual data, and if the user's feature information matches the first part, output the second audio-visual data corresponding to the second part of the first feedback text.

[0070] The feature information in response to the first audio-visual data may include feature information collected during the period from the start of outputting the first audio-visual data to the completion of the output of the first audio-visual data, or it may include user feature information collected within a period of time after the completion of the output of the first audio-visual data, or it may include both of the above.

[0071] The feature information can be in any form and can characterize the user's response to the first audio-visual data.

[0072] As examples, the feature information responding to the first audio-visual data may include at least one of the position data, velocity data, and acceleration data of one or more body parts of the user collected within the corresponding time period; or, the user's feature information may include a video data segment captured by filming the user within the corresponding time period; or, the user's feature information may include a segment of audio data obtained by capturing the user's voice through a microphone within the corresponding time period; or, the user's feature information may include several control commands input by the user within the corresponding time period for controlling the target object.

[0073] For example, the feature information in response to the first audio-visual data may include the acceleration data of the user's palm in multiple directions in space at the beginning of the time from the start of outputting the first audio-visual data to the end of the time when the first audio-visual data is output, the acceleration data of the user's palm in multiple directions in space at the end of the time, and the acceleration data of the user's palm in multiple directions in space at a certain intermediate moment (e.g., the moment when the acceleration is the greatest).

[0074] The position data, velocity data, and acceleration data of a specific body part of the user can be obtained by wearing a wearable device worn on the corresponding body part. These wearable devices can have built-in sensors capable of detecting the above data, and these wearable devices can communicate with the device executing this embodiment (e.g., Bluetooth connection), so that the device can receive at least one of the position data, velocity data, and acceleration data detected by the wearable devices through the sensors in real time.

[0075] After obtaining the feature information in response to the first audio-visual data, the device can determine whether the feature information matches the first part. If it matches, the device can continue to output the audio-visual data corresponding to the second part, that is, output the second audio-visual data corresponding to the second part.

[0076] If the feature information of the first audio-visual data does not match the first part, the device may output the corresponding information in at least one of the following output methods:

[0077] The first method involves outputting prompt information based on the first part when the user's feature information does not match the first part, prompting the user to respond to the first audio-visual data in a way that matches the first part.

[0078] The prompt message can be used to prompt the user to respond accordingly to the first audio-visual data. For example, the prompt message can be a voice message such as "Please follow the voice prompt to perform the correct action".

[0079] Alternatively, the prompt message can be consistent with the first audio-visual data. That is, if the device determines that the feature information of the user's response to the first audio-visual data does not match the first part, it can output the first audio-visual data again until the feature information of the user's response to the first audio-visual data matches the first part.

[0080] After outputting the prompt message, a certain period of time can be waited for the user to respond accordingly. During this period, characteristic information representing the user's response can be obtained. If the characteristic information obtained during this period matches the first part, the second audio-visual data can be output. If the characteristic information obtained during this period matches the first part, the prompt message can be output again, or the second audio-visual data can be output again.

[0081] The second method involves determining the matching status of the other parts of the output before the first part with the corresponding feature information. If the condition that N consecutive parts (including the first part) do not match the corresponding feature information is met, a certain period of time can be waited before continuing to output the second audio-visual data. During the waiting period, the user can be prompted that no response has been made that matches the corresponding part for N consecutive times. If the condition that N consecutive parts do not match the corresponding feature information is not met, the second audio-visual data can continue to be output.

[0082] The third method is to wait for a certain period of time before continuing to output the second audio-visual data.

[0083] In some embodiments, each time the device outputs audio-visual data corresponding to a certain part of the first feedback text, it can regard the currently output audio-visual data as the first audio-visual data and the current part as the first part, thereby sequentially outputting the audio-visual data corresponding to each part of the first feedback text in accordance with steps S102 and S103.

[0084] For example, suppose the first feedback text includes 10 parts, which are numbered from part 1 to part 10 in chronological order. After the device obtains the first feedback text, it first outputs the audio-visual data 1 corresponding to part 1. Then it determines whether the feature information of the user's response to the audio-visual data 1 matches part 1. If they match, the device continues to output the audio-visual data 2 corresponding to part 2. If the feature information of the user's response to the audio-visual data 2 matches part 2, the device continues to output the audio-visual data 3 corresponding to part 3, and so on, until the audio-visual data 10 corresponding to part 10 is output.

[0085] As an example, see Figure 2 Suppose the query information obtained in S101 is "boxing practice". The first feedback text obtained based on the query information is: "First, start training with the basic movements of boxing. The training begins now. Probing, probing, jab, side squat..."

[0086] The first feedback text can be divided into four parts, namely "probing", "probing", "jab", and "side squat". The device can identify the first "probing" as the first part, output the corresponding first audio-visual data, such as outputting the voice of "probing", and then determine whether the user's response to the voice feature information matches "probing". If they match, it is determined that the user has correctly performed the probing action.

[0087] At this point, the device can output the audio and video data corresponding to the next part of the first part, that is, output the next "probing" voice again, and obtain the feature information of the user's response to the second output voice. If the feature information obtained the second time also matches the "probing", it indicates that the user has correctly made the probing action again.

[0088] Next, the device can output the voice of "jab". If the user responds to the voice feature information of "jab" and it matches "jab", it indicates that the user has correctly performed the jab action. Then the device can continue to output the voice of "side squat". If the user responds to the voice feature information of "side squat" and it matches "side squat", it indicates that the user has correctly performed the side squat action. At this point, the boxing practice based on the first feedback text ends.

[0089] The beneficial effects of this embodiment are as follows:

[0090] After receiving the feedback text, the audio-visual data corresponding to one part of the feedback text is output first. If the user's feature information matches the part, the audio-visual data corresponding to the next part is output. In this way, the user can respond according to each output part in sequence, without having to spend time finding the next instruction from the text after each response, thus improving the convenience for the user to respond based on the feedback text.

[0091] In some embodiments, it can be done as follows Figure 3 The method shown divides the first feedback text into multiple parts.

[0092] S301, Identify the feature text contained in the first feedback text.

[0093] S302, Identify the semantic relationships between the feature texts contained in the first feedback text.

[0094] S303, based on the semantic relationship between the feature texts, the first feedback text is divided into at least a first part and a second part, each part including at least one feature text with semantic relationship.

[0095] In steps S301 and S302, the obtained first feedback text can be processed using a data processing model with semantic understanding capabilities, thereby identifying the feature text contained in the first feedback text and the semantic relationship between these feature texts.

[0096] The semantic relationships between different feature texts can vary. For example, the semantic relationship between some feature texts may be that the response of one feature text occurs before the response of another feature text, that is, there is a sequential relationship between the responses of the two feature texts; the semantic relationship between some feature texts may be that the responses of two (or more) feature texts occur at the same time.

[0097] Taking verbs as an example, suppose the first feedback text includes "Start in Mountain Pose, inhale, straighten your arms, exhale, fold forward, bend forward from the hips." In this text, "inhale," "straighten," "exhale," and "bend forward" are all part of the feature text. By identifying the semantic relationships between the feature texts, we can determine that the semantic relationship between "inhale" and "straighten" is that the two actions of inhalation and straightening occur simultaneously; the semantic relationship between "exhale" and "bend forward" is that the two actions of exhalation and bending forward occur simultaneously; and the semantic relationship between "inhale" and "exhale" is that the inhalation action occurs first, followed by the exhalation action.

[0098] In step S303, feature texts whose corresponding semantic association is that the response occurs simultaneously can be grouped into the same part, while feature texts whose corresponding semantic association is that the response occurs sequentially can be grouped into different parts.

[0099] For the non-feature text in the first feedback text, other than the feature text, it can be classified into the part to which the feature text belongs based on the semantic relationship between the non-feature text and the feature text.

[0100] Based on the previous example, by analyzing the semantics of the text, we can find that "begin in Mountain Pose" and "inhale" are related, and "fold forward" and "exhale" are related. Therefore, "begin in Mountain Pose" and "inhale" can be grouped into the same part, and "fold forward" and "exhale" can be grouped into the same part. Thus, the first feedback text of the previous example can be divided into the first part "begin in Mountain Pose, inhale, straighten arms" and the second part "exhale, fold forward, bend forward from the hips".

[0101] In some embodiments, for any portion of the first feedback text, it can be determined whether that portion matches the user's feature information in the following manner:

[0102] Obtain reference feature information for the feature text contained in this section;

[0103] Determine the similarity between the reference feature information of the feature text contained in this section and the feature information of the audio-visual data in response to this section;

[0104] If the similarity is greater than the preset matching threshold, it can be determined that this part matches the user's feature information;

[0105] If the similarity is less than or equal to the matching threshold, it can be determined that this part does not match the user's feature information.

[0106] The device can obtain a feature information database in advance, which may include multiple feature texts and reference feature information corresponding to each feature text.

[0107] Based on this, the device can find the reference feature information corresponding to the feature text contained in the first part after determining the first part.

[0108] After the device outputs the first audio-visual data and obtains the feature information of the user's response to the first audio-visual data, it can determine the similarity between the user's feature information and the reference feature information of the feature text contained in the first part. Based on whether the similarity is greater than the matching threshold, it is determined whether the feature information of the user's response to the first audio-visual data matches the first part. If the feature information of the user's response to the first audio-visual data matches the first part, it means that the user has made a correct response based on the instructions of the first part. If the feature information of the user's response to the first audio-visual data does not match the first part, it means that the user has not made a correct response based on the instructions of the first part.

[0109] The similarity between the user's feature information and the reference feature information of the feature text contained in the first part can be determined by various similarity calculation methods in the relevant technical field, which will not be elaborated here.

[0110] The reference feature information corresponding to the feature text and the user's feature information can be in the same form.

[0111] For example, if the feature information obtained in response to the first audio-visual data includes at least one of the position data, velocity data, and acceleration data of one or more body parts of the user within a certain period of time, then the reference feature information corresponding to the feature text may also include at least one of the position data, velocity data, and acceleration data of the corresponding body part of any person within a certain period of time.

[0112] The reference feature information corresponding to the feature text can be obtained in the following way:

[0113] After an individual makes a correct response based on a feature text, the feature information obtained during the response process is determined as the reference feature information corresponding to that feature text.

[0114] As an example, suppose the feature text is a verb, and the feature information includes the acceleration data of the wrist in the X, Y, and Z directions in the three-dimensional coordinate system when the response begins, when the response stops, and when the acceleration is at its maximum during the response. In this case, if you need to obtain the reference feature information corresponding to the verb "jab", you can determine the wrist acceleration data collected at the above three moments during the jab action after the individual makes the correct jab action as the reference feature information corresponding to "jab".

[0115] The reference feature information for the "jab" obtained in the above manner can be represented by Table 1 below.

[0116] Table 1

[0117] Jab X-axis 0;60;-18 Y-axis 0;40;70 Z-axis 0;30;-55

[0118] Each row in Table 1 represents the acceleration data of the human wrist on the corresponding coordinate axis at three moments during the correct jab motion, in meters per second squared. The three moments are, in order, the moment when the jab motion begins, the moment when the acceleration on the corresponding coordinate axis is the greatest during the jab motion, and the moment when the jab motion ends.

[0119] The reference feature information corresponding to the feature text can also be obtained in the following way:

[0120] Multiple similar feature texts are obtained, as well as multiple feature information when multiple users respond based on these feature texts. Then, these feature information are clustered to determine the correspondence between feature information and feature texts. Finally, the feature information corresponding to the feature text is fused to obtain the reference feature information corresponding to the feature text.

[0121] Similar feature text can include feature text that has some identical content.

[0122] As an example, multiple similar feature texts can include "open bottle cap", "open refrigerator" and "open school bag". For these three feature texts, feature information (such as the aforementioned acceleration data) can be obtained when multiple users make the action of opening bottle cap, feature information when multiple users make the action of opening refrigerator, and feature information when multiple users make the action of opening school bag.

[0123] Then, clustering is performed on the above-mentioned multiple feature information to determine which feature information corresponds to "opening the bottle cap", which feature information corresponds to "opening the refrigerator", and which feature information corresponds to "opening the schoolbag". Finally, the multiple feature information corresponding to "opening the bottle cap" is fused to obtain the reference feature information corresponding to "opening the bottle cap", the multiple feature information corresponding to "opening the refrigerator" is fused to obtain the reference feature information corresponding to "opening the refrigerator", and the multiple feature information corresponding to "opening the schoolbag" is fused to obtain the reference feature information corresponding to "opening the schoolbag".

[0124] The reference feature information corresponding to the feature text can also be obtained in the following way:

[0125] During the period when an individual continuously responds based on multiple feature texts, multiple feature information is obtained during this period. Known reference feature information is identified among the multiple feature texts. Then, based on the order of the multiple feature texts and the known reference feature information, reference feature information corresponding to the target feature text is determined among the multiple feature texts. The target feature text refers to the feature text among the multiple feature texts for which reference feature information has not yet been determined.

[0126] As an example, suppose three consecutive feature texts are "probing", "jab", and "side squat" in sequence. The reference feature information for "probing" and "jab" has already been obtained through other means. The target feature text is "side squat". After any person performs the three actions of probing, jab, and side squat in sequence according to the three feature texts, the device can determine, based on the known reference feature information for "probing" and "jab", that the feature information obtained during the first two actions corresponds to "probing" and "jab" respectively. Then, based on the order of the feature texts, it determines that the feature information obtained during the last action is the reference feature information corresponding to "side squat".

[0127] Optionally, when obtaining the first feedback text based on the user's query information, further information about the user's interaction device can be obtained. This allows for the generation of the first feedback text based on both the query information and the user's interaction device information, with the interaction device information indicating the device the user is operating. This approach improves the match between the obtained first feedback text and the user's needs, leading to first feedback text that better meets the user's requirements.

[0128] Interaction device information can be obtained based on the connection relationship between the device performing the method of this embodiment and other electronic devices. For example, when the device detects a current connection with a wearable device (such as a wristband or smartwatch), it can obtain interaction device information instructing the user to use the wearable device; when the device detects a current connection with a controller device, it can obtain interaction device information instructing the user to use the controller device.

[0129] In some embodiments, when the interactive device information instructs the user to use the wearable device, a first feedback text related to the action of a target body part can be obtained based on the user's query information. The target body part includes the body part on which the user wears the wearable device.

[0130] For example, if the wearable device being used by the user is worn on the user's wrist or arm, a first feedback text that matches the query information and is related to wrist or arm movements is obtained, such as a first feedback text related to boxing movements; if the wearable device being used by the user is worn on the user's head, a first feedback text that matches the query information and is related to head movements is obtained; if the wearable device being used by the user is worn on the user's leg, a first feedback text that matches the query information and is related to leg movements is obtained.

[0131] The body part on which a wearable device is located can be determined by the type of wearable device. For example, if the wearable device is a wristband or smartwatch, it is determined that the user wears it on their wrist; if the wearable device is a helmet or glasses, it is determined that the user wears it on their head.

[0132] In some embodiments, when the interactive device information instructs the user to use the controller device, a first feedback text related to the target application can be obtained based on the user's query information. The target application includes an application operated by the user through the controller device.

[0133] In this embodiment, one or more target applications that can be operated by the gamepad can be identified from among the multiple applications already installed on the device, or one or more target applications that the user has previously operated with the gamepad can be identified based on the user's historical operation data. Then, the first feedback text that matches the information and is related to the target application is obtained and queried.

[0134] For example, if the query information obtained is "boxing action practice", the interaction device information indicates that the user is using a gamepad, and the device has a boxing game that can be operated with a gamepad installed locally, then the first feedback text obtained by the device can be the operation guide text of the boxing game.

[0135] In some embodiments, when the interactive device information instructs the user to use a headset and microphone device, a first feedback text related to the audio can be obtained based on the query information.

[0136] For example, when the interactive device information instructs the user to use headphones and microphone, the user can search for a target song whose name matches the query information from the music stored on the device or from music that the user has previously played on the device, and then obtain first feedback text related to the target song, such as text indicating the singing style of the target song.

[0137] In some embodiments, when the interactive device information instructs the user to use a handwriting-type device, such as a graphics tablet, stylus, or other device, the user can obtain first feedback text related to drawing based on the query information.

[0138] For example, when the interactive device information instructs the user to use a graphics tablet, the user can search for and query paintings that match the information, and then obtain first feedback text related to the painting, such as first feedback text to guide the user to copy the painting.

[0139] In some alternative embodiments, if the device determines that the feature information of the user's response to the first audio-visual data does not match the first portion, the device may also perform at least one of the following processing methods.

[0140] In the first approach, if the user's feature information does not match the first part, the third part that matches the feature information is determined from the multiple parts contained in the first feedback text.

[0141] Output the audio and video data corresponding to the third part.

[0142] The second processing method involves identifying the third part that matches the feature information from among the multiple parts contained in the first feedback text, and outputting the audio-visual data corresponding to the text part located after the position in the first feedback text based on the position of the third part.

[0143] The third processing method is to obtain a second feedback text when multiple parts and feature information contained in the first feedback text do not match.

[0144] The fourth part that matches the feature information is identified in the second feedback text;

[0145] Output the audio and video data corresponding to Part 4.

[0146] In processing method one, the device can obtain reference feature information corresponding to each part of the first feedback text except for the first part, compare the reference feature information of each part with the feature information responding to the first audio-visual data, and if it is determined that the reference feature information of a certain part matches the feature information responding to the first audio-visual data, then the part is determined to be the third part, and then the audio-visual data corresponding to the third part is output.

[0147] Based on the previous example, suppose the first feedback text obtained is: "First, we will start training with the basic movements of boxing. The training begins now, tentative, tentative, jab, side squat..." When the device outputs the voice corresponding to the first part "tentative", the user responds to the voice by performing a side squat. After obtaining the feature information of the user's response to the first audio-visual data, the device can determine that the feature information does not match "tentative". Furthermore, by comparing the feature information with the reference feature information of "side squat", it finds that the feature information of the user's response to the first audio-visual data matches "side squat". Therefore, it can determine that "side squat" is the third part and output the voice of "side squat".

[0148] When executing processing method two, the device can start from the third part and sequentially output the audio and video data corresponding to each part of the first feedback text after the third part. Each time the audio and video data is output, the device can determine whether the user's response to the audio and video data and the corresponding part match in the manner described in the aforementioned embodiment, so as to determine whether to continue outputting the audio and video data corresponding to the next part.

[0149] When executing processing method two, if it is determined that the third part of the first feedback text comes before the first part, in this case, the audio and video data corresponding to the parts before the first part can be output one by one in reverse order, starting from the first part, until the audio and video data corresponding to the last part of the third part is output.

[0150] For example, suppose the first feedback text includes 10 parts, which are numbered from part 1 to part 10 in chronological order. Suppose that when the device outputs the audio-visual data corresponding to part 7, it determines that the user's response to the audio-visual data features match part 2. At this time, part 7 can be regarded as the first part, and part 4 can be regarded as the third part. The device can output the audio-visual data corresponding to part 6, part 5, part 4, and part 3 in sequence.

[0151] When performing processing method two, the device can also select several adjacent or non-adjacent parts from multiple parts located after the third part in the first feedback text, and output the audio and video data corresponding to the selected parts.

[0152] In this embodiment, after determining the third part, if the first feedback text has been output multiple times, the previous output records can be combined to determine which parts of the audio-visual data were output after the audio-visual data corresponding to the third part each time the first feedback text was output. Then, the parts that were frequently output after the third part in the previous output records are selected, and the audio-visual data corresponding to these parts are output.

[0153] For example, the first feedback text includes a fifth, sixth, and seventh part after the third part. The first feedback text has been output a total of 10 times. In 5 of these instances, the audio-visual data corresponding to the sixth part is output after the audio-visual data of the third part is output. In 3 of these instances, the audio-visual data corresponding to the seventh part is output after the audio-visual data of the third part is output. In this case, the sixth and seventh parts of the first feedback text can be selected, and the audio-visual data corresponding to the sixth part and the audio-visual data corresponding to the seventh part can be output sequentially.

[0154] Selectively outputting specific parts of the audio-visual data in the manner described above can make the output order of multiple parts of the audio-visual data in the first feedback text more in line with the user's behavioral habits.

[0155] In some alternative embodiments, if it is determined that there are multiple parts in the first feedback text that match the feature information of the user's response to the first audio-visual data, the part that is closest to the first part among the multiple parts can be identified as the third part.

[0156] In some embodiments, if the user responds to a mismatch between the feature information of the first audio-visual data and the first part, the device may also execute processing method one and processing method two in sequence. That is, the device may first output the third audio-visual data corresponding to the third part, and then output the audio-visual data corresponding to each part of the first feedback text after the third part in sequence.

[0157] When performing processing method three, the device can compare the reference feature information corresponding to each part of the first feedback text with the feature information of the user's response to the first audio-visual data. If the comparison finds that no part of the first feedback text matches the feature information of the user's response to the first audio-visual data, then the second feedback text can be obtained. The second feedback text can be text that at least meets the following conditions:

[0158] The second feedback text contains at least one part that matches the feature information of the user's response to the first audio-visual data.

[0159] After obtaining the second feedback text, the device can determine the fourth part in the second feedback text that matches the feature information of the user's response to the first audio-visual data, in the same way as the third part was determined in the first feedback text, and then output the audio-visual data corresponding to the fourth part.

[0160] The method for dividing the second feedback text into multiple parts and obtaining the audio-visual data corresponding to the multiple parts of the second feedback text can be found in the aforementioned embodiments and will not be repeated here.

[0161] In some embodiments, a second feedback text can be obtained based on query information. In this embodiment, when obtaining query information, multiple feedback texts can be obtained based on the query information, and then a first feedback text can be determined from the multiple feedback texts by combining information such as user attribute information, historical behavior information, and interaction device information, and the remaining feedback texts can be determined as the second feedback text;

[0162] Based on this, if it is found that none of the parts contained in the first feedback text match the feature information of the user's response to the first audio-visual data, a second feedback text can be selected from the remaining feedback texts.

[0163] In some embodiments, a second feedback text can be obtained based on feature information. In this embodiment, the feature information of the user's response to the first audio-visual data can be compared with each reference feature information in the aforementioned feature information database to determine the target reference feature information that matches the user's feature information. The feature text corresponding to the target reference feature information is then read, and the second feedback text is obtained based on the query information and the feature text corresponding to the target reference feature information.

[0164] As examples, the feature text corresponding to the query information and the target reference feature information can be used as search keywords to retrieve the second feedback text from the internet. Alternatively, the feature text corresponding to the query information and the target reference feature information can be input into a data processing model with text generation capabilities, which can then generate the second feedback text based on these features.

[0165] This application also provides an information output device; please refer to [link to relevant documentation]. Figure 4 This is a schematic diagram of the device, which may include the following units.

[0166] The obtaining unit 401 is used to obtain a first feedback text based on the user's query information. The first feedback text includes at least a first part and a second part. Each part of the first feedback text corresponds to audio-visual data. The audio-visual data includes at least one of audio data, video data, and image data.

[0167] The interaction unit 402 is used to output the first audio-visual data corresponding to the first part, so that the user can respond to the first audio-visual data;

[0168] The interaction unit 402 is used to acquire the feature information of the user's response to the first audio-visual data. If the user's feature information matches the first part, it outputs the second audio-visual data corresponding to the second part of the first feedback text.

[0169] Optionally, when obtaining the first feedback text based on the user's query information, the obtaining unit 401 can be used for:

[0170] The first feedback text is obtained based on the user's query information and the user's interaction device information, and the interaction device information instructs the user to operate the controlled device.

[0171] Optionally, when obtaining the first feedback text based on the user's query information and the user's interaction device information, the obtaining unit 401 may use it for at least one of the following:

[0172] When the interactive device information instructs the user to use the wearable device, the first feedback text related to the action of the target body part is obtained based on the user's query information. The target body part includes the body part of the user wearing the wearable device.

[0173] When the interactive device information instructs the user to use the controller device, the system obtains first feedback text related to the target application based on the user's query information. The target application includes the application that the user operates through the controller device.

[0174] Optionally, the obtaining unit 401 is also used for:

[0175] Identify the characteristic text contained in the first feedback text;

[0176] Identify the semantic relationships between the feature texts contained in the first feedback text;

[0177] Based on the semantic relationship between the feature texts, the first feedback text is divided into at least a first part and a second part, and each part includes at least one feature text with semantic relationship.

[0178] Optionally, the interaction unit 402 can also be used for:

[0179] If the user's feature information does not match the first part, output prompt information based on the first part to prompt the user to respond to the first audio-visual data in a way that matches the first part;

[0180] Alternatively, if the user's feature information does not match the first part, a third part that matches the feature information can be identified from among the multiple parts contained in the first feedback text.

[0181] Output the audio and video data corresponding to the third part.

[0182] Optionally, the interaction unit 402 can also be used for:

[0183] The third part that matches the feature information is identified from the multiple parts contained in the first feedback text. Based on the position of the third part in the first feedback text, the audio-visual data corresponding to the text part located after the position in the first feedback text is output.

[0184] Optionally, the interaction unit 402 can also be used for:

[0185] When multiple parts and feature information contained in the first feedback text do not match, the second feedback text is obtained through obtaining unit 401;

[0186] The fourth part that matches the feature information is identified in the second feedback text;

[0187] Output the audio and video data corresponding to Part 4.

[0188] Optionally, when obtaining the second feedback text, unit 401 can use it for at least one of the following:

[0189] The second feedback text is obtained based on the query information;

[0190] The second feedback text is obtained based on the feature information.

[0191] Optionally, each part of the first feedback text corresponds to audio-visual data, including at least one of the following:

[0192] Generate corresponding audio and video data based on each part of the first feedback text;

[0193] The corresponding audio and video data is retrieved from the database based on each part of the first feedback text.

[0194] The working principle of the information output device provided in this embodiment can be found in the relevant steps of the information output method provided in any embodiment of this application, and will not be repeated here.

[0195] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0196] For ease of description, the above systems or devices are described separately as various modules or units based on their functions. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware components.

[0197] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0198] Finally, it should be noted that in this document, relational terms such as first, second, third, and fourth are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0199] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. An information output method, comprising: The first feedback text is obtained based on the user's query information. The first feedback text includes at least a first part and a second part. Each part of the first feedback text corresponds to audio-visual data, and the audio-visual data includes at least one of audio data, video data, and image data. Output the first audio-visual data corresponding to the first part so that the user can respond to the first audio-visual data; Obtain the feature information of the user's response to the first audio-visual data. If the user's feature information matches the first part, output the second audio-visual data corresponding to the second part of the first feedback text.

2. The method according to claim 1, wherein obtaining the first feedback text based on the user's query information includes: The first feedback text is obtained based on the user's query information and the user's interaction device information, wherein the interaction device information instructs the user to operate the control device.

3. The method according to claim 2, wherein obtaining the first feedback text based on the user's query information and the user's interaction device information includes at least one of the following: When the interactive device information instructs the user to use the wearable device, a first feedback text related to the action of the target body part is obtained based on the user's query information, wherein the target body part includes the body part of the user wearing the wearable device. When the interactive device information instructs the user to use the controller device, a first feedback text related to the target application is obtained based on the user's query information. The target application includes an application that the user operates through the controller device.

4. The method according to claim 1, further comprising: Identify the characteristic text contained in the first feedback text; Identify the semantic relationships between the feature texts contained in the first feedback text; Based on the semantic correlation between the feature texts, the first feedback text is divided into at least a first part and a second part, each part including at least one feature text with semantic correlation.

5. The method according to claim 1, further comprising: If the user's feature information does not match the first part, a prompt message is output based on the first part to prompt the user to respond to the first audio-visual data in a way that matches the first part. Alternatively, if the user's feature information does not match the first part, a third part that matches the feature information can be determined from among the multiple parts contained in the first feedback text. Output the audio and video data corresponding to the third part.

6. The method according to claim 5, further comprising: Among the multiple parts contained in the first feedback text, a third part that matches the feature information is determined. Based on the position of the third part in the first feedback text, the audio-visual data corresponding to the text part located after the position in the first feedback text is output.

7. The method according to claim 1, further comprising: If multiple parts of the first feedback text do not match the feature information, a second feedback text is obtained. The fourth part that matches the feature information is identified in the second feedback text; Output the audio and video data corresponding to the fourth part.

8. The method according to claim 7, wherein obtaining the second feedback text comprises at least one of the following: The second feedback text is obtained based on the query information; The second feedback text is obtained based on the aforementioned feature information.

9. The method according to claim 1, wherein each part of the first feedback text corresponds to audio-visual data, including at least one of the following: Generate corresponding audio-visual data based on each part of the first feedback text; The corresponding audio and video data is retrieved from the database based on each part of the first feedback text.

10. An information output device, comprising: The obtaining unit is configured to obtain a first feedback text based on the user's query information. The first feedback text includes at least a first part and a second part. Each part of the first feedback text corresponds to audio-visual data, and the audio-visual data includes at least one of audio data, video data, and image data. An interactive unit is used to output the first audio-visual data corresponding to the first part, so that the user can respond to the first audio-visual data; The interaction unit is used to acquire feature information of the user's response to the first audio-visual data. If the user's feature information matches the first part, the second audio-visual data corresponding to the second part of the first feedback text is output.

Citation Information

Patent Citations

  • Interactive TV and method for realizing interaction with user by utilizing display device

    CN102724449A

  • Video query method and device, electronic device and storage medium

    CN113987271A