Audio playback method and apparatus and terminal device

By enabling the terminal device to determine the scene after component switching and using conversation information to recommend music, the problem of users needing to accurately describe their musical intentions is solved, enabling more efficient music playback.

WO2025194881A1PCT designated stage Publication Date: 2025-09-25BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/139076
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-22
Filing Date
2024-12-13
Publication Date
2025-09-25

AI Technical Summary

Technical Problem

Existing music-related AI systems require users to accurately describe their music playback intentions and music titles, which makes music playback operations highly complex.

Method used

The terminal device responds to the component switching operation, determines the current scene and obtains the associated conversation information, combines the historical conversation information, and recommends and plays music suitable for the current scene.

Benefits of technology

The accuracy of music recommendations and playback is improved, and the complexity of user input of relevant information is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024139076_25092025_PF_FP_ABST
    Figure CN2024139076_25092025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure provides an audio playback method and apparatus and a terminal device. The audio playback method comprises: in response to a switching operation on components, determining a first scenario corresponding to a currently selected first component, wherein the components are used for text dialogues, and the functions of the text dialogues of the plurality of components are different; obtaining first dialogue information associated with the first component; on the basis of the first scenario and the first dialogue information, determining recommended music among a plurality of pieces of preset music; and playing back the recommended music. Embodiments of the present disclosure can reduce the audio playback complexity.
Need to check novelty before this filing date? Find Prior Art

Description

Audio playback method, device and terminal equipment

[0001] This application claims priority to Chinese Patent Application No. 202410339302.5 filed on March 22, 2024, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field

[0002] The embodiments of the present disclosure relate to an audio playback method, apparatus, and terminal device. Background Art

[0003] With the continuous development and popularization of artificial intelligence (AI) technology, its application in the field of music is becoming more and more extensive.

[0004] Currently, music-related AI can play related music based on user input. For example, a user can input "Play Music 1 and Music 2" to the music-related AI, and the music-related AI will play Music 1 and Music 2 based on the user's input. However, in this method, the user's input to the music-related AI must accurately describe the intention of playing music and the name of the music to be played, which makes music playback more complex. Summary of the Invention

[0005] The present disclosure provides an audio playback method, apparatus, and terminal device.

[0006] In a first aspect, the present disclosure provides an audio playback method, the method comprising:

[0007] In response to a component switching operation, determining a first scene corresponding to a currently selected first component, wherein the component is used for text conversation, and functions of the text conversations between the multiple components are different;

[0008] Obtaining first conversation information associated with the first component;

[0009] According to the first scene and the first dialogue information, recommended music is determined from a plurality of preset music, and the recommended music is played.

[0010] In a second aspect, the present disclosure provides an audio playback device, comprising a first determination module, an acquisition module, a second determination module, and a playback module, wherein:

[0011] The first determining module is configured to, in response to a component switching operation, determine a first scenario corresponding to a currently selected first component, wherein the component is used for text conversation, and the functions of the text conversations between the multiple components are different;

[0012] The acquisition module is used to acquire first conversation information associated with the first component;

[0013] The second determining module is configured to determine recommended music from a plurality of preset music pieces according to the first scene and the first conversation information;

[0014] The playing module is used to play the recommended music.

[0015] In a third aspect, an embodiment of the present disclosure provides a terminal device including: a processor and a memory;

[0016] The memory stores computer-executable instructions;

[0017] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the audio playback method as described in the first aspect and various possible aspects of the first aspect.

[0018] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the audio playback method as described in the first aspect and various possible aspects of the first aspect are implemented. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, a brief introduction will be given below to the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0020] FIG1 is a schematic diagram of an application scenario provided by an embodiment of the present disclosure;

[0021] FIG2 is a schematic diagram of a flow chart of an audio playback method provided by an embodiment of the present disclosure;

[0022] FIG3 is a schematic diagram of first conversation information provided by an embodiment of the present disclosure;

[0023] FIG4 is a schematic diagram of second dialogue information provided by an embodiment of the present disclosure;

[0024] FIG5 is a schematic diagram of a process for determining recommended music according to an embodiment of the present disclosure;

[0025] FIG6 is a schematic diagram of a method for playing recommended music provided by an embodiment of the present disclosure;

[0026] FIG7 is a schematic diagram of a process of playing recommended music provided by an embodiment of the present disclosure;

[0027] FIG8 is a process diagram of an audio playback method provided by an embodiment of the present disclosure;

[0028] FIG9 is a schematic structural diagram of an audio playback device provided by an embodiment of the present disclosure; and

[0029] FIG10 is a schematic structural diagram of a terminal device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0030] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.

[0031] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0032] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the terminal device, application, server, or storage medium, that performs the operation of the disclosed technical solution, based on the prompt message.

[0033] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. In addition, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the terminal device.

[0034] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0035] To facilitate understanding, the concepts involved in the embodiments of the present disclosure are explained below.

[0036] Terminal device: is a device with wireless transceiver function. The terminal device can be deployed on land, including indoors or outdoors, handheld, wearable or vehicle-mounted. The terminal device can be a mobile phone, a tablet computer, a computer with wireless transceiver function, a virtual reality (VR) terminal device, an augmented reality (AR) terminal device, a wireless terminal in industrial control, a vehicle-mounted terminal device, a wireless terminal in self-driving, a wireless terminal device in remote medical, a wireless terminal device in smart grid, a wireless terminal device in transportation safety, a wireless terminal device in smart city, a wireless terminal device in smart home, a wearable terminal device, etc. The terminal device involved in the embodiments of the present disclosure can also be called a terminal, user equipment (UE), an access terminal device, a vehicle-mounted terminal, an industrial control terminal, a UE unit, a UE station, a mobile station, a mobile station, a remote station, a remote terminal device, a mobile device, a UE terminal device, a wireless communication device, a UE agent or a UE device, etc. Terminal devices can also be fixed or mobile.

[0037] In the related art, with the continuous development and popularization of AI technology, its application in the music field is becoming increasingly widespread. For example, AI technology can be applied in scenarios such as music creation and music recommendation. Currently, music-related AI can play related music based on user input. For example, a user can input "Play Music 1 and Music 2" by voice or text to the music-related AI, and the music-related AI will play Music 1 and Music 2 based on the user's voice or text input. However, in this method, the user's voice input to the music-related AI must accurately describe the intended music playback. For example, if the user inputs "Music 1 and Music 2" to the music-related AI, the music-related AI will retrieve information related to Music 1 and Music 2 from its knowledge base and play the relevant information, without playing Music 1 and Music 2. Secondly, the user's voice input to the music-related AI must accurately describe the name of the music to be played. For example, if the user inputs "Play upbeat music" to the music-related AI, the music-related AI will not recognize the music name and will not play the relevant music. This leads to a high level of operational complexity in music playback.

[0038] To address the technical problems in the related art, an embodiment of the present disclosure provides an audio playback method. In response to a component switching operation, a terminal device can determine a first scene corresponding to a currently selected first component and obtain first conversation information associated with the first component. The terminal device can also obtain second conversation information within a historical period, wherein the second conversation information includes information related to music. The terminal device can determine recommended music from a plurality of preset music based on the first scene, the first conversation information, and the second conversation information, and play the recommended music. In the above method, since the first conversation information is related to a first scene (e.g., fitness), the first conversation information can include specific scene information (e.g., running, cycling, etc.), and the second conversation information can accurately indicate the music that the user is interested in within the historical period. Therefore, the terminal device can accurately determine the recommended music that the user is interested in in the first scene and play the recommended music. This can improve the accuracy of music recommendations and the accuracy of music playback. In addition, during the user's conversation with the first component, the music can be played without the user inputting music-related information to the component, thereby reducing the complexity of music playback.

[0039] The application scenario of the embodiment of the present disclosure is described below with reference to FIG1 .

[0040] Figure 1 is a schematic diagram of an application scenario provided by an embodiment of the present disclosure. Referring to Figure 1 , it includes: a terminal device. The terminal device displays a page for component 1, which includes a question text a entered by the user, an answer text b output by component 1, and a voice input control. When the user switches from component 1 to component 2, the terminal device can display the page for component 2. The user can engage in a conversation with component 2 on the page for component 2. For example, the user enters question and answer text c, and component 2 outputs answer text d. After component 2 answers the question, it can generate recommended music based on the scenario corresponding to component 2 and the conversation information in component 2. The recommended music includes music 1, music 2, ..., music 10, and the terminal device can play music 1. Thus, after the terminal device switches components, when it engages in a conversation with the new component, it can determine the recommended music appropriate for the current scenario and play the recommended music, without the user having to enter music-related information into the component, thereby reducing the complexity of music playback.

[0041] It should be noted that FIG1 is only an example of an application scenario of the embodiment of the present disclosure, and does not limit the application scenario of the embodiment of the present disclosure.

[0042] The following detailed description of the technical solution of the present disclosure and how the technical solution of the present disclosure solves the above-mentioned technical problems is provided with specific embodiments. The following specific embodiments may be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. The embodiments of the present disclosure will be described below in conjunction with the accompanying drawings.

[0043] FIG2 is a flow chart of an audio playback method provided by an embodiment of the present disclosure. Referring to FIG2 , the method may include:

[0044] S201: In response to a component switching operation, determine a first scene corresponding to a currently selected first component.

[0045] The execution subject of the embodiment of the present disclosure may be a terminal device, or an audio playback device provided in the terminal device. The audio playback device may be implemented based on software, or based on a combination of software and hardware, which is not limited in the embodiment of the present disclosure.

[0046] Components can be used for text conversations. For example, the components can be components in a preset neural network model, where the preset neural network model can be used to understand and generate natural language text and process various natural language tasks (such as question-answering, conversation, and translation). The preset neural network model can include multiple components, each of which has the function of text conversation.

[0047] The functions of text conversations between multiple components are different. For example, if the text conversation function of component 1 can be a fitness recommendation function, and the text conversation function of component 2 can be a music recommendation function, then during the text conversation, component 1 will have a higher accuracy in answering fitness-related questions, while component 2 will have a higher accuracy in answering music-related questions. In this way, the differences between multiple components can improve the user experience.

[0048] It should be noted that multiple components can perform text conversations, but different components have different focuses during the training phase. For example, when training fitness-related components, the proportion of fitness-related text content in the training samples is relatively high (the training focuses on answering fitness-related questions, so the component can answer fitness-related questions with high accuracy, but the component can also answer questions in other fields). When training music-related components, the proportion of music-related text content in the training samples is relatively high (the training focuses on answering music-related questions, so the component can answer music-related questions with high accuracy, but the component can also answer questions in other fields).

[0049] The switching operation may be a component switching operation. For example, after a user has engaged in a conversation using component 1, they switch component 1 to component 2. This operation may be a component switching operation. For example, if the terminal device is currently displaying a fitness-related component page, the user may perform a component switching operation, and the terminal device may cancel the fitness-related component page and display the music-related component page.

[0050] The first component may be the currently selected component after the switching operation. For example, if the user closes component 1 and opens component 2 in the terminal device, the terminal device may determine that the first component is component 2 (the currently selected component). For example, if the page displayed on the terminal device is a page of fitness-related components, if the user closes the page of fitness-related components in the terminal device and opens the page of music-related components in the terminal device, the terminal device may determine that the first component is a music-related component.

[0051] The first scene may be a scene associated with the text conversation function of the first component. For example, if the first component is a fitness-related component, the first scene may be a fitness scene (e.g., the user is currently preparing for or is currently exercising); if the first component is a writing-related component, the first scene may be a writing scene (e.g., the user is currently preparing for or is currently writing).

[0052] Optionally, the terminal device may determine the first scenario corresponding to the currently selected first component according to the following two feasible implementations:

[0053] A possible implementation:

[0054] Obtain a correspondence between components and scenes, and determine the first component according to the first component and the correspondence between the components and scenes.

[0055] The correspondence between components and scenes may include at least one component and the scene corresponding to each component. For example, the correspondence between components and scenes may be as shown in Table 1:

[0056] Table 1

[0057] It should be noted that Table 1 is only an example of the correspondence between components and scenarios, and does not limit the correspondence between components and scenarios.

[0058] For example, if the terminal device determines that the currently selected first component is component 1, the terminal device can determine that the first scene corresponding to the first component is scene 1; if the terminal device determines that the currently selected first component is component 2, the terminal device can determine that the first scene corresponding to the first component is scene 2; if the terminal device determines that the currently selected first component is component 3, the terminal device can determine that the first scene corresponding to the first component is scene 3.

[0059] Another possible implementation:

[0060] At least one device information is obtained, and a first scenario is determined based on the device information.

[0061] The device information may include information collected by sensors. For example, the device information may include data collected by a gyroscope, data collected by an inertial sensor, images or videos captured by a camera, and data collected by an optical heart rate sensor, etc., which is not limited in the present embodiment.

[0062] It should be noted that the terminal device can obtain device information from any device such as a wearable device (such as a bracelet), a handheld device (such as a handle), etc., or it can obtain device information from the terminal device. This embodiment of the present disclosure does not limit this.

[0063] After the terminal device obtains the device information, it can process the device based on a neural network model for scene recognition (the model structure can be any feasible neural network structure) to obtain a first scene. For example, the terminal device can input the device information of at least one device into the neural network model for scene recognition, and the neural network model for scene recognition can output the scene corresponding to the device information, that is, the first scene corresponding to the first component, thereby improving the accuracy of the first scene.

[0064] It should be noted that the neural network model used to identify scenes can be a model pre-trained based on multiple groups of training samples, wherein each group of training samples can include multiple sample device information and sample scenes corresponding to the multiple sample device information. The training process of the neural network model used to identify scenes will not be repeated in the embodiments of the present disclosure.

[0065] S202: Obtain first conversation information associated with the first component.

[0066] The first conversation information may include conversation information between the user and the first component. For example, after the terminal device displays the first component, the user may ask a question in the first component, and the first component may answer the question. The first conversation information may include the user's question and the first component's answer. For example, the user and the first component may have multiple rounds of conversations, and the terminal device may obtain the text content of these multiple rounds of conversations to obtain the first conversation information.

[0067] Next, the first conversation information obtained by the terminal device is described with reference to FIG3 .

[0068] Figure 3 is a schematic diagram of a first dialogue information provided by an embodiment of the present disclosure. Please refer to Figure 3, which includes: a terminal device. The display page of the terminal device is a page related to the first component. The user can enter question text A in the page (can be based on voice input, the terminal device can convert the voice into text, and display the text), and the first component can generate an answer text a corresponding to the question text A (can be based on voice output, the terminal device can convert the voice into text, and display the text). If the user continues to enter question text B, the first component can generate an answer text b corresponding to the question text B. The terminal device can determine the first dialogue information based on multiple dialogues between the user and the first component, wherein the first dialogue information can include question text A, question text B, answer text a, and answer text b.

[0069] It should be noted that if the user uses the first component for the first time, the terminal device can obtain the first conversation information based on the new conversation content between the user and the component. If the user has used the first component multiple times, the terminal device can obtain the first conversation information based on the historical conversation content and new conversation content between the user and the component.

[0070] It should be noted that the terminal device can obtain the first conversation information corresponding to the first component according to any feasible implementation method, and the embodiment of the present disclosure is not limited to this.

[0071] S203: Determine recommended music from a plurality of preset music according to the first scene and the first dialogue information.

[0072] The preset music may be music pre-stored in the terminal device or music stored in a database, which is not limited in the embodiment of the present disclosure.

[0073] Optionally, the terminal device may determine music that matches the current scene from among multiple preset music pieces based on the first scene and the first conversation information, and obtain the recommended music. For example, if the first scene may be a fitness scene, and the first conversation information indicates that the user's fitness method is running, the terminal device may determine music suitable for running from among the multiple preset music pieces based on this information, and determine the suitable running music as the recommended music.

[0074] Optionally, the terminal device may process the first scene and the first conversation information to obtain a feature vector (embedding) of the first scene and a feature vector of the first conversation information. The terminal device may determine a feature vector for each preset music piece and, based on the multiple feature vectors, determine recommended music. For example, the terminal device may process the first scene and the first conversation information based on a neural network model composed of multiple convolutional layers (which may be a neural network model of any structure) to obtain feature vector 1, and process each preset music piece based on the neural network model to obtain feature vector 2 corresponding to each preset music piece. The terminal device may calculate the similarity between feature vector 1 and each feature vector 2, and determine recommended music based on the multiple similarities. For example, the similarity between the first scene and the first conversation information and music 1 is similarity a, and the similarity between the first scene and the first conversation information and music 2 is similarity b. If similarity a is greater than similarity b, the terminal device may determine music 1 as the recommended music piece.

[0075] Optionally, the terminal device may determine the recommended music according to the following feasible implementation method: obtaining the second conversation information within the historical period, and determining the recommended music from multiple preset music based on the first scene, the first conversation information and the second conversation information.

[0076] The second conversation information may include music-related information. For example, the second conversation information may include conversations between the user and the first component within a historical period, and the conversations include music-related information. For example, the second conversation information may include conversations between the user and other components other than the first component within a historical period, and the conversations include music-related information.

[0077] Optionally, the music-related information may include user comments on the music. For example, the music-related information may include: the user's favorite music and music genres, the user's disliked music and music genres, the user's favorite music and music genres in the first scene, and the user's disliked music and music genres in the first scene, etc., which is not limited in the present embodiment.

[0078] For example, the music-related information in the second conversation information may include: the user likes rock music, dislikes rap music, and dislikes music sung by singer 1, but likes music sung by singer 2. In rock music, the user likes music A and music B. In this way, the terminal device can accurately determine the user's preferred music type based on the second conversation information, and then accurately determine the recommended music, thereby improving the accuracy of music playback.

[0079] It should be noted that the historical period can be any set period (such as the past day, the past week, and the past month, etc.), and the embodiments of the present disclosure are not limited to this.

[0080] Next, the second dialogue information will be described with reference to FIG. 4 .

[0081] FIG4 is a schematic diagram of a second dialogue information provided by an embodiment of the present disclosure. Please refer to FIG4 , which includes: dialogue information of component 1, dialogue information of component 2, and dialogue information of component 3. Among them, the dialogue information of component 1 includes: the text "Can you recommend some music?", the text "What style of music do you like?", the text "Rock style music", and multiple recommended music. The dialogue information of component 2 includes: the text "Which rock music is good?", recommended rock music, and the text "I am working out and don't want to listen to music with a slower tempo." The dialogue information of component 3 includes: the text "What are some places suitable for visiting?", recommended places to visit, and the text "When is the best time to go to the zoo?"

[0082] Please refer to Figure 4. Since the conversation information of component 1 and the conversation information of component 2 are related to music, and the conversation information of component 3 is not related to music, the terminal device (not shown in Figure 4) can determine that the second conversation information includes the conversation information of component 1 and the conversation information of component 2 (which may include conversation information related to music in components 1 and 2, and may also include conversation information not related to music, which is not limited in this embodiment of the present disclosure). Since the conversation information of component 3 does not include conversation information related to music, the second conversation information does not include the conversation information of component 3.

[0083] In this way, based on the conversation information of component 1 and the conversation information of component 2, the terminal device can determine that the user likes rock-style music, and that the user likes music with a faster rhythm and dislikes music with a slower rhythm during fitness. Therefore, the terminal device can accurately determine the type of music the user likes based on the second conversation information, and then accurately determine the recommended music, thereby improving the accuracy and effect of music playback and improving the user experience.

[0084] It should be noted that the terminal device can obtain the second conversation information according to any feasible implementation method, and the embodiment of the present disclosure is not limited to this.

[0085] Optionally, the terminal device determines recommended music from a plurality of preset music based on the first scene, the first conversation information, and the second conversation information. For example, the terminal device may process the first scene, the first conversation information, and the second conversation information based on a neural network model composed of multiple convolutional layers (which may be a neural network model of any structure) to obtain a feature vector 1, and process each preset music based on the neural network model to obtain a feature vector 2 corresponding to each preset music. The terminal device may calculate the similarity between the feature vector 1 and each feature vector 2, and determine the recommended music based on multiple similarities. It should be noted that the method by which the terminal device determines recommended music from a plurality of preset music based on the first scene, the first conversation information, and the second conversation information is similar to the method by which the terminal device determines recommended music from a plurality of preset music based on the first scene and the first conversation information, and the embodiments of the present disclosure will not be repeated here.

[0086] Optionally, the terminal device determines recommended music from multiple preset music based on the first scene, the first conversation information and the second conversation information. Specifically, it can also be: determining scene switching information associated with the switching operation, and determining recommended music from multiple preset music according to the first scene, the first conversation information, the second conversation information and the scene switching information.

[0087] The scene switching information can be used to indicate the scene associated with the switching operation. For example, if the switching operation switches component 1 to component 2, and the scene corresponding to component 1 is scene A and the scene corresponding to component 2 is scene B, the scene switching information associated with the switching operation can include information that the user switched from scene A to scene B.

[0088] Optionally, the scene switching information may include at least one of the following:

[0089] The second scenario corresponding to the second component

[0090] an identifier of the first component and an identifier of the second component;

[0091] The text information corresponding to the switch operation.

[0092] The second component may be the component selected before the switching operation. For example, the terminal device switches the current component from component 1 to component 2 based on the switching operation. Since component 1 is the component selected before the switching operation and component 2 is the component selected after the switching operation, the terminal device may determine component 1 as the second component and component 2 as the first component.

[0093] The second scene may be a scene associated with the text conversation function of the second component. For example, if the second component is an English learning component (e.g., the user is preparing to learn English or is currently learning English), the second scene may be English learning; if the second component is a fitness component, the second scene may be a fitness scene (e.g., the user is preparing to exercise or is currently exercising).

[0094] It should be noted that the method for the terminal device to determine the second scene corresponding to the second component is the same as the method for the terminal device to determine the first scene corresponding to the first component, and the embodiments of the present disclosure will not be repeated here.

[0095] Optionally, the identifier of the first component can be a unique identifier such as the name or label of the first component, which is not limited in the embodiments of the present disclosure. The identifier of the second component can be a unique identifier such as the name or label of the second component, which is not limited in the embodiments of the present disclosure.

[0096] The text information corresponding to the switching operation is used to describe the switching operation. For example, the text information corresponding to the switching operation may be: switching from an English learning scene to a fitness scene, switching from an English learning scene to a writing scene, and so on.

[0097] In this way, because music recommendations incorporate scene switching information, the accuracy of music recommendations can be improved. For example, if the scene switching information can include that the second component is an English learning component, the first component is a fitness component, and the text corresponding to the switching operation indicates that the English learning scene is switched to a fitness scene, then when the terminal device makes music recommendations based on the first scene, the first conversation information, the second conversation information, and the scene switching information, it can recommend English music suitable for fitness that the user likes, thereby improving the accuracy of music recommendations and music playback.

[0098] It should be noted that the terminal device can obtain scene switching information based on any feasible implementation method. For example, the second scene corresponding to the second component, the identifier of the first component and the identifier of the second component can be pre-set information, and the text information corresponding to the switching operation can be pre-set information (such as pre-setting the corresponding text information for each switching operation), or it can be text information generated based on any feasible implementation method. The embodiments of the present disclosure are not limited to this.

[0099] Optionally, the terminal device determines recommended music from multiple preset music based on the first scene, the first conversation information, the second conversation information and the scene switching information. Specifically, the first scene, the first conversation information, the second conversation information and the scene switching information are input into the neural network model to obtain the first music feature, obtain the second music feature corresponding to each preset music, and determine the similarity between the first music feature and each second music feature, and determine the recommended music based on multiple similarities.

[0100] The neural network model may be a model of any structure for extracting feature vectors, which is not limited in the embodiments of the present disclosure.

[0101] The first music feature may be used to indicate the music information to be recommended. For example, the first music feature may indicate that the music information to be recommended includes rock style, Chinese, and fast tempo.

[0102] The second music feature can be used to indicate music information of the preset music. For example, if the second music feature corresponding to preset music 1 is feature A, and the second music feature corresponding to preset music 2 is feature B, feature A can indicate music information (such as style, rhythm, etc.) of music 1, and feature B can indicate music information of music 2.

[0103] It should be noted that the terminal device can process the preset music based on the neural network model for extracting audio features to obtain the second music feature corresponding to the preset music.

[0104] It should be noted that the second music feature corresponding to the preset music may be a music feature that is predetermined and stored in a database, and the terminal device may obtain the second music feature corresponding to each preset music in the database.

[0105] Optionally, the terminal device determines the similarity between the first music feature and each second music feature. Specifically, for any second music feature, the terminal device can calculate the cosine similarity between the first music feature and the second music feature, and determine the cosine similarity between the first music feature and the second music feature as the similarity between the first music feature and the second music feature.

[0106] For example, the first music feature may be feature 1, and the second music feature may include feature 2 and feature 3. The terminal device may calculate the cosine similarity between feature 1 and feature 2 to obtain the similarity between feature 1 and feature 2. The terminal device may calculate the cosine similarity between feature 1 and feature 3 to obtain the similarity between feature 1 and feature 3.

[0107] It should be noted that the terminal device may also determine the similarity between the first music feature and each second music feature based on any other feasible implementation method (such as Euclidean distance), and the embodiment of the present disclosure is not limited to this.

[0108] The terminal device may determine the recommended music based on multiple similarities. For example, the terminal device may determine N preset music corresponding to N (an integer greater than 0) second music features with the highest similarity to the first music feature as recommended music. For example, the terminal device may determine the preset music corresponding to the second music feature with the highest similarity to the first music feature as recommended music. The terminal device may also determine multiple preset music corresponding to multiple second music features with the highest similarity to the first music feature as recommended music.

[0109] The process of determining recommended music will be described below with reference to FIG. 5 .

[0110] FIG5 is a schematic diagram of a process for determining recommended music according to an embodiment of the present disclosure. Referring to FIG5 , the process includes: a first scene, first conversation information, second conversation information, scene switching information, and a neural network model. A terminal device (not shown in FIG5 ) can input the first scene, first conversation information, second conversation information, and scene switching information into the neural network model, and the neural network model can output a first music feature.

[0111] Referring to Figure 5 , the second music feature includes music feature 1 corresponding to music 1, music feature 2 corresponding to music 2, ..., and music feature n corresponding to music n. The terminal device can calculate the similarity between the first music feature and each second music feature. Here, the similarity between the first music feature and music feature 1 is similarity 1, the similarity between the first music feature and music feature 2 is similarity 2, ..., and the similarity between the first music feature and music feature n is similarity n.

[0112] Referring to Figure 5, if the terminal device recommends the 10 most similar songs, the terminal device can determine that the recommended music includes Music 1, Music 2, ..., Music 10 (the 10 songs corresponding to the top 10 second music features with the highest similarity to the first music feature). In this way, the terminal device can accurately determine the recommended music, improve the accuracy of the recommended music, and improve the accuracy of music playback.

[0113] S204: Play recommended music.

[0114] Optionally, after the terminal device determines the recommended music, it can play the recommended music in the order in which the music is recommended. For example, in the embodiment shown in FIG10 , the terminal device can play music 1, music 2, ..., music 10 in sequence.

[0115] The disclosed embodiment provides an audio playback method, wherein a terminal device can determine a first scene corresponding to a currently selected first component in response to a switching operation on a component, and obtain first conversation information associated with the first component. The terminal device can obtain second conversation information within a historical period, and determine scene switching information associated with the switching operation. The terminal device can determine recommended music from a plurality of preset music based on the first scene, the first conversation information, the second conversation information, and the scene switching information, and play the recommended music. In this way, since the terminal device can accurately determine the user's current scene based on the first scene and the first conversation information, and can accurately determine the user's favorite music type based on the second conversation information and the scene switching information, the terminal device can accurately determine the recommended music, thereby improving the accuracy of music recommendation, improving the accuracy of music playback, and improving the user experience. Moreover, during the user's conversation with the first component, the music can be played without the user inputting music-related information into the component, thereby reducing the complexity of music playback.

[0116] Based on the embodiment shown in FIG. 2 , the method for playing recommended music in the above-mentioned video playing method will be described in detail below in conjunction with FIG. 6 .

[0117] FIG6 is a schematic diagram of a method for playing recommended music provided by an embodiment of the present disclosure. Referring to FIG6 , the method process may include:

[0118] S601: Determine the opening voice before playing the recommended music.

[0119] The opening voice is used to associate the first scene with the recommended music. For example, a user may have a conversation with the first component in the current scene. When the conversation ends, the terminal device may generate recommended music. Before playing the recommended music, the terminal device may determine the opening voice and play it, thereby improving the user experience.

[0120] For example, if the first scene is a fitness scene and the first conversation information includes information about running, the terminal device can determine that the current scene is a running scene, and the opening voice generated by the terminal device can be "Start with running, and now play songs suitable for running for you." This can improve the user experience.

[0121] It should be noted that the terminal device can determine the opening voice based on any feasible implementation method, and the embodiments of the present disclosure are not limited to this.

[0122] S602: Play the opening voice and recommended music.

[0123] Optionally, the terminal device can play recommended music while playing the opening speech. For example, the terminal device can play the recommended music while playing the opening speech, and reduce the volume of the recommended music. After the opening speech ends, the volume of the recommended music can be restored to improve the music playback effect.

[0124] Optionally, the terminal device can play the recommended music after the opening voice is played. In this way, the terminal device does not need to adjust the playback volume of the recommended music, thereby improving the user experience.

[0125] Among them, the terminal device plays the recommended music, which can specifically include: obtaining music information related to the recommended music, generating a recommendation text describing the recommended music based on the music information, determining a recommended voice corresponding to the recommendation text, and playing the recommended voice when playing the recommended music.

[0126] The music information may be information related to the recommended music. For example, the music information related to the recommended music may include the music title, singer name, album name, lyrics, music-related comments, music-related background information, and any other information related to the recommended music, which is not limited in the present embodiment.

[0127] It should be noted that the terminal device can obtain music information related to the recommended music according to any feasible implementation method, and the embodiments of the present disclosure are not limited to this.

[0128] The recommendation text can be used to describe the recommended music. For example, the recommendation text can include text related to the music information of the recommended music. For example, the recommendation text can include the music title, artist name, and background information related to the music. In this way, the recommended music can be introduced based on the recommendation text, thereby improving the user experience.

[0129] It should be noted that the terminal device can process the music information based on any feasible implementation method to obtain recommended text (for example, the music information is processed based on a pre-trained neural network model, and the pre-trained neural network model can summarize the music information to obtain recommended text), and the embodiments of the present disclosure are not limited to this.

[0130] The recommended voice may be the voice corresponding to the recommended text. For example, if the recommended text is “Music 1”, the recommended voice corresponding to the recommended text may be “Music 1”.

[0131] It should be noted that the terminal device can process the recommended text based on any feasible implementation method to obtain the recommended voice (for example, processing the recommended text based on a text-to-speech model to obtain the recommended voice), and the embodiments of the present disclosure are not limited to this.

[0132] Optionally, the terminal device may play the recommended voice corresponding to the recommended music before playing the recommended music, and the terminal device may also play the recommended voice corresponding to the recommended music when starting to play the recommended music (lowering the playback volume of the recommended music and restoring the playback volume of the recommended music when the recommended voice finishes playing). The embodiments of the present disclosure are not limited to this.

[0133] Optionally, when the user is conversing with the component, since the user can input voice, the output of the component can also be voice. Therefore, when the terminal device plays the recommended voice (the recommended music is played after the recommended voice is finished), if there is currently a voice output by the component (such as the answer voice output by the component based on the question asked by the user), the terminal device can pause the playback of the recommended voice, and continue to play the recommended voice when the voice output by the component is finished. When the terminal device plays the recommended music, if there is currently a voice output by the component, the terminal device can lower the playback volume of the recommended music, and can restore the playback volume of the recommended music when the voice output by the component is finished.

[0134] Next, the process of playing recommended music will be described with reference to FIG. 7 .

[0135] Figure 7 is a schematic diagram of a process for playing recommended music provided by an embodiment of the present disclosure. Referring to Figure 7, a terminal device is included. The terminal device displays a page corresponding to the component, which may include recommended music, including Music 1, Music 2, ..., and Music 10. Before playing the recommended music, the terminal device may play an opening voiceover. After the opening voiceover ends, the recommended voiceover corresponding to Music 1 may be played.

[0136] As shown in Figure 7, after the recommended voice message corresponding to Music 1 finishes playing, the terminal device can start playing Music 1. When Music 1 finishes playing, the terminal device can stop playing Music 1 and start playing the recommended voice message corresponding to Music 2. After the recommended voice message corresponding to Music 2 finishes playing, the terminal device can start playing Music 2. In this way, the terminal device can play the recommended voice message corresponding to each recommended song while playing the recommended music, thereby improving the music playback effect and enhancing the user experience.

[0137] The disclosed embodiments provide a method for playing recommended music. A terminal device can determine the opening voice before playing the recommended music, obtain music information related to the recommended music, generate a recommendation text describing the recommended music based on the music information, determine the recommended voice corresponding to the recommended voice, and play the recommended voice when playing the recommended music. In this way, the terminal device can introduce the upcoming recommended music based on the recommended voice, thereby improving the user experience.

[0138] Based on any of the above embodiments, the process of the above audio playing method will be described below with reference to FIG. 8 .

[0139] FIG8 is a process diagram of an audio playback method provided by an embodiment of the present disclosure. Referring to FIG8 , it includes: a terminal device. The display page of the terminal device is the page of component 1 (the second component), and the page of component 1 includes the question text a input by the user, the answer text b output by component 1, and a voice input control. When the user switches component 1 to component 2, the terminal device can display the page of component 2 (the first component). The user can have a conversation with component 2 on the page of component 2.

[0140] Referring to Figure 8 , the terminal device can identify the first conversation message of component 2 and the second conversation message of component 1. The first conversation message includes conversation messages related to running warm-ups. The second conversation message includes the text "Which rock music is good?", recommended rock music, and the text "I'm working out and don't want to listen to slow music."

[0141] Referring to Figure 8 , the terminal device may determine that the first scene is a fitness scene, and the scene switching information indicates that the music recommendation scene is switched to the fitness scene. The terminal device may input the first scene, the first conversation information, the second conversation information, and the scene switching information into the neural network model, and the neural network model may output the first music feature.

[0142] Referring to FIG8 , the second music feature includes music feature 1 corresponding to music 1, music feature 2 corresponding to music 2, ..., and music feature n corresponding to music n. The terminal device can calculate the similarity between the first music feature and each second music feature. Here, the similarity between the first music feature and music feature 1 is similarity 1, the similarity between the first music feature and music feature 2 is similarity 2, ..., and the similarity between the first music feature and music feature n is similarity n.

[0143] Referring to Figure 8 , if the terminal device recommends the 10 most similar songs, the terminal device may determine that the recommended music includes Music 1, Music 2, ..., Music 10 (the 10 songs corresponding to the top 10 second music features with the highest similarity to the first music feature), and display the recommended music on the page of Component 2. The terminal device may play the opening voice before playing the recommended music. After the opening voice ends, the recommended voice corresponding to Music 1 may be played. After the recommended voice corresponding to Music 1 ends, the terminal device may start playing Music 1.

[0144] In this way, since the terminal device can accurately determine the user's current scene based on the first scene and the first conversation information, and can accurately determine the user's favorite music type based on the second conversation information and the scene switching information, the terminal device can accurately determine the recommended music, thereby improving the accuracy of music recommendations and the accuracy of music playback. In addition, the terminal device can play the recommended voice corresponding to each recommended music during the process of playing the recommended music, thereby improving the effect of music playback and the user's experience. Moreover, during the user's conversation with the first component, the music can be played without the user inputting music-related information into the component, thereby reducing the complexity of music playback.

[0145] FIG9 is a schematic diagram of the structure of an audio playback device provided by an embodiment of the present disclosure. Referring to FIG9 , the audio playback device 900 includes a first determination module 901, an acquisition module 902, a second determination module 903, and a playback module 904, wherein:

[0146] The first determining module 901 is configured to determine, in response to a component switching operation, a first scenario corresponding to a currently selected first component, wherein the component is used for text conversations, and the functions of text conversations between multiple components are different;

[0147] The acquisition module 902 is used to acquire first conversation information associated with the first component;

[0148] The second determining module 903 is configured to determine recommended music from a plurality of preset music pieces according to the first scene and the first conversation information;

[0149] The playing module 904 is used to play the recommended music.

[0150] According to one or more embodiments of the present disclosure, the second determining module 903 is specifically configured to:

[0151] Acquiring second conversation information within a historical period, where the second conversation information includes music-related information;

[0152] The recommended music is determined from the plurality of preset music based on the first scene, the first dialogue information, and the second dialogue information.

[0153] According to one or more embodiments of the present disclosure, the second determining module 903 is specifically configured to:

[0154] Determining scene switching information associated with the switching operation, where the scene switching information is used to indicate the scene associated with the switching operation;

[0155] The recommended music is determined from the plurality of preset music according to the first scene, the first dialogue information, the second dialogue information and the scene switching information.

[0156] According to one or more embodiments of the present disclosure, the scene switching information includes at least one of the following:

[0157] a second scene corresponding to a second component, the second component being the component selected before the switching operation;

[0158] an identifier of the first component and an identifier of the second component;

[0159] The text information corresponding to the switching operation.

[0160] According to one or more embodiments of the present disclosure, the second determining module 903 is specifically configured to:

[0161] Inputting the first scene, the first dialogue information, the second dialogue information, and the scene switching information into a neural network model to obtain a first music feature;

[0162] Obtaining a second music feature corresponding to each preset music, and determining a similarity between the first music feature and each second music feature;

[0163] The recommended music is determined according to multiple similarities.

[0164] According to one or more embodiments of the present disclosure, the first determining module 901 is specifically configured to:

[0165] Obtain at least one device information;

[0166] The first scenario is determined according to the device information.

[0167] According to one or more embodiments of the present disclosure, the playback module 904 is specifically configured to:

[0168] Determining to play the opening voice before the recommended music;

[0169] Play the opening voice and the recommended music.

[0170] According to one or more embodiments of the present disclosure, the playback module 904 is specifically configured to:

[0171] Acquiring music information related to the recommended music, and generating a recommendation text describing the recommended music based on the music information;

[0172] Determine a recommended voice corresponding to the recommended text, and play the recommended voice when playing the recommended music.

[0173] The audio playback device provided in the embodiment of the present disclosure can be used to implement the technical solution of the above-mentioned method embodiment. Its implementation principle and technical effects are similar and will not be described in detail in this embodiment.

[0174] FIG10 is a schematic diagram of the structure of a terminal device provided by an embodiment of the present disclosure. Please refer to FIG10 , which shows a schematic diagram of the structure of a terminal device 1000 suitable for implementing an embodiment of the present disclosure. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., as well as fixed terminals such as digital TVs, desktop computers, etc. The terminal device shown in FIG10 is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present disclosure.

[0175] As shown in FIG10 , terminal device 1000 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 1001, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 1002 or programs loaded from a storage device 1008 into a random access memory (RAM) 1003. Various programs and data required for the operation of terminal device 1000 are also stored in RAM 1003. Processing device 1001, ROM 1002, and RAM 1003 are connected to each other via a bus 1004. An input / output (I / O) interface 1005 is also connected to bus 1004.

[0176] Typically, the following devices may be connected to the I / O interface 1005: an input device 1006 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 1007 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 1008 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the terminal device 1000 to communicate with other devices wirelessly or by wire to exchange data. Although FIG10 illustrates a terminal device 1000 having various devices, it should be understood that not all of the devices shown are required to be implemented or present. More or fewer devices may alternatively be implemented or present.

[0177] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 1009, or installed from the storage device 1008, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0178] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0179] The computer-readable medium may be included in the terminal device, or may exist independently without being incorporated into the terminal device.

[0180] The computer-readable medium carries one or more programs. When the one or more programs are executed by the terminal device, the terminal device executes the method shown in the above embodiment.

[0181] An embodiment of the present disclosure provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, various methods that may be involved in the above embodiments are implemented.

[0182] An embodiment of the present disclosure provides a computer program product, including a computer program, which implements various possible methods involved in the above embodiments when executed by a processor.

[0183] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).

[0184] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0185] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."

[0186] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0187] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0188] It should be noted that the modifications of "one" and "multiple" mentioned in the present disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise clearly indicated in the context, they should be understood as "one or more".

[0189] The names of the messages or information exchanged between multiple devices in the embodiments of the present disclosure are only used for illustrative purposes and are not used to limit the scope of these messages or information.

[0190] It is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws and regulations. Data may include information, parameters and messages, such as flow switching indication information.

[0191] In a first aspect, an embodiment of the present disclosure provides an audio playback method, the audio playback method comprising:

[0192] In response to a component switching operation, determining a first scene corresponding to a currently selected first component, wherein the component is used for text conversation, and functions of the text conversations between the multiple components are different;

[0193] Obtaining first conversation information associated with the first component;

[0194] According to the first scene and the first dialogue information, recommended music is determined from a plurality of preset music, and the recommended music is played.

[0195] According to one or more embodiments of the present disclosure, determining recommended music from a plurality of preset music pieces based on the first scene and the first conversation information includes:

[0196] Acquiring second conversation information within a historical period, where the second conversation information includes music-related information;

[0197] The recommended music is determined from the plurality of preset music based on the first scene, the first dialogue information, and the second dialogue information.

[0198] According to one or more embodiments of the present disclosure, determining the recommended music from the plurality of preset music based on the first scene, the first conversation information, and the second conversation information includes:

[0199] Determining scene switching information associated with the switching operation, where the scene switching information is used to indicate the scene associated with the switching operation;

[0200] The recommended music is determined from the plurality of preset music according to the first scene, the first dialogue information, the second dialogue information and the scene switching information.

[0201] According to one or more embodiments of the present disclosure, the scene switching information includes at least one of the following:

[0202] a second scene corresponding to a second component, the second component being the component selected before the switching operation;

[0203] an identifier of the first component and an identifier of the second component;

[0204] The text information corresponding to the switching operation.

[0205] According to one or more embodiments of the present disclosure, determining the recommended music from the plurality of preset music based on the first scene, the first dialogue information, the second dialogue information, and the scene switching information includes:

[0206] Inputting the first scene, the first dialogue information, the second dialogue information, and the scene switching information into a neural network model to obtain a first music feature;

[0207] Obtaining a second music feature corresponding to each preset music, and determining a similarity between the first music feature and each second music feature;

[0208] The recommended music is determined according to multiple similarities.

[0209] According to one or more embodiments of the present disclosure, determining the first scene corresponding to the currently selected first component includes:

[0210] Obtain at least one device information;

[0211] The first scenario is determined according to the device information.

[0212] According to one or more embodiments of the present disclosure, playing the recommended music includes:

[0213] Determining to play the opening voice before the recommended music;

[0214] Play the opening voice and the recommended music.

[0215] According to one or more embodiments of the present disclosure, playing the recommended music includes:

[0216] Acquiring music information related to the recommended music, and generating a recommendation text describing the recommended music based on the music information;

[0217] Determine a recommended voice corresponding to the recommended text, and play the recommended voice when playing the recommended music.

[0218] In a second aspect, an embodiment of the present disclosure provides an audio playback device, the audio playback device including a first determination module, an acquisition module, a second determination module, and a playback module, wherein:

[0219] The first determining module is configured to, in response to a component switching operation, determine a first scenario corresponding to a currently selected first component, wherein the component is used for text conversation, and the functions of the text conversations between the multiple components are different;

[0220] The acquisition module is used to acquire first conversation information associated with the first component;

[0221] The second determining module is configured to determine recommended music from a plurality of preset music pieces according to the first scene and the first conversation information;

[0222] The playing module is used to play the recommended music.

[0223] According to one or more embodiments of the present disclosure, the second determining module is specifically configured to:

[0224] Acquiring second conversation information within a historical period, where the second conversation information includes music-related information;

[0225] The recommended music is determined from the plurality of preset music based on the first scene, the first dialogue information, and the second dialogue information.

[0226] According to one or more embodiments of the present disclosure, the second determining module is specifically configured to:

[0227] Determining scene switching information associated with the switching operation, where the scene switching information is used to indicate the scene associated with the switching operation;

[0228] The recommended music is determined from the plurality of preset music according to the first scene, the first dialogue information, the second dialogue information and the scene switching information.

[0229] According to one or more embodiments of the present disclosure, the scene switching information includes at least one of the following:

[0230] a second scene corresponding to a second component, the second component being the component selected before the switching operation;

[0231] an identifier of the first component and an identifier of the second component;

[0232] The text information corresponding to the switching operation.

[0233] According to one or more embodiments of the present disclosure, the second determining module is specifically configured to:

[0234] Inputting the first scene, the first dialogue information, the second dialogue information, and the scene switching information into a neural network model to obtain a first music feature;

[0235] Obtaining a second music feature corresponding to each preset music, and determining a similarity between the first music feature and each second music feature;

[0236] The recommended music is determined according to multiple similarities.

[0237] According to one or more embodiments of the present disclosure, the first determining module is specifically configured to:

[0238] Obtain at least one device information;

[0239] The first scenario is determined according to the device information.

[0240] According to one or more embodiments of the present disclosure, the playback module is specifically configured to:

[0241] Determining to play the opening voice before the recommended music;

[0242] Play the opening voice and the recommended music.

[0243] According to one or more embodiments of the present disclosure, the playback module is specifically configured to:

[0244] Acquiring music information related to the recommended music, and generating a recommendation text describing the recommended music based on the music information;

[0245] Determine a recommended voice corresponding to the recommended text, and play the recommended voice when playing the recommended music.

[0246] In a third aspect, an embodiment of the present disclosure provides a terminal device including: a processor and a memory;

[0247] The memory stores computer-executable instructions;

[0248] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the first aspect as described above and various methods that may be involved in the first aspect.

[0249] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the first aspect and various possible methods involved in the first aspect are implemented.

[0250] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0251] In addition, although the operations are described in a specific order, this should not be understood as requiring that these operations be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of this disclosure. Certain features described in the context of separate embodiments can also be implemented in a single embodiment in combination. Conversely, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination. Although the subject matter has been described in language specific to structural features and / or method logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims.

Claims

1. An audio playback method, comprising: In response to a component switching operation, determining a first scene corresponding to a currently selected first component, wherein the component is used for text conversation, and functions of the text conversations between the multiple components are different; Obtaining first conversation information associated with the first component; According to the first scene and the first dialogue information, recommended music is determined from a plurality of preset music, and the recommended music is played.

2. The method according to claim 1, wherein The step of determining recommended music from a plurality of preset music pieces according to the first scene and the first conversation information includes: Acquiring second conversation information within a historical period, where the second conversation information includes music-related information; The recommended music is determined from the plurality of preset music based on the first scene, the first dialogue information, and the second dialogue information.

3. The method according to claim 2, wherein: The determining the recommended music from the plurality of preset music based on the first scene, the first conversation information, and the second conversation information includes: Determining scene switching information associated with the switching operation, where the scene switching information is used to indicate the scene associated with the switching operation; The recommended music is determined from the plurality of preset music according to the first scene, the first dialogue information, the second dialogue information and the scene switching information.

4. The method according to claim 3, wherein: The scene switching information includes at least one of the following: a second scene corresponding to a second component, the second component being the component selected before the switching operation; an identifier of the first component and an identifier of the second component; The text information corresponding to the switching operation.

5. The method according to claim 3 or 4, wherein: The step of determining the recommended music from the plurality of preset music pieces according to the first scene, the first dialogue information, the second dialogue information, and the scene switching information includes: Inputting the first scene, the first dialogue information, the second dialogue information, and the scene switching information into a neural network model to obtain a first music feature; Obtaining a second music feature corresponding to each preset music, and determining a similarity between the first music feature and each second music feature; The recommended music is determined according to multiple similarities.

6. The method according to any one of claims 1 to 5, wherein: The determining the first scene corresponding to the currently selected first component includes: Obtain at least one device information; The first scenario is determined according to the device information.

7. The method according to any one of claims 1 to 6, wherein: The playing of the recommended music includes: Determining to play the opening voice before the recommended music; Play the opening voice and the recommended music.

8. The method according to claim 7, wherein: The playing of the recommended music includes: Acquiring music information related to the recommended music, and generating a recommendation text describing the recommended music based on the music information; Determine a recommended voice corresponding to the recommended text, and play the recommended voice when playing the recommended music.

9. An audio playback device, comprising a first determination module, an acquisition module, a second determination module, and a playback module, wherein: The first determining module is configured to determine, in response to a component switching operation, a first scenario corresponding to a currently selected first component, wherein the component is used for text conversation, and the functions of the text conversations between the multiple components are different; The acquisition module is configured to acquire first conversation information associated with the first component; The second determining module is configured to determine recommended music from a plurality of preset music according to the first scene and the first conversation information; The playing module is configured to play the recommended music.

10. A terminal device comprising: processor and memory, wherein The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor executes the audio playback method according to any one of claims 1 to 8.

11. A computer-readable storage medium storing computer-executable instructions, wherein: When the processor executes the computer-executable instruction, the audio playback method according to any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Music recommendation method and device and mobile terminal

    CN103605656A

  • Music recommendation method and apparatus

    CN105930429A

  • Instant messaging message sending method and electronic device

    CN110830368A

  • Knowledge graph-based recommendation method and device

    CN113449176A

  • Music recommending method, device, terminal, and storage medium

    US20200151212A1