Media resource playing method and related apparatus

Sending audio information requests and playing media resources through the calling terminal device solves the problem that users cannot interact in real time, achieving a more friendly user experience and higher fun.

WO2025140028A1PCT designated stage expired Publication Date: 2025-07-03HUAWEI TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/140932
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-29
Filing Date
2024-12-20
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In the prior art, the media resources played by the calling terminal device are configured by the media platform server, and users cannot interact in real time, resulting in a decrease in the fun and operability of the user experience.

Method used

The calling terminal device sends audio information to request to play media resources and receives and plays media resources determined by audio information, so as to enable users to flexibly control the playback of media resources through voice commands, including real-time feedback interaction of audio, pictures, animations and text information.

Benefits of technology

It improves users' interest and operability for media resources and provides a more friendly user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024140932_03072025_PF_FP_ABST
    Figure CN2024140932_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the present application are a media resource playing method and a related apparatus. The method comprises: sending a call request to a called terminal device; sending first audio information, wherein the first audio information is used for requesting the play of a media resource; and receiving a first media resource and playing same, wherein the first media resource is determined by means of the first audio information. After sending a call request to a called terminal device, a calling terminal device requests the play of a media resource by means of sending first audio information, and the calling terminal device then receives a first media resource and plays same. Therefore, by means of a voice instruction, a user can flexibly control the media resource played by the terminal device, thereby bringing more friendly and more interesting user experience to the user.
Need to check novelty before this filing date? Find Prior Art

Description

A media resource playing method and related device

[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office on December 29, 2023, with application number CN202311864928.X and invention name “A media resource playback method and related device”, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] The present application relates to the field of communication technology, and in particular to a media resource playback method and related devices. Background Art

[0003] With the continuous development of communication technology, high-definition voice over long term evolution (VOLTE) technology has gradually become part of people's lives, allowing them to enjoy a variety of media resources. For example, users can enjoy video experiences such as video ringback tone and video customer service while making a voice call, making the waiting period before the call more interesting and greatly improving the user's calling experience.

[0004] At present, the media resource playback method is usually based on the media resource playback of the communication technology (CT) domain, which can also be called the telecommunications domain. The corresponding process can be: the calling terminal device initiates a call to the called terminal device. When the called terminal device rings, the media server in the CT domain pulls the media resources corresponding to the user contract information from the media platform server based on the user contract information, and then, after receiving the confirmation playback message from the calling terminal device, instructs the media platform server to start playing the media resources for the calling terminal device.

[0005] Since the media resources played by the calling terminal device are configured by the media platform server, the user of the calling terminal device cannot interact in real time, thereby reducing the interest and operability of the user in using the media resources. Summary of the Invention

[0006] In a first aspect, an embodiment of the present application proposes a method for playing media resources, which is applied to a calling terminal device and includes: sending a call request to a called terminal device; sending a first audio message, which is used to request playing a media resource; receiving and playing a first media resource, which is determined by the first audio message.

[0007] In this embodiment of the present application, after sending a call request to the called terminal, the calling terminal device sends a first audio message requesting the playback of a media resource. The calling terminal then receives and plays the first media resource. This allows the user to flexibly control the media resource played by the terminal device through voice commands, providing a more user-friendly and engaging user experience.

[0008] In combination with the first aspect, in a possible implementation of the first aspect, the method further includes: playing a first response message, the first response message being a response message generated based on the first audio information, the first response message including audio information, picture information, animation information and / or text information.

[0009] In this embodiment of the present application, the screen (or speaker) of the calling terminal device can overlay (or play) the response information of the audio information, which includes but is not limited to: audio information, image information, animation information and / or text information. Through the above method, real-time feedback of interactive information is achieved, thereby improving the user experience.

[0010] In conjunction with the first aspect, in one possible implementation of the first aspect, the first response information includes: first text information, where the first text information is text information generated by performing speech recognition processing based on the first audio information, and the content included in the first text information corresponds to the first audio information. By feeding back the first text information corresponding to the first audio information to the calling terminal device, user operation is facilitated and the user experience is improved.

[0011] In combination with the first aspect, in a possible implementation of the first aspect, the method further includes: sending second audio information;

[0012] Receive a response to the second audio information, the response to the second audio information including: a second media resource, the second media resource is determined by the second audio information; and / or, a second response information, the second response information is a response information generated based on the second audio information, the second response information includes audio information, picture information, animation information and / or text information.

[0013] In an embodiment of the present application, users can also continue to operate media resources through audio information to further enhance the user experience.

[0014] In combination with the first aspect, in a possible implementation manner of the first aspect, the method further includes: stopping playing the first media resource; playing the second media resource; and / or playing the second response information.

[0015] In this embodiment of the present application, after the calling terminal device receives the second media resource, it can stop playing the first media resource and then play the second media resource. When the calling terminal device receives the second response information, the calling terminal device can also superimpose (or play) the second response information on the screen (or speaker), thereby improving the user experience.

[0016] In combination with the first aspect, in a possible implementation of the first aspect, the method further includes: receiving an off-hook message sent by the called terminal device; in response to the off-hook message, stopping playing the first media resource; stopping receiving audio information, and / or stopping voice recognition processing of the audio information.

[0017] In this embodiment of the present application, the calling terminal device triggers a request to play a media resource through audio information during the call-making phase. The media resource can be a video ringtone, which enhances the user experience. Furthermore, the calling terminal device can also trigger other interactive operations related to the media resource through audio information during the call-making phase. These other interactive operations include, but are not limited to, ordering and copying the media resource, or rating or liking the media resource, thereby enhancing the user experience.

[0018] In combination with the first aspect, in a possible implementation of the first aspect, sending the first audio information includes: sending the first audio information, the first audio information includes a wake-up keyword, and the wake-up keyword is used to trigger the media platform server to perform voice recognition processing on the first audio information; or, sending audio information including the wake-up keyword; sending the first audio information.

[0019] In an embodiment of the present application, the user can also input the wake-up keyword by voice to avoid misoperation.

[0020] In the second aspect, an embodiment of the present application proposes a media resource playback method, which is applied to a media platform server, and the method includes: receiving a call request sent by the calling terminal device; receiving a first audio information sent by the calling terminal device; determining a first media resource based on the first audio information; and sending the first media resource to the calling terminal device.

[0021] In this embodiment of the present application, after receiving a call request from a calling terminal device, the media platform server determines a corresponding first media resource based on the first audio information sent by the host terminal device, and then sends the first media resource to the calling terminal device. This allows users to flexibly control the media resources played by the terminal device through voice commands, providing a more user-friendly and engaging user experience.

[0022] In combination with the second aspect, in a possible implementation of the second aspect, determining the first media resource based on the first audio information includes: performing speech recognition processing based on the first audio information to generate first text information, and the content included in the first text information corresponds to the first audio information; performing semantic understanding processing based on the first text information to generate first user intent information; and determining the first media resource based on the first user intent information.

[0023] In the embodiment of the present application, the first user intention information corresponding to the first audio information is determined through voice recognition and semantic understanding, and then the first media resource is determined, thereby improving recognition accuracy.

[0024] In combination with the second aspect, in a possible implementation of the second aspect, the method also includes: generating a first response message based on the first user intention information, the first response message including audio information, picture information, animation information and / or text information; and sending the first response message to the calling terminal device.

[0025] In this embodiment of the present application, corresponding response information can also be generated based on the first user's intention information and sent to the calling terminal device. The screen (or speaker) of the calling terminal device can overlay (or play) the response information of the audio information. The response information includes but is not limited to: audio information, image information, animation information and / or text information. Through the above method, real-time feedback of interactive information is achieved, improving the user experience.

[0026] In combination with the second aspect, in a possible implementation manner of the second aspect, the first response information includes the first text information.

[0027] In combination with the second aspect, in a possible implementation of the second aspect, the first user intention information includes any one or more of the following: starting to play media resources, pausing to play media resources, switching to play media resources, switching back to play media resources, copying and ordering media resources, content feature keywords of the first audio information, or the weight of the content feature keywords.

[0028] In combination with the second aspect, in a possible implementation of the second aspect, determining the first media resource based on the first user intention information includes: determining the first media resource based on a decision recommendation model and the content feature keywords and / or the weight of the content feature keywords included in the first user intention information, wherein the decision recommendation model applies a parameter set to determine the media resource, and the parameter set includes any one or more of the following: the content feature keywords of the media resource, the weight of the content feature keywords of the media resource, the media resource tag of the media resource library, the weight of the media resource tag of the media resource library, the popularity weight of the media resource of the media resource library, the media resource release time of the media resource library, or the media resource playback rate of the media resource library, wherein the media resource library includes one or more media resources.

[0029] In combination with the second aspect, in a possible implementation of the second aspect, the method further includes: detecting whether the first audio information includes a wake-up keyword; if the first audio information includes the wake-up keyword, triggering voice recognition processing based on the first audio information; or, based on detecting that the received audio information includes the wake-up keyword, triggering voice recognition processing based on the first audio information.

[0030] In an embodiment of the present application, the user can also input the wake-up keyword by voice. The media platform server detects the wake-up keyword and initiates voice recognition processing of the first audio information after the user inputs the wake-up keyword to avoid misoperation.

[0031] In combination with the second aspect, in a possible implementation of the second aspect, the method also includes: receiving a second audio information sent by the calling terminal device; generating a response to the second audio information based on the second audio information, the response to the second audio information including: a second media resource, the second media resource is determined by the second audio information; and / or, a second response information, the second response information is a response information generated based on the second audio information, the second response information includes audio information, picture information, animation information and / or text information; sending a response to the second audio information to the calling terminal device.

[0032] In combination with the second aspect, in a possible implementation of the second aspect, a response to the second audio information is generated based on the second audio information, including: performing speech recognition processing based on the second audio information to generate second text information, wherein the content included in the second text information corresponds to the first audio information; performing semantic understanding processing based on the second text information to generate second user intent information; and generating a response to the second audio information based on the second user intent information.

[0033] In an embodiment of the present application, users can also continue to operate media resources through audio information to further enhance the user experience.

[0034] A third aspect of an embodiment of the present application provides an electronic device, comprising: a transceiver unit and a processing unit, so that the electronic device implements the method of the above-mentioned first aspect or any possible implementation of the first aspect, or enables the electronic device to implement the method of the above-mentioned second aspect or any possible implementation of the second aspect.

[0035] A fourth aspect of an embodiment of the present application provides an electronic device, comprising: a processor, the processor being coupled to a memory, the memory being used to store programs or instructions, and when the programs or instructions are executed by the processor, the electronic device implements the method of the above-mentioned first aspect or any possible implementation of the first aspect, or implements the above-mentioned second aspect or any possible implementation of the second aspect.

[0036] A fifth aspect of an embodiment of the present application provides a computer-readable medium having a computer program or instruction stored thereon. When the computer program or instruction runs on a computer, the computer executes the method of the aforementioned first aspect or any possible implementation of the first aspect, or the computer executes the method of the aforementioned second aspect or any possible implementation of the second aspect.

[0037] A sixth aspect of an embodiment of the present application provides a computer program product, which, when executed on a computer, enables the computer to execute the method of the aforementioned first aspect or any possible implementation of the first aspect, or enables the computer to execute the method of the aforementioned second aspect or any possible implementation of the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] FIG1 is a schematic diagram of a communication scenario proposed in an embodiment of the present application;

[0039] FIG2 is a schematic diagram of a flow chart of a method for playing media resources according to an embodiment of the present application;

[0040] FIG3 is a schematic diagram of another communication scenario in an embodiment of the present application;

[0041] FIG4 is a schematic diagram of a flow chart of a method for playing media resources according to an embodiment of the present application;

[0042] FIG5 is a schematic diagram of another communication scenario in an embodiment of the present application;

[0043] FIG6 is a schematic diagram of a flow chart of a method for playing media resources according to an embodiment of the present application;

[0044] FIG7 is a schematic diagram of another communication scenario in an embodiment of the present application;

[0045] FIG8 is a schematic diagram of a flow chart of a method for playing media resources according to an embodiment of the present application;

[0046] FIG9 is a schematic diagram of another communication scenario in an embodiment of the present application;

[0047] FIG10 is a schematic diagram of a flow chart of a method for playing media resources according to an embodiment of the present application;

[0048] Figures 11 to 18 are schematic diagrams of a scenario in an embodiment of the present application;

[0049] FIG19 is a schematic diagram of an application scenario in an embodiment of the present application;

[0050] FIG20 is a schematic structural diagram of an electronic device provided in an embodiment of the present application;

[0051] FIG21 is a schematic diagram of another structure of an electronic device according to an embodiment of the present application;

[0052] FIG22 is another structural diagram of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION

[0053] The terms "first", "second", etc. in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the terms used in this way can be interchanged under appropriate circumstances. This is merely a way of distinguishing when describing objects with the same properties in the embodiments of the present application. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, so that a process, method, system, product or apparatus that includes a series of units is not necessarily limited to those units, but may include other units not expressly listed or inherent to these processes, methods, products or apparatuses.

[0054] The technical solutions in the embodiments of the present application will be clearly described below in conjunction with the drawings in the embodiments of the present application. In the description of the present application, unless otherwise specified, " / " means or, for example, A / B can mean A or B; "and / or" in the present application is merely a description of the association relationship of associated objects, indicating that three relationships can exist, for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, in the description of the present application, "at least one" refers to one or more items, and "multiple items" refers to two or more items. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.

[0055] First, some of the terms used in the embodiments of the present application are explained to facilitate understanding by those skilled in the art.

[0056] (1) Terminal device: It can be a wireless terminal device that can receive network device scheduling and instruction information. The wireless terminal device can be a device that provides voice and / or data connectivity to the user, or a handheld device with wireless connection function, or other processing device connected to a wireless modem.

[0057] Terminal devices can communicate with one or more core networks or the Internet via a radio access network (RAN). Terminal devices can be mobile terminal devices, such as mobile phones (also known as "cellular" phones, mobile phones), computers, and data cards. For example, they can be portable, pocket-sized, handheld, computer-built-in, or vehicle-mounted mobile devices that exchange voice and / or data with the radio access network. Examples include personal communication service (PCS) phones, cordless phones, Session Initiation Protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), tablet computers, and computers with wireless transceiver capabilities. A wireless terminal device may also be referred to as a system, a subscriber unit, a subscriber station, a mobile station, a mobile station (MS), a remote station, an access point (AP), a remote terminal, an access terminal, a user terminal, a user agent, a subscriber station (SS), a customer premises equipment (CPE), a terminal, a user equipment (UE), a mobile terminal (MT), an unmanned aerial vehicle (UAV), etc. A terminal device may also be a wearable device or a next-generation communication system, for example, a terminal device in a 5G communication system or a terminal device in a future-evolved public land mobile network (PLMN).

[0058] (2) Internet protocol multimedia subsystem (IMS) domain.

[0059] The CT domain implements communication through the evolved packet core (EPC) and the Internet Protocol Multimedia Subsystem (IMS) domain core network. The IMS domain core network includes several application servers (ASs), such as a media platform server. The media platform server is used to provide media resource playback for terminals. For example, when providing video ringback tone services, the media platform server is also called a video ringback tone platform. The media platform server may include a media resource application server and a media resource subsystem (MRS). The media resource application server and the media resource server may be co-located or physically separate. The media resource server may also be called a ringback tone platform, a video ringback tone platform, or a ringback tone platform. The media resource server is used to provide media resources such as video ringback tone, video ringback tone, video advertising, and video customer service. For example, the media resource server creates and manages these media resources. The media application server and the media resource server may be co-located or physically separate. The media application server processes Session Initiation Protocol (SIP) signaling messages, and the media resource server provides audio streams and / or video streams to the calling terminal and / or the called terminal.

[0060] In addition, the IMS core network also includes: serving-call session control function (S-CSCF), interrogating-call session control function (I-CSCF), proxy-call session control function (P-CSCF), home subscriber server (HSS), session border controller (SBC), and several application servers, such as telephony application server (TAS), multimedia telephony application server (MMTelAS), and service continuity application server (SCCAS). The I-CSCF and S-CSCF can be combined and referred to as "I / S-CSCF". The SBC and P-CSCF can be combined and referred to as "SBC / P-CSCF". The EPC may include a packet data network gateway (PGW) device, a serving gateway (SGW) device, and a mobility management entity (MME) device.

[0061] S / P-GW devices are used to provide the functions of service gateway and packet data network gateway logical entities. SGW is the anchor point for local mobility, mainly facing the wireless access network to transmit service plane data. P-GW is the EPS anchor point, mainly facing other data networks to achieve access and interaction with multiple public data networks. SGW devices can be used to connect the IMS core network with wireless networks, and PGW devices can be used to connect the IMS core network with the Internet Protocol (IP) network. MME devices are the core devices of the EPC network and are used to provide the functions of the MME logical entity.

[0062] The above-mentioned network devices are all corresponding network devices in the wireless communication network in the prior art. They will not be described in detail here, but only briefly explained. For example: HSS devices can be used to store user subscription information and location information. SBC devices can provide secure access and media processing. MMTelAS devices provide basic multimedia telephone services and supplementary services. MME devices are the core devices of the EPC network. SGW devices can be used to connect the IMS core network with the wireless network, and PGW devices can be used to connect the IMS core network with the IP network. S-CSCF devices can be used for user registration, authentication control, session routing and service triggering control, and maintain session status information. I-CSCF devices can be used for the allocation and query of S-CSCF devices registered by users. P-CSCF devices can be used for signaling and message agents. In this application, in order to keep the description concise, CSCF devices are used to represent any one or more combinations of S-CSCF devices, I-CSCF devices, and P-CSCF devices.

[0063] (3) Media resources.

[0064] The media resources in the embodiments of the present application include, but are not limited to: audio ringback tone, video ringback tone, video advertisement, or video animation, etc.

[0065] Since the media resources currently played by the calling terminal device are configured by the media platform server, the user of the calling terminal device cannot interact in real time, thereby reducing the interest and operability of the user in using the media resources.

[0066] Based on this, the present application proposes a method for playing media resources, which includes: sending a call request to a called terminal device; sending a first audio message, where the first audio message is used to request the playback of a media resource; and receiving and playing a first media resource, where the first media resource is determined by the first audio message. After sending the call request to the called terminal device, the calling terminal device requests the playback of the media resource by sending the first audio message, and then the calling terminal device receives and plays the first media resource. This allows users to flexibly control the media resources played by the terminal device through voice commands, providing users with a more user-friendly and interesting user experience.

[0067] The following describes an embodiment of the present application in conjunction with the accompanying drawings. Please refer to Figure 1, which is a schematic diagram of a communication scenario proposed in an embodiment of the present application. A communication scenario proposed in an embodiment of the present application includes: a media platform server, a calling terminal device, and a called terminal device. The calling terminal device sends a call request to the called terminal device. The media platform server then sends a default media resource to the calling terminal device, and the calling terminal device plays the default media resource. The default media resource can be pre-ordered by the calling terminal device, or pre-allocated to the calling terminal device by the media platform server. When the calling terminal device plays the default media resource, the calling terminal device can collect the user's audio information, and then send the audio information to the media platform server. The media platform server interacts according to the audio information reported by the calling terminal device, for example, switching the playback media resource for the calling terminal device according to the audio information, or pausing the playback of the calling terminal device's media resource, or finalizing the media resource currently played by the calling terminal device.

[0068] Based on the communication scenario shown in FIG1 , please refer to FIG2 , which is a flow chart of a method for playing media resources according to an embodiment of the present application. A method for playing media resources according to an embodiment of the present application includes:

[0069] S1. The calling terminal device sends a call request to the called terminal device.

[0070] The calling terminal sends a call request to the called terminal, and the two terminals negotiate a call. After the called terminal rings, the media platform server negotiates media resources with the calling terminal. During the negotiation, the media resource transmission direction is set to allow both uplink and downlink transmission.

[0071] S2. The calling terminal device sends first audio information to the media platform server.

[0072] After the negotiation is complete, the media platform server sends a default media resource to the calling terminal device. This default media resource can be a media resource that the calling terminal device has subscribed to from the media platform server, or a media resource that the called terminal device has subscribed to from the media platform server, or a media resource that the media platform server proactively allocates to the calling terminal device or the called terminal device. The calling terminal device then plays the default media resource.

[0073] During the above process, the media platform server activates the voice interaction service. This voice interaction service allows users to send voice commands (e.g., via audio information carrying the voice message) to the media platform server through their terminal devices, and then complete the interaction based on the voice commands. In one possible implementation, the calling terminal device sends a first audio message to the media platform server, which is used to request the playback of a media resource.

[0074] In one example, the calling terminal device requests to switch the playback media resource through the first audio information. The first audio information can be a request to randomly switch to play the next media resource, or the first audio information can be a clear indication of which type of media resource to switch. Please refer to Figure 11, which is a schematic diagram of a scenario in an embodiment of the present application. In one example, the plane of the calling terminal device displays the first text information corresponding to the first audio information. The first text information is the text information obtained after the media platform server performs voice recognition based on the first audio information. The first text information corresponding to the first audio information includes: "Xiao Cai Xiao Cai, change one", and the first audio information is used to request to replace the currently playing default media resource. In another example, please refer to Figure 12, which is another schematic diagram of a scenario in an embodiment of the present application. The first text information corresponding to the first audio information includes: "Xiao Cai Xiao Cai, change one", and the first audio information is used to request to replace the currently playing default media resource and play media resources related to animals.

[0075] In another example, please refer to Figure 13, which is another scenario diagram in an embodiment of the present application. The first text information corresponding to the first audio information includes: "Xiao Cai Xiao Cai". The first audio information is used to trigger the media platform server to perform voice recognition processing on the first audio information. "Xiao Cai Xiao Cai" is used as the wake-up keyword to wake up the voice interaction service.

[0076] In another possible implementation, the calling terminal device sends a first audio message to the media platform server, and the first audio message is used to interactively process the media resource currently being played by the calling terminal device. The interactive processing includes but is not limited to: stopping the playback of the currently playing media resource, copying and subscribing to the currently playing media resource, or sharing the currently playing media resource with other users. In another example, please refer to Figure 14, which is a schematic diagram of another scenario in an embodiment of the present application. The first text message corresponding to the first audio message includes: "Xiao Cai Xiao Cai, copy this ringtone for me", and the first audio message is used to trigger the media platform server to copy and subscribe the media resource currently being played by the calling terminal device for the calling terminal device.

[0077] In another possible implementation, when the media platform server is unable to identify and process the first audio information, the media platform server may send a prompt message to the calling terminal device, and the prompt message indicates that the media platform server is unable to identify and process the first audio information. For example, please refer to Figure 15, which is another scenario diagram in an embodiment of the present application. The first audio information includes: "Xiao Cai Xiao Cai, it's raining today." After the media platform server is unable to identify and process the first audio information, it sends a prompt message to the calling terminal device. The screen of the calling terminal device displays the prompt message, and the prompt message includes "Sorry, I didn't understand. What do you want to instruct Xiao Cai to do?". Subsequently, the calling terminal device can continue to collect the user's audio information, such as the second audio information, and then the calling terminal device sends the second audio information to the media platform server, and the media platform server performs voice interaction processing based on the second audio information.

[0078] S3. The media platform server detects whether the first audio information includes a wake-up keyword.

[0079] The media platform server detects whether the user voice (audio information) of the calling terminal device contains the wake-up keyword to prevent the media platform server from mistakenly triggering voice interaction processing of the user voice. When the media platform server detects that the user voice (audio information) of the calling terminal device contains the wake-up keyword, it performs subsequent processing on the first audio information, such as voice recognition and semantic understanding.

[0080] In one possible implementation, after the media platform server receives the first audio information, the media platform server detects whether the first audio information includes a wake-up keyword, which is used to trigger the media platform server to perform voice recognition processing on the first audio information. If the first audio information includes a wake-up keyword, step S4 is entered. Exemplarily, the media platform server uses a technical method such as keyword spotting (KWS) to identify whether the first audio information includes a wake-up keyword, which can also be called a wake-up word.

[0081] In another possible implementation, after step S1, the media platform server collects audio information from the calling terminal device and identifies whether the audio information includes the wake-up keyword. If the wake-up keyword is included, the first audio information is collected and step S4 is executed.

[0082] In another possible implementation, the media platform server may also send voice interaction prompt information to the calling terminal device, and the calling terminal device displays the voice interaction prompt information on the screen of the terminal device in a real-time overlay manner. For example, please refer to Figure 16, which is another scenario diagram in an embodiment of the present application. The plane overlay of the calling terminal device displays the voice interaction prompt information, and the voice interaction prompt information includes: "Hi, hello. I am the intelligent voice assistant "Xiao Cai". You can say to me "Xiao Cai Xiao Cai, change one"". The calling terminal device continues to receive audio information input by the user.

[0083] If the media platform server does not detect that the first audio information includes the wake-up keyword, or the media platform server does not detect that the audio information of the calling terminal device includes the wake-up keyword, it is determined that the audio information currently input by the calling terminal device is not for voice interaction processing related to media resources, and the media platform server does not execute

[0084] S4. The media platform server performs speech recognition processing based on the first audio information to generate first text information.

[0085] After receiving the first audio information, the media platform server performs speech recognition processing on the first audio information to generate the first text information. For example, the media platform server recognizes the first audio information (i.e., the user voice of the calling terminal device) based on automatic speech recognition (ASR) technology and converts the first audio information into the first text information.

[0086] Optionally, the media platform server may send the first text message to the calling terminal device, and the screen of the calling terminal device may overlay and display the first text message. The first text message is shown in, for example, Figures 11 to 15 above.

[0087] S5. The media platform server performs semantic understanding processing based on the first text information to generate first user intention information.

[0088] After receiving the first text message, the media platform server performs semantic understanding processing on the first text message to generate first user intent information, where the first user intent information includes any one or more of the following: starting playback of a media resource, pausing playback of a media resource, switching playback of a media resource, rewinding playback of a media resource, copying and ordering a media resource, content feature keywords of the first audio message, or weights of the content feature keywords. For example, an artificial intelligence (AI) model can be used for semantic understanding.

[0089] After step S5, steps S6 and S9 are executed.

[0090] S6. The media platform server determines a first media resource according to the first user intention information.

[0091] The media platform server determines a corresponding first media resource based on the first user intent information. In one possible implementation, the first user intent information requests switching to a playback media resource, and the media platform server determines a media resource to be played from one or more candidate media resources as the first media resource, which may include one or more media resources.

[0092] Specifically, the media platform server determines, based on the content feature keywords and / or the weights of the content feature keywords included in the first user intent information, a media resource that meets the content feature keywords and / or the weights of the content feature keywords as the first media resource. In conjunction with the example of FIG12 , the content feature keyword included in the first user intent information is "animal." Based on the content feature keyword, the media platform server determines, from one or more candidate media resources, a media resource that includes images of animals as the first media resource.

[0093] Exemplarily, the media platform server determines the first media resource based on a decision recommendation model (or a decision recommendation algorithm, or a decision algorithm) and the content feature keywords and / or the weight of the content feature keywords included in the first user intent information, wherein the decision recommendation model applies a parameter set to determine the media resource, and the parameter set includes any one or more of the following: the content feature keywords of the media resource, the weight of the content feature keywords of the media resource, the media resource tag of the media resource library, the weight of the media resource tag of the media resource library, the popularity weight of the media resource of the media resource library, the media resource release time of the media resource library, or the media resource playback rate of the media resource library, wherein the media resource library includes one or more media resources.

[0094] Optionally, if the first user intention information determined by the media platform server requests switching to play media resources and the first user intention information does not include content feature keywords, the media platform server can randomly select a media resource for the calling terminal device as the first media resource based on the first user intention information.

[0095] S7. The media platform server sends the first media resource to the calling terminal device.

[0096] S8. The calling terminal device plays the first media resource.

[0097] After receiving the first media resource, the calling terminal device stops playing the default media resource and then plays the first media resource.

[0098] S9. The media platform server generates first response information according to the first user intention information.

[0099] The media platform server generates a first response message based on the first user intention information, and the first response message includes audio information, picture information, animation information and / or text information. The first response message may include first text information generated by voice recognition based on the first audio information. The first response message may also include temporary response information generated by preliminary processing based on the first text information. For example, the first response message in Figure 13 includes: "Xiao Cai Xiao Cai" and "Here, please speak, I'm listening..."

[0100] The first response information may also include an interactive response based on the first user intention information. For example, the first response information includes a start playback prompt: "Master, the video has been switched to playback for you", the first response information includes a processing wait prompt: "Processing, Master, please wait", the first response information includes a voice wait prompt: "Here, please speak, I'm listening", the first response information includes an interactive result prompt: "Master, the copy and download has been successfully made for you", or the first response information includes an intention not understood prompt: "Sorry, I didn't understand."

[0101] For ease of understanding, examples are provided with reference to the accompanying drawings. For example, the first response message in Figure 11 includes: "The video has been switched for you." For another example, the first response message in Figure 12 includes: "The video has been switched for you." For another example, the first response message in Figure 13 includes: "I'm here, please speak, I'm listening..." For another example, the first response message in Figure 14 includes: "Master, the copy and download has been successfully made for you." For another example, the first response message in Figure 15 includes: "Sorry, I didn't understand. What do you want to instruct Xiao Cai to do?"

[0102] S10. The media platform server sends a first response message to the calling terminal device.

[0103] S11. The calling terminal device plays a first response message.

[0104] After receiving the first response information, the calling terminal device plays the first response information in an interface for playing the first media resource or playing the default media resource.

[0105] It should be noted that the execution order of step S8 and step S11 is not limited in the embodiment of the present application. Step S8 may be executed first and then step S11, or step S11 may be executed first and then step S8, or step S8 and step S11 may be executed simultaneously. For example, the screen of the calling terminal device may superimpose the first response information while playing the first media resource.

[0106] S12. The calling terminal device sends second audio information to the media platform server.

[0107] Steps S12 to S15 are optional steps.

[0108] After the calling terminal device completes sending the first audio information to the media platform server, the calling terminal device may further send the second audio information to the media platform server.

[0109] S13. The media platform server determines a second media resource and / or generates second response information according to the second audio information.

[0110] In step S13, the media platform server processes the second audio information in a manner similar to the manner in which the media platform server processes the first audio information in steps S2 to S10, and thus is not described in detail here.

[0111] Exemplarily, when the second audio information is used to indicate the switching of the media resource to be played, the media platform service determines the second media resource to be switched based on the second user intention information corresponding to the second audio information.

[0112] It is understandable that when the user of the calling terminal device performs multiple voice interactions, the media platform server uses a similar processing method for the first audio information to perform multiple voice interactions respectively.

[0113] S14. The media platform server sends a response to the second audio information to the calling terminal device, including: a second media resource and / or second response information.

[0114] The second media resource is similar to the first media resource, and the second response information is similar to the first response information. The second response information is response information generated based on the second audio information, and the second response information includes audio information, picture information, animation information and / or text information.

[0115] S15. The calling terminal device plays the second media resource and / or the second response information.

[0116] In step S15, the calling terminal device stops playing the first media resource and / or the first response information, and then plays the second media resource and / or the second response information.

[0117] In this embodiment of the present application, after sending a call request to the called terminal, the calling terminal device sends a first audio message requesting the playback of a media resource. The calling terminal then receives and plays the first media resource. This allows the user to flexibly control the media resource played by the terminal device through voice commands, providing a more user-friendly and engaging user experience.

[0118] In combination with the foregoing embodiments, another communication scenario of the embodiment of the present application is introduced below. Please refer to Figure 3, which is a schematic diagram of another communication scenario in the embodiment of the present application. In the embodiment of the present application, the calling terminal device is connected to the IMS core network through the access network, and then connected to the media platform server through the IMS core network; the called terminal device is connected to the IMS core network through the access network, and then connected to the media platform server through the IMS core network. It can be understood that the calling terminal device and / or the called terminal device can be connected to the media platform server through other means, and the embodiment of the present application is not limited thereto.

[0119] The media platform server may include one or more logical network elements. In actual implementation, the media platform server may be deployed on a unified physical network element to implement multiple logical network elements, or may implement multiple logical network elements on multiple physical network elements according to internal characteristics. Take the media resource as an example, which is video ringback tone media (or video ringback tone, or ringback tone). In one possible implementation, the media platform server includes any one or more of the following logical network elements: video ringback tone business operation management service, artificial intelligence model, video ringback tone intelligent semantic understanding service, video ringback tone interactive response processing service, video ringback tone playback decision service, video ringback tone intelligent interactive service, video ringback tone application service, or video ringback tone media service. The following is a detailed description:

[0120] The Video Ringback Tone (VRBT) application service, also known as the signaling access and processing unit for video RBT, implements logical functions such as media negotiation and media playback control during the ringing phase of a video RBT user's call. The RBT application service supports invoking the playback decision service based on user intent information to obtain the media resource to be switched, notifying the calling terminal device to switch to the new media resource, and overlaying interactive feedback information.

[0121] 2. The Video Ringback Ringback Tone (RBT) media service is used to implement media playback of video RBT during the ringing phase of a call, sending the video RBT content to the calling terminal device via the network in the form of an audio and video media stream for playback. The Video Ringback Ringback Tone (RBT) media service supports uplink processing and intelligent recognition of the audio media stream of the calling terminal device's audio information (user voice), and real-time identification of wake-up keywords in the audio information. In response to detecting the wake-up keyword, the service generates corresponding text information for the calling terminal device's audio information, and supports the playback of the text information as subtitles overlaid on the video RBT media playing on the calling terminal device's screen to provide interactive real-time feedback.

[0122] 3. The Video Ringback Tone (RBT) Playback Decision Service selects the specific RBT to be played for a video ringback tone user (i.e., the user of the calling terminal device) during each call, based on the RBT media subscribed to by the calling terminal device. RBT content is selected in real time based on keywords in the user's intended content. When the semantics of the audio information from the calling terminal device indicate a specific type of RBT, a decision algorithm is used to determine the RBT content that matches the intended content.

[0123] 4. The video ringback tone intelligent interactive service is used to accurately understand the semantics of the audio information of the calling terminal device, obtain the user interaction intention of the calling terminal device, and make corresponding interactive responses.

[0124] 5. Artificial intelligence model, used to pre-train the video ringback tone intelligent interactive service, to achieve accurate semantic understanding and interactive feedback of the text information generated by the calling terminal device based on the audio information in the video ringback tone voice interaction scenario.

[0125] 6. Video ringback tone business operation and management services are used to provide users with video ringback tone business management services, including activating video material functions and ordering video ringback tone tones.

[0126] In conjunction with the media platform server illustrated in FIG3 , please refer to FIG4 , which is a flow chart illustrating an embodiment of a method for playing media resources in accordance with an embodiment of the present application. When the media platform server includes: a video ringback tone application service, a video ringback tone media service, a video ringback tone playback decision service, and a video ringback tone intelligent interactive service, a method for playing media resources proposed in an embodiment of the present application includes:

[0127] D1. The calling terminal device sends a call request to the called terminal device.

[0128] D2. The calling terminal device and the called terminal device conduct call negotiation and media resource (color ring tone) negotiation.

[0129] D3. The video ringback tone application service notifies the video ringback tone media service to play the media resource, which carries the voice recognition identifier.

[0130] In step D3, in a possible implementation, the video ringback ring application service sends voice recognition identification information to the video ringback ring media service, and the voice recognition identification information instructs the video ringback ring media service to start receiving the uplink audio media stream and voice recognition of the wake-up keyword.

[0131] In another possible implementation, the video ringback tone application service sends a voice interaction prompt to the video ringback tone media service. This voice interaction prompt notifies the calling terminal device to display the voice interaction prompt, which includes prompting the user to enter a wake-up keyword. The calling terminal device then displays the voice interaction prompt in real time as an overlay on the interface playing the default media resource.

[0132] It should be noted that the video ringback tone application service sends voice recognition identification information and voice interaction prompt information to the video ringback tone media service.

[0133] D4. The video ringback tone media service sends default media resources and voice interaction prompt information to the calling terminal device.

[0134] In step D4, the default media resource is played on the screen of the calling terminal device, and the voice interaction prompt information is superimposed and played in real time.

[0135] D5. The calling terminal device sends the first audio information to the video ringback tone media service.

[0136] D6. The video ringback tone media service detects whether the first audio information includes a wake-up keyword.

[0137] The video ringback tone media service parses the first audio message from the calling terminal device and, using techniques such as keyword recognition, determines whether the first audio message includes the wake-up keyword. If so, the process proceeds to step D7. If not, the process proceeds to step D7. Otherwise, the service either does not process the message or provides a message indicating that the message was not understood, such as "Sorry, I didn't understand."

[0138] D7. The video ringback tone media service performs speech recognition processing on the first audio information to generate a first text message.

[0139] Exemplarily, the video ringback tone media service recognizes the first audio information based on automatic speech recognition technology and converts the first audio information into the first text information.

[0140] D8. The video ringback tone media service sends the first text message and temporary response information to the calling terminal device.

[0141] The video ring back tone media service superimposes the first text message and the corresponding temporary response message into the video ring back tone media stream (default media resource) and sends it to the calling terminal device. The first text message and the temporary response message are superimposed and played in real time on the screen of the calling terminal device.

[0142] D9. The video ringback tone media service reports the first text message to the video ringback tone intelligent interactive service.

[0143] D10. The video ringback tone intelligent interactive service performs semantic understanding processing on the first text information to generate first user intention information, and determines first response information based on the first user intention information.

[0144] D11. The video ringback tone intelligent interactive service sends an interactive processing request to the video ringback tone application service. The interactive processing request is used to request to play a new media resource. The interactive processing request includes first user intention information, first response information and / or first text information.

[0145] The video ring back tone intelligent interaction service sends an interaction processing request to the video ring back tone application service (or video ring back tone voice management service) based on the first user intention information. The interaction processing request is used to request the media ring back tone application service to perform corresponding service processing on the first user intention information. The interaction processing request includes the first user intention information, the first response information, and / or the first text message.

[0146] Exemplarily, the first user intention information includes but is not limited to: interactive response instructions, such as switching playback media resources, rewinding playback media resources, or copying and ordering media resources, etc., or content feature keywords (the content feature keywords can also be called video content tags) and the weight of content feature keywords, etc.

[0147] D12. The video ringback tone application service sends a media resource query request to the video ringback tone playback decision service. The request carries the first user intention information.

[0148] After the video ringback ring application service determines that the first user intention information is to switch the media resource to be played, the video ringback ring application service sends a media resource query request to the video ringback ring playback decision service. The request carries the first user intention information. For example, the request includes: content feature keywords and content feature keyword weights, etc.

[0149] D13. The video ringback tone playback decision service determines the first media resource according to the media resource query request.

[0150] The video ringback tone playback decision service processes the media resource query request based on the specified decision recommendation model, decides to select the media resource (i.e., video ringback tone) that best meets the corresponding content feature keywords, and returns it to the video ringback tone application service. The decision recommendation model applies a parameter set to determine the media resource. The parameter set includes any one or more of the following: content feature keywords of the media resource, weights of the content feature keywords of the media resource, media resource tags of the media resource library, weights of the media resource tags of the media resource library, popularity weights of the media resources of the media resource library, media resource release time of the media resource library, or media resource playback rate of the media resource library, wherein the media resource library includes one or more media resources.

[0151] D14. The video ringback tone playback decision service sends identification information of the first media resource to the video ringback tone application service.

[0152] D15. The video ringback tone application service notifies the video ringback tone media service to switch to playing the first media resource and to play the first response information.

[0153] D16. The video ringback tone media service sends the first media resource and the first response information to the calling terminal device.

[0154] The calling terminal device plays the first media resource and the first response information on the screen.

[0155] D17. The called terminal device goes off-hook.

[0156] D18. The video ringback tone application service instructs the video ringback tone media service to stop playing the media resource and stop receiving the audio.

[0157] D19: The calling terminal device and the called terminal device renegotiate to connect the call.

[0158] In this embodiment, intelligent voice interaction is implemented in the video ringback tone playback scenario. During the video ringback tone playback process, real-time interaction based on the user's voice is achieved, providing users with a more friendly and interesting user experience. Because this solution does not rely on the terminal and network, it can be implemented on the video ringback tone platform side, which facilitates promotion and expands the application scope of the service.

[0159] In conjunction with the foregoing embodiments, another communication scenario in accordance with an embodiment of the present application will now be described. Please refer to Figure 5, which is a schematic diagram of another communication scenario in accordance with an embodiment of the present application. In another possible implementation, the media platform server in accordance with an embodiment of the present application includes any one or more of the following logical network elements: a video ringback tone playback decision service, a video ringback tone intelligent interaction service, a video ringback tone application service, or a video ringback tone media service.

[0160] In conjunction with the media platform server illustrated in FIG5 , please refer to FIG6 , which is a flow chart illustrating an embodiment of a method for playing media resources in accordance with an embodiment of the present application. When the media platform server includes: a video ringback tone application service, a video ringback tone media service, a video ringback tone playback decision service, or a video ringback tone intelligent interactive service, a method for playing media resources proposed in an embodiment of the present application includes:

[0161] F1. The calling terminal device sends a call request to the called terminal device.

[0162] F2: The calling terminal device and the called terminal device perform call negotiation and media resource (color ring tone) negotiation.

[0163] F3. The video ringback tone application service notifies the video ringback tone media service to play the media resource, which carries the voice recognition identifier.

[0164] F4. The video ringback tone media service sends default media resources and voice interaction prompt information to the calling terminal device.

[0165] Accordingly, the calling terminal device plays the default media resource (ie, the default video ringback tone) and simultaneously displays the voice interaction prompt information superimposed on the default video ringback tone. The voice interaction prompt information is shown in FIG15 .

[0166] F5. The calling terminal device sends the first audio information to the video ringback tone media service.

[0167] F6. The video ringback tone media service detects whether the first audio information includes a wake-up keyword.

[0168] Optionally, if a wake-up keyword is included, the call is marked as a voice wake-up state, and the complete audio information of the calling terminal device is voice recognized and converted into text information.

[0169] F7. The video ringback tone media service performs speech recognition processing on the first audio information to generate a first text message.

[0170] F8. The video ringback tone media service sends the first text message and temporary response information to the calling terminal device.

[0171] Accordingly, the calling terminal device plays the default media resource (i.e., the default video ringback tone) and simultaneously displays a temporary response message superimposed on the default video ringback tone. An example of this temporary response message is shown in FIG17 , which is another schematic diagram of another scenario in an embodiment of the present application. The temporary response message includes "Processing, please wait a moment..."

[0172] F9. The video ringback tone media service reports the first text message to the video ringback tone intelligent interactive service.

[0173] F10. The video ringback tone intelligent interactive service performs semantic understanding processing on the first text information to generate first user intention information, and determines first response information based on the first user intention information.

[0174] F11. The video ringback tone intelligent interactive service sends an interactive processing request to the video ringback tone application service. The interactive processing request is used to request the playback of new media resources. The interactive processing request includes first user intention information, first response information and / or first text information.

[0175] F12. The video ringback tone application service sends a media resource query request to the video ringback tone playback decision service. The request carries the first user intention information.

[0176] F13. The video ringback tone playback decision service determines the first media resource according to the media resource query request.

[0177] F14. The video ringback tone playback decision service sends identification information of the first media resource to the video ringback tone application service.

[0178] In another possible implementation, the video ringback ring application service selects the next video ringback ring (the next video ringback ring serves as the first media resource) according to the content in the existing play rule list, and notifies the video ringback ring media service to switch to playing the first media resource.

[0179] F15. The video ringback tone application service notifies the video ringback tone media service to switch to playing the first media resource and playing the first response information.

[0180] F16. The video ringback tone media service sends the first media resource and the first response information to the calling terminal device.

[0181] Accordingly, the first media resource is played on the screen of the calling terminal device, and the first response information is superimposed and displayed. The first response information is shown in Figure 18, which is another scenario diagram in an embodiment of the present application. The first response information includes "Owner, the video playback has been switched for you."

[0182] F17. The calling terminal device sends the second audio information to the video ringback tone media service.

[0183] F18. The video ringback tone media service detects whether the second audio information includes a wake-up keyword.

[0184] Optionally, step F18 may not be performed and step F19 may be performed directly, because the current call has been marked as the voice wake-up state in step F6.

[0185] F19. The video ringback tone media service performs voice recognition processing on the second audio information to generate second text information.

[0186] F20. The video ringback tone media service sends a second text message and temporary response information to the calling terminal device.

[0187] F21. The video ringback tone media service reports the second text information to the video ringback tone intelligent interactive service.

[0188] F22. The video ringback tone intelligent interactive service performs semantic understanding processing on the second text information to generate second user intention information, and determines second response information based on the second user intention information.

[0189] F23. The video ringback tone intelligent interactive service sends an interactive processing request to the video ringback tone application service. The interactive processing request is used to request the playback of new media resources. The interactive processing request includes the second user intention information, the second response information and / or the second text information.

[0190] F24. The video ringback tone application service sends a media resource query request to the video ringback tone playback decision service. The request carries the second user intention information.

[0191] F25. The video ringback tone playback decision service determines the second media resource according to the media resource query request.

[0192] The video ringback tone AS selects the next song based on the content in the existing play rule list, notifies the media service to switch playback, and displays the voice interaction result response message to the terminal user on the video.

[0193] F26. The video ringback tone playback decision service sends identification information of the second media resource to the video ringback tone application service.

[0194] F27. The video ringback tone application service notifies the video ringback tone media service to switch to playing the second media resource and to play the second response information.

[0195] F28. The video ringback tone media service sends a second media resource and a second response message to the calling terminal device.

[0196] Correspondingly, after receiving the second media resource, the calling terminal device stops playing the first media resource, and then plays the second media resource and superimposes and displays the second response information.

[0197] F29: The called terminal device goes off-hook.

[0198] F30: The video ringback tone application service instructs the video ringback tone media service to stop playing media resources and stop receiving audio.

[0199] F31: The calling terminal device and the called terminal device renegotiate to reconnect the call.

[0200] In combination with the foregoing embodiments, another communication scenario of the embodiment of the present application is introduced below. Please refer to Figure 7, which is a schematic diagram of another communication scenario in the embodiment of the present application. In another possible implementation, the media platform server in the embodiment of the present application includes any one or more of the following logical network elements: video ringback tone playback decision service, video ringback tone application service, video ringback tone intelligent interactive service or video ringback tone media service, wherein the video ringback tone intelligent interactive service includes: video ringback tone interactive response processing service and video ringback tone intelligent semantic understanding service, and the video ringback tone media service includes: video ringback tone interactive voice recognition service and video ringback tone media playback service.

[0201] In conjunction with the media platform server illustrated in FIG7 , please refer to FIG8 , which is a flow chart illustrating an embodiment of a method for playing media resources in accordance with an embodiment of the present application. When the media platform server includes: a video ringback tone playback decision service, a video ringback tone application service, a video ringback tone intelligent interactive service, or a video ringback tone media service, a method for playing media resources proposed in an embodiment of the present application includes:

[0202] G1. The calling terminal device sends a call request to the called terminal device.

[0203] G2: The calling terminal device and the called terminal device conduct call negotiation and media resource (color ring tone) negotiation.

[0204] G3. The video ringback tone application service identifies whether the calling terminal device is allowed to perform intelligent voice interaction.

[0205] G4. The video ringback tone application service notifies the video ringback tone media playback service to play the media resource, which carries the voice recognition identifier.

[0206] G5. The video ringback tone media playback service sends default media resources and voice interaction prompt information to the calling terminal device.

[0207] G6. The calling terminal device sends the fourth audio information (without carrying the wake-up keyword) to the video ringback tone media playback service.

[0208] G7. The video ringback tone media playback service recognizes that the fourth audio information does not carry the wake-up keyword.

[0209] G8. The calling terminal device sends a third audio message (carrying a wake-up keyword) to the video ringback tone media playback service.

[0210] G9. The video ringback tone media playback service recognizes that the third audio information carries the wake-up keyword.

[0211] G10. The video ringback tone media playback service requests the video ringback tone interactive voice recognition service to wake up the voice recognition service.

[0212] The video ringback tone media playback service analyzes the voice media stream (third audio information) of the calling terminal device and identifies the wake-up keyword. It then notifies the video ringback tone voice recognition service to start the service and uploads the third audio information to the video ringback tone voice recognition service. It then marks the call as voice wake-up state.

[0213] G11. The video ringback tone media playback service sends the third audio information to the video ringback tone interactive voice recognition service.

[0214] G12. The video ringback tone interactive voice recognition service performs voice recognition processing on the third audio information to generate third text information.

[0215] G13. The video ringback tone interactive voice recognition service sends a third text message and a temporary answer message to the calling terminal device.

[0216] G14. The video ringback tone interactive voice recognition service reports the third text information to the video ringback tone intelligent semantic understanding service.

[0217] G15. The video ringback tone intelligent semantic understanding service performs semantic understanding processing on the third text information to generate third user intention information.

[0218] G16. The video ringback tone intelligent semantic understanding service sends the third user intention information to the video ringback tone interactive response processing service.

[0219] G17. The video ringback tone interactive response processing service determines third response information based on the third user intention information.

[0220] G18. The video ringback tone interactive response processing service sends a first interactive processing request to the video ringback tone application service. The first interactive processing request includes third user intention information, third response information and / or third text information. The third response information includes a waiting voice prompt.

[0221] G19. The video ringback tone application service sends a third response message to the video ringback tone media playback service. The waiting voice prompt carried in the third response message is used to notify the waiting user of voice interaction input.

[0222] G20. The video ringback tone media playback service sends a third response message to the calling terminal device.

[0223] Since the third audio message only includes the wake-up keyword, the video ringback tone media playback service sends a third response message to the calling terminal device. The waiting voice prompt carried in the third response message is used to notify the waiting user of the voice interaction input. The waiting voice prompt is an example of FIG13 .

[0224] G21. The calling terminal device sends the first audio information to the video ringback tone media playback service.

[0225] G22. The video ringback tone media playback service sends the first audio information to the video ringback tone interactive voice recognition service.

[0226] G23. The video ringback tone interactive voice recognition service performs voice recognition processing on the first audio information to generate a first text message.

[0227] G24. The video ringback tone interactive voice recognition service sends a first text message and a temporary answer message to the calling terminal device.

[0228] G25. The video ringback tone interactive voice recognition service reports the first text information to the video ringback tone intelligent semantic understanding service.

[0229] G26. The video ringback tone intelligent semantic understanding service performs semantic understanding processing on the first text information to generate first user intention information.

[0230] G27. The video ringback tone intelligent semantic understanding service sends the first user intention information to the video ringback tone interactive response processing service.

[0231] G28. The video ringback tone interactive response processing service determines the first response information according to the first user intention information.

[0232] G29. The video ringback tone interactive response processing service sends a second interactive processing request to the video ringback tone application service. The second interactive processing request includes the first user intention information, the first response information and / or the first text information. The first user intention information includes switching the playback media resources.

[0233] G30. The video ringback tone application service notifies the video ringback tone media playback service to switch to playing the first media resource.

[0234] G31. The video ringback tone media playback service sends a first media resource and a first response message to the calling terminal device.

[0235] G32. The called terminal device goes off-hook.

[0236] G33. The video ringback tone application service instructs the video ringback tone media playback service to stop playing media resources and stop receiving audio.

[0237] G34: The calling terminal device and the called terminal device renegotiate to reconnect the call.

[0238] In conjunction with the foregoing embodiments, another communication scenario in accordance with an embodiment of the present application will now be described. Please refer to FIG9 , which is a schematic diagram of another communication scenario in accordance with an embodiment of the present application. In another possible implementation, the media platform server in accordance with an embodiment of the present application includes any one or more of the following logical network elements: a video ringback tone operation and management service, or a video ringback tone call playback platform, wherein the video ringback tone call playback platform includes: a video ringback tone intelligent interactive service, a video ringback tone playback decision service, and a video ringback tone application service.

[0239] In conjunction with the media platform server shown in FIG9 , please refer to FIG10 , which is a flow chart of a method for playing media resources in accordance with an embodiment of the present application. When the media platform server includes a video ringback tone operation management service and a video ringback tone call playback platform, a method for playing media resources proposed in an embodiment of the present application includes:

[0240] H1. The calling terminal device sends a call request to the called terminal device.

[0241] H2: The calling terminal device and the called terminal device perform call negotiation and media resource (color ring tone) negotiation.

[0242] H3. The video ringback tone application service identifies whether the calling terminal device is allowed to perform intelligent voice interaction.

[0243] H4. The video ringback tone application service notifies the video ringback tone media service to play the media resource, which carries the voice recognition identifier.

[0244] H5 and video ringback tone media services send default media resources and voice interaction prompt information to the calling terminal device.

[0245] H6. The calling terminal device sends the first audio information (carrying the wake-up keyword and the ringback tone copy subscription instruction) to the video ringback tone media service.

[0246] H7. The video ringback tone media service detects whether the first audio information includes a wake-up keyword.

[0247] H8. The video ringback tone media service performs speech recognition processing on the first audio information to generate a first text message.

[0248] H9. The video ringback tone media service sends the first text message and temporary response information to the calling terminal device.

[0249] H10. The video ringback tone media service reports the first text message to the video ringback tone intelligent interactive service.

[0250] H11. The video ringback tone intelligent interactive service performs semantic understanding processing on the first text information to generate first user intention information, and determines first response information based on the first user intention information.

[0251] H12. The video ringback tone intelligent interactive service sends the first user intention information to the video ringback tone operation management service.

[0252] H13. The video ringback tone operation management service copies and orders the currently playing media resources for the calling terminal device based on the first user intention information.

[0253] H14. The video ringback tone intelligent interactive service sends an interactive processing request to the video ringback tone application service. The interactive processing request carries first response information, and the first response information includes a media resource subscription result.

[0254] H15. The video ringback tone application service notifies the video ringback tone media service to play the first response information.

[0255] H16. The video ringback tone media service sends a first response message to the calling terminal device.

[0256] Correspondingly, the screen of the calling terminal device plays the first response information, and the first response information is shown in, for example, FIG14 .

[0257] H17. The calling terminal device sends the second audio information to the video ringback tone media service.

[0258] H18. The video ringback tone media service detects whether the second audio information includes a wake-up keyword.

[0259] H19. The video ringback tone media service performs speech recognition processing on the second audio information to generate second text information.

[0260] H20. The video ringback tone media service sends a second text message and temporary response information to the calling terminal device.

[0261] H21. The video ringback tone media service reports the second text information to the video ringback tone intelligent interactive service.

[0262] H22. The video ringback tone intelligent interactive service performs semantic understanding processing on the second text information to generate second user intention information, and determines second response information based on the second user intention information, where the second response information indicates that the second audio information is not understood.

[0263] H23. The video ringback tone intelligent interactive service sends an interactive processing request to the video ringback tone application service. The interactive processing request includes the second user intention information, the second response information and / or the second text information. The second response information indicates that the second audio information is not understood.

[0264] H24. The video ringback tone application service notifies the video ringback tone media service to play the second response information.

[0265] H25. The video ringback tone media service sends a second response message to the calling terminal device.

[0266] Correspondingly, the screen of the calling terminal device plays the second response information, and the first response information is shown in, for example, FIG16 .

[0267] H26. The called terminal device goes off-hook.

[0268] H27. The video ringback tone application service instructs the video ringback tone media service to stop playing media resources and stop receiving audio.

[0269] H28: The calling terminal device and the called terminal device renegotiate to connect the call.

[0270] In conjunction with the above embodiments, the application scenario of the method provided in the embodiment of the present application is described below. The application scenario of the method provided in the embodiment of the present application can be shown in Figure 19. Figure 19 is a schematic diagram of an application scenario in the embodiment of the present application, which includes: user 001 and electronic device 002.

[0271] Among them, user 001: The user can interact with electronic device 002 through gestures / voice, etc. to open an application on the electronic device (or a functional module in the application, or a mini program, quick application, etc.).

[0272] Electronic device 002: Equipped with an operating system, the operating system has built-in system-level apps, and users can install / uninstall apps as needed. Electronic device 002 has a screen that displays to user 001, through which user 001 can operate applications on electronic device 002. Electronic device 002 typically has a large screen, such as a tablet computer or a foldable phone.

[0273] In addition to the aforementioned tablet computers (pads) or foldable screen mobile phones, the electronic devices in the embodiments of the present application may also be non-foldable screen mobile phones, smart watches, smart glasses, smart bracelets, portable game consoles, personal digital assistants (PDAs), laptop computers, ultra mobile personal computers (UMPCs), handheld computers, netbooks, in-vehicle media playback devices, wearable electronic devices (e.g., watches, bracelets, glasses), virtual reality (VR) terminal devices, augmented reality (AR) terminal devices, and other digital display products. The electronic device 002 may be the calling terminal device and / or the called terminal device in the embodiments of the present application.

[0274] Please refer to Figure 20, which is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. As shown in Figure 20, the electronic device may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, an earphone interface 270D, a sensor module 280, a button 290, a motor 291, an indicator 292, a camera 293, a display 294, and a subscriber identification module (SIM) card interface 295. Among them, the sensor module 280 can include a pressure sensor 280A, a gyroscope sensor 280B, an air pressure sensor 280C, a magnetic sensor 280D, an acceleration sensor 280E, a distance sensor 280F, a proximity light sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, etc.

[0275] It should be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device. In other embodiments, the electronic device may include more or fewer components than shown, or may combine or separate certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0276] The processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors.

[0277] The controller can be the nerve center and command center of the electronic device. The controller can generate operation control signals based on the instruction opcode and timing signals to complete the control of instruction fetching and execution.

[0278] Processor 210 may also include a memory for storing instructions and data. In some embodiments, the memory in processor 210 is a cache memory. This memory can store instructions or data that have just been used or are being recycled by processor 210. If processor 210 needs to use the same instruction or data again, it can directly access the memory. This avoids duplicate accesses, reduces processor 210 latency, and thus improves system efficiency.

[0279] In some embodiments, the processor 210 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface.

[0280] It is understood that the interface connection relationship between the modules illustrated in this embodiment is only for illustrative purposes and does not constitute a structural limitation on the electronic device. In other embodiments, the electronic device may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.

[0281] The charging management module 240 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 240 can receive charging input from the wired charger via the USB interface 230. In some wireless charging embodiments, the charging management module 240 can receive wireless charging input via the electronic device's wireless charging coil. While charging the battery 242, the charging management module 240 can also power the electronic device through the power management module 241.

[0282] The power management module 241 is used to connect the battery 242, the charging management module 240, and the processor 210. The power management module 241 receives input from the battery 242 and / or the charging management module 240 and provides power to the processor 210, the internal memory 221, the external memory, the display 294, the camera 293, and the wireless communication module 260. The power management module 241 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 241 can also be set in the processor 210. In other embodiments, the power management module 241 and the charging management module 240 can also be set in the same device.

[0283] The wireless communication function of the electronic device can be implemented through antenna 1, antenna 2, mobile communication module 250, wireless communication module 260, modem processor and baseband processor.

[0284] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in an electronic device can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.

[0285] The mobile communication module 250 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to electronic devices. The mobile communication module 250 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 250 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 250 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In some embodiments, at least some of the functional modules of the mobile communication module 250 can be set in the processor 210. In some embodiments, at least some of the functional modules of the mobile communication module 250 can be set in the same device as at least some of the modules of the processor 210.

[0286] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 270A, the receiver 270B, etc.) or displays an image or video through the display screen 294. In some embodiments, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 210 and be set in the same device as the mobile communication module 250 or other functional modules.

[0287] The wireless communication module 260 can provide wireless communication solutions for electronic devices, including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR), etc. The wireless communication module 260 can be one or more devices that integrate at least one communication processing unit. The wireless communication module 260 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 210. The wireless communication module 260 can also receive the signal to be sent from the processor 210, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.

[0288] In some embodiments, antenna 1 of the electronic device is coupled to mobile communication module 250, and antenna 2 is coupled to wireless communication module 260, so that the electronic device can communicate with a network and other devices via wireless communication technologies. The wireless communication technologies may include global system for mobile communications (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), time-division code division multiple access (TDSCDMA), long term evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include global positioning system (GPS), global navigation satellite system (GLONASS), Beidou navigation satellite system (BDS), quasi-zenith satellite system (QZSS) and / or satellite-based augmentation system (SBAS).

[0289] The electronic device implements display functionality through a GPU, display screen 294, and an application processor. A GPU is a microprocessor for image processing that connects display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 210 may include one or more GPUs that execute program instructions to generate or modify display information.

[0290] Display screen 294 is used to display images, videos, etc. Display screen 294 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a MiniLED, a MicroLED, a Micro-oLed, or a quantum dot light-emitting diode (QLED).

[0291] Optionally, the display screen 294 displays the running interface of the application on the electronic device in the form of a visual window. In the embodiment of the present application, the display screen displays the running interface of a specific type of application (such as SMS and instant messaging software) in a dual-page display mode, showing two interfaces on one display screen, and the two interfaces can be parent and child pages of each other.

[0292] It is understood that if the electronic device is a foldable screen mobile phone, the display screen 294 can also be called a foldable screen. It can be folded into two left and right screens along the longitudinal folding edge of the foldable screen, or it can be folded into two upper and lower screens along the horizontal folding edge of the foldable screen, etc.

[0293] It should be noted that the at least two screens formed after the electronic device in the embodiment of the present application (including the electronic device that folds inward and the electronic device that folds outward) is folded can be multiple independent screens, or a complete screen with an integrated structure, which is folded into at least two parts.

[0294] For example, the foldable screen can be a flexible foldable screen. The flexible foldable screen includes a folding edge made of a flexible material. Part or all of the flexible foldable screen is made of a flexible material. When the flexible foldable screen is folded, the at least two screens formed are a complete integrated screen, but are folded into at least two parts.

[0295] For another example, the foldable screen of the electronic device may be a multi-screen foldable screen. The multi-screen foldable screen may include multiple (two or more) screens. These multiple screens are multiple independent display screens. These multiple screens may be connected in sequence via folding axes. Each screen may rotate about the folding axis connected to it, thereby achieving folding of the multi-screen foldable screen.

[0296] The electronic device can realize the shooting function through the ISP, the camera 293, the video codec, the GPU, the display 294 and the application processor.

[0297] The ISP processes data fed back by camera 293. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then passed to the ISP for processing and converted into a visible image. The ISP can also perform algorithmic optimization on image noise, brightness, and skin tone. The ISP can also optimize parameters such as exposure and color temperature of the captured scene. In some embodiments, the ISP can be located within camera 293.

[0298] The camera 293 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In some embodiments, the electronic device may include 1 or N cameras 293, where N is a positive integer greater than 1.

[0299] Camera 293 can also be used to provide the user with a personalized, contextualized service experience based on the external environment and user actions perceived by the electronic device. Specifically, camera 293 can acquire rich and accurate information, enabling the electronic device to perceive the external environment and user actions. Specifically, in an embodiment of the present application, camera 293 can be used to identify whether the user of the electronic device is the first user or the second user.

[0300] Digital signal processors (DSPs) are used to process digital signals. Besides digital image signals, they can also process other digital signals. For example, when an electronic device selects a frequency, the DSP performs a Fourier transform on the frequency energy.

[0301] Video codecs are used to compress or decompress digital video. Electronic devices may support one or more video codecs. This allows them to play or record videos in a variety of encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.

[0302] The NPU is a neural network (NN) computing processor. Drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain, it rapidly processes input information and can continuously self-learn. The NPU enables intelligent cognitive applications in electronic devices, such as image recognition, face recognition, speech recognition, and text comprehension.

[0303] The external memory interface 220 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor 210 via the external memory interface 220 to implement data storage functions. For example, files such as music and videos can be stored on the external memory card.

[0304] The internal memory 221 can be used to store computer executable program codes, which include instructions. The processor 210 executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory 221. For example, in an embodiment of the present application, the processor 210 can display the corresponding display content on the display screen in response to the user's operation on the display screen 294 by executing the instructions stored in the internal memory 221. The internal memory 221 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc. The data storage area may store data created during the use of the electronic device (such as audio data, a phone book, etc.), etc. In addition, the internal memory 221 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0305] The electronic device can implement audio functions such as music playback and recording through the audio module 270, the speaker 270A, the receiver 270B, the microphone 270C, the headphone jack 270D, and the application processor.

[0306] The audio module 270 is used to convert digital audio information into analog audio signals for output, and also to convert analog audio input into digital audio signals. The audio module 270 can also be used to encode and decode audio signals. In some embodiments, the audio module 270 can be located within the processor 210, or some of its functional modules can be located within the processor 210. The speaker 270A, also known as the "speaker," is used to convert electrical audio signals into sound signals. The electronic device can listen to music or make hands-free calls through the speaker 270A. The receiver 270B, also known as the "earpiece," is used to convert electrical audio signals into sound signals. When the electronic device receives a call or voice message, the user can hold the receiver 270B close to their ear to hear the voice. The microphone 270C, also known as the "microphone," is used to convert sound signals into electrical signals. When making a call, sending a voice message, or triggering a function in the electronic device through a voice assistant, the user can speak into the microphone 270C by placing their mouth close to it. An electronic device can be equipped with at least one microphone 270C. In other embodiments, the electronic device may be provided with two microphones 270C, which can not only collect sound signals but also implement noise reduction. In other embodiments, the electronic device may be provided with three, four, or more microphones 270C, which can collect sound signals, reduce noise, identify sound sources, implement directional recording, and so on.

[0307] The headphone jack 270D is used to connect a wired headphone and can be a USB interface 230 or a 3.5mm open mobile terminal platform (OMTP) standard interface or a cellular telecommunications industry association of the USA (CTIA) standard interface.

[0308] Pressure sensor 280A is used to sense pressure signals and convert them into electrical signals. In some embodiments, pressure sensor 280A can be located on display screen 294. There are many types of pressure sensors 280A, such as resistive, inductive, and capacitive. A capacitive pressure sensor can include at least two parallel plates made of conductive material. When force is applied to pressure sensor 280A, the capacitance between the electrodes changes. The electronic device determines the intensity of the pressure based on this change in capacitance. When a touch operation is applied to display screen 294, the electronic device detects the intensity of the touch operation based on pressure sensor 280A. The electronic device can also calculate the location of the touch based on the detection signal from pressure sensor 280A. In some embodiments, touch operations applied to the same touch location but with different touch intensities can correspond to different operation instructions. For example, when a touch operation with an intensity less than a pressure threshold is applied to a short message application icon, a command to view short messages is executed. When a touch operation with an intensity greater than or equal to the pressure threshold is applied to a short message application icon, a command to create a new short message is executed.

[0309] The gyroscope sensor 280B can be used to determine the motion posture of the electronic device. In some embodiments, the angular velocity of the electronic device around three axes (i.e., x, y, and z axes) can be determined by the gyroscope sensor 280B. The gyroscope sensor 280B can be used for anti-shake shooting. For example, when the shutter is pressed, the gyroscope sensor 280B detects the angle of the electronic device's shaking, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to offset the shaking of the electronic device through reverse movement to achieve anti-shake. The gyroscope sensor 280B can also be used for navigation and somatosensory game scenes. In addition, the gyroscope sensor 280B can also be used to measure the rotation amplitude or movement distance of the electronic device.

[0310] The air pressure sensor 280C is used to measure air pressure. In some embodiments, the electronic device calculates the altitude using the air pressure value measured by the air pressure sensor 280C to assist in positioning and navigation.

[0311] Magnetic sensor 280D includes a Hall effect sensor. The electronic device can use magnetic sensor 280D to detect the opening and closing of a flip case. In some embodiments, when the electronic device is a flip phone, the electronic device can detect the opening and closing of the flip cover based on magnetic sensor 280D. Based on the detected opening and closing status of the case or flip cover, features such as automatic unlocking of the flip cover can be configured.

[0312] Accelerometer 280E can detect the magnitude of an electronic device's acceleration in all directions (generally three axes). When the electronic device is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the electronic device's posture, for applications such as landscape and portrait screen switching and pedometers. Furthermore, accelerometer 280E can also be used to measure the electronic device's orientation (i.e., the orientation vector).

[0313] Distance sensor 280F is used to measure distance. The electronic device can measure distance using infrared or laser. In some embodiments, when shooting a scene, the electronic device can use distance sensor 280F to measure distance to achieve fast focusing.

[0314] The proximity light sensor 280G may include, for example, a light emitting diode (LED) and a light detector, such as a photodiode. The light emitting diode may be an infrared light emitting diode. The electronic device emits infrared light outward through the light emitting diode. The electronic device uses a photodiode to detect infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device. When insufficient reflected light is detected, the electronic device can determine that there is no object near the electronic device. The electronic device can use the proximity light sensor 280G to detect when the user holds the electronic device close to the ear to talk, so as to automatically turn off the screen to save power. The proximity light sensor 280G can also be used in leather case mode and pocket mode to automatically unlock and lock the screen.

[0315] The ambient light sensor 280L senses ambient light levels. The electronic device can adaptively adjust the brightness of the display screen 294 based on the perceived ambient light levels. The ambient light sensor 280L can also be used to automatically adjust the white balance when taking photos. The ambient light sensor 280L can also work with the proximity sensor 280G to detect whether the electronic device is in a pocket to prevent accidental touches.

[0316] The fingerprint sensor 280H is used to collect fingerprints. Electronic devices can use the collected fingerprint characteristics to unlock the device, access application locks, take photos with the fingerprint, answer calls with the fingerprint, and more.

[0317] The temperature sensor 280J is used to detect temperature. In some embodiments, the electronic device uses the temperature detected by the temperature sensor 280J to implement a temperature processing strategy. For example, when the temperature reported by the temperature sensor 280J exceeds a threshold, the electronic device reduces the performance of the processor located near the temperature sensor 280J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is lower than another threshold, the electronic device heats the battery 242 to prevent the electronic device from shutting down abnormally due to low temperature. In other embodiments, when the temperature is lower than another threshold, the electronic device boosts the output voltage of the battery 242 to prevent abnormal shutdown due to low temperature.

[0318] Touch sensor 280K, also known as a "touch panel," can be disposed on display screen 294. The touch sensor 280K and display screen 294 form a touch screen, also known as a "touch screen." Touch sensor 280K is used to detect touch operations applied thereto or in the vicinity thereof. The touch sensor can transmit the detected touch operations to an application processor to determine the type of touch event. Visual output related to the touch operations can be provided via display screen 294. In other embodiments, touch sensor 280K can also be disposed on the surface of the electronic device, at a location different from that of display screen 294.

[0319] The bone conduction sensor 280M can acquire vibration signals. In some embodiments, the bone conduction sensor 280M can acquire vibration signals from the vibrating bones of the human body's vocal cords. The bone conduction sensor 280M can also contact the human pulse to receive blood pressure signals. In some embodiments, the bone conduction sensor 280M can also be set in headphones to form bone conduction headphones. The audio module 270 can parse the voice signal based on the vibration signal of the vibrating bones of the vocal cords acquired by the bone conduction sensor 280M to implement voice functions. The application processor can parse heart rate information based on the blood pressure signals acquired by the bone conduction sensor 280M to implement heart rate detection functions.

[0320] Keys 290 include a power button, a volume button, and the like. Keys 290 may be mechanical keys or touch-sensitive keys. The electronic device may receive key inputs and generate key signal inputs related to user settings and function control of the electronic device.

[0321] The electronic device uses various sensors in the sensor module 280 , buttons 290 , and / or cameras 293 .

[0322] Motor 291 can generate vibration prompts. Motor 291 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. For touch operations acting on different areas of the display screen 294, motor 291 can also correspond to different vibration feedback effects. Different application scenarios (for example: time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.

[0323] The indicator 292 may be an indicator light, which may be used to indicate the charging status, power level change, messages, missed calls, notifications, etc.

[0324] The SIM card interface 295 is used to connect a SIM card. The SIM card can be connected to and separated from the electronic device by inserting it into or removing it from the SIM card interface 295. The electronic device can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 295 can support Nano SIM cards, Micro SIM cards, SIM cards, and the like. Multiple cards can be inserted into the same SIM card interface 295 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 295 can also be compatible with different types of SIM cards. The SIM card interface 295 can also be compatible with external memory cards. Electronic devices interact with the network through SIM cards to implement functions such as calls and data communications. In some embodiments, the electronic device uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device and cannot be separated from the electronic device.

[0325] The methods in the aforementioned embodiments can all be implemented in the electronic device 002 having the aforementioned hardware structure. For ease of understanding, the following description is schematically based on an example in which the electronic device is a mobile phone including a foldable screen, and the foldable screen is in a fully unfolded state. Whether the mobile phone includes a foldable screen, the form of the foldable screen of the mobile phone, and the number of screens after the foldable screen is folded are not limited here.

[0326] The above describes the media resource playback method in the embodiment of the present application. The following describes the electronic device in the embodiment of the present application. Please refer to Figure 21. Figure 21 shows another embodiment of the electronic device in the embodiment of the present application, including:

[0327] In one example, the electronic device is applied to a calling terminal device, and the electronic device includes:

[0328] The transceiver unit 2101 is used to send a call request to the called terminal device;

[0329] The transceiver unit 2101 is further configured to send a first audio message, where the first audio message is used to request playback of a media resource;

[0330] The transceiver unit 2101 is further configured to receive and play a first media resource, where the first media resource is determined by the first audio information.

[0331] In one possible implementation,

[0332] The transceiver unit 2101 is further configured to play a first response message, where the first response message is response information generated based on the first audio information, and the first response message includes audio information, picture information, animation information and / or text information.

[0333] In one possible implementation, the first response information includes:

[0334] The first text information is text information generated by performing speech recognition processing based on the first audio information, and the content included in the first text information corresponds to the first audio information.

[0335] In one possible implementation,

[0336] The transceiver unit 2101 is further configured to send second audio information;

[0337] The transceiver unit 2101 is further configured to receive a response to the second audio information, where the response to the second audio information includes:

[0338] a second media resource, wherein the second media resource is determined by the second audio information;

[0339] And / or, second response information, the second response information is response information generated based on the second audio information, the second response information includes audio information, picture information, animation information and / or text information.

[0340] In one possible implementation,

[0341] The transceiver unit 2101 is further configured to stop playing the first media resource;

[0342] The transceiver unit 2101 is further configured to play the second media resource;

[0343] And / or, play the second response information.

[0344] In one possible implementation,

[0345] The transceiver unit 2101 is further configured to receive an off-hook message sent by the called terminal device;

[0346] The processing unit 2102 is configured to stop playing the first media resource in response to the off-hook message;

[0347] The processing unit 2102 is further configured to stop receiving the audio information and / or stop the speech recognition processing of the audio information.

[0348] In one possible implementation,

[0349] The transceiver unit 2101 is further configured to send the first audio information, where the first audio information includes a wake-up keyword, and the wake-up keyword is used to trigger the media platform server to perform voice recognition processing on the first audio information;

[0350] or,

[0351] The transceiver unit 2101 is further configured to send audio information including the wake-up keyword;

[0352] The transceiver unit 2101 is further configured to send the first audio information.

[0353] In another possible implementation, the electronic device is applied to a media platform server, including:

[0354] The transceiver unit 2101 is further configured to receive a call request sent by the calling terminal device;

[0355] The transceiver unit 2101 is further configured to receive the first audio information sent by the calling terminal device;

[0356] The processing unit 2102 is further configured to determine a first media resource based on the first audio information;

[0357] The transceiver unit 2101 is further configured to send the first media resource to the calling terminal device.

[0358] In one possible implementation,

[0359] The processing unit 2102 is further configured to perform speech recognition processing on the first audio information to generate first text information, where the first text information includes content corresponding to the first audio information;

[0360] The processing unit 2102 is further configured to perform semantic understanding processing based on the first text information to generate first user intention information;

[0361] The processing unit 2102 is further configured to determine the first media resource according to the first user intention information.

[0362] In one possible implementation,

[0363] The processing unit 2102 is further configured to generate first response information according to the first user intention information, where the first response information includes audio information, image information, animation information, and / or text information;

[0364] The transceiver unit 2101 is further configured to send the first response information to the calling terminal device.

[0365] In a possible implementation, the first response information includes the first text information.

[0366] In one possible implementation, the first user intention information includes any one or more of the following:

[0367] Start playing a media resource, pause playing a media resource, switch playing a media resource, switch back playing a media resource, copy and order a media resource, content feature keywords of the first audio information, or the weight of the content feature keywords.

[0368] In one possible implementation,

[0369] The processing unit 2102 is further configured to determine the first media resource based on the decision recommendation model and the content feature keywords and / or the weights of the content feature keywords included in the first user intent information, wherein:

[0370] The decision recommendation model applies a parameter set to determine media resources, and the parameter set includes any one or more of the following: content feature keywords of the media resources, weights of the content feature keywords of the media resources, media resource tags of the media resource library, weights of the media resource tags of the media resource library, popularity weights of the media resources of the media resource library, release time of the media resources of the media resource library, or playback rate of the media resources of the media resource library, wherein the media resource library includes one or more media resources.

[0371] In one possible implementation,

[0372] The processing unit 2102 is further configured to detect whether the first audio information includes a wake-up keyword;

[0373] The processing unit 2102 is further configured to trigger speech recognition processing based on the first audio information if the first audio information includes the wake-up keyword;

[0374] Alternatively, the processing unit 2102 is further configured to trigger speech recognition processing based on the first audio information based on detecting that the received audio information includes the wake-up keyword.

[0375] In one possible implementation,

[0376] The transceiver unit 2101 is further configured to receive the second audio information sent by the calling terminal device;

[0377] The processing unit 2102 is further configured to generate a response to the second audio information according to the second audio information.

[0378] The response to the second audio information includes:

[0379] a second media resource, wherein the second media resource is determined by the second audio information;

[0380] and / or, second response information, where the second response information is response information generated based on the second audio information, and the second response information includes audio information, picture information, animation information and / or text information;

[0381] The transceiver unit 2101 is further configured to send a response to the second audio information to the calling terminal device.

[0382] In one possible implementation,

[0383] The processing unit 2102 is further configured to perform speech recognition processing based on the second audio information to generate second text information, where the second text information includes content corresponding to the first audio information;

[0384] The processing unit 2102 is further configured to perform semantic understanding processing based on the second text information to generate second user intention information;

[0385] The processing unit 2102 is further configured to generate a response to the second audio information according to the second user intention information.

[0386] Referring to FIG. 22 , this application provides a schematic diagram of the structure of another electronic device. The electronic device may include a processor 2201, a memory 2202, and a communication port 2203. The processor 2201, the memory 2202, and the communication port 2203 are interconnected via a circuit. The memory 2202 stores program instructions and data.

[0387] The memory 2202 stores program instructions and data corresponding to the steps executed by the calling terminal device and / or the media platform server in the corresponding implementation methods shown in Figures 2 to 18 above.

[0388] The processor 2201 is configured to execute the steps performed by the calling terminal device and / or the media platform server as shown in any of the embodiments shown in FIG. 2 to FIG. 18 .

[0389] The communication port 2203 can be used to receive and send data, and to execute the steps related to acquisition, sending, and receiving in any of the embodiments shown in Figures 2 to 18 above.

[0390] In one implementation, the electronic device may include more or fewer components relative to those in FIG. 22 . This application is merely an illustrative description and does not limit this.

[0391] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.

[0392] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0393] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in whole or in part through software, hardware, firmware, or any combination thereof.

[0394] When software is used to implement the integrated unit, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, hard disk, tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).

Claims

1. A method for playing media resources, characterized in that, The method is applied to the calling terminal device, and the method includes: Sending a call request to the called terminal device; Sending first audio information, where the first audio information is used to request playing a media resource; Receiving and playing a first media resource, where the first media resource is determined by the first audio information.

2. The method according to claim 1, wherein The method further includes: Playing a first response message, where the first response message is a response message generated based on the first audio information, and the first response message includes audio information, picture information, animation information, and / or text information.

3. The method according to any one of claims 1 or 2, characterized in that, The first response message includes: First text information, where the first text information is text information generated by performing speech recognition processing on the first audio information, and the content included in the first text information corresponds to the first audio information.

4. The method according to any one of claims 1 to 3, characterized in that The method further includes: Sending second audio information; Receiving a response to the second audio information, where the response to the second audio information includes: A second media resource, where the second media resource is determined by the second audio information; And / or, a second response message, where the second response message is a response message generated based on the second audio information, and the second response message includes audio information, picture information, animation information, and / or text information.

5. The method according to claim 4, characterized in that The method further includes: Stopping playing the first media resource; Playing the second media resource; And / or, playing the second response message.

6. The method according to any one of claims 1-5, characterized in that, The method further includes: Receiving an off-hook message sent by the called terminal device; In response to the off-hook message, stopping playing the first media resource; Stopping receiving audio information, and / or stopping speech recognition processing on the audio information.

7. The method according to any one of claims 1-6, characterized in that, Sending the first audio information includes: Sending the first audio information, where the first audio information includes a wake-up keyword, and the wake-up keyword is used to trigger the media platform server to perform speech recognition processing on the first audio information; Or, Sending audio information including the wake-up keyword; Sending the first audio information.

8. A method for playing media resources, characterized in that, The method is applied to the media platform server, and the method includes: Receiving a call request sent by the calling terminal device; Receiving first audio information sent by the calling terminal device; Determining a first media resource according to the first audio information; Sending the first media resource to the calling terminal device.

9. The method according to claim 8, wherein The determining the first media resource according to the first audio information includes: Performing speech recognition processing on the first audio information to generate first text information, where the content included in the first text information corresponds to the first audio information; Performing semantic understanding processing on the first text information to generate first user intent information; Determining the first media resource according to the first user intent information.

10. The method according to claim 9, characterized in that The method further includes: Generating a first response message according to the first user intent information, where the first response message includes audio information, picture information, animation information, and / or text information; Sending the first response message to the calling terminal device.

11. The method according to claim 9 or 10, characterized in that The first response message includes the first text information.

12. The method according to any one of claims 9-11, characterized in that, The first user intent information includes any one or more of the following: Start playing the media resource, pause playing the media resource, switch the played media resource, rewind the played media resource, copy and subscribe to the media resource, the content feature keywords of the first audio information, or the weights of the content feature keywords.

13. The method according to claim 12, characterized in that, Determining the first media resource according to the first user intent information includes: Determining the first media resource according to a decision recommendation model and the content feature keywords and / or the weights of the content feature keywords included in the first user intent information, where The decision recommendation model determines media resources using a parameter set, and the parameter set includes any one or more of the following: the content feature keywords of the media resources, the weights of the content feature keywords of the media resources, the media resource tags of the media resource library, the weights of the media resource tags of the media resource library, the popularity weights of the media resources in the media resource library, the release time of the media resources in the media resource library, or the play rates of the media resources in the media resource library, where the media resource library includes one or more media resources.

14. The method according to any one of claims 9-13, characterized in that, The method further includes: Detecting whether the first audio information includes a wake-up keyword; If the first audio information includes the wake-up keyword, triggering speech recognition processing according to the first audio information; Or, triggering speech recognition processing according to the first audio information based on detecting that the received audio information includes the wake-up keyword.

15. The method according to any one of claims 8-14, characterized in that, The method further includes: Receiving second audio information sent by the calling terminal device; Generating a response to the second audio information according to the second audio information, The response to the second audio information includes: A second media resource determined by the second audio information; And / or, a second response message, which is a response message generated based on the second audio information, and the second response message includes audio information, picture information, animation information, and / or text information; Sending the response to the second audio information to the calling terminal device.

16. The method according to any one of claims 8-15, characterized in that, Generating a response to the second audio information according to the second audio information includes: Performing speech recognition processing according to the second audio information to generate second text information, and the content included in the second text information corresponds to the first audio information; Performing semantic understanding processing according to the second text information to generate second user intent information; Generating a response to the second audio information according to the second user intent information.

17. An electronic device, characterized in that, Includes: A transceiver unit and a processing unit, enabling the electronic device to execute the method according to any one of claims 1-7 or 8-16.

18. An electronic device, characterized in that, Includes: A processor, the processor is coupled to a memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the electronic device is enabled to execute the method according to any one of claims 1-7 or 8-16.

19. A computer storage medium, characterized in that, Includes computer instructions, and when the computer instructions run on a terminal device, the terminal device is enabled to execute the method according to any one of claims 1-7 or 8-16.

20. A computer program product, characterized in that, When the computer program product runs on a computer, the computer is caused to execute the method according to any one of claims 1-7 or 8-16.

Citation Information

Patent Citations

  • Method and system for implementing VOLTE (Voice Over Long Term Evolution) coloring ringback tone

    CN106027817A

  • Method for controlling video polyphonic ringtone playing and related device

    CN111049778A

  • Media resource playing method, related device and system

    CN114125163A

  • Providing communication services using sets of I / O devices

    EP3912312A1

  • Methods and systems for processing call establishment request

    US20170126901A1