Media resource playing method and related device
By sending audio information in the calling terminal device to request playback of media resources, and performing voice recognition and semantic understanding on the media platform server side, and dynamically adjusting the playback of media resources, the problem that users cannot interact in real time in the prior art is solved, and the user experience is improved.
Patent Information
- Application Number
- CN202311864928.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
The existing media resource playback method cannot realize user interaction in the calling terminal device, reducing the fun and operability of the user experience.
By sending audio information in the calling terminal device to request playback of media resources, and performing voice recognition and semantic understanding on the media platform server side, the user's intention is determined to dynamically adjust the playback of media resources.
It realizes that users can flexibly control the playback of media resources through voice commands, improving the friendliness and fun of the user experience.
Smart Images

Figure CN120238524A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of communication technology, and in particular to a media resource playback method and related devices. Background Art
[0002] With the continuous development of communication technology, high-definition voice over long term evolution (VOLTE) technology has gradually entered people's lives, and people can enjoy various types of media resource experiences. For example, users can enjoy video experience while making a voice call, such as video ringback tone and video customer service, making the waiting stage before the call more interesting and greatly improving the user's calling experience.
[0003] At present, the media resource playback method is usually based on the media resource playback in the communication technology (CT) domain, which can also be called the telecommunications domain. The corresponding process can be: the calling terminal device initiates a call to the called terminal device. When the called terminal device rings, the media server in the CT domain pulls the media resources corresponding to the user contract information from the media platform server based on the user contract information, and then after receiving the confirmation playback message from the calling terminal device, instructs the media platform server to start playing the media resources for the calling terminal device.
[0004] Since the media resources played by the calling terminal device are configured by the media platform server, the user of the calling terminal device cannot interact in real time, thus reducing the interest and operability of the user in using the media resources. Summary of the invention
[0005] In the first aspect, an embodiment of the present application proposes a method for playing media resources, which is applied to a calling terminal device, and the method includes: sending a call request to a called terminal device; sending a first audio information, wherein the first audio information is used to request playing a media resource; receiving and playing a first media resource, wherein the first media resource is determined by the first audio information.
[0006] In the embodiment of the present application, after sending a call request to the called terminal device, the calling terminal device sends a first audio message to request to play the media resource, and then the calling terminal device receives and plays the first media resource. This allows the user to flexibly control the media resource played by the terminal device through voice commands, providing the user with a more friendly and interesting user experience.
[0007] In combination with the first aspect, in a possible implementation manner of the first aspect, the method further includes: playing a first response message, where the first response message is a response message generated based on the first audio message, and the first response message includes audio information, picture information, animation information, and / or text information.
[0008] In the embodiments of the present application, a response message of the audio information can be superimposed and displayed (or played) on the screen (or speaker) of the calling terminal device. The response message includes, but is not limited to: audio information, picture information, animation information, and / or text information. Through the above method, real-time feedback of interactive information is realized, and the user experience is improved.
[0009] In combination with the first aspect, in a possible implementation manner of the first aspect, the first response message includes: first text information, where the first text information is text information generated by performing speech recognition processing on the first audio information, and the content included in the first text information corresponds to the first audio information. By feeding back the first text information corresponding to the first audio information to the calling terminal device, it is convenient for the user to operate and improves the user experience.
[0010] In combination with the first aspect, in a possible implementation manner of the first aspect, the method further includes: sending a second audio message;
[0011] Receiving a response to the second audio message, where the response to the second audio message includes: a second media resource, where the second media resource is determined by the second audio message; and / or, a second response message, where the second response message is a response message generated based on the second audio message, and the second response message includes audio information, picture information, animation information, and / or text information.
[0012] In the embodiments of the present application, the user can also continue to operate the media resource through the audio information, further improving the user experience.
[0013] In combination with the first aspect, in a possible implementation manner of the first aspect, the method further includes: stopping playing the first media resource; playing the second media resource; and / or, playing the second response message.
[0014] In the embodiments of the present application, when the calling terminal device receives the second media resource, it can stop playing the first media resource and then play the second media resource. When the calling terminal device receives the second response message, the calling terminal device can also superimpose and display (or play) the second response message on the screen (or speaker), improving the user experience.
[0015] In combination with the first aspect, in a possible implementation of the first aspect, the method further includes: receiving an off-hook message sent by the called terminal device; in response to the off-hook message, stopping playing the first media resource; stopping the reception of audio information, and / or stopping the speech recognition processing of the audio information.
[0016] In the embodiments of the present application, the calling terminal device triggers a request to play a media resource through audio information during the call stage. This media resource can be a video ringtone, which enhances the operation interest. In addition, the calling terminal device can also trigger other interactive operations related to the media resource through audio information during the call stage. These other interactive operations include, but are not limited to: subscribing to copy the media resource, or rating or liking the media resource, etc., which enhances the operation interest.
[0017] In combination with the first aspect, in a possible implementation of the first aspect, sending the first audio information includes: sending the first audio information, where the first audio information includes a wake-up keyword, and the wake-up keyword is used to trigger the media platform server to perform speech recognition processing on the first audio information; or, sending audio information including the wake-up keyword; sending the first audio information.
[0018] In the embodiments of the present application, the user can also input the wake-up keyword by voice to avoid misoperation.
[0019] In a second aspect, an embodiment of the present application provides a method for playing a media resource. The method is applied to a media platform server and includes: receiving a call request sent by the calling terminal device; receiving first audio information sent by the calling terminal device; determining a first media resource according to the first audio information; and sending the first media resource to the calling terminal device.
[0020] In the embodiments of the present application, after the media platform server receives the call request from the calling terminal device, it determines the corresponding first media resource according to the first audio information sent by the host terminal device, and then sends the first media resource to the calling terminal device. This enables the user to flexibly control the media resource played by the terminal device through voice commands, bringing a more friendly and interesting user experience to the user.
[0021] In combination with the second aspect, in a possible implementation of the second aspect, the determining the first media resource according to the first audio information includes: performing speech recognition processing on the first audio information to generate first text information, where the content included in the first text information corresponds to the first audio information; performing semantic understanding processing on the first text information to generate first user intent information; and determining the first media resource according to the first user intent information.
[0022] In the embodiments of the present application, the first user intent information corresponding to the first audio information is determined through speech recognition and semantic understanding, and then the first media resource is determined, thereby improving the recognition accuracy.
[0023] Combined with the second aspect, in a possible implementation manner of the second aspect, the method further includes: generating a first response message according to the first user intent information, where the first response message includes audio information, picture information, animation information, and / or text information; and sending the first response message to the calling terminal device.
[0024] In the embodiments of the present application, a corresponding response message may also be generated according to the first user intent information and sent to the calling terminal device. The response message of the audio information may be superimposed and displayed (or played) on the screen (or speaker) of the calling terminal device, and the response message includes, but is not limited to, audio information, picture information, animation information, and / or text information. Through the above method, real-time feedback of interactive information is realized, and the user experience is improved.
[0025] Combined with the second aspect, in a possible implementation manner of the second aspect, the first response message includes the first text information.
[0026] Combined with the second aspect, in a possible implementation manner of the second aspect, the first user intent information includes any one or more of the following: start playing a media resource, pause playing a media resource, switch the playing media resource, rewind the playing media resource, copy and order a media resource, the content feature keywords of the first audio information, or the weights of the content feature keywords.
[0027] Combined with the second aspect, in a possible implementation manner of the second aspect, determining the first media resource according to the first user intent information includes: determining the first media resource according to a decision recommendation model and the content feature keywords and / or the weights of the content feature keywords included in the first user intent information, where the decision recommendation model determines a media resource by applying a parameter set, and the parameter set includes any one or more of the following: the content feature keywords of the media resource, the weights of the content feature keywords of the media resource, the media resource tags of the media resource library, the weights of the media resource tags of the media resource library, the popularity weights of the media resources in the media resource library, the release time of the media resources in the media resource library, or the play rate of the media resources in the media resource library, where the media resource library includes one or more media resources.
[0028] In combination with the second aspect, in a possible implementation manner of the second aspect, the method further includes: detecting whether the first audio information includes a wake-up keyword; if the first audio information includes the wake-up keyword, triggering speech recognition processing based on the first audio information; or, triggering speech recognition processing based on the first audio information according to detecting that the received audio information includes the wake-up keyword.
[0029] In the embodiments of the present application, the user can also input the wake-up keyword by voice. The media platform server detects the wake-up keyword and triggers the speech recognition processing of the first audio information after the user inputs the wake-up keyword, avoiding misoperations.
[0030] In combination with the second aspect, in a possible implementation manner of the second aspect, the method further includes: receiving second audio information sent by the calling terminal device; generating a response to the second audio information according to the second audio information, where the response to the second audio information includes: a second media resource determined by the second audio information; and / or, a second response message, where the second response message is a response message generated based on the second audio information, and the second response message includes audio information, picture information, animation information, and / or text information; sending the response to the second audio information to the calling terminal device.
[0031] In combination with the second aspect, in a possible implementation manner of the second aspect, generating a response to the second audio information according to the second audio information includes: performing speech recognition processing according to the second audio information to generate second text information, where the content included in the second text information corresponds to the first audio information; performing semantic understanding processing according to the second text information to generate second user intent information; generating a response to the second audio information according to the second user intent information.
[0032] In the embodiments of the present application, the user can also continue to operate the media resource through audio information, further improving the user experience.
[0033] A third aspect of the embodiments of the present application provides an electronic device, including: a transceiver unit and a processing unit, enabling the electronic device to implement the method in the first aspect or any possible implementation manner of the first aspect, or enabling the electronic device to implement the method in the second aspect or any possible implementation manner of the second aspect.
[0034] The fourth aspect of the embodiments of the present application provides an electronic device, including: a processor, the processor is coupled to a memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the electronic device implements the methods in the above-mentioned first aspect or any possible implementation manner of the first aspect, or enables the electronic device to implement the methods in the above-mentioned second aspect or any possible implementation manner of the second aspect.
[0035] The fifth aspect of the embodiments of the present application provides a computer-readable medium, on which computer programs or instructions are stored. When the computer programs or instructions are run on a computer, the computer executes the methods in the foregoing first aspect or any possible implementation manner of the first aspect, or enables the computer to execute the methods in the foregoing second aspect or any possible implementation manner of the second aspect.
[0036] The sixth aspect of the embodiments of the present application provides a computer program product. When the computer program product is executed on a computer, the computer executes the methods in the foregoing first aspect or any possible implementation manner of the first aspect, or enables the computer to execute the methods in the foregoing second aspect or any possible implementation manner of the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a schematic diagram of a communication scenario proposed by the embodiments of the present application;
[0038] Figure 2 It is a schematic flowchart of an embodiment of a media resource playing method in the embodiments of the present application;
[0039] Figure 3 It is another schematic diagram of a communication scenario in the embodiments of the present application;
[0040] Figure 4 It is a schematic flowchart of an embodiment of a media resource playing method in the embodiments of the present application;
[0041] Figure 5 It is another schematic diagram of a communication scenario in the embodiments of the present application;
[0042] Figure 6 It is a schematic flowchart of an embodiment of a media resource playing method in the embodiments of the present application;
[0043] Figure 7 It is another schematic diagram of a communication scenario in the embodiments of the present application;
[0044] Figure 8 It is a schematic flowchart of an embodiment of a media resource playing method in the embodiments of the present application;
[0045] Figure 9 It is another schematic diagram of a communication scenario in the embodiments of the present application;
[0046] Figure 10 This is a schematic flowchart of an embodiment of a media resource playing method in an embodiment of the present application;
[0047] Figures 11 to 18 This is a schematic diagram of a scenario in an embodiment of the present application;
[0048] Figure 19 This is a schematic diagram of an application scenario in an embodiment of the present application;
[0049] Figure 20 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0050] Figure 21 Another schematic diagram of the structure of the electronic device in an embodiment of the present application;
[0051] Figure 22 Another schematic diagram of the structure of the electronic device in an embodiment of the present application. Detailed implementation manners
[0052] Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings of the present application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing embodiments of the present application. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not necessarily have to be limited to those units, but may include other units that are not clearly listed or are inherent to these processes, methods, products or devices.
[0053] Next, the technical solutions in the embodiments of the present application will be clearly described in conjunction with the accompanying drawings in the embodiments of the present application. In the description of the present application, unless otherwise specified, " / " means "or". For example, A / B may represent A or B; "and / or" in the present application is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, in the description of the present application, "at least one item" means one or more items, and "multiple items" means two or more items. "At least one of the following (items)" or its similar expressions refer to any combination of these items, including any combination of single (item) or plural (items). For example, at least one of a, b, or c may represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c may be single or multiple.
[0054] First, some terms in the embodiments of the present application are explained to facilitate understanding by those skilled in the art.
[0055] (1) Terminal device: It can be a wireless terminal device capable of receiving scheduling and indication information from a network device. The wireless terminal device can be a device that provides voice and / or data connectivity to a user, or a handheld device with a wireless connection function, or other processing devices connected to a wireless modem.
[0056] The terminal device can communicate with one or more core networks or the Internet via a radio access network (RAN). The terminal device can be a mobile terminal device, such as a mobile phone (or a "cellular" phone, mobile phone), a computer, and a data card. For example, it can be a portable, pocket-sized, handheld, computer-integrated, or vehicle-mounted mobile device that exchanges voice and / or data with the wireless access network. For example, personal communication service (PCS) phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistant (PDA), tablet (Pad), computers with wireless transceiver functions, etc. The wireless terminal device can also be referred to as a system, subscriber unit, subscriber station, mobile station, mobile station (MS), remote station, access point (AP), remote terminal device, access terminal device, user terminal device, user agent, subscriber station (SS), customer premises equipment (CPE), terminal, user equipment (UE), mobile terminal (MT), drone, etc. The terminal device can also be a wearable device and the next-generation communication system. For example, the terminal device in a 5G communication system or the terminal device in a future evolved public land mobile network (PLMN), etc.
[0057] (2) Internet Protocol Multimedia Subsystem (IMS) domain.
[0058] The CT domain realizes communication through the evolved packet core (EPC) and the core network of the Internet Protocol Multimedia Subsystem (IMS) domain, etc. The core network of the IMS domain includes several application servers (AS), such as a media platform server, which is used to provide media resource playback for terminals. For example, when providing a video ringback tone service, the media platform server is also called a video ringback tone platform. The media platform server may include a media resource application server and a media resource server (MRS). The media resource application server and the media resource server may be co-located or physically separated. The media resource server may also be called a ringback tone platform, a video ringback tone platform, or a call waiting tone platform, and is used to provide media resources such as video ringback tones, video call vibrations, video advertisements, and video customer services. For example, the media resource server produces and manages the above media resources. The media application server and the media resource server may be co-located or physically separated. The media application server processes Session Initiation Protocol (SIP) signaling messages, and the media resource server provides an audio stream and / or a video stream for the calling terminal and / or the called terminal.
[0059] In addition, the IMS domain core network further includes: a serving-call session control function (S-CSCF) device, an interrogating-call session control function (I-CSCF) device, a proxy-call session control function (P-CSCF) device, a home subscriber server (HSS) device, a session border controller (SBC) device, and several application servers, such as a telephony application server (TAS), a multimedia telephony application server (MMTelAS), a server centralization and continuity application server (SCCAS), etc. Among them, the I-CSCF device and the S-CSCF device can be co-located and can be abbreviated as the "I / S-CSCF" device. The SBC device and the P-CSCF device can be co-located and can be abbreviated as the "SBC / P-CSCF" device. The EPC may include a packet data network gateway (PGW) device, a serving gateway (SGW) device, and a mobile management entity (MME) device.
[0060] The S / P-GW device is used to provide the functions of the serving gateway and the packet data network gateway logical entities. The SGW is the anchor point for local mobility, mainly facing the radio access network for the transmission of service plane data. The P-GW is the EPS anchor point, mainly facing other data networks to realize access interaction with multiple public data networks. The SGW device can be used for the connection between the IMS core network and the wireless network, and the PGW device can be used for the connection between the IMS core network and the Internet Protocol (IP) network. The MME device is the core device of the EPC network and is used to provide the functions of the MME logical entity.
[0061] The above network devices are corresponding network devices in a wireless communication network in the prior art, which will not be described in detail here and will only be briefly introduced. For example: The HSS device can be used to store user subscription information and location information. The SBC device can provide secure access and media processing. The MMTelAS device provides basic multimedia telephone services and supplementary services. The MME device is the core device of the EPC network. The SGW device can be used for the connection between the IMS core network and the wireless network, and the PGW device can be used for the connection between the IMS core network and the IP network. The S-CSCF device can be used for user registration, authentication control, session routing, and service triggering control, and maintain session state information. The I-CSCF device can be used for the assignment and query of the S-CSCF device for user registration. The P-CSCF device can be used for signaling and message proxy. In this application, for the sake of simplicity of description, the CSCF device is used to represent any one or more combinations of the S-CSCF device, the I-CSCF device, and the P-CSCF device.
[0062] (3), Media resources.
[0063] The media resources in the embodiments of this application include, but are not limited to: audio ringtones, video ringtones, video advertisements, or video animations, etc.
[0064] Since the media resources played by the current calling terminal device are configured by the media platform server, the user of the calling terminal device cannot interact in real time, thus reducing the interest and operability of the user using the media resources.
[0065] Based on this, this application proposes a media resource playing method, which includes: sending a call request to the called terminal device; sending first audio information, where the first audio information is used to request playing media resources; receiving and playing first media resources, where the first media resources are determined by the first audio information. After the calling terminal device sends a call request to the called terminal device, it requests to play media resources by sending first audio information, and then the calling terminal device receives and plays the first media resources. This enables the user to flexibly control the media resources played by the terminal device through voice commands, bringing a more friendly and interesting user experience to the user.
[0066] The following introduces the embodiments of this application with reference to the accompanying drawings. Please refer to Figure 1 , Figure 1A schematic diagram of a communication scenario proposed in an embodiment of the present application. A communication scenario proposed in an embodiment of the present application includes: a media platform server, a calling terminal device, and a called terminal device. The calling terminal device sends a call request to the called terminal device. Then, the media platform server sends default media resources to the calling terminal device, and the calling terminal device plays the default media resources. The default media resources can be pre-ordered by the calling terminal device or pre-allocated by the media platform server for the calling terminal device. When the calling terminal device plays the default media resources, the calling terminal device can collect the user's audio information and then send the audio information to the media platform server. The media platform server performs interactions based on the audio information reported by the calling terminal device, such as switching the media resources played by the calling terminal device according to the audio information, or pausing the media resources played by the calling terminal device, or finalizing the media resources played by the current calling terminal device.
[0067] Based on the Figure 1 schematic communication scenario shown, please refer to Figure 2 , Figure 2 which is a schematic flowchart of the implementation of a media resource playback method in an embodiment of the present application. A media resource playback method proposed in an embodiment of the present application includes:
[0068] S1. The calling terminal device sends a call request to the called terminal device.
[0069] The calling terminal device sends a call request to the called terminal device, and a call negotiation is carried out between the calling terminal device and the called terminal device. After the called terminal device rings, the media platform server and the calling terminal device complete the negotiation of media resources. During the negotiation process, the transmission direction of the media resources is set to allow both upstream and downstream transmissions.
[0070] S2. The calling terminal device sends the first audio information to the media platform server.
[0071] After the negotiation is completed, the media platform server sends default media resources to the calling terminal device. The default media resources can be media resources ordered by the calling terminal device from the media platform server, or media resources ordered by the called terminal device from the media platform server, or media resources actively allocated by the media platform server to the calling terminal device or the called terminal device. Then, the calling terminal device plays the default media resources.
[0072] In the above process, the media platform server enables the voice interaction service. This voice interaction service allows users to send voice commands to the media platform server through a terminal device (for example, carrying the voice command through audio information), and then complete the interaction according to the voice command. In one possible implementation, the calling terminal device sends the first audio information to the media platform server, and the first audio information is used to request the playback of media resources.
[0073] In one example, the calling terminal device requests to switch the playback of media resources through the first audio information. The first audio information can be a request to randomly switch to play the next media resource, or the first audio information can also clearly indicate which type of media resource to switch to. Please refer to Figure 11 , Figure 11 which is a schematic diagram of a scenario in an embodiment of this application. In one example, the flat surface of the calling terminal device displays the first text information corresponding to the first audio information, and the first text information is the text information obtained by the media platform server through voice recognition based on the first audio information. The first text information corresponding to the first audio information includes: "Xiaocai Xiaocai, change one", and this first audio information is used to request to replace the currently played default media resource. In another example, please refer to Figure 12 , Figure 12 which is another schematic diagram of a scenario in an embodiment of this application. The first text information corresponding to the first audio information includes: "Xiaocai Xiaocai, I want to watch animals", and this first audio information is used to request to replace the currently played default media resource and play media resources related to animals.
[0074] In another example, please refer to Figure 13 , Figure 13 which is another schematic diagram of a scenario in an embodiment of this application. The first text information corresponding to the first audio information includes: "Xiaocai Xiaocai", and this first audio information is used to trigger the media platform server to perform voice recognition processing on the first audio information. "Xiaocai Xiaocai" is used as the wake-up keyword to wake up the voice interaction service.
[0075] In another possible implementation, the calling terminal device sends the first audio information to the media platform server, and the first audio information is used to perform interaction processing on the media resource currently played by the calling terminal device. This interaction processing includes but is not limited to: stopping the playback of the currently played media resource, copying and subscribing to the currently played media resource, or sharing the currently played media resource with other users. In another example, please refer to Figure 14 , Figure 14 which is another schematic diagram of a scenario in an embodiment of this application. The first text information corresponding to the first audio information includes: "Xiaocai Xiaocai, copy this ringtone for me", and this first audio information is used to trigger the media platform server to copy and subscribe to the media resource currently played by the calling terminal device for the calling terminal device.
[0076] In another possible implementation, when the media platform server fails to recognize and process the first audio information, the media platform server may send a prompt message to the calling terminal device, and the prompt message indicates that the media platform server fails to recognize and process the first audio information. Exemplarily, please refer to Figure 15 , Figure 15 , which is another scenario schematic diagram in the embodiments of the present application. The first audio information includes: "Xiaocai, Xiaocai, it's raining today". After the media platform server fails to recognize and process the first audio information, it sends a prompt message to the calling terminal device. The screen of the calling terminal device displays the prompt message, and the prompt message includes "Sorry, I didn't understand. What do you want to instruct Xiaocai to do?". Subsequently, the calling terminal device may continue to collect the user's audio information, for example, the second audio information, and then the calling terminal device sends the second audio information to the media platform server, and the media platform server performs voice interaction processing based on the second audio information.
[0077] S3. The media platform server detects whether the first audio information includes a wake-up keyword.
[0078] The media platform server prevents the media platform server from mistakenly triggering voice interaction processing of the user's voice by detecting whether the user's voice (audio information) of the calling terminal device includes a wake-up keyword. When the media platform server detects that the user's voice (audio information) of the calling terminal device includes a wake-up keyword, subsequent processing is performed on the first audio information, such as speech recognition and semantic understanding.
[0079] In a possible implementation, when the media platform server receives the first audio information, the media platform server detects whether the first audio information includes a wake-up keyword, and the wake-up keyword is used to trigger the media platform server to perform speech recognition processing on the first audio information. If the first audio information includes a wake-up keyword, step S4 is entered. Exemplarily, the media platform server uses technical methods such as keyword spotting (KWS) to identify whether the first audio information includes a wake-up keyword, and the wake-up keyword may also be referred to as a wake-up word.
[0080] In another possible implementation, after step S1, the media platform server collects the audio information of the calling terminal device and identifies whether the audio information includes a wake-up keyword. If a wake-up keyword is included, the first audio information is continuously collected and then step S4 is executed.
[0081] In another possible implementation, the media platform server may also send a voice interaction prompt message to the calling terminal device, and the calling terminal device displays the voice interaction prompt message on the screen of the terminal device in a real-time overlay manner. Exemplarily, please refer to Figure 16, Figure 16 This is another schematic diagram of a scenario in the embodiment of the present application. The voice interaction prompt information is superimposed and displayed on the plane of the calling terminal device. The voice interaction prompt information includes: "Hi, hello. I'm the intelligent voice assistant 'Xiaocai'. You can say to me 'Xiaocai Xiaocai, change one'". The calling terminal device continues to receive the audio information input by the user.
[0082] If the media platform server does not detect that the first audio information includes a wake-up keyword, or the media platform server does not detect that the audio information of the calling terminal device includes a wake-up keyword, it is determined that the currently input audio information of the calling terminal device is not for voice interaction processing related to media resources, and the media platform server does not execute
[0083] S4. The media platform server performs speech recognition processing on the first audio information to generate first text information.
[0084] After receiving the first audio information, the media platform server performs speech recognition processing on the first audio information to generate first text information. Exemplarily, the media platform server recognizes the first audio information (i.e., the user voice of the calling terminal device) based on Automatic Speech Recognition (ASR) technology and converts the first audio information into first text information.
[0085] Optionally, the media platform server may send the first text information to the calling terminal device, and the first text information is superimposed and displayed on the screen of the calling terminal device. The first text information is as described above Figures 11 to 15 shown.
[0086] S5. The media platform server performs semantic understanding processing on the first text information to generate first user intention information.
[0087] After receiving the first text information, the media platform server performs semantic understanding processing on the first text information to generate first user intention information. The first user intention information includes any one or more of the following: start playing media resources, pause playing media resources, switch playing media resources, rewind playing media resources, copy and order media resources, content feature keywords of the first audio information, or weights of the content feature keywords. Exemplarily, artificial intelligence (AI) models can be used for semantic understanding.
[0088] After step S5, step S6 and step S9 are executed.
[0089] S6. The media platform server determines a first media resource according to the first user intention information.
[0090] The media platform server determines the corresponding first media resource according to the first user intent information. In a possible implementation, if the first user intent information requests to switch the played media resource, the media platform server determines the media resource to be played from one or more candidate media resources as the first media resource, and the first media resource may include one or more media resources.
[0091] Specifically, the media platform server determines the media resource that conforms to the content feature keyword and / or the weight of the content feature keyword according to the content feature keyword included in the first user intent information and / or the weight of the content feature keyword as the first media resource. Combining Figure 12 with the example, the content feature keyword included in the first user intent information is "animal", and according to this content feature keyword, the media platform server determines the media resource including animal pictures from one or more candidate media resources as the first media resource.
[0092] Exemplarily, the media platform server determines the first media resource according to the decision recommendation model (or decision recommendation algorithm, or decision algorithm) and the content feature keyword and / or the weight of the content feature keyword included in the first user intent information, where the decision recommendation model determines the media resource by applying a parameter set, and the parameter set includes any one or more of the following: the content feature keyword of the media resource, the weight of the content feature keyword of the media resource, the media resource label of the media resource library, the weight of the media resource label of the media resource library, the popularity weight of the media resources in the media resource library, the release time of the media resources in the media resource library, or the play rate of the media resources in the media resource library, where the media resource library includes one or more media resources.
[0093] Optionally, if the first user intent information determined by the media platform server requests to switch the played media resource and the first user intent information does not include a content feature keyword, the media platform server may randomly select a media resource as the first media resource for the calling terminal device according to the first user intent information.
[0094] S7. The media platform server sends the first media resource to the calling terminal device.
[0095] S8. The calling terminal device plays the first media resource.
[0096] After receiving the first media resource, the calling terminal device stops playing the default media resource and then plays the first media resource.
[0097] S9. The media platform server generates the first response information according to the first user intent information.
[0098] The media platform server generates a first response message according to the first user intent information, and the first response message includes audio information, picture information, animation information, and / or text information. The first response message may include first text information generated by performing speech recognition on the first audio information. The first response message may also include a temporary response message generated by performing preliminary processing on the first text information. For example, Figure 13 the first response message in
[0099] includes: "Xiaocai Xiaocai" and "Here, please speak, I'm listening...". The first response message may further include an interactive response based on the first user intent information. For example, the first response message includes a start-playing prompt: "Master, I have switched to playing the video for you", the first response message includes a processing-waiting prompt: "Processing, please wait a moment, master", the first response message includes a voice-waiting prompt: "Here, please speak, I'm listening", the first response message includes an interactive-result prompt: "Master, the copy and download for you have been successful", or the first response message includes an intent-not-understood prompt: "Sorry, I didn't understand".
[0100] For ease of understanding, an example is given in combination with the accompanying drawings. For example, Figure 11 the first response message in Figure 12 includes: "I have switched the video for you". Another example, Figure 13 the first response message in Figure 14 includes: "Here, please speak, I'm listening...". Another example, Figure 15 the first response message in
[0101] S10. The media platform server sends the first response message to the calling terminal device.
[0102] S11. The calling terminal device plays the first response message.
[0103] After receiving the first response message, the calling terminal device superimposes and plays the first response message on the interface for playing the first media resource or the default media resource.
[0104] It should be noted that the application embodiment does not limit the execution order of step S8 and step S11. It may be that step S8 is executed first and then step S11, or step S11 is executed first and then step S8, or step S8 and step S11 may be executed simultaneously. For example, the screen of the calling terminal device superimposes and displays the first response message while playing the first media resource.
[0105] S12. The calling terminal device sends the second audio information to the media platform server.
[0106] Steps S12 - S15 are optional steps.
[0107] After the calling terminal device finishes sending the first audio information to the media platform server, the calling terminal device may also send the second audio information to the media platform server.
[0108] S13. The media platform server determines the second media resource and / or generates the second response information according to the second audio information.
[0109] In step S13, the processing method of the media platform server for the second audio information is similar to the processing method of the media platform server for the first audio information in the foregoing steps S2 - S10, which will not be elaborated here.
[0110] Exemplarily, when the second audio information is used to indicate the media resource to be switched for playback, the media platform service determines the second media resource to be switched according to the second user intention information corresponding to the second audio information.
[0111] It can be understood that when the user of the calling terminal device conducts multiple voice interactions, the media platform server respectively uses a similar processing method for the first audio information for multiple voice interactions.
[0112] S14. The media platform server sends a response to the second audio information to the calling terminal device, including: the second media resource and / or the second response information.
[0113] Regarding the second media resource, it is similar to the first media resource, and regarding the second response information, it is similar to the first response information. The second response information is the response information generated based on the second audio information, and the second response information includes audio information, picture information, animation information, and / or text information.
[0114] S15. The calling terminal device plays the second media resource and / or the second response information.
[0115] In step S15, the calling terminal device stops playing the first media resource and / or the first response information. Then, the calling terminal device plays the second media resource and / or the second response information.
[0116] In the embodiments of the present application, after the calling terminal device sends a call request to the called terminal device, it requests to play a media resource by sending the first audio information, and then the calling terminal device receives and plays the first media resource. This enables the user to flexibly control the media resource played by the terminal device through voice commands, bringing a more friendly and interesting user experience to the user.
[0117] Combined with the foregoing embodiments, the following describes another communication scenario of the embodiments of the present application. Please refer to Figure 3 , Figure 3 which is a schematic diagram of another communication scenario in the embodiments of the present application. In the embodiments of the present application, the calling terminal device is connected to the IMS core network through the access network, and then is connected to the media platform server through the IMS core network; the called terminal device is connected to the IMS core network through the access network, and then is connected to the media platform server through the IMS core network. It can be understood that the calling terminal device and / or the called terminal device may be connected to the media platform server in other ways, which is not limited in the embodiments of the present application.
[0118] The media platform server may include one or more logical network elements. In actual implementation, the media platform server may be deployed on a unified physical network element to implement multiple logical network elements, or may implement multiple logical network elements on multiple physical network elements according to internal characteristics. Taking the media resource as the video ringback tone media (or called video ringback tone, or ringback tone) as an example for illustration. In a possible implementation manner, the media platform server includes any one or more of the following logical network elements: video ringback tone service operation and management service, artificial intelligence model, video ringback tone intelligent semantic understanding service, video ringback tone interactive response processing service, video ringback tone playback decision service, video ringback tone intelligent interactive service, video ringback tone application service, or video ringback tone media service. The following is a specific description:
[0119] 1. The video ringback tone application service may also be referred to as the signaling access and processing unit of the video ringback tone. In the ringing stage of the call of the video ringback tone user, it realizes logical functions such as media negotiation and media playback control of the video ringback tone. The media ringback tone application service supports calling the playback decision service according to the user intention information to obtain the media resources that need to be switched for playback, and notifying the calling terminal device to switch to play new media resources and overlay and display the response information of the interactive feedback, etc.
[0120] 2. The video ringback tone media service is used to realize the media playback of the video ringback tone in the ringing stage of the call, and send the video ringback tone content to the calling terminal device for playback in the form of an audio-visual media stream through the network. The video ringback tone media service supports the uplink processing and intelligent recognition of the audio information (user voice) of the calling terminal device, and real-time identification of the wake-up keywords of the audio information. In response to detecting the wake-up keywords, it generates corresponding text information for the audio information of the calling terminal device, and supports overlaying and playing the text information on the video ringback tone media played on the screen of the calling terminal device to realize interactive real-time feedback.
[0121] 3. The video ringback tone playback decision service is used to select and decide the specific video ringback tone to be played for the user according to the video ringback tones subscribed by the calling terminal device user (i.e., the user of the calling terminal device) in each call. It selects the corresponding video ringback tone content in real time according to the content feature keywords of the user intention information. When the semantic meaning of the audio information of the calling terminal device hopes to play a certain type of characteristic video ringback tone, it makes a real-time decision for the calling terminal device on the video ringback tone content that conforms to the user intention information based on the decision algorithm.
[0122] 4. The video ringback tone intelligent interaction service is used to accurately understand the semantics of the audio information of the calling terminal device, obtain the user interaction intention of the calling terminal device, and perform corresponding interaction responses.
[0123] 5. The artificial intelligence model is used to pre-train the video ringback tone intelligent interaction service to achieve accurate semantic understanding and interaction feedback of the text information generated by the calling terminal device based on the audio information in the video ringback tone voice interaction scenario.
[0124] 6. The video ringback tone service operation and management service is used to provide the service management of video ringback tones for users, including opening the video material function and subscribing to video ringback tone sounds, etc.
[0125] Combined with Figure 3 the schematic media platform server, please refer to Figure 4 , Figure 4 which is the schematic flowchart of the implementation process of a media resource playback method in an embodiment of this application. When the media platform server includes: the video ringback tone application service, the video ringback tone media service, the video ringback tone playback decision service, and the video ringback tone intelligent interaction service, a media resource playback method proposed in an embodiment of this application includes:
[0126] D1. The calling terminal device sends a call request to the called terminal device.
[0127] D2. Call negotiation and media resource (ringback tone) negotiation are carried out between the calling terminal device and the called terminal device.
[0128] D3. The video ringback tone application service notifies the video ringback tone media service to play the media resource, carrying the voice recognition identifier.
[0129] In step D3, in a possible implementation manner, the video ringback tone application service sends voice recognition identifier information to the video ringback tone media service, and this voice recognition identifier information instructs the video ringback tone media service to turn on the reception of the uplink audio media stream and the voice recognition of the wake-up keyword.
[0130] In another possible implementation, the video ringback tone application service sends a voice interaction prompt message to the video ringback tone media service. The voice interaction prompt message is used to notify the calling terminal device to display the voice interaction prompt message, and the voice interaction prompt message includes prompting the user to input a wake-up keyword. The calling terminal device displays the voice interaction prompt message in a real-time overlay manner on the interface for playing the default media resource.
[0131] It should be noted that the video ringback tone application service sends a voice recognition identification message and a voice interaction prompt message to the video ringback tone media service.
[0132] D4. The video ringback tone media service sends the default media resource and the voice interaction prompt message to the calling terminal device.
[0133] In step D4, the default media resource is played on the screen of the calling terminal device, and the voice interaction prompt message is played in a real-time overlay manner.
[0134] D5. The calling terminal device sends the first audio message to the video ringback tone media service.
[0135] D6. The video ringback tone media service detects whether the first audio message includes a wake-up keyword.
[0136] The video ringback tone media service analyzes the first audio message of the calling terminal device, and uses technical methods such as keyword recognition to identify whether the first audio message includes a wake-up keyword. If it includes a wake-up keyword, it proceeds to step D7; if not, it does nothing or feeds back a prompt indicating that the intention is not understood to the calling terminal device, for example: "Sorry, I didn't understand."
[0137] D7. The video ringback tone media service performs voice recognition processing on the first audio message to generate the first text message.
[0138] Exemplarily, the video ringback tone media service recognizes the first audio message based on automatic speech recognition technology and converts the first audio message into the first text message.
[0139] D8. The video ringback tone media service sends the first text message and a temporary response message to the calling terminal device.
[0140] The video ringback tone media service superimposes the first text message and the corresponding temporary response message onto the video ringback tone media stream (default media resource) and sends them to the calling terminal device. The first text message and the temporary response message are played in a real-time overlay manner on the screen of the calling terminal device.
[0141] D9. The video ringback tone media service reports the first text message to the video ringback tone intelligent interaction service.
[0142] D10. The video ringback tone intelligent interaction service performs semantic understanding processing on the first text information to generate the first user intention information, and determines the first response information according to the first user intention information.
[0143] D11. The video ringback tone intelligent interaction service sends an interaction processing request to the video ringback tone application service. The interaction processing request is used to request the playback of a new media resource, and the interaction processing request includes the first user intention information, the first response information, and / or the first text information.
[0144] The video ringback tone intelligent interaction service sends an interaction processing request to the video ringback tone application service (or the video ringback tone voice management service) based on the first user intention information. The interaction processing request is used to request the media ringback tone application service to perform corresponding service processing on the first user intention information. The interaction processing request includes the first user intention information, the first response information, and / or the first text information.
[0145] Exemplarily, the first user intention information includes, but is not limited to: interaction response instructions, such as switching the playback of media resources, reverting to the previous media resource for playback, or copying and subscribing to media resources, etc., or, content feature keywords (the content feature keywords can also be referred to as video content tags) and the weights of the content feature keywords, etc.
[0146] D12. The video ringback tone application service sends a media resource query request to the video ringback tone playback decision service, and the request carries the first user intention information.
[0147] After the video ringback tone application service determines that the first user intention information is to switch the media resource for playback, the video ringback tone application service sends a media resource query request to the video ringback tone playback decision service, and the request carries the first user intention information. For example, the request includes: content feature keywords and the weights of the content feature keywords, etc.
[0148] D13. The video ringback tone playback decision service determines the first media resource according to the media resource query request.
[0149] The video ringback tone playback decision service processes the media resource query request based on a specified decision recommendation model, and decides to select the media resource (i.e., the video ringback tone) that best matches the corresponding content feature keywords, and returns it to the video ringback tone application service. The decision recommendation model determines the media resource by applying a parameter set, and the parameter set includes any one or more of the following: content feature keywords of the media resource, weights of the content feature keywords of the media resource, media resource tags in the media resource library, weights of the media resource tags in the media resource library, popularity weights of the media resources in the media resource library, release times of the media resources in the media resource library, or, playback rates of the media resources in the media resource library, where the media resource library includes one or more media resources.
[0150] D14. The video ringback tone playback decision service sends the identification information of the first media resource to the video ringback tone application service.
[0151] D15. The video ringback tone application service notifies the video ringback tone media service to switch to play the first media resource and play the first response message.
[0152] D16. The video ringback tone media service sends the first media resource and the first response message to the calling terminal device.
[0153] The calling terminal device plays the first media resource and the first response message on the screen.
[0154] D17. The called terminal device picks up the phone.
[0155] D18. The video ringback tone application service instructs the video ringback tone media service to stop playing the media resource and stop receiving sound.
[0156] D19. The calling terminal device and the called terminal device renegotiate to connect the call.
[0157] In the embodiments of the present application, interaction is realized based on intelligent voice in the video ringback tone playback scenario. During the video ringback tone playback process, real-time interaction based on the user's voice is realized, bringing a more friendly and interesting user experience to the user. Since this solution does not depend on the terminal and the network and can be realized relying on the video ringback tone platform side, it is convenient for promotion and improves the application scope of the service.
[0158] Combined with the foregoing embodiments, another communication scenario of the embodiments of the present application will be introduced next. Please refer to Figure 5 , Figure 5 which is a schematic diagram of another communication scenario in the embodiments of the present application. In another possible implementation manner, the media platform server in the embodiments of the present application includes any one or more of the following logical network elements: video ringback tone playback decision service, video ringback tone intelligent interaction service, video ringback tone application service, or video ringback tone media service.
[0159] Combined with Figure 5 the media platform server shown, please refer to Figure 6 , Figure 6 which is a schematic flowchart of an embodiment of a media resource playback method in the embodiments of the present application. When the media platform server includes: video ringback tone application service, video ringback tone media service, video ringback tone playback decision service, or video ringback tone intelligent interaction service, a media resource playback method proposed in the embodiments of the present application includes:
[0160] F1. The calling terminal device sends a call request to the called terminal device.
[0161] F2. Call negotiation and media resource (ringback tone) negotiation are carried out between the calling terminal device and the called terminal device.
[0162] F3. The video ringback tone application service notifies the video ringback tone media service to play the media resource, carrying a voice recognition identifier.
[0163] F4. The video ringback tone media service sends the default media resource and voice interaction prompt information to the calling terminal device.
[0164] Correspondingly, the calling terminal device plays the default media resource (i.e., the default video ringback tone), and at the same time, the voice interaction prompt information is superimposed and displayed on the default video ringback tone. The voice interaction prompt information is, for example Figure 15 as shown.
[0165] F5. The calling terminal device sends the first audio information to the video ringback tone media service.
[0166] F6. The video ringback tone media service detects whether the first audio information includes a wake-up keyword.
[0167] Optionally, if the wake-up keyword is included, the current call is marked as the voice wake-up state, and the complete audio information of the calling terminal device is voice-recognized and converted into text information.
[0168] F7. The video ringback tone media service performs voice recognition processing on the first audio information to generate the first text information.
[0169] F8. The video ringback tone media service sends the first text information and a temporary response message to the calling terminal device.
[0170] Correspondingly, the calling terminal device plays the default media resource (i.e., the default video ringback tone), and at the same time, the temporary response message is superimposed and displayed on the default video ringback tone. The temporary response message is, for example Figure 17 as shown, Figure 17 which is another scenario schematic diagram in the embodiments of the present application. The temporary response message includes "Processing, please wait a moment, master...".
[0171] F9. The video ringback tone media service reports the first text information to the video ringback tone intelligent interaction service.
[0172] F10. The video ringback tone intelligent interaction service performs semantic understanding processing on the first text information to generate the first user intention information, and determines the first response information according to the first user intention information.
[0173] F11. The video ringback tone intelligent interaction service sends an interaction processing request to the video ringback tone application service. The interaction processing request is used to request to play a new media resource, and the interaction processing request includes the first user intention information, the first response information, and / or the first text information.
[0174] F12. The video ringback tone application service sends a query media resource request to the video ringback tone playing decision service, and this request carries the first user intention information.
[0175] F13. The video ringback tone playing decision service determines the first media resource according to the query media resource request.
[0176] F14. The video ringback tone playing decision service sends the identification information of the first media resource to the video ringback tone application service.
[0177] In another possible implementation manner, the video ringback tone application service selects the next video ringback tone (this next video ringback tone is used as the first media resource) according to the content in the existing playing rule list, and notifies the video ringback tone media service to switch to play the first media resource.
[0178] F15. The video ringback tone application service notifies the video ringback tone media service to switch to play the first media resource and play the first response information.
[0179] F16. The video ringback tone media service sends the first media resource and the first response information to the calling terminal device.
[0180] Correspondingly, the first media resource is played on the screen of the calling terminal device, and the first response information is superimposed and displayed. This first response information is, for example Figure 18 as shown, Figure 18 which is another scenario schematic diagram in the embodiments of the present application. The first response information includes "Master, the video has been switched to play for you".
[0181] F17. The calling terminal device sends the second audio information to the video ringback tone media service.
[0182] F18. The video ringback tone media service detects whether the second audio information includes a wake-up keyword.
[0183] Optionally, step F18 may not be executed, and step F19 may be directly executed because the current call has been marked as the voice wake-up state in step F6.
[0184] F19. The video ringback tone media service performs speech recognition processing on the second audio information to generate the second text information.
[0185] F20. The video ringback tone media service sends the second text information and the temporary response information to the calling terminal device.
[0186] F21. The video ringback tone media service reports the second text information to the video ringback tone intelligent interaction service.
[0187] The video ringback tone intelligent interaction service performs semantic understanding processing on the second text information to generate second user intention information, and determines a second response message according to the second user intention information.
[0188] F23. The video ringback tone intelligent interaction service sends an interaction processing request to the video ringback tone application service. The interaction processing request is used to request the playback of a new media resource, and the interaction processing request includes the second user intention information, the second response message, and / or the second text information.
[0189] F24. The video ringback tone application service sends a query media resource request to the video ringback tone playback decision service, and the request carries the second user intention information.
[0190] F25. The video ringback tone playback decision service determines a second media resource according to the query media resource request.
[0191] The video ringback tone AS selects the next one according to the content in the existing playback rule list, notifies the media service to switch the playback, and superimposes and displays the voice interaction result response message on the video for the end user.
[0192] F26. The video ringback tone playback decision service sends the identification information of the second media resource to the video ringback tone application service.
[0193] F27. The video ringback tone application service notifies the video ringback tone media service to switch the playback to the second media resource and play the second response message.
[0194] F28. The video ringback tone media service sends the second media resource and the second response message to the calling terminal device.
[0195] Correspondingly, after receiving the second media resource, the calling terminal device stops playing the first media resource, then plays the second media resource and superimposes and displays the second response message.
[0196] F29. The called terminal device picks up the phone.
[0197] F30. The video ringback tone application service instructs the video ringback tone media service to stop playing the media resource and stop receiving audio.
[0198] F31. The calling terminal device and the called terminal device renegotiate to connect the call.
[0199] Combined with the foregoing embodiments, another communication scenario of the embodiments of the present application will be introduced next. Please refer to Figure 7 , Figure 7This is another schematic diagram of a communication scenario in the embodiments of the present application. In another possible implementation, the media platform server in the embodiments of the present application includes any one or more of the following logical network elements: video ringback tone playback decision service, video ringback tone application service, video ringback tone intelligent interaction service, or video ringback tone media service. Among them, the video ringback tone intelligent interaction service includes: video ringback tone interaction response processing service and video ringback tone intelligent semantic understanding service, and the video ringback tone media service includes: video ringback tone interactive voice recognition service and video ringback tone media playback service.
[0200] Combined with Figure 7 the indicated media platform server, please refer to Figure 8 , Figure 8 This is a schematic flowchart of the implementation of a media resource playback method in the embodiments of the present application. When the media platform server includes: video ringback tone playback decision service, video ringback tone application service, video ringback tone intelligent interaction service, or video ringback tone media service, a media resource playback method proposed in the embodiments of the present application includes:
[0201] G1. The calling terminal device sends a call request to the called terminal device.
[0202] G2. Call negotiation and media resource (ringback tone) negotiation are carried out between the calling terminal device and the called terminal device.
[0203] G3. The video ringback tone application service identifies whether to allow the calling terminal device to perform intelligent voice interaction.
[0204] G4. The video ringback tone application service notifies the video ringback tone media playback service to play the media resource, carrying a voice recognition identifier.
[0205] G5. The video ringback tone media playback service sends the default media resource and voice interaction prompt information to the calling terminal device.
[0206] G6. The calling terminal device sends the fourth audio information (without carrying a wake-up keyword) to the video ringback tone media playback service.
[0207] G7. The video ringback tone media playback service identifies that the fourth audio information does not carry a wake-up keyword.
[0208] G8. The calling terminal device sends the third audio information (carrying a wake-up keyword) to the video ringback tone media playback service.
[0209] G9. The video ringback tone media playback service identifies that the third audio information carries a wake-up keyword.
[0210] G10. The video ringback tone media playback service requests to wake up the voice recognition service from the video ringback tone interactive voice recognition service.
[0211] When the video ringback tone media playback service analyzes the voice media stream (the third audio information) of the calling terminal device and recognizes the wake-up keyword, it notifies the video ringback tone voice recognition service to start the service and uploads the third audio information to the video ringback tone voice recognition service. Then, this call is marked as the voice wake-up state.
[0212] G11. The video ringback tone media playback service sends the third audio information to the video ringback tone interactive voice recognition service.
[0213] G12. The video ringback tone interactive voice recognition service performs voice recognition processing on the third audio information to generate the third text information.
[0214] G13. The video ringback tone interactive voice recognition service sends the third text information and the temporary response information to the calling terminal device.
[0215] G14. The video ringback tone interactive voice recognition service reports the third text information to the video ringback tone intelligent semantic understanding service.
[0216] G15. The video ringback tone intelligent semantic understanding service performs semantic understanding processing on the third text information to generate the third user intent information.
[0217] G16. The video ringback tone intelligent semantic understanding service sends the third user intent information to the video ringback tone interactive response processing service.
[0218] G17. The video ringback tone interactive response processing service determines the third response information according to the third user intent information.
[0219] G18. The video ringback tone interactive response processing service sends a first interactive processing request to the video ringback tone application service. The first interactive processing request includes the third user intent information, the third response information, and / or the third text information. The third response information includes a waiting voice prompt.
[0220] G19. The video ringback tone application service sends the third response information to the video ringback tone media playback service. The waiting voice prompt carried by the third response information is used to notify the waiting user of the voice interaction input.
[0221] G20. The video ringback tone media playback service sends the third response information to the calling terminal device.
[0222] Since the third audio information only includes the wake-up keyword, the video ringback tone media playback service sends the third response information to the calling terminal device. The waiting voice prompt carried by the third response information is used to notify the waiting user of the voice interaction input. This waiting voice prompt is, for example Figure 13 an example.
[0223] G21. The calling terminal device sends the first audio information to the video ringback tone media playback service.
[0224] G22. The video ringback tone media playback service sends the first audio information to the video ringback tone interactive voice recognition service.
[0225] G23. The video ringback tone interactive voice recognition service performs voice recognition processing on the first audio information to generate the first text information.
[0226] G24. The video ringback tone interactive voice recognition service sends the first text information and the temporary response information to the calling terminal device.
[0227] G25. The video ringback tone interactive voice recognition service reports the first text information to the video ringback tone intelligent semantic understanding service.
[0228] G26. The video ringback tone intelligent semantic understanding service performs semantic understanding processing on the first text information to generate the first user intent information.
[0229] G27. The video ringback tone intelligent semantic understanding service sends the first user intent information to the video ringback tone interactive response processing service.
[0230] G28. The video ringback tone interactive response processing service determines the first response information according to the first user intent information.
[0231] G29. The video ringback tone interactive response processing service sends a second interactive processing request to the video ringback tone application service. The second interactive processing request includes the first user intent information, the first response information, and / or the first text information. The first user intent information includes switching to play a media resource.
[0232] G30. The video ringback tone application service notifies the video ringback tone media playback service to switch to play the first media resource.
[0233] G31. The video ringback tone media playback service sends the first media resource and the first response information to the calling terminal device.
[0234] G32. The called terminal device picks up the phone.
[0235] G33. The video ringback tone application service instructs the video ringback tone media playback service to stop playing the media resource and stop receiving audio.
[0236] G34. The calling terminal device and the called terminal device renegotiate to connect the call.
[0237] Combined with the foregoing embodiments, next, another communication scenario of the embodiments of the present application will be introduced. Please refer to Figure 9 , Figure 9This is another schematic diagram of a communication scenario in the embodiments of the present application. In another possible implementation manner, the media platform server in the embodiments of the present application includes any one or more of the following logical network elements: video ringback tone operation management service, or video ringback tone call playback platform, where the video ringback tone call playback platform includes: video ringback tone intelligent interaction service, video ringback tone playback decision-making service, and video ringback tone application service.
[0238] Combined with Figure 9 the media platform server shown in the schematic diagram, please refer to Figure 10 , Figure 10 This is a schematic flowchart of an embodiment of a media resource playback method in the embodiments of the present application. When the media platform server includes: video ringback tone operation management service and video ringback tone call playback platform, a media resource playback method proposed in the embodiments of the present application includes:
[0239] H1. The calling terminal device sends a call request to the called terminal device.
[0240] H2. Call negotiation and media resource (ringback tone) negotiation are performed between the calling terminal device and the called terminal device.
[0241] H3. The video ringback tone application service identifies whether to allow the calling terminal device to perform intelligent voice interaction.
[0242] H4. The video ringback tone application service notifies the video ringback tone media service to play the media resource, carrying a voice recognition identifier.
[0243] H5. The video ringback tone media service sends the default media resource and voice interaction prompt information to the calling terminal device.
[0244] H6. The calling terminal device sends the first audio information (carrying a wake-up keyword and a ringback tone copy subscription instruction) to the video ringback tone media service.
[0245] H7. The video ringback tone media service detects whether the first audio information includes a wake-up keyword.
[0246] H8. The video ringback tone media service performs voice recognition processing on the first audio information to generate the first text information.
[0247] H9. The video ringback tone media service sends the first text information and a temporary response message to the calling terminal device.
[0248] H10. The video ringback tone media service reports the first text information to the video ringback tone intelligent interaction service.
[0249] H11. The video ringback tone intelligent interaction service performs semantic understanding processing on the first text information to generate the first user intention information, and determines the first response information according to the first user intention information.
[0250] H12. The video ringback tone intelligent interaction service sends the first user intent information to the video ringback tone operation and management service.
[0251] H13. The video ringback tone operation and management service copies and subscribes to the currently played media resource for the calling terminal device according to the first user intent information.
[0252] H14. The video ringback tone intelligent interaction service sends an interaction processing request to the video ringback tone application service. The interaction processing request carries the first response information, and the first response information includes the media resource subscription result.
[0253] H15. The video ringback tone application service notifies the video ringback tone media service to play the first response information.
[0254] H16. The video ringback tone media service sends the first response information to the calling terminal device.
[0255] Correspondingly, the screen of the calling terminal device plays the first response information. The first response information is as follows, for example Figure 14 as shown.
[0256] H17. The calling terminal device sends the second audio information to the video ringback tone media service.
[0257] H18. The video ringback tone media service detects whether the second audio information includes a wake-up keyword.
[0258] H19. The video ringback tone media service performs speech recognition processing on the second audio information to generate the second text information.
[0259] H20. The video ringback tone media service sends the second text information and the temporary response information to the calling terminal device.
[0260] H21. The video ringback tone media service reports the second text information to the video ringback tone intelligent interaction service.
[0261] H22. The video ringback tone intelligent interaction service performs semantic understanding processing on the second text information to generate the second user intent information, and determines the second response information according to the second user intent information. The second response information indicates that the second audio information is not understood.
[0262] H23. The video ringback tone intelligent interaction service sends an interaction processing request to the video ringback tone application service. The interaction processing request includes the second user intent information, the second response information, and / or the second text information. The second response information indicates that the second audio information is not understood.
[0263] H24. The video ringback tone application service notifies the video ringback tone media service to play the second response information.
[0264] H25. The video ringback tone media service sends a second response message to the calling terminal device.
[0265] Correspondingly, the second response message is played on the screen of the calling terminal device, and the first response message is as follows Figure 16 as shown.
[0266] H26. The called terminal device picks up the call.
[0267] H27. The video ringback tone application service instructs the video ringback tone media service to stop playing the media resource and stop voice recording.
[0268] H28. The calling terminal device and the called terminal device renegotiate to connect the call.
[0269] Combined with the foregoing embodiments, the application scenarios of the method provided in the embodiments of the present application will be described next. The application scenarios of the method provided in the embodiments of the present application can be as Figure 19 shown. Figure 19 This is a schematic diagram of an application scenario in the embodiments of the present application. This application scenario includes: User 001 and electronic device 002.
[0270] Among them, User 001: The user can interact with the electronic device 002 by means of gestures / voices, etc., to open an application program (or a function module in the application program, or a small program, a quick application, etc.) on the electronic device.
[0271] Electronic device 002: It is equipped with an operating system, and a system-level APP is built in the operating system. The user can also install / uninstall the APP according to his own needs. The electronic device 002 has a screen for displaying to the user 001, and the user 001 can operate the application program on the electronic device 002 through the screen. The electronic device 002 usually has a relatively large screen, such as a tablet computer, a folding screen mobile phone, etc.
[0272] In addition to the tablet computer (pad) or folding screen mobile phone mentioned above, the electronic device in the embodiments of the present application can also be a non-folding screen mobile phone, a smart watch, smart glasses, a smart bracelet, a portable game console, a personal digital assistant (PDA), a notebook computer, an ultramobile personal computer (UMPC), a handheld computer, a netbook, an in-vehicle media playback device, a wearable electronic device (such as: a watch, a bracelet, glasses), a virtual reality (VR) terminal device, an augmented reality (AR) terminal device and other digital display products. The electronic device 002 can be the calling terminal device and / or the called terminal device in the embodiments of the present application.
[0273] Please refer to Figure 20 , Figure 20 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 20 shown, the electronic device may include a processor 210, an external memory interface 220, an internal memory 221, a universal serial bus (USB) interface 230, a charging management module 240, a power management module 241, a battery 242, an antenna 1, an antenna 2, a mobile communication module 250, a wireless communication module 260, an audio module 270, a speaker 270A, a receiver 270B, a microphone 270C, a headphone interface 270D, a sensor module 280, a button 290, a motor 291, an indicator 292, a camera 293, a display screen 294, and a subscriber identification module (SIM) card interface 295, etc. Among them, the sensor module 280 may include a pressure sensor 280A, a gyroscope sensor 280B, a barometric pressure sensor 280C, a magnetic sensor 280D, an acceleration sensor 280E, a distance sensor 280F, a proximity light sensor 280G, a fingerprint sensor 280H, a temperature sensor 280J, a touch sensor 280K, an ambient light sensor 280L, a bone conduction sensor 280M, etc.
[0274] It can be understood that the structure illustrated in this embodiment does not constitute a specific limitation on the electronic device. In other embodiments, the electronic device may include more or fewer components than shown, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0275] The processor 210 may include one or more processing units. For example, the processor 210 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0276] The controller can be the nerve center and command center of an electronic device. The controller can generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
[0277] A memory can also be set in the processor 210 to store instructions and data. In some embodiments, the memory in the processor 210 is a cache memory. This memory can save the instructions or data that the processor 210 has just used or recycled. If the processor 210 needs to use the instruction or data again, it can directly call it from the said memory. This avoids repeated accesses, reduces the waiting time of the processor 210, and thus improves the efficiency of the system.
[0278] In some embodiments, the processor 210 can include one or more interfaces. The interfaces can include an inter-integrated circuit (I2C) interface, an inter-integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a subscriber identity module (SIM) interface, and / or a universal serial bus (USB) interface, etc.
[0279] It can be understood that the interface connection relationships among the modules illustrated in this embodiment are only illustrative and do not constitute a structural limitation on the electronic device. In some other embodiments, the electronic device can also adopt different interface connection methods in the above embodiments, or a combination of multiple interface connection methods.
[0280] The charging management module 240 is used to receive charging input from a charger. Among them, the charger can be a wireless charger or a wired charger. In some embodiments of wired charging, the charging management module 240 can receive the charging input of the wired charger through the USB interface 230. In some embodiments of wireless charging, the charging management module 240 can receive the wireless charging input through the wireless charging coil of the electronic device. While charging the battery 242, the charging management module 240 can also supply power to the electronic device through the power management module 241.
[0281] The power management module 241 is used to connect the battery 242, the charging management module 240, and the processor 210. The power management module 241 receives inputs from the battery 242 and / or the charging management module 240 and supplies power to the processor 210, the internal memory 221, the external memory, the display screen 294, the camera 293, the wireless communication module 260, etc. The power management module 241 can also be used to monitor parameters such as the battery capacity, the number of battery charge cycles, and the battery health status (leakage, impedance). In some other embodiments, the power management module 241 can also be disposed in the processor 210. In some other embodiments, the power management module 241 and the charging management module 240 can also be disposed in the same device.
[0282] The wireless communication function of the electronic device can be implemented by the antenna 1, the antenna 2, the mobile communication module 250, the wireless communication module 260, the modulation and demodulation processor, and the baseband processor, etc.
[0283] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas. For example, the antenna 1 can be multiplexed as the diversity antenna of the wireless local area network. In some other embodiments, the antenna can be used in combination with a tuning switch.
[0284] The mobile communication module 250 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc. applied to the electronic device. The mobile communication module 250 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc. The mobile communication module 250 can receive electromagnetic waves by the antenna 1, perform filtering, amplification, etc. on the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 250 can also amplify the signal modulated by the modulation and demodulation processor and convert it into electromagnetic waves through the antenna 1 for radiation. In some embodiments, at least some functional modules of the mobile communication module 250 can be disposed in the processor 210. In some embodiments, at least some functional modules of the mobile communication module 250 and at least some modules of the processor 210 can be disposed in the same device.
[0285] The modulation and demodulation processor may include a modulator and a demodulator. Among them, the modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. Subsequently, the demodulator transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is transmitted to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 270A, the receiver 270B, etc.), or displays an image or video through the display screen 294. In some embodiments, the modulation and demodulation processor may be an independent device. In other embodiments, the modulation and demodulation processor may be independent of the processor 210 and be provided in the same device as the mobile communication module 250 or other functional modules.
[0286] The wireless communication module 260 may provide solutions for wireless communications applied to the electronic device, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. The wireless communication module 260 may be one or more devices integrating at least one communication processing unit. The wireless communication module 260 receives electromagnetic waves via the antenna 2, performs frequency modulation and filtering processing on the electromagnetic wave signal, and transmits the processed signal to the processor 210. The wireless communication module 260 may also receive the signal to be transmitted from the processor 210, perform frequency modulation and amplification on it, and convert it into electromagnetic waves through the antenna 2 and radiate it out.
[0287] In some embodiments, antenna 1 of the electronic device is coupled to the mobile communication module 250, and antenna 2 is coupled to the wireless communication module 260, enabling the electronic device to communicate with the network and other devices through wireless communication technologies. The wireless communication technologies may include Global System for Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Synchronous Code Division Multiple Access (TDSCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology, etc. The GNSS may include Global Positioning System (GPS), Global Navigation Satellite System (GLONASS), BeiDou Navigation Satellite System (BDS), Quasi-Zenith Satellite System (QZSS), and / or Satellite Based Augmentation Systems (SBAS).
[0288] The electronic device implements the display function through the GPU, the display screen 294, and the application processor, etc. The GPU is a microprocessor for image processing, connected to the display screen 294 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. The processor 210 may include one or more GPUs, which execute program instructions to generate or change display information.
[0289] The display screen 294 is used to display images, videos, etc. The display screen 294 includes a display panel. The display panel can adopt a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light-emitting diode (QLED), etc.
[0290] Optionally, the display screen 294 displays the running interface of the application on the electronic device in the form of a visualization window. In the embodiment of the present application, the display screen displays the running interface of specific types of applications (such as text message applications and instant messaging software) in a double-page display manner, and two interfaces are displayed on one display screen, and the two interfaces can be parent-child pages of each other.
[0291] It can be understood that if the electronic device is a folding-screen mobile phone, then the display screen 294 can also be called a folding screen. It can be folded into two screens on the left and right along the longitudinal folding edge of the folding screen, or it can be folded into two screens on the top and bottom along the transverse folding edge of the folding screen, etc.
[0292] It should be noted that at least two screens formed after the folding of the electronic device (including the electronic device that folds inward and the electronic device that folds outward) in the embodiment of the present application can be multiple independent screens, or can be a complete screen with an integrated structure, but only folded into at least two parts.
[0293] For example, the folding screen can be a flexible folding screen. The flexible folding screen includes a folding edge made of a flexible material. Part or all of the flexible folding screen is made of a flexible material. At least two screens formed after the folding of the flexible folding screen are a complete screen with an integrated structure, but only folded into at least two parts.
[0294] For another example, the folding screen of the electronic device can be a multi-screen folding screen. The multi-screen folding screen can include multiple (two or more) screens. These multiple screens are multiple separate display screens. These multiple screens can be sequentially connected by a folding axis. Each screen can rotate around the folding axis connected to it to realize the folding of the multi-screen folding screen.
[0295] An electronic device can implement a shooting function through an ISP, a camera 293, a video codec, a GPU, a display screen 294, an application processor, etc.
[0296] The ISP is used to process the data fed back by the camera 293. For example, when taking a photo, the shutter is opened, and light passes through the lens and is transmitted to the camera's photosensitive element. The optical signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye. The ISP can also optimize the noise and brightness of the image through algorithms. The ISP can also optimize parameters such as the exposure and color temperature of the shooting scene. In some embodiments, the ISP can be provided in the camera 293.
[0297] The camera 293 is used to capture still images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard format such as RGB or YUV. In some embodiments, the electronic device can include one or N cameras 293, where N is a positive integer greater than 1.
[0298] The camera 293 can also be used to provide a personalized and contextual service experience to the user according to the perceived external environment and the user's actions. Among them, the camera 293 can obtain rich and accurate information so that the electronic device can perceive the external environment and the user's actions. Specifically, in the embodiments of the present application, the camera 293 can be used to identify whether the user of the electronic device is the first user or the second user.
[0299] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy, etc.
[0300] The video codec is used to compress or decompress digital videos. The electronic device can support one or more video codecs. In this way, the electronic device can play or record videos in multiple encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, MPEG4, etc.
[0301] The NPU is a neural-network (NN) computing processor. By drawing on the structure of biological neural networks, such as the transmission pattern between human brain neurons, it can quickly process input information and can also continuously self-learn. Through the NPU, applications such as intelligent cognition of electronic devices can be realized, such as: image recognition, face recognition, speech recognition, text understanding, etc.
[0302] The external memory interface 220 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device. The external memory card communicates with the processor 210 through the external memory interface 220 to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0303] The internal memory 221 can be used to store computer-executable program code, and the executable program code includes instructions. The processor 210 executes various functional applications and data processing of the electronic device by running the instructions stored in the internal memory 221. For example, in the embodiment of the present application, the processor 210 can, by executing the instructions stored in the internal memory 221, in response to the operation of the user on the display screen 294, display corresponding display content on the display screen. The internal memory 221 can include a storage program area and a storage data area. Among them, the storage program area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The storage data area can store data created during the use of the electronic device (such as audio data, phone book, etc.). In addition, the internal memory 221 can include high-speed random access memory and can also include non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0304] The electronic device can implement audio functions through the audio module 270, the speaker 270A, the receiver 270B, the microphone 270C, the headphone jack 270D, and the application processor, etc. Such as music playback, recording, etc.
[0305] The audio module 270 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 270 can also be used for encoding and decoding audio signals. In some embodiments, the audio module 270 can be disposed in the processor 210, or some functional modules of the audio module 270 can be disposed in the processor 210. The speaker 270A, also referred to as a "loudspeaker", is used to convert an audio electrical signal into a sound signal. The electronic device can listen to music or hands-free calls through the speaker 270A. The receiver 270B, also referred to as an "earpiece", is used to convert an audio electrical signal into a sound signal. When the electronic device answers a call or a voice message, the voice can be listened to by placing the receiver 270B close to the human ear. The microphone 270C, also referred to as a "microphone" or "transmitter", is used to convert a sound signal into an electrical signal. When making a call, sending a voice message, or when it is necessary to trigger the electronic device to perform certain functions through a voice assistant, the user can make a sound by bringing the mouth close to the microphone 270C to input the sound signal into the microphone 270C. The electronic device can be provided with at least one microphone 270C. In some other embodiments, the electronic device can be provided with two microphones 270C, which can not only collect sound signals but also implement a noise reduction function. In some other embodiments, the electronic device can also be provided with three, four or more microphones 270C to implement functions such as collecting sound signals, noise reduction, identifying the sound source, and implementing a directional recording function.
[0306] The headphone jack 270D is used to connect a wired headphone. The headphone jack 270D can be a USB interface 230, or a 3.5mm open mobile terminal platform (OMTP) standard interface, or a cellular telecommunications industry association of the USA (CTIA) standard interface.
[0307] The pressure sensor 280A is used to sense the pressure signal and can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor 280A can be set on the display screen 294. There are many types of pressure sensors 280A, such as resistive pressure sensors, inductive pressure sensors, capacitive pressure sensors, etc. The capacitive pressure sensor can be a parallel plate including at least two conductive materials. When a force acts on the pressure sensor 280A, the capacitance between the electrodes changes. The electronic device determines the intensity of the pressure based on the change in capacitance. When a touch operation acts on the display screen 294, the electronic device detects the touch operation intensity according to the pressure sensor 280A. The electronic device can also calculate the touch position based on the detection signal of the pressure sensor 280A. In some embodiments, touch operations acting on the same touch position but with different touch operation intensities can correspond to different operation instructions. For example: when a touch operation with a touch operation intensity less than a pressure threshold acts on a short message application icon, an instruction to view the short message is executed. When a touch operation with a touch operation intensity greater than or equal to the pressure threshold acts on a short message application icon, an instruction to create a new short message is executed.
[0308] The gyro sensor 280B can be used to determine the motion posture of the electronic device. In some embodiments, the angular velocity of the electronic device around three axes (i.e., x, y, and z axes) can be determined by the gyro sensor 280B. The gyro sensor 280B can be used for anti-shake shooting. For example, when the shutter is pressed, the gyro sensor 280B detects the angle of the electronic device shaking, calculates the distance that the lens module needs to compensate based on the angle, and allows the lens to offset the shaking of the electronic device through reverse movement to achieve anti-shake. The gyro sensor 280B can also be used for navigation and somatosensory game scenes. In addition, the gyro sensor 280B can also be used to measure the rotation amplitude or moving distance of the electronic device.
[0309] The air pressure sensor 280C is used to measure air pressure. In some embodiments, the electronic device calculates the altitude through the air pressure value measured by the air pressure sensor 280C to assist in positioning and navigation.
[0310] The magnetic sensor 280D includes a Hall sensor. The electronic device can use the magnetic sensor 280D to detect the opening and closing of the flip leather case. In some embodiments, when the electronic device is a flip phone, the electronic device can detect the opening and closing of the flip cover according to the magnetic sensor 280D. Then, according to the detected opening and closing state of the leather case or the opening and closing state of the flip cover, the flip cover can be automatically unlocked.
[0311] The acceleration sensor 280E can detect the magnitude of the acceleration of the electronic device in various directions (generally three axes). When the electronic device is stationary, it can detect the magnitude and direction of gravity. It can also be used to identify the posture of the electronic device and is applied to applications such as horizontal and vertical screen switching and pedometers. In addition, the acceleration sensor 280E can also be used to measure the orientation of the electronic device (i.e., the direction vector of the orientation).
[0312] The distance sensor 280F is used to measure distance. The electronic device can measure distance through infrared or laser. In some embodiments, when shooting a scene, the electronic device can use the distance sensor 280F to measure the distance to achieve fast focusing.
[0313] The proximity light sensor 280G may include, for example, a light-emitting diode (LED) and a light detector, such as a photodiode. The light-emitting diode can be an infrared light-emitting diode. The electronic device emits infrared light outward through the light-emitting diode. The electronic device uses the photodiode to detect the infrared reflected light from nearby objects. When sufficient reflected light is detected, it can be determined that there is an object near the electronic device. When insufficient reflected light is detected, the electronic device can determine that there is no object near the electronic device. The electronic device can use the proximity light sensor 280G to detect when the user holds the electronic device close to the ear for a call, so as to automatically turn off the screen to save power. The proximity light sensor 280G can also be used for automatic unlocking and locking of the leather case mode and pocket mode.
[0314] The ambient light sensor 280L is used to sense the ambient light brightness. The electronic device can adaptively adjust the brightness of the display screen 294 according to the sensed ambient light brightness. The ambient light sensor 280L can also be used to automatically adjust the white balance when taking pictures. The ambient light sensor 280L can also cooperate with the proximity light sensor 280G to detect whether the electronic device is in the pocket to prevent accidental touch.
[0315] The fingerprint sensor 280H is used to collect fingerprints. The electronic device can use the collected fingerprint characteristics to achieve fingerprint unlocking, access to application locks, fingerprint photography, fingerprint answering of incoming calls, etc.
[0316] The temperature sensor 280J is used to detect temperature. In some embodiments, the electronic device uses the temperature detected by the temperature sensor 280J to execute a temperature processing strategy. For example, when the temperature reported by the temperature sensor 280J exceeds the threshold, the electronic device reduces the performance of the processor near the temperature sensor 280J to reduce power consumption and implement thermal protection. In other embodiments, when the temperature is lower than another threshold, the electronic device heats the battery 242 to avoid abnormal shutdown of the electronic device caused by low temperature. In still other embodiments, when the temperature is lower than yet another threshold, the electronic device boosts the output voltage of the battery 242 to avoid abnormal shutdown caused by low temperature.
[0317] The touch sensor 280K, also known as the "touch panel". The touch sensor 280K can be disposed on the display screen 294, and the touch sensor 280K and the display screen 294 form a touch screen, also known as the "touch control screen". The touch sensor 280K is used to detect touch operations acting thereon or nearby. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 294. In some other embodiments, the touch sensor 280K can also be disposed on the surface of the electronic device, at a different position from where the display screen 294 is located.
[0318] The bone conduction sensor 280M can acquire vibration signals. In some embodiments, the bone conduction sensor 280M can acquire vibration signals of the vibrating bone mass of the human vocal part. The bone conduction sensor 280M can also contact the human pulse and receive blood pressure pulsation signals. In some embodiments, the bone conduction sensor 280M can also be disposed in the earphone to form a bone conduction earphone. The audio module 270 can parse out voice signals based on the vibration signals of the vibrating bone mass of the human vocal part acquired by the bone conduction sensor 280M to implement the voice function. The application processor can parse out heart rate information based on the blood pressure pulsation signals acquired by the bone conduction sensor 280M to implement the heart rate detection function.
[0319] The keys 290 include a power-on key, volume keys, etc. The keys 290 can be mechanical keys. They can also be touch keys. The electronic device can receive key inputs and generate key signal inputs related to the user settings and function controls of the electronic device.
[0320] Among them, the electronic device is through various sensors in the sensor module 280, the keys 290, and / or the camera 293, etc.
[0321] The motor 291 can generate vibration prompts. The motor 291 can be used for incoming call vibration prompts and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playing, etc.) can correspond to different vibration feedback effects. For touch operations acting on different regions of the display screen 294, the motor 291 can also correspond to different vibration feedback effects. Different application scenarios (such as: time reminder, receiving information, alarm clock, game, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.
[0322] The indicator 292 can be an indicator light and can be used to indicate the charging state, power change, and can also be used to indicate messages, missed calls, notifications, etc.
[0323] The SIM card interface 295 is used to connect to a SIM card. The SIM card can be inserted into or removed from the SIM card interface 295 to achieve contact and separation from the electronic device. The electronic device can support one or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 295 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 295 simultaneously. The types of multiple cards can be the same or different. The SIM card interface 295 can also be compatible with different types of SIM cards. The SIM card interface 295 can also be compatible with external memory cards. The electronic device interacts with the network through the SIM card to implement functions such as calls and data communication. In some embodiments, the electronic device uses an eSIM, that is, an embedded SIM card. The eSIM card can be embedded in the electronic device and cannot be separated from the electronic device.
[0324] The methods in the foregoing embodiments can all be implemented in the electronic device 002 with the above hardware structure. For ease of understanding, the following uses a mobile phone including a folding screen and with the folding screen in a fully unfolded state as an example for illustrative description. The presence or absence of a folding screen in the mobile phone, the form of the folding screen of the mobile phone, and the number of screens after the folding screen is folded are not limited herein.
[0325] The media resource playback method in the embodiments of the present application has been described above. Next, the electronic device in the embodiments of the present application will be described. Please refer to Figure 21 , Figure 21 Another embodiment of the electronic device in the embodiments of the present application includes:
[0326] In one example, the electronic device is applied to a calling terminal device, and the electronic device includes:
[0327] A transceiver unit 2101, configured to send a call request to a called terminal device;
[0328] The transceiver unit 2101 is further configured to send first audio information, where the first audio information is used to request the playback of media resources;
[0329] The transceiver unit 2101 is further configured to receive and play first media resources, where the first media resources are determined by the first audio information.
[0330] In a possible implementation manner,
[0331] The transceiver unit 2101 is further configured to play first response information, where the first response information is response information generated based on the first audio information, and the first response information includes audio information, picture information, animation information, and / or text information.
[0332] In a possible implementation, the first response information includes:
[0333] First text information, which is the text information generated by performing speech recognition processing on the first audio information, and the content included in the first text information corresponds to the first audio information.
[0334] In a possible implementation,
[0335] The transceiver unit 2101 is further configured to send a second audio information;
[0336] The transceiver unit 2101 is further configured to receive a response to the second audio information, and the response to the second audio information includes:
[0337] A second media resource, which is determined by the second audio information;
[0338] And / or, a second response information, which is a response information generated based on the second audio information, and the second response information includes audio information, picture information, animation information, and / or text information.
[0339] In a possible implementation,
[0340] The transceiver unit 2101 is further configured to stop playing the first media resource;
[0341] The transceiver unit 2101 is further configured to play the second media resource;
[0342] And / or, play the second response information.
[0343] In a possible implementation,
[0344] The transceiver unit 2101 is further configured to receive an off-hook message sent by the called terminal device;
[0345] The processing unit 2102 is configured to stop playing the first media resource in response to the off-hook message;
[0346] The processing unit 2102 is further configured to stop receiving the audio information and / or stop the speech recognition processing on the audio information.
[0347] In a possible implementation,
[0348] The transceiver unit 2101 is further configured to send the first audio information, and the first audio information includes a wake-up keyword, and the wake-up keyword is used to trigger the media platform server to perform speech recognition processing on the first audio information;
[0349] Or,
[0350] The transceiver unit 2101 is further configured to send audio information including the wake-up keyword.
[0351] The transceiver unit 2101 is further configured to send the first audio information.
[0352] In another possible implementation, the electronic device is applied to a media platform server and includes:
[0353] The transceiver unit 2101 is further configured to receive a call request sent by the calling terminal device.
[0354] The transceiver unit 2101 is further configured to receive the first audio information sent by the calling terminal device.
[0355] The processing unit 2102 is further configured to determine a first media resource according to the first audio information.
[0356] The transceiver unit 2101 is further configured to send the first media resource to the calling terminal device.
[0357] In a possible implementation,
[0358] The processing unit 2102 is further configured to perform speech recognition processing according to the first audio information to generate first text information, and the content included in the first text information corresponds to the first audio information.
[0359] The processing unit 2102 is further configured to perform semantic understanding processing according to the first text information to generate first user intent information.
[0360] The processing unit 2102 is further configured to determine the first media resource according to the first user intent information.
[0361] In a possible implementation,
[0362] The processing unit 2102 is further configured to generate a first response message according to the first user intent information, and the first response message includes audio information, picture information, animation information, and / or text information.
[0363] The transceiver unit 2101 is further configured to send the first response message to the calling terminal device.
[0364] In a possible implementation, the first response message includes the first text information.
[0365] In a possible implementation, the first user intent information includes any one or more of the following:
[0366] Start playing the media resource, pause playing the media resource, switch the played media resource, rewind the played media resource, copy and order the media resource, the content feature keywords of the first audio information, or the weights of the content feature keywords.
[0367] In a possible implementation,
[0368] The processing unit 2102 is further configured to determine the first media resource according to the decision recommendation model and the content feature keywords and / or the weights of the content feature keywords included in the first user intention information, where
[0369] The decision recommendation model determines the media resource by applying a parameter set, and the parameter set includes any one or more of the following: the content feature keywords of the media resource, the weights of the content feature keywords of the media resource, the media resource tags of the media resource library, the weights of the media resource tags of the media resource library, the popularity weights of the media resources in the media resource library, the release time of the media resources in the media resource library, or the play rates of the media resources in the media resource library, where the media resource library includes one or more media resources.
[0370] In a possible implementation,
[0371] The processing unit 2102 is further configured to detect whether the first audio information includes a wake-up keyword;
[0372] The processing unit 2102 is further configured to, if the first audio information includes the wake-up keyword, trigger speech recognition processing according to the first audio information;
[0373] Alternatively, the processing unit 2102 is further configured to trigger speech recognition processing according to the first audio information based on detecting that the received audio information includes the wake-up keyword.
[0374] In a possible implementation,
[0375] The transceiver unit 2101 is further configured to receive second audio information sent by the calling terminal device;
[0376] The processing unit 2102 is further configured to generate a response to the second audio information according to the second audio information,
[0377] The response to the second audio information includes:
[0378] A second media resource, which is determined by the second audio information;
[0379] And / or, a second response message, which is a response message generated based on the second audio message, and the second response message includes audio message, picture message, animation message, and / or text message;
[0380] The transceiver unit 2101 is further configured to send a response to the second audio message to the calling terminal device.
[0381] In a possible implementation manner,
[0382] The processing unit 2102 is further configured to perform speech recognition processing on the second audio message to generate a second text message, and the content included in the second text message corresponds to the first audio message;
[0383] The processing unit 2102 is further configured to perform semantic understanding processing on the second text message to generate a second user intention message;
[0384] The processing unit 2102 is further configured to generate a response to the second audio message according to the second user intention message.
[0385] Refer to Figure 22 , a schematic structural diagram of another electronic device provided in this application. The electronic device may include a processor 2201, a memory 2202, and a communication port 2203. The processor 2201, the memory 2202, and the communication port 2203 are interconnected by lines. Among them, program instructions and data are stored in the memory 2202.
[0386] The program instructions and data corresponding to the steps executed by the calling terminal device and / or the media platform server in the corresponding embodiments shown above are stored in the memory 2202. Figures 2 to 18
[0387] The processor 2201 is configured to execute the steps executed by the calling terminal device and / or the media platform server in any of the embodiments shown above. Figures 2 to 18
[0388] The communication port 2203 can be used for receiving and sending data, and is used to execute the steps related to acquisition, sending, and receiving in any of the embodiments shown above. Figures 2 to 18
[0389] In an implementation manner, the electronic device may include more or fewer components relative to Figure 22 This is only an exemplary illustration in this application and is not limited.
[0390] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0391] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0392] In addition, in each embodiment of this application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in whole or in part through software, hardware, firmware, or any combination thereof.
[0393] When implementing the integrated units using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from a website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more available media integrated. The available medium can be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a solid state disk (SSD)), etc.
Claims
1. A method for playing media resources, characterized in that The method is applied to the calling terminal device, and the method includes: Sending a call request to the called terminal device; Sending first audio information, where the first audio information is used to request playing a media resource; Receiving and playing a first media resource, where the first media resource is determined by the first audio information.
2. The method according to claim 1, wherein The method further includes: Playing a first response message, where the first response message is a response message generated based on the first audio information, and the first response message includes audio information, picture information, animation information, and / or text information.
3. The method according to any one of claims 1 or 2, characterized in that The first response message includes: First text information, where the first text information is text information generated by performing speech recognition processing on the first audio information, and the content included in the first text information corresponds to the first audio information.
4. The method according to any one of claims 1 to 3, characterized in that The method further includes: Sending second audio information; Receiving a response to the second audio information, where the response to the second audio information includes: A second media resource, where the second media resource is determined by the second audio information; And / or, a second response message, where the second response message is a response message generated based on the second audio information, and the second response message includes audio information, picture information, animation information, and / or text information.
5. The method according to claim 4, characterized in that The method further includes: Stopping playing the first media resource; Playing the second media resource; And / or, playing the second response message.
6. The method according to any one of claims 1-5, characterized in that, The method further includes: Receiving an off-hook message sent by the called terminal device; In response to the off-hook message, stopping playing the first media resource; Stopping receiving audio information, and / or stopping speech recognition processing on the audio information.
7. The method according to any one of claims 1 to 6, characterized in that, Sending the first audio information includes: Sending the first audio information, where the first audio information includes a wake-up keyword, and the wake-up keyword is used to trigger the media platform server to perform speech recognition processing on the first audio information; Or, Sending audio information including the wake-up keyword; Sending the first audio information.
8. A method for playing media resources, characterized in that, The method is applied to the media platform server, and the method includes: Receiving a call request sent by the calling terminal device; Receiving first audio information sent by the calling terminal device; Determining a first media resource according to the first audio information; Sending the first media resource to the calling terminal device.
9. The method according to claim 8, wherein The determining the first media resource according to the first audio information includes: Performing speech recognition processing on the first audio information to generate first text information, where the content included in the first text information corresponds to the first audio information; Performing semantic understanding processing on the first text information to generate first user intent information; Determining the first media resource according to the first user intent information.
10. The method according to claim 9, wherein The method further includes: Generating a first response message according to the first user intent information, where the first response message includes audio information, picture information, animation information, and / or text information; Sending the first response message to the calling terminal device.
11. The method according to claim 9 or 10, characterized in that The first response message includes the first text information.
12. The method according to any one of claims 9 - 11, characterized in that The first user intent information includes any one or more of the following: Start playing media resources, pause playing media resources, switch to play media resources, rewind to play media resources, copy and subscribe to media resources, the content feature keywords of the first audio information, or the weights of the content feature keywords.
13. The method according to claim 12, wherein Determine the first media resource according to the first user intent information, including: Determine the first media resource according to the decision recommendation model and the content feature keywords and / or the weights of the content feature keywords included in the first user intent information, where The decision recommendation model determines media resources by applying a parameter set, and the parameter set includes any one or more of the following: the content feature keywords of the media resources, the weights of the content feature keywords of the media resources, the media resource tags of the media resource library, the weights of the media resource tags of the media resource library, the popularity weights of the media resources in the media resource library, the release time of the media resources in the media resource library, or the play rates of the media resources in the media resource library, where the media resource library includes one or more media resources.
14. The method according to any one of claims 9 - 13, characterized in that The method further includes: Detect whether the first audio information includes a wake-up keyword; If the first audio information includes the wake-up keyword, trigger speech recognition processing according to the first audio information; Or, trigger speech recognition processing according to the first audio information based on detecting that the received audio information includes the wake-up keyword.
15. The method according to any one of claims 8 - 14, characterized in that, The method further includes: Receive second audio information sent by the calling terminal device; Generate a response to the second audio information according to the second audio information, The response to the second audio information includes: A second media resource, which is determined by the second audio information; And / or, a second response message, which is a response message generated based on the second audio information, and the second response message includes audio information, picture information, animation information, and / or text information; Send the response to the second audio information to the calling terminal device.
16. The method according to any one of claims 8-15, characterized in that, Generate a response to the second audio information according to the second audio information, including: Perform speech recognition processing according to the second audio information to generate second text information, and the content included in the second text information corresponds to the first audio information; Perform semantic understanding processing according to the second text information to generate second user intent information; Generate a response to the second audio information according to the second user intent information.
17. An electronic device, characterized in that, Includes: A transceiver unit and a processing unit, such that the electronic device executes the method according to any one of claims 1-7 or 8-16.
18. An electronic device, characterized in that, Includes: A processor, the processor is coupled to a memory, and the memory is used to store programs or instructions. When the programs or instructions are executed by the processor, the electronic device executes the method according to any one of claims 1-7 or 8-16.
19. A computer storage medium, characterized in that, Includes computer instructions, and when the computer instructions run on a terminal device, the terminal device executes the method according to any one of claims 1-7 or 8-16.
20. A computer program product, characterized in that, When the computer program product runs on a computer, the computer is caused to execute the method according to any one of claims 1-7 or 8-16.