Voice control method, multimedia system, vehicle, and storage medium
Semantic analysis and application recommendation in in-vehicle systems address the issue of default responses, enhancing user experience and retention by aligning audio command responses with user preferences.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- BYD CO LTD
- Filing Date
- 2023-05-23
- Publication Date
- 2026-04-27
AI Technical Summary
Existing in-vehicle multimedia systems respond to audio commands using default applications, often not preferred by the user, leading to a poor user experience.
Perform semantic analysis on voice commands to identify user preferences, and use a recommendation policy to select and respond with the appropriate application from a list, considering foreground, background, recently used, or preset applications, and update applications to match content types.
Improves user experience by ensuring responses align with user preferences, increasing user retention through targeted application usage.
Smart Images

Figure 0007852090000001 
Figure 0007852090000002 
Figure 0007852090000003
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This disclosure claims priority to Chinese Patent Application No. 202210759453.7, filed on June 30, 2022. The entire content of the above - referenced application is incorporated herein by reference.
[0002] This disclosure relates to the field of audio processing technology, and more particularly, to an audio control method, a multimedia system, a vehicle, and a storage medium.
Background Art
[0003] Due to the rapid development of the Internet and vehicle technology, users can surf the Internet using an in - vehicle multimedia system. A plurality of applications of the same type (e.g., KuGou and Kuwo) are installed in existing in - vehicle multimedia systems. When a user generates an audio command, the system directly calls the default application to respond, but the default application is not the application preferred by the user, thus resulting in a poor user experience.
Summary of the Invention
[0004] Embodiments of this disclosure provide an audio control method, a multimedia system, a vehicle, and a storage medium to solve the problem that an audio control response cannot be implemented by using the application preferred by the user.
[0005] Embodiments of this disclosure provide an audio control method including the following.
[0006] Semantic analysis is performed on the received target audio command to obtain a semantic analysis result.
[0007] If the semantic analysis result matches a preset keyword, the target voice command is responded to by using the application corresponding to the preset keyword.
[0008] If the semantic analysis results do not match the preset keywords, the target service type is determined based on the semantic analysis results, and a list of applications corresponding to the target service type is obtained.
[0009] By using the first recommendation policy, the recommended application is determined from the application list.
[0010] By using the recommended application, the target voice command will be responded to.
[0011] According to one embodiment of the present disclosure, determining a recommended application from an application list by using a first recommendation policy includes the following:
[0012] If the application list includes a foreground application, the foreground application will be determined as the recommended application.
[0013] If the application list does not include a foreground application but includes a background application, the background application will be determined as the recommended application.
[0014] If the application list does not include foreground and background applications, but includes recently used applications, the recommended application will be determined from the application list by using the second recommendation policy.
[0015] If the application list does not include foreground applications, background applications, or recently used applications, a preset application will be selected as the recommended application.
[0016] Recently used applications are those used within a preset period prior to the current time.
[0017] According to one embodiment of the present disclosure, determining a recommended application from an application list by using a second recommendation policy includes the following:
[0018] Based on the semantic analysis results, the target content type is determined, and it is decided whether recently used applications contain content corresponding to the target content type.
[0019] If a recently used application contains content that corresponds to the target content type, that recently used application will be determined as a recommended application.
[0020] If recently used applications do not contain content corresponding to the target content type, it is determined whether there are candidate applications with content corresponding to the target content type.
[0021] If there are candidate applications that contain content corresponding to the target content type, those candidate applications will be selected as recommended applications.
[0022] If no candidate applications containing content corresponding to the target content type exist, a preset application will be selected as the recommended application.
[0023] According to an embodiment of the present disclosure, responding to a target voice command by using a recommended application includes the following.
[0024] Based on the semantic analysis result, a target content type is determined, and it is judged whether the recommended application includes content corresponding to the target content type.
[0025] If the recommended application includes content corresponding to the target content type, the target voice command is responded to by using the recommended application.
[0026] If the recommended application does not include content corresponding to the target content type, a candidate application including the target content type is updated to be the recommended application, and the target voice command is responded to by using the updated recommended application.
[0027] According to an embodiment of the present disclosure, updating a candidate application including the target content type to be the recommended application includes the following.
[0028] A candidate application including the target content type is determined as the first application, and the dominant content type of the first application is obtained.
[0029] The first application whose dominant content type is the target content type is updated to be the recommended application.
[0030] According to an embodiment of the present disclosure, the semantic analysis result includes the target content type.
[0031] Responding to a target voice command by using a recommended application includes the following.
[0032] The system acquires dominant content types corresponding to the recommended applications, and then matches these dominant content types with the target content types.
[0033] If the preferred content type for the recommended application matches the target content type, the target voice command will be responded to by using the recommended application.
[0034] If the preferred content type corresponding to the recommended application does not match the target content type, preferred content recommendation information is obtained, the target voice command is responded to by using the recommended application, and preferred content recommendation information is prompted.
[0035] According to one embodiment of this disclosure, obtaining preferential content recommendation information includes the following:
[0036] A candidate application other than the recommended application in the application list is selected as the second application, and the dominant content type corresponding to the second application is acquired.
[0037] Based on the preferred content types and target content types corresponding to the second application, preferred content recommendation information is obtained.
[0038] According to one embodiment of the present disclosure, obtaining preferential content recommendation information based on a preferential content type and a target content type corresponding to a second application includes the following:
[0039] Similarity calculations are performed on the dominant content type and target content type corresponding to the second application to obtain the content type similarity for the second application.
[0040] The second application with the best content type similarity is selected as the target application.
[0041] Based on the target application, preferential content recommendations are obtained.
[0042] In one embodiment, the application list includes at least one candidate application, and each candidate application includes at least one current content type.
[0043] Before the target voice command is responded to by using the recommended application, the voice control method further includes the following:
[0044] Current traffic and user ratings are obtained for each current content type within the candidate application.
[0045] An overall rating is determined for each current content type based on the current traffic and user ratings associated with that content type.
[0046] The dominant content type for a candidate application is determined based on its overall rating for at least one current content type.
[0047] One embodiment of the present disclosure provides a multimedia system including memory, a processor, and a computer program stored in the memory and running on the processor. The voice control method is performed when the processor executes the computer program.
[0048] One embodiment of this disclosure provides a vehicle including a multimedia system.
[0049] One embodiment of the present disclosure provides a computer-readable storage medium for storing a computer program. The voice control method is performed when the computer program is executed by a processor.
[0050] According to the voice control method, multimedia system, vehicle, and storage medium, when the semantic analysis result matches a preset keyword, user preferences can be intuitively reflected. Therefore, an application corresponding to the preset keyword may be used to respond to the target voice command. This helps improve the user experience and further increase user retention. When the semantic analysis result does not match a preset keyword, a recommended application is determined from an application list corresponding to the target service type, and that recommended application is used to respond to the target voice command. This can satisfy user preferences to some extent, improve the user experience, and further increase user retention.
[0051] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings necessary to illustrate the embodiments of this disclosure are briefly described below. Obviously, the accompanying drawings in the following description illustrate only some embodiments of this disclosure, and those skilled in the art can derive other drawings from these without creative effort. [Brief explanation of the drawing]
[0052] [Figure 1] This is a schematic diagram of the application environment for a voice control method according to one embodiment of the present disclosure. [Figure 2] This is a flowchart of a voice control method according to one embodiment of the present disclosure. [Figure 3] This is another flowchart of a voice control method according to one embodiment of the present disclosure. [Figure 4] This is another flowchart of a voice control method according to one embodiment of the present disclosure. [Figure 5] This is another flowchart of a voice control method according to one embodiment of the present disclosure. [Figure 6] This is another flowchart of a voice control method according to one embodiment of the present disclosure. [Figure 7] This is another flowchart of a voice control method according to one embodiment of the present disclosure. [Figure 8] This is another flowchart of a voice control method according to one embodiment of the present disclosure. [Figure 9] This is another flowchart of a voice control method according to one embodiment of the present disclosure. [Figure 10] This is another flowchart of a voice control method according to one embodiment of the present disclosure. [Figure 11] This is another flowchart of a voice control method according to one embodiment of the present disclosure. [Modes for carrying out the invention]
[0053] The following describes the technical solutions in the embodiments of this disclosure with reference to the accompanying drawings of the embodiments. Clearly, the embodiments described are not all of the embodiments of this disclosure, but rather a selection of them. All other embodiments that can be obtained by those skilled in the art without creative effort based on the embodiments of this disclosure are within the scope of this disclosure.
[0054] Embodiments of this disclosure provide a voice control method. The voice control method may be applied to the application environment shown in Figure 1. In particular, the voice control method is applied to a voice control system (hereinafter simply referred to as the "system"). The voice control system may be a multimedia system in a vehicle (i.e., an in-vehicle infotainment system), or another system capable of implementing voice control. The voice control system can call applications that satisfy user preferences and improve the user experience in order to respond based on collected target voice commands entered by the user.
[0055] In one embodiment, a voice control method is provided. As shown in Figure 1, the voice control method includes the following steps.
[0056] S101: Semantic analysis is performed on the received target voice command, and the semantic analysis results are obtained.
[0057] S102: If the semantic analysis result matches a preset keyword, the target voice command is responded to by using the application corresponding to the preset keyword.
[0058] S103: If the semantic analysis results do not match the preset keywords, the target service type is determined based on the semantic analysis results, and a list of applications corresponding to the target service type is obtained.
[0059] S104: The recommended application is determined from the application list by using the first recommendation policy.
[0060] S105: The target voice command is responded to by using the recommended application.
[0061] A target voice command is a command entered by the user via voice, instructing a specific application to be executed. The semantic analysis result is the result of semantic analysis performed on the target voice command.
[0062] In one example, in step S101, the multimedia system can receive a target voice command entered by the user, invoke the system's preset voice recognition tool, perform semantic analysis on the target voice command, and obtain the semantic analysis result. The target voice command processing process includes a voice recognition process and a semantic analysis process. The voice recognition tool may be any tool for performing semantic analysis, including but not limited to the iFLYTEK voice recognition tool. For example, the multimedia system can receive a target voice command entered by the user in voice format, such as "play crosstalk," "listen to economic news," "listen to ghost stories," "pop music," or "play a movie," perform semantic analysis on the target voice command, and obtain the semantic analysis result in text format.
[0063] Preset keywords are keywords that are bound to a specific application.
[0064] In one example, after obtaining the semantic analysis result in step S102, the multimedia system can perform matching between the semantic analysis result and preset keywords. When the semantic analysis result matches a preset keyword, the application corresponding to the preset keyword may be used directly to respond to the target voice command. For example, if the multimedia system obtains the semantic analysis result "I want to open Tencent Video" and this result matches the preset keyword "Tencent Video" pre-configured in the system, the multimedia system will respond to the target voice command by using the application "Tencent Video APP". Since user preferences can be intuitively reflected when the semantic analysis result matches a preset keyword, directly using the application corresponding to the preset keyword to respond to the target voice command satisfies user preferences, improves the user experience, and further increases user retention.
[0065] The target service type is the corresponding service type that reflects the function to which the user's utterance belongs in the semantic analysis results. The application list is a list formed by at least one candidate application. A candidate application is an application that can display content corresponding to the target service type.
[0066] In one example, after obtaining the semantic analysis result in step S103, the multimedia system can perform matching between the semantic analysis result and preset keywords. If the semantic analysis result does not match the preset keywords, the multimedia system can determine the target service type based on the semantic analysis result, and then query the system for at least one candidate application that includes that target service type based on that target service type, and obtain an application list based on at least one candidate application. In this example, the multimedia system obtains field information corresponding to the service type field from the semantic analysis result and determines the field information as the target service type. For example, semantic analysis results obtained by analysis using a speech recognition tool include field information corresponding to the service field, and the field information corresponding to the service field is determined as the target service type. The service field is used to reflect the service type corresponding to the function to which the user's utterance belongs, and includes, but is not limited to, "air conditioner=Aircontrol", "navigation=Map", "music=Music", and "radio station=internetRadio".
[0067] The first recommendation policy is a preset policy used to determine recommended applications based on the service type. Recommended applications are those that are recommended for use by the user.
[0068] In one example, after determining a list of applications corresponding to the target service type based on the target service type in step S104, the multimedia system can call a pre-configured first recommendation policy to select the application that best matches the user's preferences from the application list and determine that application as the recommended application. In this example, the first recommendation policy may be used to analyze user preferences based on user behavior or other information and determine the processing policy for the recommended application that matches the user's preferences.
[0069] In one example, after determining the recommended application in step S105, the multimedia system can respond to the target voice command by using the recommended application, for example, by playing audio based on the recommended application. Since the recommended application satisfies user preferences to some extent, it can be understood that using the recommended application to respond to the target voice command can improve the user experience and further increase user retention.
[0070] As shown in Figure 10, after receiving a target voice command, the multimedia system can invoke a speech recognition tool to perform semantic analysis on the target voice command and obtain the semantic analysis results. Next, the semantic analysis results are matched with preset keywords. If the semantic analysis results match the preset keywords, the application corresponding to the preset keywords is used to respond to the target voice command. If the semantic analysis results do not match the preset keywords, the content of the service field in the semantic analysis results is determined as the target service type, and then a list of applications corresponding to that target service type is obtained. A recommended application is then determined from the application list, and that recommended application is used to respond to the target voice command. This allows the target voice command response process to satisfy user preferences, improve the user experience, and further increase user retention.
[0071] In this embodiment, when the semantic analysis results match preset keywords, user preferences can be intuitively reflected. Therefore, an application corresponding to the preset keywords may be used to respond to the target voice command. This helps improve the user experience and further increase user retention. When the semantic analysis results do not match preset keywords, a recommended application is determined from a list of applications corresponding to the target service type, and that recommended application is used to respond to the target voice command. This can satisfy user preferences to some extent, improve the user experience, and further increase user retention.
[0072] In one embodiment, the application list includes at least one candidate application corresponding to the same target service type.
[0073] Step S104, specifically, determining a recommended application from the application list by using a first recommendation policy, includes determining a recommended application from at least one candidate application by using a first recommendation policy.
[0074] An application list is a list formed by at least one candidate application. A candidate application is an application that can display content corresponding to the target service type. In this example, the application list contains at least one candidate application corresponding to the same target service type.
[0075] In one example, after determining a list of applications corresponding to the target service type based on semantic analysis results, the multimedia system determines at least one candidate application from the application list, then analyzes the at least one candidate application by executing a pre-configured first recommendation policy, selects the recommended application that best suits the user's preferences from the at least one candidate application, and determines that recommended application as the recommended application. By using the recommended application, the system can respond to target voice commands, thereby allowing the user to view content of interest, and thus improving user retention.
[0076] In this embodiment, the application list corresponding to a target service type includes at least one candidate application corresponding to the same target service type. The candidate application is an application that satisfies the basic user requirements (i.e., the target service type) extracted from the semantic analysis results. Then, from the at least one candidate application, a recommended application that best matches the user's preferences is selected, thereby better matching user preferences and improving user retention when that recommended application is used to respond to target voice commands.
[0077] In one embodiment, as shown in Figure 2, step S104, specifically, determining a recommended application from the application list by using a first recommendation policy, includes the following steps:
[0078] S201: If the application list includes a foreground application, the foreground application is determined to be the recommended application.
[0079] S202: If the application list does not include a foreground application but includes a background application, the background application will be determined as the recommended application.
[0080] S203: If the application list does not include foreground and background applications, but includes recently used applications, the recommended application is determined by analyzing recently used applications using a second recommendation policy.
[0081] S204: If the application list does not include foreground applications, background applications, or recently used applications, a preset application is determined as the recommended application.
[0082] Recently used applications are those used within a preset time period prior to the current time.
[0083] A foreground application is an application that is currently running in the foreground; in other words, an application that is running and visible to the user. A background application is an application that is currently running in the background; in other words, an application that can continue to run related services for a short time after the application has been closed. Recently used applications are applications that were used within a preset period prior to the current time. The preset period is a user-defined period, which may be one day or one week. Preset applications are applications that are set by default, other than foreground applications, background applications, and recently used applications, and may be applications provided by the system.
[0084] In one example, after obtaining the application list in step S201, the multimedia system needs to determine whether the application list includes a foreground application. If the application list includes a foreground application, this indicates that the candidate application is currently running in the foreground and the user is browsing the foreground application. This may reflect the user's preference for the foreground application. Therefore, the foreground application may be determined as the recommended application. In this way, the recommended application with the highest user preference is obtained.
[0085] For example, in step S202, if the application list does not contain any foreground applications, the multimedia system determines whether the application list contains background applications. If the application list contains background applications, this indicates that there are no candidate applications currently running in the foreground, but that candidate applications are currently running in the background. Background applications can be understood as applications that were opened earlier and have not been completely closed. This may also reflect that the user has a certain degree of preference for background applications. Therefore, background applications may be determined as recommended applications. In this way, recommended applications with a high degree of user preference are obtained.
[0086] The second recommendation policy is a preset policy used to analyze recently used applications in order to determine the recommended application.
[0087] In one example, in step S203, if the application list does not include foreground and background applications, the multimedia system determines whether the application list includes recently used applications. If the application list includes recently used applications, this indicates that there are no candidate applications currently running in the foreground or background, but that recently used applications ran earlier and their execution history is buffered. This may reflect that the user has viewed recently used applications and prefers them. Therefore, a second recommendation policy is used to analyze recently used applications to determine recommended applications. In this way, recommended applications with common user preferences are obtained. In this example, the second recommendation policy is used to analyze recently used applications, and recently used applications that meet preset criteria may be determined as recommended applications. Hereinafter, preset criteria are pre-set conditions that may be determined based on user preferences.
[0088] For example, in step S204, if the application list does not include foreground applications, background applications, and recently used applications, the multimedia system can determine a preset application, which is set by default by the system, as the recommended application. The preset application is one of the candidate applications and corresponds to the target service type. In this way, a recommended application that satisfies the basic user requirements is obtained.
[0089] As shown in Figure 10, the multimedia system first determines whether the application list contains foreground applications. If it does, the foreground application is used as the recommended application to respond to the target voice command. If it does not contain foreground applications, it is determined whether the application list contains background applications. If it does, the background application is used as the recommended application to respond to the target voice command. If it does not contain background applications, it is determined whether the application list contains recently used applications. If it contains recently used applications, the recommended application needs to be further determined by using a second recommendation policy. If it does not contain recently used applications, a preset application is determined as the recommended application to respond to the target voice command.
[0090] In this embodiment, recommended applications are determined by sequentially selecting from foreground applications, background applications, recently used applications, and preset applications in order of user preference. Therefore, using the recommended application to respond to a target voice command better matches user preferences and further improves user retention.
[0091] In one embodiment, as shown in Figure 3, step S203, specifically, the determination of a recommended application from the application list by using a second recommendation policy, includes the following steps:
[0092] S301: Based on the semantic analysis results, the target content type is determined, and it is decided whether the recently used application contains content corresponding to the target content type.
[0093] S302: If a recently used application contains content that corresponds to the target content type, the recently used application is determined to be the recommended application.
[0094] S303: If the most recently used application does not contain content corresponding to the target content type, it is determined whether there are candidate applications with content corresponding to the target content type.
[0095] S304: If a candidate application exists that contains content corresponding to the target content type, that candidate application is selected as the recommended application.
[0096] S305: If no candidate application containing content corresponding to the target content type exists, a preset application is selected as the recommended application.
[0097] The target content type is the content type recognized in the semantic analysis results, in other words, the content type recognized based on the target voice command. The semantic analysis results obtained by analysis using a voice recognition tool include not only the target service type corresponding to the service field, but also the target content type corresponding to the semantic field. Target service type analysis may be performed based on the target service type and target content type. Semantic fields are fields used to distinguish utterances of different subcategories under the service field and are used to implement different content types. For example, the service field internetRadio is further divided into several semantic fields, namely the program field, tag field, category field, and presenter field. Different user utterances in the semantic analysis results can be understood as indicating different target content types corresponding to the semantic fields in the semantic analysis results. For example, if the semantic analysis result is "listen to a comedy radio station", the category field is entered. As another example, if the semantic analysis result is "Listen to crosstalk by XX", the moderator field will be populated.
[0098] For example, if in step S301 the multimedia system needs to analyze recently used applications by using a second recommended policy, the multimedia system must first determine the target content type based on the semantic analysis results, and then determine whether the recently used applications contain content corresponding to the target content type, in order to determine whether the recently used applications can satisfy the user requirements.
[0099] For example, in step S302, if the recently used application contains content of the target content type in the multimedia system, this indicates that the recently used application can display content of the target content type requested by the user. Furthermore, the recently used application is an application used by the user within a preset period prior to the present time and reflects the user's preferences to some extent. Therefore, the recently used application may be determined as a recommended application, thereby better aligning with the user's preferences to subsequently use the recommended application to respond to target voice commands.
[0100] For example, in step S303, if the recently used application does not contain content of the target content type in the multimedia system, this indicates that the recently used application cannot display content of the target content type requested by the user. In this case, it is necessary to determine whether any candidate applications other than the recently used application contain content of the target content type from among the at least one candidate application corresponding to the target service type determined based on the semantic analysis results, and a recommended application is determined based on the result of this determination.
[0101] For example, in step S304, if a candidate application containing content corresponding to the target content type exists within the multimedia system, this indicates that the service type of the candidate application includes the target service type determined based on the semantic analysis results, and that the candidate application contains content of the target content type determined based on the semantic analysis results, and therefore content that satisfies the user requirement can be displayed. Thus, the candidate application may be determined as the recommended application.
[0102] For example, in step S305, if no candidate application containing content corresponding to the target content type exists within the multimedia system, this indicates that no candidate application satisfies both the target service type and the target content type. In this case, a preset application may be determined as the recommended application. The preset application is one of the candidate applications and corresponds to the target service type. In this way, a recommended application that satisfies the basic user requirements is obtained.
[0103] In this embodiment, to determine a recommended application, selections are made sequentially from recently used applications, candidate applications containing the target content type, and preset applications, based on whether the application contains content of the target content type. Therefore, the determined recommended application can best satisfy user preferences. In this way, using the recommended application to respond to target voice commands better matches user preferences and further improves user retention.
[0104] In one embodiment, as shown in Figure 4, step S105, specifically, the response of the target voice command by using the recommended application, includes the following steps:
[0105] S401: Based on the semantic analysis results, the target content type is determined, and it is decided whether the recommended application contains content corresponding to the target content type.
[0106] S402: If the recommended application contains content that corresponds to the target content type, the target voice command will be responded to by using the recommended application.
[0107] S403: If the recommended application does not contain content corresponding to the target content type, the candidate application containing the target content type is updated to become the recommended application, and the target voice command is responded to by using the updated recommended application.
[0108] The target content type is the content type recognized in the semantic analysis results, in other words, the content type recognized based on the target speech command. The semantic analysis results obtained by analysis using a speech recognition tool include not only the target service type corresponding to the service field, but also the target content type corresponding to the semantic field. Target service type analysis may be performed based on the target service type and target content type. The semantic field is a field used to distinguish utterances of different subcategories under the service field and is used to implement different content types.
[0109] For example, in step S401, when the multimedia system responds to a target voice command by using a recommended application, the multimedia system may first determine the target content type based on the semantic analysis results, and then determine whether the recommended application contains content of the target content type, and whether the recommended application can satisfy the content of the target content type requested by the user.
[0110] For example, in step S402, if the recommended application contains content corresponding to the target content type, the multimedia system assumes that the recommended application can display the content requested by the user. Furthermore, the recommended application is an application with a high user preference. Therefore, the recommended application can be used to respond to the target voice command in order to display content of the target content type, thereby better satisfying user preferences and further improving user retention. For example, the target content type determined based on the target voice command is "Play song 1". If the recommended application contains song 1, that recommended application may be used to play song 1. In this way, when content requested by the user is displayed, the recommended application that satisfies user preferences may be used for display, thereby better matching user preferences and further improving user retention.
[0111] For example, in step S403, if the recommended application does not contain content corresponding to the target content type, the multimedia system assumes that the recommended application cannot display the content requested by the user. In this way, the multimedia system updates the candidate application that includes the target content type to become the recommended application, and by using the updated recommended application, it can respond to the target voice command, thereby enabling the updated recommended application to display content of the target content type and further improving user retention. For example, the target content type determined based on the target voice command is "Play song 1". If the recommended application does not include song 1, a candidate application containing song 1 is determined from at least one candidate application in the application list, the candidate application containing song 1 is updated to become the recommended application, and by using the updated recommended application, song 1 is played, thus enabling the updated recommended application to display the content requested by the user and further improving user retention.
[0112] In this embodiment, it is determined whether the recommended application needs to be updated based on whether the recommended application contains content of the target content type, ensuring that the recommended application can display the content requested by the user, thereby allowing the user request to be fulfilled when the system responds, and further improving user retention rates.
[0113] In one embodiment, as shown in Figure 5, step S403, specifically updating candidate applications that include a target content type to become recommended applications, includes the following steps:
[0114] S501: A candidate application containing the target content type is determined to be the first application, and the first application acquires the dominant content type.
[0115] S502: The first application whose dominant content type is the target content type will be updated to become the recommended application.
[0116] The first application is a candidate application that includes the target content type; in other words, the first application is an application that matches both the target service type and the target content type determined based on the semantic analysis results. The dominant content type is a content type with high user ratings and can be understood as a content type with high user ratings or frequent user access.
[0117] In one example, in step S501, the multimedia system can determine from the application list at least one candidate application containing the target content type as the first application, ensuring that the first application matches the target service type and target content type determined based on the semantic analysis results, and that the first application can display the content requested by the user. Furthermore, after determining the first application, the multimedia system needs to obtain the dominant content type of the first application. For example, the multimedia system can determine the dominant content type of the first application by querying a pre-configured dominant content type information table.
[0118] In one example, after acquiring the dominant content type of at least one first application in step S502, the multimedia system performs a matching of the dominant content type with the target content type determined based on the semantic analysis results, updates the first application whose dominant content type is the target content type to become a recommended application, and can use the updated recommended application to respond to a target voice command and display content corresponding to the target content type. Since the target content type is the dominant content type of the updated recommended application, this indicates that the target content type has a high user rating, which in turn makes it more attractive to users and thus improves user retention and user experience.
[0119] In this example, if there are at least two first applications where the dominant content type is the target content type, user ratings for at least two first applications will be obtained for the target content type, and the first application with the highest user rating will be determined as the recommended application, ensuring that users are more likely to be engaged when the updated recommended application displays content of the target content type, thus improving user retention and user experience.
[0120] In this embodiment, at least one candidate application containing the target content type is determined as the first application, and the first application whose dominant content type is the target content type is updated to become the recommended application, thereby making the updated recommended application more engaging to users when displaying content of the target content type, and thus improving user retention and user experience.
[0121] In one embodiment, the semantic analysis results include the target content type.
[0122] As shown in Figure 6, step S105, specifically, the response of the target voice command by using the recommended application, includes the following steps:
[0123] S601: A superior content type corresponding to the recommended application is acquired, and matching is performed between the superior content type corresponding to the recommended application and the target content type.
[0124] S602: If the preferred content type corresponding to the recommended application matches the target content type, the target voice command will be responded to by using the recommended application.
[0125] S603: If the preferred content type corresponding to the recommended application does not match the target content type, preferred content recommendation information is obtained, the target voice command is responded to by using the recommended application, and preferred content recommendation information is prompted.
[0126] The target content type is the content type recognized in the semantic analysis results, in other words, the content type recognized based on the target speech command. The semantic analysis results obtained by analysis using a speech recognition tool include not only the target service type corresponding to the service field, but also the target content type corresponding to the semantic field. Target service type analysis may be performed based on the target service type and target content type. The semantic field is a field used to distinguish utterances of different subcategories under the service field and is used to implement different content types. The dominant content type is a content type with a high user rating, which can be understood as a content type whose user rating meets preset criteria.
[0127] In one example, in step S601, when the multimedia system responds to a target voice command by using a recommended application, particularly by using a recently used application, the multimedia system needs to query a pre-configured dominant content type information table to determine the dominant content type of the recommended application. The dominant content type can be understood as a high-quality program for the recommended application, such as a comedy show. In this example, the multimedia system can further perform matching between the dominant content type of the recommended application and the target content type in the semantic analysis results, and based on the analysis results, determine what content should be displayed in the recommended application.
[0128] For example, in step S602, when the dominant content type matches the target content type, the multimedia system can determine that the target content type of interest to the user is the dominant content type in the recommended application, for example, a high-quality program in the recommended application. In this way, the multimedia system responds to the target voice command by using the recommended application, thereby making the user more engaging and improving user retention.
[0129] Advantageous content recommendation information is relevant information for applications where the advantageous content type is the target content type.
[0130] In one example, in step S603, if the dominant content type does not match the target content type, the multimedia system determines that the target content type of interest to the user is not a dominant content type in the recommended application, for example, not a high-quality program in the recommended application. In this way, the multimedia system can compare the dominant content types of candidate applications other than the recommended application with the target content type to obtain dominant content recommendation information, and then respond to the target voice command by using the recommended application and prompt for the dominant content recommendation information. In one embodiment, the target voice command can be responded to by using the recommended application, thereby making the user more engaging and improving user retention. In another embodiment, the user can also be made aware of the dominant content recommendation information and determine whether an application whose dominant content type matches the target content type should be used to respond to the target voice command, further improving user retention.
[0131] As shown in Figure 10, the multimedia system can determine a recently used application as recommended application APP1 and perform matching between the target content type (content in the semantic field) in the semantic analysis results and the dominant content type of recommended application APP1. If the target content type matches the dominant content type of recommended application APP1, recommended application APP1 is used to respond to the target voice command. If the target content type does not match the dominant content type of recommended application APP1, matching is performed between the target content type (content in the semantic field) and the dominant content type of another candidate application in the application list. If there is a candidate application whose dominant content type matches the target content type, that candidate application is determined as target application APP2, dominant content recommendation information (e.g., APP2 has higher quality programming) is generated, recommended application APP1 is controlled to respond to the target voice command, and the dominant content recommendation information (e.g., APP2 has higher quality programming) is displayed. If no candidate applications exist whose dominant content type matches the target content type, the recommended application APP1 is invoked to respond to the target voice command, and there is no need to be prompted for dominant content recommendation information.
[0132] In the voice control method provided in this embodiment, based on a comparison between the preferred content type of the recommended application and the target content type of the semantic analysis results, it is determined whether preferred content recommendation information should be displayed in addition to using the recommended application to respond to the target voice command, thereby satisfying user preferences and improving user retention.
[0133] In one embodiment, as shown in Figure 7, step S603, specifically the acquisition of preferential content recommendation information, includes the following steps.
[0134] S701: A candidate application other than the recommended application in the application list is determined as the second application, and a dominant content type corresponding to the second application is acquired.
[0135] S702: Based on the preferred content type and target content type corresponding to the second application, preferred content recommendation information is obtained.
[0136] In one example, in step S701, the multimedia system can determine all candidate applications in the application list other than the recommended application as the second application, and then query a pre-configured dominant content type information table in the system to determine the dominant content type of the second application.
[0137] In one example, after acquiring the dominant content type corresponding to the second application in step S702, the multimedia system can analyze and determine the second application corresponding to the dominant content type most similar to the target content type, based on the dominant content type corresponding to the second application and the target content type determined based on the semantic analysis results. Based on the second application with the most similar dominant content type, it can generate corresponding dominant content recommendation information. The dominant content recommendation information is used to display information about the dominant content type in the second application that is most similar to the target content type, thereby enabling the user to know the dominant content recommendation information and reminding them whether the second application with the most similar dominant content type should be used to respond to the target voice command, which can further improve user retention.
[0138] In this embodiment, when the dominant content type corresponding to the recommended application does not match the target content type, a candidate application other than the recommended application may be determined as the second application to avoid repeated calculations. This helps to conserve computational resources and improve processing efficiency. Dominant content recommendation information is generated based on the target content type and the dominant content type corresponding to the second application, thereby enabling the user to know the dominant content recommendation information and determine whether the second application with the most similar dominant content type should be used to respond to the target voice command, which can further improve user retention.
[0139] In one embodiment, as shown in Figure 8, step S702, specifically, obtaining preferential content recommendation information based on the preferential content type and target content type corresponding to the second application, includes the following steps:
[0140] S801: Similarity calculations are performed on the dominant content type and target content type corresponding to the second application to obtain the content type similarity for the second application.
[0141] S802: The second application with the best content-type similarity is determined to be the target application.
[0142] S803: Based on the target application, preferential content recommendation information is obtained.
[0143] In one example, after acquiring all second applications in step S801, the multimedia system can obtain content type similarity for each second application by performing similarity calculations between the dominant content type and target content type corresponding to the second application, using a cosine similarity algorithm, but not limited to one. The content type similarity is used to reflect the similarity between the target content type and the dominant content type.
[0144] In one example, after obtaining content type similarity for all second applications in step S802, the multimedia system can compare all content type similarities and determine the second application with the best content type similarity as the target application. In this specification, the target application can be understood as a candidate application whose dominant content type is most similar to the target content type. In other words, the target application can satisfy the basic user requirements corresponding to the target service type and ensure that the target content type of the target application is the dominant content type. This helps to improve user preference and further enhance user retention.
[0145] In one example, after determining the target application in step S803, the multimedia system can obtain preferential content recommendation information by filling in the program name or another unique identifier corresponding to the target application in a pre-configured recommendation information template. This allows the user to know the preferential content recommendation information and determine whether an application whose preferential content type matches the target content type should be used to respond to the target voice command, thereby further improving user retention.
[0146] In this embodiment, a second application having the best content type similarity and dominant content type corresponding to the target content type is selected and determined as the target application, generating dominant content recommendation information that thereby satisfies the basic user requirements corresponding to the target service type and also ensures that the target content type of the target application is a dominant content type. This helps to improve user preference and further increase user retention.
[0147] In one embodiment, the application list includes at least one candidate application, and each candidate application includes at least one current content type.
[0148] As shown in Figure 9, before step S105, specifically before the target voice command is responded to by using the recommended application, the voice control method further includes the following steps:
[0149] S901: Current traffic and current user ratings are obtained for each current content type within the candidate application.
[0150] S902: Based on the current traffic and user ratings corresponding to each current content type, an overall rating is obtained for the current content type.
[0151] S903: The dominant content type for a candidate application is determined based on the overall rating corresponding to at least one current content type.
[0152] A candidate application is an application in the application list that corresponds to the target service type. The current content type is the content type currently set by the candidate application, and may include, but is not limited to, crosstalk, sketch, or another content type.
[0153] For example, in step S901, the multimedia system can further acquire current traffic and current user ratings for each current content type in each candidate application through real-time statistical collection. Current traffic is traffic within a preset period prior to the present time and can reflect the amount of users who accessed the current content type within that preset period and can reflect, to some extent, the user preference for the content of the current content type. Current user ratings are the rating values at the present time, or may be limited to rating values within a preset period prior to the present time and reflect the user preference for the content of the current content type.
[0154] For example, after obtaining the current traffic and current user ratings corresponding to each current content type in step S902, the multimedia system can perform weighting or other calculations on the current traffic and current user ratings to determine the overall rating corresponding to the current content type. For instance, the multimedia system can first perform normalization on the current traffic and current user ratings to obtain traffic normalized values and rating normalized values, respectively, and then perform weighting by referencing pre-configured traffic weights and pre-configured rating weights to obtain the overall rating corresponding to each current content type. Thus, the overall rating can reflect the user preference for the content of the current content type, or the degree of positive acceptance of the content of the current content type.
[0155] In one example, after receiving an overall rating corresponding to at least one current content type in step S903, the multimedia system can process the overall rating using a preset priority rating criterion to determine the dominant content type for the application being processed from at least one current content type. For example, the multimedia system may determine a current content type whose overall rating is higher than a preset rating as the dominant content type for the application being processed. The preset rating as used herein is a pre-set rating used to evaluate whether the dominant content type criterion is met. In another example, the multimedia system may, alternatively, determine the first N (N≧1) current content types with high overall ratings as the dominant content types for the application being processed.
[0156] In this embodiment, before a recommended application is used to respond to a target voice command, the overall rating of each current content type can be determined based on the current traffic and current user rating corresponding to the current content type, and the dominant content type corresponding to each application being processed can be updated. Therefore, when an application being processed (including, but not limited to, a recommended application) is used to respond to a target voice command based on the dominant content type of the application being processed, user preferences are better satisfied. This helps to improve user preference and further increase user retention.
[0157] It should be understood that the step sequence numbers do not represent the execution sequence in the embodiments described above. The execution sequence of a process should be determined based on the function and internal logic of the process and should not impose any limitations on the implementation process of the embodiments of this disclosure.
[0158] In one embodiment, a multimedia system is provided. The multimedia system includes memory, a processor, and a computer program stored in memory and running on the processor. The voice control method in the above-described embodiment, for example, steps S101 to S105 shown in Figure 1, or steps shown in Figures 2 to 8, are performed when the processor executes the computer program. To avoid repetition, further details are not described herein.
[0159] In one embodiment, a vehicle is provided. The vehicle includes the multimedia system of the embodiment described above. The multimedia system may perform the voice control method of the embodiment described above, for example, steps S101-S105 shown in Figure 1, or steps shown in Figures 2-8. To avoid repetition, further details are not described here.
[0160] In one embodiment, a computer-readable storage medium is provided. The computer-readable storage medium stores a computer program. The voice control method in the above-described embodiment, for example, steps S101 to S105 shown in Figure 1, or steps shown in Figures 2 to 8, are performed when the computer program is executed by the processor. To avoid repetition, further details are not described herein.
[0161] Those skilled in the art will understand that all or part of the steps of the methods in the embodiments described above may be performed by a computer program instructing the relevant hardware. The computer program may be stored on a non-volatile computer-readable storage medium. When the computer program is executed, it may include the steps of the embodiments of the methods described above. All references to memory, storage, databases, or other media used in the embodiments provided in this disclosure may include non-volatile memory and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random-access memory (RAM) or external cache. For illustrative purposes only, RAM can be acquired in multiple forms, including but not limited to static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data-rate SDRAM (DDRSDRAM), extended SDRAM (ESDRAM), synchlink DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct Rambus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM).
[0162] Those skilled in the art will clearly understand that, for the sake of simplicity and clarity, the division into functional units or modules described above is used merely as an example for illustrative purposes. In actual application, the aforementioned functions may be allocated to various functional units or modules for implementation based on requirements. In other words, the internal structure of the device is divided into various functional units or modules to implement all or some of the functions described above.
[0163] The embodiments described above are intended solely to illustrate the technical solutions of the Disclosure and are not intended to limit the Disclosure. While the Disclosure is described in detail with reference to the embodiments described above, it should be understood by those skilled in the art that the technical solutions described in the embodiments above may still be modified, or some of the technical features may be replaced with equivalent substitutions. Such modifications or substitutions shall not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the Disclosure and shall remain within the scope of the Disclosure.
Claims
1. Perform semantic analysis on the received target voice command and obtain the semantic analysis results, and If the semantic analysis result matches a preset keyword, the system responds to the target voice command by using the application corresponding to the preset keyword. or Perform semantic analysis on the received target voice command and obtain the semantic analysis results. If the semantic analysis result does not match the preset keyword, the target service type is determined based on the semantic analysis result, and a list of applications corresponding to the target service type is obtained. The first recommendation policy is used to determine the recommended application from the application list, and Responding to the target voice command by using the recommended application. Equipped with, By using the first recommendation policy, the recommended applications can be determined from the application list. If the application list includes a foreground application, the foreground application is determined to be the recommended application. If the application list does not include the foreground application, but includes the background application, the background application shall be determined as the recommended application. If the application list does not include the foreground application and the background application, and the application list includes recently used applications, then the recommended application is determined from the application list by using a second recommendation policy, and If the application list does not include the foreground application, the background application, and the recently used application, a preset application is determined to be the recommended application. Equipped with, The recently used application is an application that was used within a preset period prior to the present time. By using the second recommendation policy, the recommended applications can be determined from the application list. Based on the semantic analysis results, determine the target content type and determine whether the recently used application contains content corresponding to the target content type. If the recently used application contains the content corresponding to the target content type, the recently used application is determined to be the recommended application, or If the recently used application does not have the content corresponding to the target content type, determine whether there is a candidate application for the content corresponding to the target content type, and If a candidate application exists that has the content corresponding to the target content type, the candidate application is determined to be the recommended application, or If no candidate application exists that contains the content corresponding to the target content type, the preset application is determined to be the recommended application. A voice control method comprising the following features.
2. By using the aforementioned recommended application, the response to the target voice command is made Based on the semantic analysis results, determine the target content type, and determine whether the recommended application has the content corresponding to the target content type, and If the recommended application includes the content corresponding to the target content type, then using the recommended application will result in responding to the target voice command, or If the recommended application does not have the content corresponding to the target content type, update the candidate application that has the target content type to become the recommended application, and respond to the target voice command by using the updated recommended application. The voice control method according to claim 1, comprising:
3. Updating the candidate application having the target content type to become the recommended application is The candidate application having the target content type is determined as the first application, and the dominant content type of the first application is obtained, and Updating the first application, whose dominant content type is the target content type, to become the recommended application. The voice control method according to claim 2, comprising:
4. The semantic analysis results include the target content type, By using the aforementioned recommended application, the response to the target voice command is made To acquire a superior content type corresponding to the recommended application, and to perform matching between the superior content type corresponding to the recommended application and the target content type, and If the preferred content type corresponding to the recommended application matches the target content type, the recommended application will be used to respond to the target voice command, or If the preferred content type corresponding to the recommended application does not match the target content type, obtain preferred content recommendation information, respond to the target voice command by using the recommended application, and prompt for the preferred content recommendation information. The voice control method according to claim 1, comprising:
5. To obtain the aforementioned superior content recommendation information, Select a candidate application other than the recommended application in the application list as a second application, acquire a dominant content type corresponding to the second application, and To obtain the preferred content recommendation information based on the preferred content type and the target content type corresponding to the second application. The voice control method according to claim 4, comprising:
6. The acquisition of the preferred content recommendation information based on the preferred content type and the target content type corresponding to the second application is as follows: Similarity calculations are performed on the preferred content type and the target content type corresponding to the second application to obtain the content type similarity corresponding to the second application. Determining the second application having the best content type similarity as the target application, and Based on the aforementioned target application, obtain the aforementioned preferential content recommendation information. The voice control method according to claim 5, comprising:
7. The application list comprises at least one candidate application, and each of the candidate applications comprises at least one current content type. Before responding to the target voice command by using the recommended application, the voice control method To obtain current traffic and current user ratings corresponding to each current content type within the candidate application, Based on the current traffic and current user ratings corresponding to each current content type, to obtain an overall rating corresponding to the current content type, and Determining the dominant content type of the candidate application based on the overall rating corresponding to at least one current content type. The voice control method according to claim 4, further comprising:
8. A multimedia system comprising memory, a processor, and a computer program stored in the memory and running on the processor, wherein when the processor executes the computer program, the voice control method described in claim 1 is performed.
9. A vehicle equipped with the multimedia system described in claim 8.
10. A computer-readable storage medium for storing a computer program, wherein when the computer program is executed by a processor, the voice control method described in claim 1 is performed.
Citation Information
Patent Citations
Developer voice actions system
JP2019144598A
Voice control of media playback system
JP2019168696A
Multiple Voice Services
JP2019533182A