Voice control method, multimedia system, automobile and storage medium

By performing semantic parsing and recommendation strategy selection on the voice commands of the in-vehicle multimedia system, the problem of user preference application response was solved, thereby improving user experience and retention rate.

CN117373440BActive Publication Date: 2026-05-01BYD CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BYD CO LTD
Filing Date
2022-06-30
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing in-vehicle multimedia systems cannot accurately respond to user-preferred applications when receiving voice commands, resulting in a poor user experience.

Method used

By performing semantic parsing on the received voice commands, obtaining the semantic parsing results, and matching preset keywords or determining the target business type based on the results, a recommendation strategy is adopted to select a suitable application from the application list to respond to the voice commands.

Benefits of technology

It improved the user experience, increased user retention of the in-vehicle multimedia system, and met users' personalized needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117373440B_ABST
    Figure CN117373440B_ABST
Patent Text Reader

Abstract

The application discloses a voice control method, a multimedia system, a car and a storage medium. The method comprises the following steps: performing semantic analysis on a received target voice instruction to obtain a semantic analysis result; if the semantic analysis result matches a preset keyword, an application program corresponding to the preset keyword is used to respond to the target voice instruction; if the semantic analysis result does not match the preset keyword, a target business type is determined according to the semantic analysis result, an application program list corresponding to the target business type is obtained, a first recommendation strategy is used to determine a recommended application program from the application program list, and the recommended application program is used to respond to the target voice instruction. The method can use the recommended application program to respond to the target voice instruction, meet user preferences and improve user retention rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of voice processing technology, and more particularly to a voice control method, a multimedia system, an automobile, and a storage medium. Background Technology

[0002] With the rapid development of internet and automotive technologies, users can now surf the internet using in-vehicle multimedia systems. However, existing in-vehicle multimedia systems often have multiple similar applications (such as Kugou and Kuwo). When a user issues a voice command, the system directly calls the default application, which is not the user's preferred application, resulting in a poor user experience. Summary of the Invention

[0003] This invention provides a voice control method, a multimedia system, a vehicle, and a storage medium to solve the problem of not being able to respond to voice control using user-preferred applications.

[0004] This invention provides a voice control method, including:

[0005] Perform semantic parsing on the received target voice command and obtain the semantic parsing result;

[0006] If the semantic parsing result matches a preset keyword, then the application corresponding to the preset keyword is used to respond to the target voice command;

[0007] If the semantic parsing result does not match the preset keywords, the target business type is determined based on the semantic parsing result, and the application list corresponding to the target business type is obtained;

[0008] A first recommendation strategy is adopted to determine recommended applications from the list of applications;

[0009] The recommended application is used to respond to the target voice command.

[0010] Preferably, the step of employing a first recommendation strategy to determine recommended applications from the application list includes:

[0011] If the application list includes a foreground application, then the foreground application is identified as a recommended application;

[0012] If the application list does not include foreground applications, but the application list includes background applications, then the background applications are identified as recommended applications.

[0013] If the application list does not include foreground and background applications, but includes recently used applications, then a second recommendation strategy is adopted to determine recommended applications from the application list.

[0014] If the application list does not include foreground applications, background applications, and recently used applications, then the default application will be determined as the recommended application.

[0015] The recently used application refers to an application that was used within a preset period prior to the current moment.

[0016] Preferably, a second recommendation strategy is adopted to determine recommended applications from the application list;

[0017] Based on the semantic parsing results, the target content type is determined, and it is determined whether the recently used application contains content corresponding to the target content type.

[0018] If the recently used application contains content corresponding to the target content type, then the recently used application is identified as a recommended application;

[0019] If the recently used applications do not contain content corresponding to the target content type, then determine whether there are candidate applications with content corresponding to the target content type;

[0020] If there are candidate applications that contain content corresponding to the target content type, then the candidate applications are determined as recommended applications;

[0021] If no candidate application contains content corresponding to the target content type, the preset application will be selected as the recommended application.

[0022] Preferably, the step of using the recommended application to respond to the target voice command includes:

[0023] Based on the semantic parsing results, the target content type is determined, and it is determined whether the recommendation application contains content corresponding to the target content type.

[0024] If the recommended application contains content corresponding to the target content type, then the recommended application is used to respond to the target voice command;

[0025] If the recommended application does not contain content corresponding to the target content type, then the candidate application that contains the target content type will be updated as the recommended application, and the updated recommended application will be used to respond to the target voice command.

[0026] Preferably, updating the candidate applications containing the target content type to recommended applications includes:

[0027] The candidate applications containing the target content type are identified as the first application, and the advantageous content types of the first application are obtained;

[0028] The first application whose advantageous content type is the target content type is updated to a recommended application.

[0029] Preferably, the semantic parsing result includes the target content type;

[0030] The step of using the recommended application to respond to the target voice command includes:

[0031] Obtain the advantageous content type corresponding to the recommended application, and match the advantageous content type corresponding to the recommended application with the target content type;

[0032] If the recommended application's content type matches the target content type, then the recommended application is used to respond to the target voice command;

[0033] If the advantageous content type corresponding to the recommended application does not match the target content type, then the advantageous content recommendation information is obtained, the recommended application is used to respond to the target voice command, and the advantageous content recommendation information is displayed.

[0034] Preferably, obtaining the recommended content information includes:

[0035] The candidate applications in the application list, excluding the recommended applications, are identified as the second applications, and the advantageous content types corresponding to the second applications are obtained.

[0036] Based on the advantageous content type corresponding to the second application and the target content type, advantageous content recommendation information is obtained.

[0037] Preferably, obtaining the recommended content information based on the advantageous content type corresponding to the second application and the target content type includes:

[0038] The similarity between the advantageous content type corresponding to the second application and the target content type is calculated to obtain the content type similarity of the second application.

[0039] The second application with the highest content type similarity is identified as the target application;

[0040] Based on the target application, obtain the recommended information for the advantageous content.

[0041] In one embodiment, the application list includes at least one candidate application, and each candidate application includes at least one current content type;

[0042] Before employing the recommended application and responding to the target voice command, the voice control command further includes:

[0043] Obtain the current number of visits and the current user rating for each of the current content types in the candidate applications;

[0044] Based on the current number of visits and the current user rating for each of the current content types, obtain the comprehensive rating corresponding to the current content type;

[0045] The advantageous content types of the candidate applications are determined based on the comprehensive score corresponding to at least one current content type.

[0046] This invention provides a multimedia system including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the above-described voice control method.

[0047] This invention provides a car that includes the aforementioned multimedia system.

[0048] This invention provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned voice control method.

[0049] When the semantic parsing results match the preset keywords, the aforementioned voice control methods, multimedia systems, automobiles, and storage media can intuitively reflect user preferences. Therefore, the application corresponding to the preset keywords can be used to respond to the target voice command, which helps to improve the user experience and thus increase user retention. When the semantic parsing results do not match the preset keywords, recommended applications are determined from the list of applications corresponding to the target business type. Using the recommended applications to respond to the target voice command can, to some extent, satisfy user preferences, improve the user experience, and thus increase user retention. Attached Figure Description

[0050] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 This is a schematic diagram of an application environment of the voice control method in one embodiment of the present invention;

[0052] Figure 2 This is a flowchart of a voice control method according to an embodiment of the present invention;

[0053] Figure 3 This is another flowchart of the voice control method in one embodiment of the present invention;

[0054] Figure 4 This is another flowchart of the voice control method in one embodiment of the present invention;

[0055] Figure 5 This is another flowchart of the voice control method in one embodiment of the present invention;

[0056] Figure 6 This is another flowchart of the voice control method in one embodiment of the present invention;

[0057] Figure 7 This is another flowchart of the voice control method in one embodiment of the present invention;

[0058] Figure 8 This is another flowchart of the voice control method in one embodiment of the present invention;

[0059] Figure 9 This is another flowchart of the voice control method in one embodiment of the present invention;

[0060] Figure 10 This is another flowchart of the voice control method in one embodiment of the present invention;

[0061] Figure 11 This is another flowchart of the voice control method in one embodiment of the present invention. Detailed Implementation

[0062] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0063] The voice control method provided in this embodiment of the invention can be applied to, for example... Figure 1The application environment shown is as follows. Specifically, this voice control method is applied in a voice control system (hereinafter referred to as "the system"). The voice control system can be a multimedia system in a car (i.e., a vehicle infotainment system) or other systems that can realize voice control. It can call up applications that meet the user's preferences to respond based on the target voice command collected by the user, thereby improving the user experience.

[0064] In one embodiment, a voice control method is provided, such as... Figure 1 As shown, the voice control method includes:

[0065] S101: Perform semantic parsing on the received target speech command and obtain the semantic parsing result;

[0066] S102: If the semantic parsing result matches the preset keywords, then the application corresponding to the preset keywords is used to respond to the target voice command;

[0067] S103: If the semantic parsing result does not match the preset keywords, then determine the target business type based on the semantic parsing result and obtain the application list corresponding to the target business type;

[0068] S104: Using the first recommendation strategy, determine recommended applications from the application list;

[0069] S105: Uses a recommended application to respond to target voice commands.

[0070] The target voice command is the instruction input by the user via voice to instruct a specific application to function. The semantic parsing result is the result of semantic parsing of the target voice command.

[0071] As an example, in step S101, the multimedia system can receive a target voice command input by the user, call a pre-set speech recognition tool, and perform semantic parsing on the target voice command. This process includes speech recognition and semantic parsing to obtain the semantic parsing result. The speech recognition tool can be any tool that performs semantic parsing, including but not limited to iFlytek speech recognition tools. For example, the multimedia system can receive target voice commands input by the user in voice form, such as "play crosstalk," "I want to listen to financial news," "I want to listen to ghost stories," "popular music," and "play a movie," perform semantic parsing on the target voice command, and obtain the semantic parsing result in text form.

[0072] Among them, preset keywords refer to keywords that are pre-set and bound to a specific application.

[0073] As an example, in step S102, after obtaining the semantic parsing result, the multimedia system can match the result with preset keywords. When the semantic parsing result matches the preset keywords, the application corresponding to the preset keywords can be directly used to respond to the target voice command. For example, if the multimedia system obtains the semantic parsing result "I want to open Tencent Video," and this result matches the preset keyword "Tencent Video," then the Tencent Video APP will be used to respond to the target voice command. Since the semantic parsing result matches the preset keywords, it can intuitively reflect user preferences. Directly using the application corresponding to the preset keywords to respond to the target voice command satisfies user preferences, improves user experience, and thus increases user retention.

[0074] The target service type refers to the service type corresponding to the function of the user's dialogue in the semantic parsing results. The application list refers to a list consisting of at least one candidate application. A candidate application is an application that can display content corresponding to the target service type.

[0075] As an example, in step S103, after obtaining the semantic parsing result, the multimedia system can match the semantic parsing result with preset keywords. If the semantic parsing result does not match the preset keywords, the system can determine the target service type based on the semantic parsing result, and then query the system for at least one candidate application containing the target service type. Based on at least one candidate application, the system obtains an application list. In this example, the multimedia system obtains the field information corresponding to the service type field from the semantic parsing result and determines it as the target service type. For example, the semantic parsing result obtained by the speech recognition tool includes the field information corresponding to the Service field. The field information corresponding to the Service field is determined as the target service type. The Service field is used to reflect the service type corresponding to the function to which the user's speech belongs, including but not limited to Air control, Map, Music, and Internet Radio.

[0076] The first recommendation strategy is a pre-set strategy used to determine the recommended applications based on the business type. Recommended applications are those applications recommended to users.

[0077] As an example, in step S104, after the multimedia system determines the corresponding application list based on the target service type, it can invoke a pre-set first recommendation strategy to select an application from the application list that best matches the user's preferences and designate it as a recommended application. In this example, the first recommendation strategy can analyze user preferences based on user behavior or other information to determine the processing strategy for recommending applications that match the user's preferences.

[0078] As an example, in step S105, after determining the recommended application, the multimedia system uses the recommended application to respond to the target voice command; for example, it can play audio based on the recommended application. Understandably, since the recommended application satisfies user preferences to some extent, using the recommended application to respond to the target voice command can improve the user experience and thus increase user retention.

[0079] like Figure 10 As shown, after receiving a target voice command, the multimedia system can call a speech recognition tool to perform semantic analysis on the target voice command and obtain the semantic analysis result. Then, the semantic analysis result is matched with preset keywords. If the match is successful, the application corresponding to the preset keyword is used to respond to the target voice command. If the match fails, the content of the Service field in the semantic analysis result is determined as the target business type, and then a list of applications corresponding to the target business type is obtained. Recommended applications are selected from the application list and used to respond to the target voice command. This ensures that the response process of the target voice command meets user preferences, which can improve the user experience and thus increase user retention.

[0080] In this embodiment, when the semantic parsing result matches the preset keywords, it can intuitively reflect user preferences. Therefore, the application corresponding to the preset keywords can be used to respond to the target voice command, which helps to improve the user experience and thus improve the user retention rate. When the semantic parsing result does not match the preset keywords, a recommended application is determined from the application list corresponding to the target business type. The recommended application is used to respond to the target voice command, which can satisfy user preferences to a certain extent, improve the user experience, and thus improve the user retention rate.

[0081] In one embodiment, the application list includes at least one candidate application corresponding to the same target service type;

[0082] Step S104, namely, using the first recommendation strategy to determine recommended applications from the application list, includes: using the first recommendation strategy to determine recommended applications from at least one candidate application.

[0083] The application list refers to a list consisting of at least one candidate application. A candidate application is an application that can display content corresponding to the target business type. In this example, the application list includes at least one candidate application corresponding to the same target business type.

[0084] As an example, after determining the application list corresponding to the target business type based on the semantic parsing results, the multimedia system can identify at least one candidate application from the application list, and then execute a pre-set first recommendation strategy to analyze the at least one candidate application. From the at least one candidate application, the system can select the recommended application that best meets the user's preferences and identify it as the recommended application. In order to use the recommended application to respond to the target voice command, the user can browse the content they are interested in through the recommended application, thereby improving the user retention rate.

[0085] In this embodiment, the application list corresponding to the target business type includes at least one candidate application corresponding to the same target business type. All of these applications meet the basic user needs (i.e., the target business type) extracted from the semantic parsing results. Then, the recommended application that best meets the user's preferences is selected from the at least one candidate application. This makes the recommended application more in line with the user's preferences when responding to the target voice command, thereby improving the user retention rate.

[0086] In one embodiment, such as Figure 2 As shown, step S104, which involves using the first recommendation strategy to determine recommended applications from the application list, includes:

[0087] S201: If the application list includes a foreground application, then the foreground application is identified as a recommended application;

[0088] S202: If the application list does not include foreground applications, but the application list includes background applications, then the background applications will be identified as recommended applications.

[0089] S203: If the application list does not include foreground and background applications, but includes recently used applications, then the second recommendation strategy is adopted to analyze the recently used applications and determine the recommended applications.

[0090] S204: If the application list does not include foreground applications, background applications, and recently used applications, then the default application will be selected as the recommended application.

[0091] Among them, recently used applications refer to applications used within a preset period prior to the current moment.

[0092] In this context, foreground applications refer to applications currently running in the foreground, meaning applications that are visible to the user. Background applications are applications currently running in the background, meaning applications that continue to provide minor services even after being closed. Recently used applications refer to applications used within a preset period prior to the current moment. This preset period is user-defined and can be one day or one week. Default applications are applications that are set by default, excluding foreground, background, and recently used applications; these can be system-installed applications.

[0093] As an example, in step S201, after the multimedia system obtains the application list, it needs to determine whether the application list includes a foreground application. If the application list includes a foreground application, it means that there is a candidate application running in the foreground at the current moment, indicating that the user is browsing the foreground application, which reflects the user's preference for the foreground application. Therefore, the foreground application can be identified as a recommended application to obtain the recommended application with the highest user preference.

[0094] As an example, in step S202, when the multimedia system does not include foreground applications in the application list, it determines whether the application list includes background applications. If the application list includes background applications, it means that although there are no candidate applications running in the foreground at the current moment, there are candidate applications running in the background. The background application can be understood as an application that was opened before the current moment and has not been completely closed. It can also reflect that the user prefers the background application to a certain extent. Therefore, the background application can be identified as a recommended application to obtain a recommended application that meets the user's preference to a higher degree.

[0095] The second recommendation strategy is a pre-set strategy used to analyze recently used applications to determine recommended applications.

[0096] As an example, in step S203, when the multimedia system does not include foreground and background applications in the application list, it determines whether the application list includes recently used applications. If the application list includes recently used applications, it means that no candidate applications are currently running in the foreground or background, but recently used applications were running before the current time and have cached running records, reflecting that the user has browsed recently used applications and that user preferences exist. Therefore, the second recommendation strategy is adopted to analyze recently used applications, determine recommended applications, and obtain recommended applications with a moderate degree of user preference. In this example, the second recommendation strategy is used to analyze recently used applications, which can determine recently used applications that meet preset conditions as recommended applications. These preset conditions are pre-set conditions that can be determined based on user preferences.

[0097] As an example, in step S204, when the multimedia system does not include foreground applications, background applications, and recently used applications in the application list, it can determine the system's default preset applications as recommended applications. These preset applications are any one of the candidate applications, corresponding to the target business type, to obtain recommended applications that meet the user's basic needs.

[0098] like Figure 10 As shown, the multimedia system first determines whether the application list contains a foreground application. If it does, the foreground application is used as the recommended application to respond to the target voice command. If it does not contain a foreground application, the system then determines whether the application list contains a background application. If it does, the background application is used as the recommended application to respond to the target voice command. If it does not contain a background application, the system then determines whether the application list contains recently used applications. If it does, a second recommendation strategy is used to further determine the recommended application. If it does not contain recently used applications, the default application is determined as the recommended application to respond to the target voice command.

[0099] In this embodiment, based on the user's preference level from high to low, foreground applications, background applications, recently used applications, and preset applications are selected as recommended applications in sequence. This makes the recommended applications respond to target voice commands in a way that better matches the user's preferences, thereby improving user retention rate.

[0100] In one embodiment, such as Figure 3 As shown, step S203, which involves using the second recommendation strategy to determine recommended applications from the application list, includes:

[0101] S301: Determine the target content type based on the semantic parsing results, and determine whether the recently used application contains content corresponding to the target content type;

[0102] S302: If a recently used application contains content corresponding to the target content type, then the recently used application will be identified as a recommended application.

[0103] S303: If the recently used application does not contain content corresponding to the target content type, determine whether there is a candidate application with content corresponding to the target content type.

[0104] S304: If there are candidate applications that contain content corresponding to the target content type, then the candidate applications will be identified as recommended applications.

[0105] S305: If there is no candidate application containing content corresponding to the target content type, the default application will be selected as the recommended application.

[0106] The target content type refers to the content type identified in the semantic parsing results; in other words, it's the content type identified from the target voice command. The semantic parsing results obtained using speech recognition tools include not only the target business type corresponding to the Service field but also the target content type corresponding to the Semantic field. Target business type analysis can be performed based on both the target business type and the target content type. The Semantic field is used under the Service field to distinguish different subcategories of speech, thus implementing different content types. For example, under the InternetRadio Service field, it is further divided into Semantic fields such as program, tags, category, and presenter. Understandably, different user speech in the semantic parsing results will result in different target content types corresponding to the Semantic fields. For example, if the semantic parsing result is "listen to comedy radio," then the category field will be filled; and if the semantic parsing result is "listen to XX's crosstalk," then the presenter field will be filled.

[0107] As an example, in step S301, when the multimedia system needs to use the second recommendation strategy to analyze recently used applications, it first needs to determine the target content type based on the semantic parsing results, and then determine whether the recently used applications contain content corresponding to the target content type, so as to determine whether the recently used applications can meet the user's needs.

[0108] As an example, in step S302, when the multimedia system finds that a recently used application contains content of the target content type, it indicates that the recently used application can display content of the target content type required by the user. Since the recently used application is an application that the user has used within a preset period before the current moment, it reflects the user's preferences to a certain extent. Therefore, the recently used application can be identified as a recommended application so that the use of the recommended application to respond to the target voice command is more in line with the user's preferences.

[0109] As an example, in step S303, when the multimedia system finds that the recently used application does not contain content of the target content type, it indicates that the recently used application cannot display the target content type required by the user. In this case, it is necessary to determine whether other candidate applications besides the recently used application contain content of the target content type from at least one candidate application corresponding to the target business type determined by the semantic parsing structure, so as to determine the recommended application based on the judgment result.

[0110] As an example, in step S304, when the multimedia system has a candidate application that contains content corresponding to the target content type, it indicates that the business type of the candidate application contains the target business type determined by the semantic parsing result, and the candidate application contains content of the target content type determined by the semantic parsing result, so that content that meets the user's needs can be displayed. Therefore, the candidate application can be determined as a recommended application.

[0111] As an example, in step S305, if the multimedia system does not have a candidate application that contains content corresponding to the target content type, it indicates that the candidate application does not simultaneously meet the requirements of the target business type and the target content type. In this case, a preset application can be determined as the recommended application. This preset application is any one of the candidate applications that corresponds to the target business type, in order to obtain a recommended application that meets the user's basic needs.

[0112] In this embodiment, based on whether the application contains content of the target content type, recently used applications, candidate applications containing the target content type, and preset applications are selected in sequence to determine the recommended application. This ensures that the determined recommended application can meet user preferences to the greatest extent, so that the recommended application responds to the target voice command in a way that better matches user preferences, thereby improving user retention rate.

[0113] In one embodiment, such as Figure 4 As shown, step S105, which involves using the recommended application to respond to the target voice command, includes:

[0114] S401: Determine the target content type based on the semantic parsing results, and determine whether the recommended application contains content corresponding to the target content type;

[0115] S402: If the recommended application contains content corresponding to the target content type, then the recommended application shall be used to respond to the target voice command;

[0116] S403: If the recommended application does not contain content corresponding to the target content type, then the candidate application that contains the target content type will be updated to the recommended application, and the updated recommended application will be used to respond to the target voice command.

[0117] The target content type refers to the content type identified in the semantic parsing results; in other words, it's the content type identified from the target voice command. The semantic parsing results obtained using speech recognition tools include not only the target business type corresponding to the Service field but also the target content type corresponding to the Semantic field. Target business type analysis can be performed based on both the target business type and the target content type. The Semantic field, under the Service field, is used to distinguish different subcategories of speech, thus implementing different content types.

[0118] As an example, in step S401, when the multimedia system uses a recommended application to respond to the target content type, it can first determine the target content type based on the semantic parsing results, and then determine whether the recommended application contains content of the target content type, so as to determine whether the recommended application can meet the user's needs for the target content type.

[0119] As an example, in step S402, when the multimedia system determines that the recommended application contains content corresponding to the target content type, it considers the recommended application to be able to display the content needed by the user. Since the recommended application is an application with a high user preference, it can be used to respond to the target voice command and display content of the target content type, making it more in line with user preferences and thus improving user retention. For example, if the target content type determined based on the target voice command is to play song 1, and the recommended application contains song 1, then playing song 1 can be used to display the content needed by the user. Using a recommended application that meets user preferences is more in line with user preferences and thus improves user retention.

[0120] As an example, in step S403, when the multimedia system determines that the recommended application cannot display the content required by the user if it does not contain content corresponding to the target content type, it can update the candidate application containing the target content type to the recommended application. The updated recommended application then responds to the target voice command, enabling it to display the target content type, thereby improving user retention. For instance, if the target content type determined by the target voice command is to play song 1, and the recommended application does not contain song 1, then from at least one candidate application in the application list, the candidate application containing song 1 is determined, updated to the recommended application, and then the updated recommended application is used to play song 1, thus displaying the content required by the user and improving user retention.

[0121] In this embodiment, it is determined whether the recommended application needs to be updated based on whether the recommended application contains content of the target content type, so as to ensure that the recommended application can display the content required by the user, so that the system can meet the user's needs when responding, thereby improving the user retention rate.

[0122] In one embodiment, such as Figure 5 As shown, step S403 updates the candidate applications containing the target content type to recommended applications, including:

[0123] S501: Identify the candidate application containing the target content type as the first application and obtain the advantageous content type of the first application;

[0124] S502: Update the first application whose preferred content type is the target content type to the recommended application.

[0125] The first application refers to the candidate application that contains the target content type; in other words, the first application is the application that matches both the target business type and the target content type determined by the semantic analysis results. The dominant content type refers to content types with high user ratings, which can be understood as content types with high user ratings or frequent user visits.

[0126] As an example, in step S501, the multimedia system can select at least one candidate application containing the target content type from the application list as the first application, ensuring that the first application matches the target business type and target content type determined by the semantic parsing results, thus guaranteeing that the first application can display the content required by the user. Furthermore, after determining the first application, the multimedia system also needs to obtain the advantageous content types of the first application; for example, it can query a pre-set advantageous content type information table to determine the advantageous content types of the first application.

[0127] As an example, in step S502, after obtaining the advantageous content type of at least one first application, the multimedia system can match the advantageous content type with the target content type determined by the semantic parsing result. This updates the first application whose advantageous content type is the target content type to a recommended application, allowing the updated recommended application to respond to target voice commands and display content corresponding to the target content type. Since the target content type is the advantageous content type of the updated recommended application, it indicates a higher user rating, making it more attractive to users and improving user retention and experience.

[0128] In this example, if there are at least two first applications whose dominant content type is the target content type, then at least two user ratings of the first applications in the target content type are obtained, and the first application with the highest user rating is selected as the recommended application. This ensures that when the updated recommended application displays content of the target content type, it is easier to attract users and improve user retention and experience.

[0129] In this embodiment, at least one candidate application containing the target content type is identified as the first application. The first application, whose dominant content type is the target content type, is updated as the recommended application. This makes it easier for the updated recommended application to attract users when displaying content of the target content type, thereby improving user retention and user experience.

[0130] In one embodiment, the semantic parsing result includes the target content type;

[0131] like Figure 6 As shown, step S105, which involves using the recommended application to respond to the target voice command, includes:

[0132] S601: Obtain the advantageous content type corresponding to the recommended application, and match the advantageous content type corresponding to the recommended application with the target content type;

[0133] S602: If the advantageous content type of the recommended application matches the target content type, then the recommended application is used to respond to the target voice command;

[0134] S603: If the advantageous content type corresponding to the recommended application does not match the target content type, then obtain the advantageous content recommendation information, adopt the recommended application, respond to the target voice command, and prompt the advantageous content recommendation information.

[0135] The target content type refers to the content type identified in the semantic parsing results; in other words, it's the content type identified from the target voice command. The semantic parsing results obtained using speech recognition tools include not only the target business type corresponding to the Service field but also the target content type corresponding to the Semantic field. Target business type analysis can be performed based on both the target business type and the target content type. The Semantic field, under the Service field, is used to distinguish different subcategories of speech and to implement different content types. The advantageous content type refers to the content type with high user ratings; it can be understood as the content type whose user ratings meet preset standards.

[0136] As an example, in step S601, when the multimedia system responds to a target voice command using a recommended application, especially a recently used application, it needs to query a pre-set advantageous content type information table to determine the advantageous content type of the recommended application. This advantageous content type can be understood as high-quality content from the recommended application, such as comedy shows. In this example, the multimedia system can also match the advantageous content type of the recommended application with the target content type in the semantic parsing results to determine the content to be displayed in the recommended application based on the analysis results.

[0137] As an example, in step S602, when the advantageous content type and the target content type match, the multimedia system can determine that the target content type that the user is interested in is the advantageous content type in the recommended application, such as a high-quality program in the recommended application. By responding to the target voice command through the recommended application, it is easier to attract the user's attention and improve the user retention rate.

[0138] Among them, the advantageous content recommendation information refers to the relevant information of applications whose advantageous content type is the target content type.

[0139] As an example, in step S603, when the multimedia system determines that the target content type and the advantageous content type do not match, the target content type that the user is interested in is not an advantageous content type in the recommended application. For example, it may not be a high-quality program in the recommended application. The system can then compare the advantageous content types of other candidate applications (excluding the recommended application) with the target content type to obtain advantageous content recommendation information. The recommended application can then respond to the target voice command and display the advantageous content recommendation information. On the one hand, responding to the target voice command through the recommended application can attract user attention and improve user retention. On the other hand, it can also let users understand the advantageous content recommendation information so that they can determine whether they need to respond to the target voice command through an application whose advantageous content type matches the target content type, which can further improve user retention.

[0140] like Figure 10 As shown, the multimedia system can identify recently used applications as recommended application APP1, and match the target content type (the content of the Semantic field) in the semantic parsing result with the advantageous content type of recommended application APP1. If the match is successful, recommended application APP1 responds to the target voice command. If the match fails, the target content type (the content of the Semantic field) is matched with the advantageous content type of other candidate applications in the application list. If a candidate application is successfully matched, the target application APP2 is identified, advantageous content recommendation information is generated (e.g., APP2 has better programs), recommended application APP1 is controlled to respond to the target voice command, and advantageous content recommendation information (e.g., APP2 has better programs) is displayed. If no candidate application is successfully matched, recommended application APP1 is invoked to respond to the target voice command without prompting advantageous content recommendation information.

[0141] In the voice control method provided in this embodiment, based on the comparison results between the advantageous content types of the recommended application and the target content types of the semantic parsing results, it is determined whether, in addition to using the recommended application to respond to the target voice command, it is also necessary to display advantageous content recommendation information to meet user preferences and improve user retention.

[0142] In one embodiment, such as Figure 7 As shown, step S603, which involves obtaining recommended content information, includes:

[0143] S701: Select the candidate applications from the application list other than the recommended applications as the second application, and obtain the advantageous content type corresponding to the second application;

[0144] S702: Based on the advantageous content type and target content type corresponding to the second application, obtain advantageous content recommendation information.

[0145] As an example, in step S701, the multimedia system can identify all candidate applications in the application list other than the recommended applications as the second application; then, it can query the system's pre-set advantageous content type information table to determine the advantageous content type of the second application.

[0146] As an example, in step S702, after obtaining the advantageous content type corresponding to the second application, the multimedia system can analyze and determine the second application corresponding to the advantageous content type most similar to the target content type based on the advantageous content type corresponding to the second application and the target content type determined by semantic parsing results. This allows for the formation of corresponding advantageous content recommendation information based on the most similar second application. This advantageous content recommendation information displays the information most similar to the target content type within the second application, enabling users to understand the recommendation information and prompting them whether they need to use the most similar second application to respond to the target voice command. It can also further improve user retention.

[0147] In this embodiment, when the advantageous content type corresponding to the recommended application does not match the target content type, a second application can be determined from the candidate applications other than the recommended application to avoid redundant calculations, which helps to save computing resources and improve processing efficiency. Based on the target content type and the advantageous content type corresponding to the second application, advantageous content recommendation information is formed. Users can understand the advantageous content recommendation information so as to determine which second application is most similar to respond to the target voice command, which can also further improve user retention rate.

[0148] In one embodiment, such as Figure 8 As shown, step S102, which involves obtaining recommended content information based on the advantageous content type and target content type corresponding to the second application, includes:

[0149] S801: Calculate the similarity between the advantageous content type and the target content type corresponding to the second application to obtain the content type similarity of the second application;

[0150] S802: Identify the second application with the highest content type similarity as the target application;

[0151] S803: Based on the target application, obtain recommended content information.

[0152] As an example, in step S801, after acquiring all the second applications, the multimedia system may use, but is not limited to, a cosine similarity algorithm to calculate the similarity between the dominant content type and the target content type corresponding to each second application, thereby obtaining the content type similarity for each second application. This content type similarity is used to reflect the degree of similarity between the target content type and the dominant content type.

[0153] As an example, in step S802, after obtaining the content type similarity of all second applications, the multimedia system can compare all content type similarities and determine the second application with the highest content type similarity as the target application. The target application here can be understood as the candidate application whose advantageous content type is most similar to the target content type. This means it meets the basic user needs corresponding to the target business type and also ensures that its target content type is an advantageous content type, which helps improve user preference and thus increases user retention.

[0154] As an example, in step S803, after the multimedia system determines the target application, it can fill the program name or other unique identifier corresponding to the target application into the pre-set recommendation information template to obtain advantageous content recommendation information, so that users can understand the advantageous content recommendation information and thus determine whether to respond to the target voice command through an application whose advantageous content type matches the target content type, which can also further improve the user retention rate.

[0155] In this embodiment, the second application with the highest similarity between the target content type and the advantageous content type is selected as the target application, thereby forming advantageous content recommendation information. This can meet the basic needs of users corresponding to the target business type and ensure that the target content type is an advantageous content type, which helps to improve user preference and thus improve user retention rate.

[0156] In one embodiment, the application list includes at least one candidate application, and each candidate application includes at least one current content type;

[0157] like Figure 9 As shown, before step S105, before employing the recommended application and responding to the target voice command, the voice control command further includes:

[0158] S901: Get the current number of visits and the current user rating for each current content type in the candidate applications;

[0159] S902: Based on the current number of visits and the current user rating for each current content type, obtain the comprehensive score corresponding to the current content type;

[0160] S903: Determine the advantageous content types of the candidate applications based on the comprehensive score corresponding to at least one current content type.

[0161] Among them, the candidate application refers to the application in the application list, which is the application corresponding to the target business type. The current content type refers to the content type set by the candidate application at the current moment, such as including but not limited to crosstalk, skits, or other content types.

[0162] As an example, in step S901, the multimedia system can also obtain in real time the current access volume and current user rating for each current content type in each candidate application. The current access volume is the access volume within a preset period prior to the current moment, reflecting how many users accessed the current content type within that preset period, and to some extent, reflecting users' liking for the content of the current content type. The current user rating refers to the rating value at the current moment, or it can be limited to the rating value within a preset period prior to the current moment, reflecting users' liking for the content of the current content type.

[0163] As an example, in step S902, the multimedia system, after obtaining the current number of visits and the current user rating for each current content type, can perform weighted processing or other calculations on the current number of visits and the current user rating to determine the comprehensive rating corresponding to the current content type. For example, the multimedia system can first normalize the current number of visits and the current user rating to obtain the normalized visit value and the normalized rating value respectively, and then perform weighted processing in combination with the pre-set visit weight and rating weight to obtain the comprehensive rating corresponding to each current content type, so that the comprehensive rating can reflect the user's liking for the content of the current content type or the degree of its positive reception.

[0164] As an example, in step S903, after receiving a comprehensive score corresponding to at least one current content type, the multimedia system can process the comprehensive score using pre-set advantage evaluation conditions to determine the advantageous content type of the candidate application from at least one current content type. For example, the multimedia system can determine the current content type with a comprehensive score greater than a preset score as the advantageous content type of the candidate application. Here, the preset score refers to a pre-set score used to evaluate whether the standard for an advantageous content type is met. As another example, the multimedia system can also determine the top N current content types with the largest comprehensive scores as the advantageous content types of the candidate application, where N≥1.

[0165] In this embodiment, before the recommended application responds to the target voice command, its comprehensive score can be determined based on the current access volume and current user rating corresponding to each current content type. This is used to update the advantageous content type corresponding to each candidate application, so that when each candidate application (including but not limited to the recommended application) responds to the target voice command, it is more in line with user preferences, which helps to improve the degree of user preference and thus improve user retention rate.

[0166] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0167] In one embodiment, a multimedia system is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the voice control method described in the above embodiment, for example... Figure 1 As shown in S101-S105, or Figures 2 to 8 As shown in the figure, to avoid repetition, it will not be repeated here.

[0168] In one embodiment, a vehicle is provided, including the multimedia system described in the above embodiments, which is capable of executing the voice control method described in the above embodiments, for example... Figure 1 As shown in S101-S105, or Figures 2 to 8 As shown in the figure, to avoid repetition, it will not be repeated here.

[0169] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a processor, the computer program implements the voice control method described in the above embodiment, for example... Figure 1 As shown in S101-S105, or Figures 2 to 8 As shown in the figure, to avoid repetition, it will not be repeated here.

[0170] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0171] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0172] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A voice control method, characterized in that, include: Perform semantic parsing on the received target voice command and obtain the semantic parsing result; If the semantic parsing result matches a preset keyword, the application corresponding to the preset keyword is used to respond to the target voice command; the preset keyword refers to a keyword that is pre-set and bound to a specific application, and the preset keyword can intuitively reflect user preferences; If the semantic parsing result does not match the preset keywords, the target business type is determined based on the semantic parsing result, and the application list corresponding to the target business type is obtained; If the application list includes a foreground application, then the foreground application is identified as a recommended application; if the application list does not include a foreground application, but includes a background application, then the background application is identified as a recommended application; if the application list does not include both foreground and background applications, but includes recently used applications, then a second recommendation strategy is adopted to determine recommended applications from the application list; if the application list does not include foreground, background, and recently used applications, then a preset application is identified as a recommended application; wherein, recently used applications refer to applications used within a preset period prior to the current moment. The recommended application is used to respond to the target voice command.

2. The voice control method as described in claim 1, characterized in that, The second recommendation strategy is adopted to determine recommended applications from the application list; Based on the semantic parsing results, the target content type is determined, and it is determined whether the recently used application contains content corresponding to the target content type. If the recently used application contains content corresponding to the target content type, then the recently used application is identified as a recommended application; If the recently used applications do not contain content corresponding to the target content type, then determine whether there are candidate applications with content corresponding to the target content type; If there are candidate applications that contain content corresponding to the target content type, then the candidate applications are determined as recommended applications; If no candidate application contains content corresponding to the target content type, the preset application will be selected as the recommended application.

3. The voice control method as described in any one of claims 1-2, characterized in that, The step of using the recommended application to respond to the target voice command includes: Based on the semantic parsing results, the target content type is determined, and it is determined whether the recommendation application contains content corresponding to the target content type. If the recommended application contains content corresponding to the target content type, then the recommended application is used to respond to the target voice command; If the recommended application does not contain content corresponding to the target content type, then the candidate application that contains the target content type will be updated as the recommended application, and the updated recommended application will be used to respond to the target voice command.

4. The voice control method as described in claim 3, characterized in that, The step of updating candidate applications containing the target content type to recommended applications includes: The candidate applications containing the target content type are identified as the first application, and the advantageous content types of the first application are obtained; The first application whose advantageous content type is the target content type is updated to a recommended application.

5. The voice control method as described in any one of claims 1-2, characterized in that, The semantic parsing results include the target content type; The step of using the recommended application to respond to the target voice command includes: Obtain the advantageous content type corresponding to the recommended application, and match the advantageous content type corresponding to the recommended application with the target content type; If the recommended application's content type matches the target content type, then the recommended application is used to respond to the target voice command; If the advantageous content type corresponding to the recommended application does not match the target content type, then the advantageous content recommendation information is obtained, the recommended application is used to respond to the target voice command, and the advantageous content recommendation information is displayed.

6. The voice control method as described in claim 5, characterized in that, The acquisition of recommended content information includes: The candidate applications in the application list, excluding the recommended applications, are identified as the second applications, and the advantageous content types corresponding to the second applications are obtained. Based on the advantageous content type corresponding to the second application and the target content type, advantageous content recommendation information is obtained.

7. The voice control method as described in claim 6, characterized in that, The step of obtaining advantageous content recommendation information based on the advantageous content type corresponding to the second application and the target content type includes: The similarity between the advantageous content type corresponding to the second application and the target content type is calculated to obtain the content type similarity of the second application. The second application with the highest content type similarity is identified as the target application; Based on the target application, obtain the recommended information for the advantageous content.

8. The voice control method as described in claim 5, characterized in that, The application list includes at least one candidate application, and each candidate application includes at least one current content type; Before employing the recommended application and responding to the target voice command, the voice control method further includes: Obtain the current number of visits and the current user rating for each of the current content types in the candidate applications; Based on the current number of visits and the current user rating for each of the current content types, obtain the comprehensive rating corresponding to the current content type; The advantageous content types of the candidate applications are determined based on the comprehensive score corresponding to at least one current content type.

9. A multimedia system, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the voice control method as described in any one of claims 1 to 8.

10. A car, characterized in that, Including the multimedia system as described in claim 9.

11. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the voice control method as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Application program pushing method and device and terminal

    CN109543091A