Model training method, voice recommendation method and electronic equipment

By obtaining user data from multiple non-vehicle voice data sources, performing clustering and transfer learning, and training user preference prediction models, the problem of poor personalization of the vehicle voice recommendation system is solved, and personalized voice recommendation is achieved.

CN120338098APending Publication Date: 2025-07-18GUANGZHOU AUTOMOBILE GROUP CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510373466.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

There is a problem of poor personalization in the vehicle voice recommendation system, especially when user information is limited, it is difficult to achieve personalized voice recommendation.

Method used

By obtaining user data from multiple non-vehicle voice data sources, clustering processing is carried out to establish user tag system data, using transfer learning to migrate user preference data to vehicle voice scenarios, train user preference prediction models, predict user preference data in different voice conversation scenarios, and make personalized recommendations.

Benefits of technology

It realizes personalized voice recommendation in the vehicle voice system, improves the accuracy and efficiency of recommendations, solves the problem of sparse data, and adapts to the personalized needs of new users and multiple scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120338098A_ABST
    Figure CN120338098A_ABST
Patent Text Reader

Abstract

The invention discloses a model training method, a voice recommendation method and electronic equipment. The model training method comprises the following steps: clustering user data of a target user obtained from a plurality of data sources to obtain user tag system data; according to the user tag system data, obtaining corresponding user preference data of the target user in the plurality of data sources; and migrating the user preference data of the target user corresponding to the plurality of data sources to the vehicle-mounted voice scene through migration learning to obtain sample data, training the initial prediction model through the sample data to obtain a user preference prediction model for predicting the user preference data in different voice conversation scenes in the vehicle-mounted voice scene, and predicting the user preference data in the vehicle-mounted voice scene according to the user preference prediction model. The user preference data of the user in different vehicle-mounted voice conversation scenes are predicted through the user preference prediction model, voice recommendation is carried out based on the user preference data, and personalized vehicle-mounted voice recommendation is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to a model training method, a voice recommendation method, and an electronic device. Background Art

[0002] With the development of science and technology, automobiles are becoming more and more intelligent. For example, a vehicle can have an artificial intelligence conversation with a user through an in-vehicle voice system; among them, the in-vehicle voice system can recommend information of interest to the user. However, due to the limited information left by users in the voice field, there is a problem of weak personalization in in-vehicle voice recommendation. Summary of the Invention

[0003] In view of this, embodiments of the present application propose a model training method, a voice recommendation method, and an electronic device to improve the above problems.

[0004] In a first aspect, an embodiment of the present application provides a model training method, the method including: obtaining user tag system data corresponding to a target user, where the user tag system data is obtained by clustering user data of the target user obtained from multiple data sources, and the multiple data sources are data sources corresponding to non-vehicle-mounted voice data; obtaining user preference data corresponding to the target user in the multiple data sources according to the user tag system data; through transfer learning, migrating the user preference data corresponding to the target user in the multiple data sources to an in-vehicle voice scenario to obtain sample data, where the sample data includes sample preference data in different voice conversation scenarios in the in-vehicle voice scenario; training an initial prediction model with the sample data to obtain a user preference prediction model, where the user preference prediction model is used to predict user preference data in different voice conversation scenarios in the in-vehicle voice scenario.

[0005] In a second aspect, an embodiment of the present application provides a voice recommendation method, the method including: obtaining the current voice conversation scenario; inputting the voice conversation scenario into a user preference prediction model to obtain target preference data output by the user preference prediction model in the voice conversation scenario, where the user preference prediction model is obtained by training an initial prediction model with sample data, and the sample data is obtained by migrating user preference data corresponding to a target user in multiple data sources to an in-vehicle voice scenario; determining target voice recommendation content according to the target preference data, and performing voice recommendation based on the target voice recommendation content.

[0006] In a third aspect, an embodiment of the present application provides a model training device, which includes: a user label system data acquisition module, a user preference data acquisition module, a transfer learning module, and a user preference prediction model acquisition module. Among them, the user label system data acquisition module is used to acquire user label system data corresponding to a target user, where the user label system data is obtained by clustering the user data of the target user obtained from multiple data sources, and the multiple data sources are data sources corresponding to non-vehicle-mounted voice data; the user preference data acquisition module is used to acquire user preference data corresponding to the target user in the multiple data sources according to the user label system data; the transfer learning module is used to transfer the user preference data corresponding to the target user in the multiple data sources to the vehicle-mounted voice scenario through transfer learning to obtain sample data, where the sample data includes sample preference data in different voice conversation scenarios in the vehicle-mounted voice scenario; the user preference prediction model acquisition module is used to train an initial prediction model through the sample data to obtain a user preference prediction model, and the user preference prediction model is used to predict user preference data in different voice conversation scenarios in the vehicle-mounted voice scenario.

[0007] In a fourth aspect, an embodiment of the present application provides a voice recommendation device, which includes: a voice conversation scenario acquisition module, a target preference data acquisition module, and a voice recommendation module. Among them, the voice conversation scenario acquisition module is used to acquire the current voice conversation scenario; the target preference data acquisition module is used to input the voice conversation scenario into a user preference prediction model to obtain target preference data output by the user preference prediction model in the voice conversation scenario, where the user preference prediction model is obtained by training an initial prediction model through sample data, and the sample data is obtained by transferring the user preference data corresponding to a target user in multiple data sources to the vehicle-mounted voice scenario; the voice recommendation module is used to determine target voice recommendation content according to the target preference data and perform voice recommendation based on the target voice recommendation content.

[0008] In a fifth aspect, an embodiment of the present application provides an electronic device, including a memory and a processor, the memory is coupled to the processor, and the memory stores instructions, and when the instructions are executed by the processor, the processor executes the model training method provided in the first aspect and the voice recommendation method provided in the second aspect.

[0009] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, in which program code is stored, and the program code can be called by a processor to execute the model training method provided in the first aspect and the voice recommendation method provided in the second aspect.

[0010] In the solution of this application, the electronic device obtains the user label system data obtained by clustering the user data of the target user obtained from the data sources corresponding to multiple non-vehicle voice data. According to the user label system data, the user preference data of the target user corresponding to multiple data sources is obtained. By clustering the user data of the target user in multiple data sources, a comprehensive user attribute label system is established; and according to the user label system data, the user preference data of the target user corresponding to multiple data sources is obtained. Through transfer learning, the user preference data of the target user corresponding to multiple data sources is migrated to the vehicle-mounted voice scenario, and sample data including sample preference data in different voice conversation scenarios in the vehicle-mounted voice scenario is obtained. And the initial prediction model is trained with the sample data to obtain a user preference prediction model for predicting the user preference data in different voice conversation scenarios in the vehicle-mounted voice scenario, so as to migrate the user preference data corresponding to the data source of the non-vehicle voice data to the vehicle-mounted voice scenario as the samples for model training, avoiding data sparsity in model training, and performing voice recommendation based on the user preference data predicted by the user preference prediction model for the user in different vehicle-mounted voice conversation scenarios, realizing personalized vehicle-mounted voice recommendation. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] In order to more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of this application. For those skilled in the art, without creative efforts, other drawings can also be obtained according to these drawings.

[0012] Figure 1 FIG. shows a schematic diagram of an application scenario of a model training method provided by an embodiment of this application;

[0013] Figure 2 FIG. shows a schematic flowchart of a model training method provided by an embodiment of this application;

[0014] Figure 3 FIG. shows a schematic flowchart of a voice recommendation method provided by an embodiment of this application;

[0015] Figure 4 FIG. shows a schematic flowchart of a voice recommendation method provided by an embodiment of this application;

[0016] Figure 5 FIG. shows a block diagram of the structure of a voice recommendation system provided by an embodiment of this application;

[0017] Figure 6 FIG. shows a block diagram of the modules of a model training device provided by an embodiment of this application;

[0018] Figure 7The block diagram of a voice recommendation device provided by an embodiment of the present application is shown;

[0019] Figure 8 The block diagram of an electronic device provided by an embodiment of the present application is shown;

[0020] Figure 9 The storage unit for storing or carrying the program code for implementing the model training method and the voice recommendation method according to the embodiments of the present application is shown. Detailed implementation manners

[0021] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application.

[0022] The in-vehicle voice system is an intelligent system that controls various functions of the vehicle through voice commands; it uses technologies such as speech recognition, semantic understanding, and speech synthesis to achieve intelligent interaction with the vehicle. Users can complete operations such as making a call, controlling the volume, adjusting the air conditioner, querying the route, opening and closing the window, and playing music only by voice commands, which greatly improves the convenience and safety of driving.

[0023] In the related art, the in-vehicle voice system mainly focuses on personalized recommendation in a single voice domain, that is, according to the interest characteristics of a single voice domain in the user's voice conversation, information that the user is interested in is recommended to the user. However, most of the personalized recommendations in a single voice domain are based on the collaborative filtering recommendation method, which makes there are some problems and limitations in voice recommendation. For example, the personalized recommendation in a single voice domain often faces the problem of data sparsity, and it is difficult to obtain a good recommendation effect through training samples. That is, due to the limited information left by users in a single voice domain, it is difficult to achieve true personalization in the personalized recommendation in a single voice domain.

[0024] In addition, the related art proposes a cross-domain recommendation method based on collaborative filtering, which can find a set of users with similar interests to the target user through the collaborative filtering algorithm, and recommend the items that the users in the set like and do not appear in the interest list of the target user to the user after sorting with a certain weight. Among them, the core of collaborative filtering is based on historical data. Therefore, there is a "cold start" problem for new users, and the recommendation effect depends on the amount and accuracy of the user's historical preference data. For some special users, good recommendations cannot be given, and affected by popularity, its long-tail performance is not good. Therefore, in the related art, due to the limited information left by users in the voice domain, there is a problem of weak personalization in in-vehicle voice recommendation.

[0025] In view of the above problems, through long-term research, the inventors have found and proposed the model training method, voice recommendation method, and electronic device provided in the embodiments of the present application. By cross-domain migrating the user preference data corresponding to the data source of non-vehicle voice data to the vehicle voice scenario as samples for model training, the data sparsity in model training is avoided, and voice recommendation is performed based on the user preference data predicted by the user preference prediction model in different vehicle voice conversation scenarios, realizing personalized vehicle voice recommendation. Among them, the specific control method for vehicle video monitoring will be described in detail in the subsequent embodiments.

[0026] The embodiments of the present application will be described below with reference to the accompanying drawings.

[0027] Please refer to Figure 1 , Figure 1 which shows a schematic flowchart of the model training method provided by an embodiment of the present application. In a specific embodiment, the model training method can be applied to a model training device 200 as shown in Figure 6 and an electronic device 100 configured with the model training device 200 ( Figure 8 ). The following will take the electronic device as an example to illustrate the specific process of this embodiment. Of course, it can be understood that the electronic device to which this embodiment is applied can include devices such as vehicles, desktop computers, laptop computers, and vehicle-mounted terminals, which are not limited herein. The following will elaborate on the Figure 1 shown process in detail. The model training specifically may include the following steps:

[0028] Step S110: Obtain the user tag system data corresponding to the target user, where the user tag system data is obtained by clustering the user data of the target user obtained from multiple data sources, and the multiple data sources are data sources corresponding to non-vehicle voice data.

[0029] In some embodiments, the electronic device can obtain the user tag system data corresponding to the target user from an associated cloud or electronic device. Among them, the target user can be the owner of the vehicle or the user represented by the account logged in to the vehicle's in-vehicle computer. Among them, the user tag system data can be obtained by clustering the user data of the target user obtained from multiple data sources; among them, the multiple data sources are data sources corresponding to non-vehicle voice data.

[0030] Optionally, multiple data sources may include data sources such as vehicle applications, in-vehicle TBOX, social platforms, etc. Herein, a data source can be understood as the source of data or the field to which the data belongs. Among them, different data sources can correspond to different fields; it can include hardware devices or software applications, which are not limited herein. Exemplarily, the data source is a mobile phone. Correspondingly, the user data of the target user obtained by the electronic device from this data source may include the user data of non-vehicle voice usage scenarios on the mobile phone, the user data on non-vehicle voice applications, etc.

[0031] As an implementable manner, the user data of the target user obtained from multiple data sources may include: the behavior data generated by the target user using the vehicle application in the vehicle application, the data of the target user using the vehicle in the in-vehicle TBOX, and the social attribute data generated by the target user using the leisure and entertainment platform and the social platform in the social platform. Exemplarily, the behavior data generated by the target user using the vehicle application in the vehicle application can be obtained from the vehicle owner APP, and may include various behavior data such as the user browsing the owner APP, operating the owner APP, and the duration of staying on the owner APP; among them, the electronic device can perform user feature analysis in aspects such as vehicle remote control, vehicle services, and automotive life based on this behavior data.

[0032] Exemplarily, the data of the target user using the vehicle in the in-vehicle TBOX can be obtained from the in-vehicle TBOX of the vehicle, and may include various data of using the vehicle such as the number of vehicle drives, the driving time of the vehicle, the driving speed of the vehicle, the opening and closing of the air conditioner in the vehicle, and the control of the windows and seats; among them, the electronic device can perform user feature analysis in aspects such as vehicle driving, function usage, and vehicle maintenance based on this data of using the vehicle.

[0033] Exemplarily, the social attribute data generated by the target user using the leisure and entertainment platform and the social platform in the social platform can be obtained from the leisure and entertainment platform and the social platform, and can include social attribute information of the target user such as demographic information, social circle, and leisure and entertainment places; among them, the electronic device can perform user feature analysis on aspects such as the social relationship, living habits, and hobbies of the target user based on this social attribute information. In some embodiments, the electronic device can be pre-set with clustering algorithms (such as, k-means clustering algorithm, mean shift clustering algorithm, density-based clustering algorithm, etc.). Among them, after the electronic device obtains the user data of the target user obtained from multiple data sources from the associated cloud or the electronic device, it can perform clustering processing on the user data of the target user obtained from multiple data sources based on the clustering algorithm and the user label system to obtain user label system data. Among them, the user label system can be established based on a preset label dimension, the number of the preset label dimensions can be one or more, and the preset label dimension can be set by the user independently or obtained through third-party experimental data, which is not limited herein.

[0034] Exemplarily, please refer to Table 1, which shows a table of user label system data provided in an embodiment of the present application. Among them, the user label system data can be obtained by performing clustering processing on the user data of the target user obtained from multiple data sources based on the user label system; the number of preset label dimensions included in the user label system is 9, and the 9 preset label dimensions can include nine dimensions of basic attribute labels, geographical location labels, car usage habit labels, interest hobby labels, social network labels, psychological feature labels, consumption behavior labels, travel habit labels, and voice usage labels of the user. Among them, the descriptions corresponding to each preset label dimension together form a comprehensive set of user label system data.

[0035] Table 1

[0036]

[0037] In some embodiments, an electronic device may obtain a user dataset X based on user data of a target user obtained from multiple data sources; wherein, the user data may be samples in the user dataset X, and n pieces of user data correspond to n samples. Among them, K user label clusters may be preset in the electronic device, where the user label clusters can be understood as preset label dimensions; among them, initial center points may be preset in the K user label cluster groups. The electronic device may calculate the distances between the n samples {X1, X2, X3, …, Xn} and the K center points. For example, the distances between X1 and the K center points are {d1, …, dk} respectively. Accordingly, the electronic device may assign each sample to the center point closest to it according to the distances between each sample and the K center points, cluster the n samples according to the center points, calculate the center points of each cluster group based on the samples in each cluster group, and perform iterative screening processing on the samples in each cluster group until a user label system data in dimensions such as user basic attributes, interest preferences, and driving habits converges.

[0038] Step S120: Obtain user preference data of the target user corresponding to the multiple data sources according to the user label system data.

[0039] In some embodiments, after the electronic device obtains the user label system data corresponding to the target user, it may obtain the user preference data of the target user corresponding to the multiple data sources according to the user label system data. Exemplarily, the electronic device performs clustering processing on the user data of the target user obtained from multiple data sources (such as vehicle applications, in-vehicle TBOX, social platforms, etc.) to obtain the user label system data; based on this, the electronic device may obtain the user preference data of the target user corresponding to the data sources of vehicle applications, the data sources driven by in-vehicle TBOX, and the data sources of social platforms according to the user label system data.

[0040] In some embodiments, a convolutional neural network CNN may be preset in the electronic device, and the CNN may be used to extract user features; among them, the electronic device may further include an MLSD learner, and the MLSD learner may fuse the user features extracted by the CNN to obtain user preference data. Accordingly, the electronic device may input the user label system data into the CNN, extract user features through the CNN, and may input the user features extracted by the CNN into the learner to obtain the user preference data of the target user corresponding to the multiple data sources output by the learner.

[0041] Step S130: Through transfer learning, migrate the user preference data of the target user corresponding to the multiple data sources to the in-vehicle voice scenario to obtain sample data, where the sample data includes sample preference data in different voice conversation scenarios in the in-vehicle voice vehicle.

[0042] In some embodiments, after the electronic device obtains the user preference data of the target user corresponding to multiple data sources, through transfer learning, the user preference data of the target user corresponding to multiple data sources can be migrated to the in-vehicle voice scenario to obtain sample data. Among them, the sample data may include sample preference data in different voice dialogue scenarios in the in-vehicle voice scenario. Among them, in the process of migrating the user preference data of the target user corresponding to multiple data sources to the in-vehicle voice scenario by the electronic device, some specific preferences of the target user in different in-vehicle voice dialogue scenarios can be screened out to improve the diversity of the sample data and enhance the personalization of the model obtained by training based on the sample data.

[0043] Exemplarily, please refer to Table 2, which shows a table of sample data provided by an embodiment of the present application. Among them, the user label system data corresponds to 9 preset label dimensions of the user's basic attribute labels, interest preference labels, driving habit labels, hobbies labels, social network labels, psychological characteristic labels, consumption behavior labels, travel habit labels, and voice usage labels. Among them, the electronic device can obtain the user preference data of the target user corresponding to multiple data sources according to the user label system data (such as destination preference, route preference, air conditioner preference, seat preference, music preference, movie preference, frequently contacted people, parking preference, etc.), and can migrate the user preference data of the target user corresponding to multiple data sources to the in-vehicle voice scenario through transfer learning and preference screening to obtain sample data. Among them, the sample data includes sample preference data in different voice dialogue scenarios in the in-vehicle voice scenario. Among them, different voice dialogue scenarios in the in-vehicle voice scenario may include navigating to a destination, selecting a navigation route, adjusting the air conditioner, adjusting the seat, playing music, playing a movie / TV series, making a phone call, selecting a parking location, etc.

[0044] Table 2

[0045]

[0046] Step S140: Train the initial prediction model with the sample data to obtain a user preference prediction model, where the user preference prediction model is used to predict the user preference data in different voice dialogue scenarios in the in-vehicle voice scenario.

[0047] In some embodiments, an initial prediction model may be pre-set in an electronic device. Among them, the initial prediction model may correspond to multiple voice conversation scenarios. After the electronic device obtains sample data, it can train the initial prediction model through the sample data to obtain a user preference prediction model. The user preference prediction model can be used to predict user preference data in different voice conversation scenarios in the in-vehicle voice scenario, so as to transfer the user's preferences to the in-vehicle voice scenario through transfer learning, and screen out some specific preferences of the user in different in-vehicle voice conversation scenarios to obtain sample data for multi-task and multi-scenario algorithm model training, so as to obtain the user's preferences in different in-vehicle voice conversation scenarios, and actively recommend the content of the voice conversation according to the user's preferences to meet the user's personalized needs.

[0048] In some embodiments, in the process of the electronic device training the initial prediction model through sample data to obtain a user preference prediction model, the electronic device can input the voice conversation scenario in the sample data into the initial prediction model, obtain the user preference data output by the initial prediction model in this voice conversation scenario, and adjust the parameters of the initial prediction model according to the similarity between the user preference data and the sample preference data corresponding to this voice conversation scenario in the sample data until the initial prediction model converges to obtain the user preference prediction model.

[0049] Among them, the electronic device can use a neural network to implement the intervention of the multi-scenario model, extract the invariant representations in the voice conversation scenario, and can use decoupled representation learning to eliminate the influence of the voice conversation scenario information. Among them, the electronic device can optimize based on gradient orthogonality and the MAML meta-learning paradigm to solve the type conflict of two different voice conversation scenarios and alleviate the influence of the unique representations of the voice conversation scenario on the estimation tasks of other voice conversation scenarios; and can, through the transfer of user preference data, make the user-item representations that appear in a single voice conversation scenario interact with the information of different voice conversation scenarios respectively. In the process of the electronic device training the initial prediction model through sample data, it can use the voice conversation scenario-related features to learn the user behavior association. For example, use the multi-head self-attention mechanism to process the user behavior sequence:

[0050]

[0051] Among them, head i represents each head, E h represents the input sequence, Q represents the query vector, K represents the key vector, V represents the value vector, d i represents the dimension of the query and the key, and respectively represent learnable weight matrices.

[0052] Among them, the information of the target item and the voice dialogue scenario can be included in the interaction between the items of the attention mechanism, and the outputs of the multiple heads will be finally associated together as the output of the final voice dialogue scenario representation:

[0053] head = Concat(head1, head2,..., headn)Wout,

[0054] Among them, head represents the number of heads, Concat represents concatenating the results of all heads, and Wout represents the learnable weight matrix.

[0055] Among them, when the electronic device trains the initial prediction model through sample data to obtain the user preference prediction model, the model can be evaluated through the accuracy, recall rate, and mean variance of the model to reduce the difference between the model prediction value and the true value.

[0056] Among them, the accuracy of the model refers to the proportion of correctly classified samples in the total number of samples:

[0057]

[0058] Among them, the recall rate of the model refers to the proportion of positive samples that are predicted as positive samples among the truly positive samples:

[0059]

[0060] Among them, the mean squared error is the average of the squares of the differences between the predicted values and the true values:

[0061]

[0062] Among them, TP represents the true positive example, that is, the number of positive samples predicted as positive samples; TN represents the true negative example, that is, the number of negative samples predicted as negative samples; FP represents the false positive example, that is, the number of negative samples predicted as positive samples; FN represents the false negative example, that is, the number of positive samples predicted as negative samples. Among them, n is the number of samples, y i is the true value, is the predicted value. Among them, the smaller the MES, the smaller the error between the result predicted by the model and the true value, and the better the performance of the model.

[0063] It can be understood that in this embodiment, after the electronic device obtains a user preference prediction model for predicting user preference data in different voice conversation scenarios in the in-vehicle voice scenario, the electronic device can perform user preference prediction based on the user preference prediction model. Optionally, the electronic device can perform user action prediction based on the user preference prediction model and the change of the vehicle's usage scenario, and can update the voice recommendation content in real time based on the user preference data in different voice conversation scenarios predicted by the user preference prediction model. Exemplarily, the electronic device determines that the vehicle's usage scenario is that the user gets in the car. Based on this, the electronic device can predict a series of subsequent actions of the user as seat adjustment, air conditioner adjustment, navigation setting, music playing, parking, etc.; correspondingly, the electronic device can determine the user preference data of the voice conversation scenario corresponding to the user's subsequent actions based on the user preference prediction model, and can perform voice content recommendation based on the user preference data of each voice conversation scenario.

[0064] The model training method provided by an embodiment of the present application obtains user label system data obtained by clustering the user data of a target user obtained from data sources corresponding to multiple non-vehicle voice data. According to the user label system data, the user preference data of the target user corresponding to multiple data sources is obtained. By clustering the user data of the target user in multiple data sources, a comprehensive user attribute label system is established; and according to the user label system data, the user preference data of the target user corresponding to multiple data sources is obtained. Through transfer learning, the user preference data of the target user corresponding to multiple data sources is migrated to the in-vehicle voice scenario to obtain sample data including sample preference data in different voice conversation scenarios in the in-vehicle voice scenario, and the initial prediction model is trained with the sample data to obtain a user preference prediction model for predicting user preference data in different voice conversation scenarios in the in-vehicle voice scenario, so as to migrate the user preference data corresponding to the data source of the non-vehicle voice data to the in-vehicle voice scenario as the sample for model training, avoiding data sparsity in model training, and performing voice recommendation based on the user preference data of the user in different in-vehicle voice conversation scenarios predicted by the user preference prediction model, realizing personalized in-vehicle voice recommendation.

[0065] Please refer to Figure 2 , Figure 2 which shows a schematic flowchart of the model training method provided by an embodiment of the present application. This method is applied to the above-mentioned electronic device. The following will elaborate in detail on the Figure 2 process shown. The model training method may specifically include the following steps:

[0066] Step S210: Obtain the user tag system data corresponding to the target user. The user tag system data is obtained by clustering the user data of the target user obtained from multiple data sources. The multiple data sources are the data sources corresponding to non-vehicle voice data. The user data of the target user obtained from the multiple data sources includes: the behavior data generated by the target user using the vehicle application in the vehicle application, the data of the vehicle used by the target user during vehicle driving, and the social attribute data generated by the target user using the leisure and entertainment platform and the social platform in the social platform.

[0067] In some embodiments, the user data of the target user obtained by the electronic device from multiple data sources may include: the behavior data generated by the target user using the vehicle application in the vehicle application, the data of the vehicle used by the target user during vehicle driving, and the social attribute data generated by the target user using the leisure and entertainment platform and the social platform in the social platform. Optionally, the electronic device may collect the user behavior data of the vehicle APP by means of front-end buried points, such as the behavior data of the user's clicks, shares, collections, etc. Optionally, the electronic device may collect the data of the target user using the vehicle reported by the vehicle-side T-BOX by means of periodic collection and trigger collection. Optionally, the electronic device may collect the social attribute data of the target user from the leisure and entertainment platform and the social media platform through API calls, data scraping, etc.

[0068] Among them, in the process of the electronic device collecting the data of the target user using the vehicle reported by the vehicle-side T-BOX based on the periodic collection and trigger collection method, there may be cases of missing or incorrect collection. Based on this, in this embodiment, the electronic device may store the collected data of the target user using the vehicle in a preset database for accumulation, so as to store a large amount of resources for obtaining user data in the field of vehicle driving and improve the reliability of user data.

[0069] In some embodiments, after an electronic device obtains user data of a target user from multiple data sources, it can perform data cleaning, transcoding, and other processing on the user data of the target user corresponding to multiple data sources collected from different channels to ensure the uniformity of the user data format, and can convert the user data into features that can be understood by machines and algorithms; among them, the user data can include data such as dates, numerical values, full-width and half-width characters, genders, mobile phone numbers, etc. Exemplarily, the electronic device can view the missing values of the user data based on pandas and numpy and fill in the missing values; and can process the outliers in the user data. For example, according to the normal distribution, values exceeding 3σ standard deviations are defined as outliers, and outlier detection and deletion processing are performed; among them, the electronic device can also perform standardization processing, normalization processing, etc. on the user data of the target user in multiple fields to unify the format and content of the user data corresponding to multiple data sources, merge the user data from different channels, and perform relevance verification to obtain user label system data.

[0070] Step S220: Extract features from the user label system data to obtain user features of the target user.

[0071] In some embodiments, the process by which an electronic device extracts features from user label system data to obtain user features of a target user may include extracting features from the user label system data based on a first network model to obtain user features of the target user. Among them, the first network model may at least include a parameter-sharing expert network, a task expert network, and a learner expert network.

[0072] Exemplarily, a first network model may be pre-set in the electronic device, and the first network model can perform multi-level task adaptive representation extraction. Exemplarily, the first network model may include a parameter-sharing expert network, a task expert network, and a learner expert network. The first network model can extract cross-task global representations, task-aware local representations, and multi-view task-aware representations through the parameter-sharing expert network, the task expert network, and the learner expert network; among them, the representations extracted by the first network model can be understood as user features. Among them, the first network model may also include a gating network; among them, the first network model can be fused through a gating network to output the final speech dialogue representation:

[0073]

[0074] g l (x) = softmax(W l x),

[0075] where, e l (x) idenotes the embedded features generated and spliced by the expert, k denotes the total number of experts, and f l (x) represents the sum of the products of the embedded features of each tasker. Among them, W l denotes the parameters of the L-th learner, and g l (x) represents weight normalization of the features. Among them, the electronic device can construct multiple learners for each data source separately to extract user features of the voice conversation scenario from multiple perspectives. Among them, the electronic device can input user label data into the first network model, perform multi-level task adaptive representation extraction through a parameter-sharing expert network, a task expert network, and a learner expert network, and can be fused through a gating network to output the final representation (i.e., user features).

[0076] Step S230: Fuse the user features of the target user to obtain the user preference data of the target user corresponding to the multiple data sources.

[0077] In some embodiments, when the electronic device obtains the user features of the target user, it can fuse the user features of the target user to obtain the user preference data of the target user corresponding to the multiple data sources.

[0078] In some embodiments, a self-distillation learner can be pre-set in the electronic device. Correspondingly, the electronic device can fuse the user features of the target user based on the self-distillation learner to obtain the user preference data of the target user corresponding to the multiple data sources. Among them, the electronic device can use the mean of different learners as the user preference output of the fusion learner, that is, the user preference data output of the target user corresponding to the multiple data sources. Optionally, on this basis, the electronic device can also construct a regularized loss to constrain the distance between each learner and the fusion learner to achieve sharing of user preference data between different learners.

[0079] Exemplarily, the electronic device can adopt an MLSD self-distillation learner, integrate the pre-estimation values of multiple learners, and obtain a fusion learner; and can construct a self-distillation loss based on the fusion learner to improve the robustness of the sub-learners:

[0080]

[0081] Among them, l s denotes the loss function value, k denotes the scale parameter, d ctr denotes the distance between the learner and the fusion learner, and V ctr denotes the features of the subset.

[0082] Step S240: Through transfer learning, transfer the user preference data corresponding to the target user in the multiple data sources to the in-vehicle voice scenario to obtain sample data. The sample data includes sample preference data in different voice conversation scenarios in the in-vehicle voice scenario, and the voice conversation scenarios include at least one of getting on the vehicle scenario, driving scenario, multimedia playing scenario, and call scenario.

[0083] In some embodiments, different voice conversation scenarios may be preset in the electronic device, such as getting on the vehicle scenario, driving scenario, multimedia playing scenario, call scenario, etc. Among them, the getting on the vehicle scenario may include voice conversation scenarios such as seat adjustment and air conditioner adjustment, the driving scenario may include voice conversation scenarios such as navigation destination setting, navigation route selection, and parking position selection, the multimedia playing scenario may include voice conversation scenarios such as playing music and playing movies / TV series, and the call scenario may include voice conversation scenarios such as making a call. Among them, the electronic device may transfer the user preference data corresponding to the target user in the multiple data sources to the preset voice conversation scenario to obtain sample preference data in different voice conversation scenarios.

[0084] It should be noted that the goal of transfer learning is to use the knowledge learned in other fields to help the learning tasks in the new field, so as to make full use of the knowledge and experience in the source field and improve the accuracy and efficiency of recommendation in the target field. And the cross-domain recommendation algorithm based on transfer learning requires reasonable feature transformation and model design, otherwise it will lead to information loss and model overfitting. Among them, negative transfer is the main challenge faced by transfer learning: negative transfer refers to the scenario where transferring knowledge from the source field to the target field does not bring any improvement, but instead leads to a decline in the overall performance of the target field. Based on this, in this embodiment, the user's interest preference data can be transferred from the non-in-vehicle voice scenario to the in-vehicle voice scenario based on the method of clustering and identifying relevance before transfer learning, so as to avoid negative transfer and improve the rationality and usability of the sample data.

[0085] Step S250: Train the initial prediction model with the sample data to obtain a user preference prediction model, which is used to predict the user preference data in different voice conversation scenarios in the in-vehicle voice scenario.

[0086] For the specific description of step S250, please refer to the description of step S140 above and will not be elaborated here.

[0087] The model training method provided by an embodiment of this application, compared with Figure 1The model training method shown. In this embodiment, the user data of the target user obtained from multiple data sources may include: the behavior data generated by the target user using the vehicle application in the vehicle application, the data of the vehicle used by the target user during vehicle driving, and the social attribute data generated by the target user using the leisure and entertainment platform and the social platform in the social platform. Thus, by intelligently mining and analyzing the user data of the target user collected in the owner APP, vehicle driving, and social platform based on the clustering algorithm, a comprehensive user label system data is established, ensuring the diversity of user data. At the same time, in this embodiment, the voice dialogue scenario may include at least one of the getting-on-the-vehicle scenario, driving scenario, multimedia playback scenario, and call scenario. Thus, the user preference data of the target user corresponding to multiple data sources is migrated to different voice dialogue scenarios in the in-vehicle voice scenario, obtaining the sample data for model training, improving the utilization rate of the user data obtained from multiple data sources, and improving the accuracy and efficiency of model training. In addition, this embodiment can also perform feature extraction on the user label system data to obtain the user features of the target user; fuse the user features of the target user to obtain the user preference data corresponding to the target user in multiple data sources. Thus, through the extraction of user features and the cross-domain transfer learning method, the interest preference data of the user is migrated from the non-vehicle voice scenario to the vehicle voice scenario to train the user preference prediction model for multiple scenarios, improving the accuracy of model training.

[0088] Please refer to Figure 3 , Figure 3 shows a schematic flowchart of a voice recommendation method provided by an embodiment of the present application. In a specific embodiment, the voice recommendation method can be applied to a voice recommendation device 300 as shown in Figure 7 and an electronic device 100 configured with the voice recommendation device 300 ( Figure 8 ). Below, taking the electronic device as an example, the specific process of this embodiment will be described. Of course, it can be understood that the electronic device to which this embodiment is applied may include devices such as vehicles, desktop computers, laptop computers, and in-vehicle terminals, which are not limited herein. Below, the Figure 3 shown process will be elaborated in detail. The voice recommendation method may specifically include the following steps:

[0089] Step S310: Obtain the current voice dialogue scenario.

[0090] In some embodiments, a visual sensor may be provided in the vehicle, and the visual sensor may collect images in the vehicle in real time; accordingly, the electronic device may obtain images in the vehicle collected by the visual sensor, and may obtain the actions of the user in the vehicle based on the images, and may determine the current voice conversation scene of the vehicle based on the actions of the user. Exemplarily, the electronic device may determine the actions of the user in the vehicle by analyzing the images in the vehicle collected by the visual sensor, such as getting in the vehicle, adjusting the seat, adjusting the air conditioner, etc.; accordingly, after the electronic device determines the actions of the user in the vehicle, it may determine the current voice conversation scene of the vehicle based on the actions of the user.

[0091] Among them, the electronic device may be pre-set with a correspondence between the user's actions and the vehicle's voice dialogue scene; wherein the correspondence between the user's actions and the vehicle's voice dialogue scene may include a one-to-one relationship, a one-to-many relationship, and a many-to-one relationship, which is not limited here. Among them, the electronic device may obtain the current voice dialogue scene of the vehicle based on the mapping relationship and the user's actions in the determined vehicle. Exemplarily, the mapping relationship between the user's actions and the vehicle's voice dialogue scene may include, getting on the vehicle corresponds to the getting on scene, seat adjustment corresponds to the driving scene, air conditioning adjustment corresponds to the driving scene, navigation destination setting corresponds to the driving scene, navigation route setting corresponds to the driving scene, parking space setting corresponds to the driving scene, playing music corresponds to the multimedia playback scene, playing video corresponds to the multimedia playback scene, making a phone call corresponds to the call scene, etc.

[0092] In some embodiments, an infrared sensor may be provided in the vehicle, and the infrared sensor may collect infrared data in the vehicle in real time; accordingly, the electronic device may obtain infrared data in the vehicle collected by the infrared sensor, and may determine the current voice conversation scene of the vehicle based on the infrared data. For example, if it is determined based on the infrared data that a user has boarded the vehicle, the current scene of the vehicle may be determined to be a boarding scene; for another example, if it is determined based on the infrared data that the temperature in the vehicle exceeds a preset temperature, the current scene of the vehicle may be determined to be a driving scene (e.g., an air conditioning adjustment scene).

[0093] In some embodiments, a pressure sensor may be provided in the vehicle, and the pressure sensor may be provided at a preset position of the vehicle (e.g., a seat, a physical button, etc.), and may collect the pressure at the preset position in the vehicle in real time; accordingly, the electronic device may obtain the pressure at the preset position in the vehicle collected by the pressure sensor, and may determine the current voice conversation scene of the vehicle based on the pressure at the preset position. For example, if it is determined that the pressure change rate of the driver's seat is greater than the preset pressure change rate, then the current scene of the vehicle may be determined to be the boarding scene; for another example, if it is determined that the pressure of the physical button for adjusting the air conditioner is greater than the preset pressure, then the current scene of the vehicle may be determined to be the driving scene (air conditioner adjustment scene).

[0094] Step S320: Input the voice conversation scenario into the user preference prediction model to obtain the target preference data under the voice conversation scenario output by the user preference prediction model. The user preference prediction model is obtained by training an initial prediction model with sample data, and the sample data is obtained by migrating the user preference data of the target user corresponding to multiple data sources to the in-vehicle voice scenario.

[0095] In some embodiments, a user preference prediction model may be pre-set in the electronic device. The user preference prediction model can be obtained by training an initial prediction model with sample data, and the sample data can be obtained by migrating the user preference data of the target user corresponding to multiple data sources to the in-vehicle voice scenario. Among them, the user preference data of the target user corresponding to multiple data sources can be obtained by clustering the user data of the target user obtained from multiple data sources, and the multiple data sources are data sources corresponding to non-vehicle voice data. Optionally, the electronic device can obtain the user data of the target user from multiple data sources in real time, and can obtain the user preference prediction model according to the user data of the target user obtained from multiple data sources, and can update the user preference prediction model in real time.

[0096] In some embodiments, after the electronic device obtains the voice conversation scenario where the vehicle is currently located, it can input the voice conversation scenario into the user preference prediction model to obtain the target preference data under the voice conversation scenario output by the user preference prediction model.

[0097] Step S330: Determine the target voice recommendation content according to the target preference data, and perform voice recommendation based on the target voice recommendation content.

[0098] In some embodiments, after the electronic device obtains the target preference data under the voice conversation scenario where the vehicle is currently located, it can determine the target voice recommendation content according to the target preference data, and can perform voice recommendation based on the target voice recommendation content.

[0099] In some embodiments, the electronic device may be pre-set with a voice content generation model; among them, the electronic device can input the target preference data into the voice content generation model to obtain the target voice recommendation content output by the voice content generation model. Optionally, the user preference prediction model can be a multi-task and multi-scenario model, and the electronic device can also generate voice content based on the user preference prediction model.

[0100] Among them, the electronic device can use the network structure of EPNET + feature-level dynamic weight to depict the information of the voice conversation scenario; among them, the electronic device can use the activation function sigmoid to convert the input of the voice content generation model into a probability value:

[0101] V = sigmoid(W o head out + b o )

[0102] where W o represents the learnable weight matrix, head out represents the final multi - head output, and b o represents the scaling factor; thus, end - to - end learning is performed on the user preference prediction model and the speech content generation model. DomainNet is used to process all features, output feature weights, which act on features other than the speech dialogue scenario features, and finally integrated into a global vector representation. The generated representation is used to generate speech dialogue recommendation content.

[0103] In some embodiments, the electronic device can update the speech dialogue recommendation content in real time according to the changes in the user preference prediction model and the speech dialogue scenario. Exemplarily, a user preference prediction model can be preset in the electronic device to perform user preference prediction. Among them, the electronic device can predict a series of subsequent actions of the user according to the changes in the user preference prediction model and the scenario, and update the speech dialogue recommendation content in real time (e.g., user gets in the car scenario → the speech dialogue recommendation content is to recommend adjusting the seat → the speech dialogue recommendation content is to recommend turning on the air conditioner to the common temperature and wind speed → the speech dialogue recommendation content is to recommend the frequently visited destinations → the speech dialogue recommendation content is to recommend the frequently traveled routes → the speech dialogue recommendation content is to recommend the frequently listened - to music tracks → arriving at the destination scenario → the speech dialogue recommendation content is to recommend the frequently used parking spaces).

[0104] In some embodiments, the process by which the electronic device determines the target speech recommendation content according to the target preference data and performs speech recommendation based on the target speech recommendation content may include determining the initial speech recommendation content according to the target preference data and performing collaborative filtering processing on the initial speech recommendation content to obtain the target speech recommendation content. It can be understood that collaborative filtering has no special requirements for the recommended objects, can process unstructured complex objects, the calculated recommendations are open, can share others' experiences, and can well support users in discovering potential interest preferences.

[0105] Among them, the electronic device can use the collaborative filtering algorithm to filter, rank, etc. the initial recommended speech dialogue content, remove the dialogue content that does not conform to the in - vehicle speech dialogue scenario and the content that the user is not interested in, so as to generate the final recommendation content of the speech engine and improve the accuracy of the prediction of the speech content generation model.

[0106] In some embodiments, the electronic device may record the user's historical selections of target voice recommendation content in different in-vehicle voice conversation scenarios, and may perform collaborative filtering processing on the target voice recommendation content determined according to the target preference data, so as to achieve personalized in-vehicle voice recommendation and improve the accuracy of voice recommendation.

[0107] Exemplarily, the electronic device may obtain a similarity matrix between voice conversations by calculating the similarity of the voice conversation column vectors in the collinearity matrix, and then find a preset number of voice conversations with the top similarity rankings of the user's historical positive feedback voice conversations to form a set of similar voice conversations. Among them, the electronic device may sort the voice conversation content based on the similarity scores to generate a final recommendation list. Among them, the electronic device may also adopt an adjusted cosine similarity calculation method to eliminate the influence of different user scoring habits by subtracting the average score of the user's voice conversation scoring:

[0108]

[0109] Among them, R u,i represents the score of user u for voice conversation i, represents the average score of user u's scoring, and R u,j represents the score of user u for voice conversation j. Among them, the electronic device may predict the voice conversations that the user has not scored according to the calculated similarity between voice conversations. For example, a new value may be re-estimated by means of linear regression:

[0110]

[0111] Among them, voice conversation N is a similar conversation to conversation i, and α and β are obtained by performing linear regression calculation on the scoring vectors of voice conversations N and i, and ε is the error of the regression model.

[0112] Exemplarily, please refer to Table 3, which shows a table of target voice recommendation content provided by an embodiment of the present application. Among them, the user label system data corresponds to 9 preset label dimensions of the user's basic attribute labels, interest preference labels, driving habit labels, hobby labels, social network labels, psychological characteristic labels, consumption behavior labels, travel habit labels, and voice usage labels. Among them, the electronic device can migrate the user preference data of the target user corresponding to multiple data sources according to the user label system data to the in-vehicle voice scenario to obtain sample data (for example, the sample preference data of the voice conversation scenario of navigating to the destination is the destination preference, and the sample preference data of the voice conversation scenario of selecting a navigation route is the route preference, etc.). Correspondingly, the electronic device can train a user preference prediction model based on the sample data, and can predict the user preference data in different voice conversation scenarios in the in-vehicle voice scenario based on the user preference prediction model, so as to determine the recommended content of the voice conversation based on the user preference data. Exemplarily, the user preference prediction model can determine the user preference data according to the current voice conversation scenario of the vehicle. Further, the electronic device can determine the target voice recommendation content in the voice conversation scenario based on the user preference data.

[0113] Table 3

[0114]

[0115] Exemplarily, please refer to Figure 4 and Figure 5 , where Figure 4 shows a schematic flowchart of the voice recommendation method provided by an embodiment of the present application. Figure 5The architecture block diagram of a voice recommendation system provided by an embodiment of the present application is shown. Among them, the voice recommendation system may include a data collection module, a data analysis module, a recommendation algorithm module, a voice system module, and a voice recommendation module. Among them, the voice recommendation system may obtain user data of a target user from data sources such as a car owner APP, vehicle driving, and social platforms through the data collection module. Among them, the voice recommendation system may perform data analysis and processing on the user data of the target user in multiple fields through the data analysis module and the recommendation algorithm module. For example, clustering processing is performed on the user data of the target user obtained from multiple data sources based on the K-means algorithm to obtain user label system data (such as user data in the user attribute label dimension, user data in the interest preference label dimension, user data in the behavior habit label dimension, and user data in the behavior feature label dimension), and user feature extraction and feature fusion processing may be performed on the user label system data based on a network model including an MLSD learner to obtain user preference data corresponding to the target user in multiple data sources, and the user preference data corresponding to the target user in multiple data sources may be migrated to the in-vehicle voice scenario based on the method of transfer learning to obtain sample data for model training, and the initial prediction model may be trained through the sample data to obtain a user preference prediction model for predicting user preference data in different voice conversation scenarios in the in-vehicle voice scenario, and the initial voice recommendation content may be determined based on the user preference data predicted by the user preference prediction model, and the initial voice recommendation content may be filtered and ranked based on the collaborative filtering algorithm to obtain the target voice recommendation content. Among them, the voice recommendation system may output the target voice recommendation content through the voice system module based on the cloud engine, in-vehicle engine, HMI interaction, and voice skills. Among them, the voice recommendation system may output the target voice recommendation content through the voice recommendation module (such as frequently visited destinations, air conditioner temperature adjustment, frequently listened to music, driving model)

[0116] In this embodiment, the electronic device can collect user data of a user from multiple data sources such as the car owner APP, vehicle driving, and social platforms, perform intelligent mining and analysis based on the k-means clustering algorithm in machine learning, establish a comprehensive set of user label system data, and can extract user features and perform feature fusion on the user label system data to obtain user preference data corresponding to the target user in multiple data sources. It can also use cross-domain transfer learning to transfer the user preference data of the user from a non-vehicle-mounted voice scenario to a vehicle-mounted voice scenario, perform multi-task and multi-scenario algorithm model training, output vehicle-mounted voice recommendation content, and can implement more accurate personalized voice recommendation content through a collaborative filtering algorithm. It can be understood that through the cross-domain recommendation method, data from multiple data sources such as the car owner APP, vehicle driving, and social platforms are combined, and the transfer learning algorithm is applied to the vehicle-mounted voice recommendation system. Even if a user uses the vehicle-mounted voice less frequently, the system can understand the user's expected needs and perform true personalized recommendations. In addition, the voice recommendation system can also predict a series of subsequent actions of the user according to the change of the driving scenario, match the user preference data, and update the voice dialogue recommendation content in real time, which can alleviate the problem that it is difficult to obtain good recommendation effects through training samples due to the limited information left by the user in the vehicle-mounted voice scenario; and can transform the vehicle-mounted voice system from a simple dialogue machine with question-and-answer to an anthropomorphic intelligent system that can understand the user's expected needs and actively initiate interactions, providing an emotional and personalized vehicle-mounted voice usage experience for the user and supporting the intelligent development of the vehicle.

[0117] The voice recommendation method provided by an embodiment of the present application includes obtaining the current voice dialogue scenario; inputting the voice dialogue scenario into a user preference prediction model to obtain target preference data output by the user preference prediction model in the voice dialogue scenario, where the user preference prediction model is obtained by training an initial prediction model with sample data, and the sample data is obtained by migrating the user preference data corresponding to the target user in multiple data sources to the vehicle-mounted voice scenario; determining target voice recommendation content according to the target preference data, and performing voice recommendation based on the target voice recommendation content, so as to predict the user preference data in the current voice dialogue scenario through the user preference prediction model trained based on the user data corresponding to the target user in multiple data sources, and determine the voice recommendation content according to the user preference data for voice recommendation, realizing personalized vehicle-mounted voice recommendation.

[0118] Please refer to Figure 6 , Figure 6 which shows a block diagram of a model training device provided by an embodiment of the present application. The model training device 200 is applied to the above-mentioned electronic device. The following will be directed to Figure 6The following describes the process shown in detail. The model training device 200 includes: a user label system data acquisition module 210, a user preference data acquisition module 220, a transfer learning module 230, and a user preference prediction model acquisition module 240, where:

[0119] The user label system data acquisition module 210 is configured to acquire user label system data corresponding to a target user. The user label system data is obtained by clustering the user data of the target user obtained from multiple data sources, and the multiple data sources are data sources corresponding to non-vehicle-mounted voice data.

[0120] The user preference data acquisition module 220 is configured to acquire user preference data corresponding to the target user in the multiple data sources according to the user label system data.

[0121] The transfer learning module 230 is configured to transfer the user preference data corresponding to the target user in the multiple data sources to the vehicle-mounted voice scenario through transfer learning to obtain sample data, and the sample data includes sample preference data in different voice conversation scenarios in the vehicle-mounted voice scenario.

[0122] The user preference prediction model acquisition module 240 is configured to train an initial prediction model through the sample data to obtain a user preference prediction model, and the user preference prediction model is used to predict user preference data in different voice conversation scenarios in the vehicle-mounted voice scenario.

[0123] Further, the voice conversation scenario includes at least one of a getting-on-the-vehicle scenario, a driving scenario, a multimedia playing scenario, and a call scenario.

[0124] Further, the user preference data acquisition module 220 may include: a feature extraction unit and a user feature fusion unit, where:

[0125] The feature extraction unit is configured to extract features from the user label system data to obtain user features of the target user.

[0126] The user feature fusion unit is configured to fuse the user features of the target user to obtain user preference data corresponding to the target user in the multiple data sources.

[0127] Further, the feature extraction unit may include: a feature extraction subunit, where:

[0128] The feature extraction subunit is configured to extract features from the user label system data based on a first network model to obtain user features of the target user, where the first network model includes at least a parameter-sharing expert network, a task expert network, and a learner expert network.

[0129] Further, the user feature fusion unit may include: a user feature fusion subunit, where:

[0130] The user feature fusion subunit is configured to fuse the user features of the target user based on the self-distillation learner to obtain the user preference data of the target user corresponding to the multiple data sources.

[0131] Further, the user data of the target user obtained from the multiple data sources includes: behavior data generated by the target user using the vehicle application in the vehicle application, data of the target user using the vehicle during vehicle driving, and social attribute data generated by the target user using the leisure and entertainment platform and the social platform in the social platform.

[0132] Please refer to Figure 7 , Figure 7 which shows a block diagram of a voice recommendation device provided by an embodiment of the present application. The voice recommendation device 300 is applied to the above-mentioned electronic device. The following will elaborate in detail on the Figure 7 shown process. The voice recommendation device 300 includes: a voice dialogue scenario acquisition module 310, a target preference data acquisition module 320, and a voice recommendation module 330, where:

[0133] The voice dialogue scenario acquisition module 310 is configured to acquire the current voice dialogue scenario.

[0134] The target preference data acquisition module 320 is configured to input the voice dialogue scenario into the user preference prediction model to obtain the target preference data output by the user preference prediction model in the voice dialogue scenario. The user preference prediction model is obtained by training an initial prediction model with sample data, and the sample data is obtained by migrating the user preference data of the target user corresponding to multiple data sources to the in-vehicle voice scenario.

[0135] The voice recommendation module 330 is configured to determine target voice recommendation content according to the target preference data and perform voice recommendation based on the target voice recommendation content.

[0136] Further, the voice recommendation module 330 may include: an initial voice recommendation content determination unit and a target voice recommendation content acquisition unit, where:

[0137] The initial voice recommendation content determination unit is configured to determine initial voice recommendation content according to the target preference data.

[0138] The target voice recommendation content acquisition unit is configured to perform collaborative filtering processing on the initial voice recommendation content to obtain the target voice recommendation content.

[0139] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described devices and modules can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0140] In several embodiments provided in the present application, the coupling between modules can be electrical, mechanical, or other forms of coupling.

[0141] In addition, in each embodiment of the present application, each functional module can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules.

[0142] Please refer to Figure 8 , which shows a structural block diagram of an electronic device provided in an embodiment of the present application. The electronic device 100 can be a device with processing capabilities such as a core network, a cloud platform, etc. Among them, the electronic device 100 can be understood as a server. The electronic device 100 in the present application can include one or more of the following components: a processor 110, a memory 120, and one or more application programs, where one or more application programs can be stored in the memory 120 and configured to be executed by one or more processors 110, and one or more programs are configured to execute the methods described in the foregoing method embodiments.

[0143] Among them, the processor 110 may include one or more processing cores. The processor 110 connects various parts within the entire electronic device 100 through various interfaces and lines, and performs various functions of the electronic device 100 and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 120, and by calling data stored in the memory 120. Optionally, the processor 110 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 110 may integrate a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing the content to be displayed; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor 110 and may be implemented separately through a communication chip.

[0144] The memory 120 may include random access memory (RAM) and may also include read-only memory. The memory 120 can be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for implementing at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created during the use of the electronic device 100 (such as phone book, audio and video data, chat record data, etc.).

[0145] Please refer to Figure 9 , which shows a structural block diagram of a computer-readable storage medium provided by an embodiment of the present application. Program code is stored in the computer-readable medium 400, and the program code can be called by a processor to execute the methods described in the above method embodiments.

[0146] The computer-readable storage medium 400 can be an electronic memory such as a flash memory, EEPROM (Electrically Erasable Programmable Read-Only Memory), EPROM, a hard disk, or a ROM. Optionally, the computer-readable storage medium 400 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 400 has a storage space for program code 410 that executes any of the method steps in the above-described method. These program codes can be read out from or written into one or more computer program products. The program code 410 can be compressed in an appropriate form, for example.

[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.

Claims

1. A model training method, characterized in that, The method comprises: Acquire user tag system data corresponding to the target user, wherein the user tag system data is obtained by clustering user data of the target user obtained from multiple data sources, and the multiple data sources are data sources corresponding to the non-vehicle voice data; According to the user tag system data, obtaining the user preference data of the target user corresponding to the multiple data sources; By transfer learning, the user preference data of the target user corresponding to the multiple data sources is transferred to the in-vehicle voice scene to obtain sample data, wherein the sample data includes sample preference data in different voice dialogue scenes in the in-vehicle voice scene; The initial prediction model is trained by the sample data to obtain a user preference prediction model, wherein the user preference prediction model is used to predict user preference data in different voice dialogue scenarios in the in-vehicle voice scenario.

2. The method according to claim 1, wherein The voice dialogue scenario includes at least one of a boarding scenario, a driving scenario, a multimedia playback scenario, and a call scenario.

3. The method according to claim 1, characterized in that, The acquiring, according to the user tag system data, user preference data corresponding to the target user in the multiple data sources includes: Extracting features from the user tag system data to obtain user features of the target user; The user features of the target user are integrated to obtain the user preference data of the target user corresponding to the multiple data sources.

4. The method according to claim 3, wherein The extracting features of the user tag system data to obtain user features of the target user includes: Based on the first network model, feature extraction is performed on the user label system data to obtain user features of the target user, wherein the first network model at least includes a parameter sharing expert network, a task expert network and a learner expert network.

5. The method according to claim 3, wherein The fusing the user features of the target user to obtain the user preference data of the target user corresponding to the multiple data sources includes: The user features of the target user are fused based on the self-distillation learner to obtain the user preference data of the target user corresponding to the multiple data sources.

6. The method according to any one of claims 1-5, characterized in that, The user data of the target user obtained from multiple data sources includes: behavioral data generated by the target user using the vehicle application in the vehicle application, data on the target user using the vehicle during vehicle driving, and social attribute data generated by the target user using the leisure and entertainment platform and the social platform in the social platform.

7. A voice recommendation method, characterized in that, The method comprises: Get the current voice dialogue scene; Inputting the voice conversation scenario into a user preference prediction model, and obtaining target preference data in the voice conversation scenario output by the user preference prediction model, wherein the user preference prediction model is obtained by training an initial prediction model with sample data, and the sample data is obtained by migrating user preference data corresponding to a target user from multiple data sources to an in-vehicle voice scenario; Target voice recommendation content is determined according to the target preference data, and voice recommendation is performed based on the target voice recommendation content.

8. The method according to claim 7, characterized in that The determining target voice recommendation content according to the target preference data includes: Determining initial voice recommendation content according to the target preference data; The initial voice recommendation content is subjected to collaborative filtering processing to obtain the target voice recommendation content.

9. A model training device, characterized in that, The device comprises: A user label system data acquisition module, used to acquire user label system data corresponding to a target user, wherein the user label system data is obtained by clustering user data of the target user obtained from multiple data sources, wherein the multiple data sources are data sources corresponding to the non-vehicle voice data; A user preference data acquisition module, used to acquire the user preference data of the target user corresponding to the multiple data sources according to the user tag system data; A transfer learning module, configured to transfer the user preference data of the target user corresponding to the multiple data sources to the in-vehicle voice scene through transfer learning to obtain sample data, wherein the sample data includes sample preference data under different voice dialogue scenes in the in-vehicle voice scene; The user preference prediction model acquisition module is used to train the initial prediction model through the sample data to obtain the user preference prediction model, and the user preference prediction model is used to predict the user preference data in different voice dialogue scenarios in the in-vehicle voice scenario.

10. A voice recommendation device, characterized in that, The device comprises: A voice dialogue scene acquisition module is used to obtain the current voice dialogue scene; a target preference data acquisition module, configured to input the voice dialogue scenario into a user preference prediction model, and obtain target preference data in the voice dialogue scenario output by the user preference prediction model, wherein the user preference prediction model is obtained by training an initial prediction model with sample data, and the sample data is obtained by migrating user preference data corresponding to a target user from multiple data sources to an in-vehicle voice scenario; The voice recommendation module is used to determine target voice recommendation content according to the target preference data, and perform voice recommendation based on the target voice recommendation content.

11. An electronic device, characterized in that, include: one or more processors; Memory; One or more application programs, wherein the one or more application programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the method according to any one of claims 1-8.

Citation Information

Cited By

  • Recommendation method and system based on big data tags

    CN120994907A

  • Cabin scene semantic graph construction method and device and program product

    CN121480661A