Recommendation method and device, electronic equipment and computer readable storage medium

By adaptively adjusting the initial characteristics of the user's voice signal, accurate tone and voiceprint characteristics are obtained, which solves the problem of acoustic feature variability in voiceprint recognition and achieves accurate personalized recommendations.

CN119964575APending Publication Date: 2025-05-09SHENZHEN TCL DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510099799.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-21
Publication Date
2025-05-09

AI Technical Summary

Technical Problem

In voiceprint recognition technology, the acoustic characteristics of user voice are highly varied, making it difficult to accurately identify the user's identity and realize personalized recommendations.

Method used

By obtaining the initial characteristics of the user's voice signal, adaptive adjustments are made to obtain the tone characteristics, and then the accurate voiceprint characteristics are determined. Based on the matching degree of voiceprint characteristics and reference users, the target user is determined and personalized recommendations are made based on the information of the target user.

Benefits of technology

By adaptively adjusting the initial features, correcting the variability of acoustic features, achieving accurate voiceprint feature recognition and personalized recommendations, improving the accuracy and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119964575A_ABST
    Figure CN119964575A_ABST
Patent Text Reader

Abstract

Embodiments of the invention disclose a recommendation method and apparatus, an electronic device and a computer readable storage medium. The method comprises the steps of obtaining an initial feature of a voice signal of a user; performing adaptive adjustment on the initial features to obtain timbre features of the voice signals; determining voiceprint features of the user based on the timbre features; determining a target user from the reference users according to the matching degree between the voiceprint feature of the user and the voiceprint feature of each reference user; acquiring information of the target user; and according to the information of the target user, carrying out personalized recommendation on the user. Thus, according to the scheme, the initial features of the voice signals of the user can be adaptively adjusted, accurate voiceprint features are obtained, and then accurate personalized recommendation is achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of recommendation technology, and specifically to a recommendation method, device, electronic device, and computer-readable storage medium. Background Art

[0002] Voiceprint recognition is a type of biometric technology that can convert sound signals into electrical signals and then use computers to identify them. Related technologies use voiceprints to identify users and make personalized content recommendations for identified users. However, the acoustic characteristics of each person's voice are both relatively stable and variable; the variability of acoustic characteristics is caused by a variety of factors such as physiology, pathology, psychology, simulation, disguise or environmental interference. The variability of acoustic characteristics makes it difficult to accurately identify the identity of a user based on their voice, thereby achieving accurate personalized recommendations. Summary of the invention

[0003] The embodiments of the present application provide a recommendation method, device, electronic device and computer-readable storage medium, which can adaptively adjust the initial features of the user's voice signal to obtain accurate voiceprint features, thereby achieving accurate personalized recommendations.

[0004] In a first aspect, an embodiment of the present application provides a recommendation method, comprising:

[0005] Acquire initial features of a user's voice signal;

[0006] Adaptively adjusting the initial features to obtain timbre features of the speech signal;

[0007] Determining a voiceprint feature of the user based on the timbre feature;

[0008] Determining a target user from the reference users according to a matching degree between the voiceprint feature of the user and the voiceprint features of each reference user;

[0009] Acquire information of the target user;

[0010] According to the information of the target user, personalized recommendations are made to the user.

[0011] In one embodiment, the adaptively adjusting the initial features to obtain the timbre features of the speech signal includes:

[0012] Inputting the initial features into a timbre extraction model to obtain an adaptive weight assigned by the timbre extraction model to the initial features;

[0013] The initial feature is adjusted according to the adaptive weight to obtain the timbre feature output by the timbre extraction model.

[0014] In one embodiment, the training step of the timbre extraction model at least includes:

[0015] Acquire a plurality of speech signal samples and a timbre feature sample of each of the speech signal samples;

[0016] Acquire initial feature samples of each of the speech signal samples;

[0017] Inputting a plurality of the initial feature samples into an initial timbre extraction model to obtain adaptive weight samples assigned by the initial timbre extraction model to the initial feature samples;

[0018] Adjusting the initial feature sample according to the adaptive weight sample to obtain the predicted timbre feature output by the initial timbre extraction model;

[0019] Based on the matching degree between the timbre feature samples of each of the speech signal samples and the predicted timbre feature, the initial timbre extraction model is trained to obtain the trained timbre extraction model.

[0020] In one embodiment, determining the voiceprint feature of the user based on the timbre feature includes:

[0021] Extracting the sound features of the speech signal based on the timbre features;

[0022] Determine the voiceprint feature of the user according to the timbre feature and the sound feature.

[0023] In one embodiment, the initial features include cepstral graphs and Mel-frequency cepstral coefficients;

[0024] The obtaining of initial features of the user's voice signal comprises:

[0025] Performing discrete Fourier transform on the user's voice signal to obtain a frequency domain signal;

[0026] Extracting the amplitude of the speech signal from the frequency domain signal, and determining the square of the amplitude as a power spectrum;

[0027] Taking the logarithm of the power spectrum to obtain the logarithm of the power spectrum;

[0028] Performing an inverse discrete Fourier transform on the logarithm of the power spectrum to obtain a cepstrum of the user's speech signal;

[0029] Filtering the power spectrum through a Mel frequency filter bank to obtain power spectrum energy in Mel frequency scale;

[0030] Taking the logarithm of the power spectrum energy of the Mel frequency scale to obtain the logarithm of the power spectrum of the Mel frequency scale;

[0031] Perform discrete cosine transform on the logarithm of the power spectrum of the mel-frequency scale to obtain mel-frequency cepstrum coefficients of the user's speech signal.

[0032] In one of the embodiments, before determining the target user from the reference users based on the matching degree between the voiceprint feature of the user and the voiceprint features of each reference user, the method further includes:

[0033] Acquire voiceprint features of each of the reference users;

[0034] Acquire information of each of the reference users;

[0035] The voiceprint features and information of each reference user are correspondingly stored in the server.

[0036] In one embodiment, the method further comprises:

[0037] When the voiceprint feature of the reference user does not match the voiceprint feature of the user, obtaining information of the user;

[0038] The user's information and voiceprint features are stored in the server accordingly.

[0039] In one embodiment, the method further comprises:

[0040] Obtaining information of the user from each device;

[0041] The user information stored in the server is updated using the user information obtained from each of the devices; the user information stored in the server is used to make personalized recommendations to the user logged in to each of the devices.

[0042] In a second aspect, an embodiment of the present application provides a recommendation device, including:

[0043] A feature acquisition module, used to acquire initial features of a user's voice signal;

[0044] A feature adjustment module, used for adaptively adjusting the initial feature to obtain the timbre feature of the speech signal;

[0045] A feature determination module, used to determine the voiceprint feature of the user based on the timbre feature;

[0046] A user determination module, configured to determine a target user from the reference users according to a matching degree between the voiceprint feature of the user and the voiceprint features of each reference user;

[0047] An information acquisition module, used to acquire the information of the target user;

[0048] The personalized recommendation module is used to make personalized recommendations to the user based on the information of the target user.

[0049] In one embodiment, the feature adjustment module includes:

[0050] An input unit, used for inputting the initial feature into a timbre extraction model to obtain an adaptive weight assigned by the timbre extraction model to the initial feature;

[0051] An adjustment unit is used to adjust the initial feature according to the adaptive weight to obtain the timbre feature output by the timbre extraction model.

[0052] In one embodiment, the training step of the timbre extraction model at least includes:

[0053] Acquire a plurality of speech signal samples and a timbre feature sample of each of the speech signal samples;

[0054] Acquire initial feature samples of each of the speech signal samples;

[0055] Inputting a plurality of the initial feature samples into an initial timbre extraction model to obtain adaptive weight samples assigned by the initial timbre extraction model to the initial feature samples;

[0056] Adjusting the initial feature sample according to the adaptive weight sample to obtain the predicted timbre feature output by the initial timbre extraction model;

[0057] Based on the matching degree between the timbre feature samples of each of the speech signal samples and the predicted timbre feature, the initial timbre extraction model is trained to obtain the trained timbre extraction model.

[0058] In one embodiment, the feature determination module includes:

[0059] An extraction unit, configured to extract a sound feature of the speech signal based on the timbre feature;

[0060] A voiceprint feature determination unit is used to determine the voiceprint feature of the user according to the timbre feature and the sound feature.

[0061] In one embodiment, the initial features include cepstral graphs and Mel-frequency cepstral coefficients; the feature acquisition module includes:

[0062] A transform unit, configured to perform a discrete Fourier transform on the user's voice signal to obtain a frequency domain signal;

[0063] an amplitude extraction unit, configured to extract the amplitude of the speech signal from the frequency domain signal, and determine the square of the amplitude as a power spectrum;

[0064] A first logarithm taking unit, used for taking the logarithm of the power spectrum to obtain the logarithm of the power spectrum;

[0065] an inverse transform unit, configured to perform an inverse discrete Fourier transform on the logarithm of the power spectrum to obtain a cepstrum of the user's speech signal;

[0066] A filtering unit, configured to filter the power spectrum through a Mel frequency filter bank to obtain power spectrum energy in a Mel frequency scale;

[0067] A second logarithm taking unit, used for taking the logarithm of the power spectrum energy of the Mel frequency scale to obtain the logarithm of the power spectrum of the Mel frequency scale;

[0068] The cosine transform unit is used to perform discrete cosine transform on the logarithm of the power spectrum of the Mel-frequency scale to obtain Mel-frequency cepstrum coefficients of the user's speech signal.

[0069] In one of the embodiments, before determining the target user from the reference users according to the matching degree between the voiceprint feature of the user and the voiceprint features of each reference user, the device further comprises:

[0070] A voiceprint feature acquisition module, used to acquire the voiceprint features of each of the reference users;

[0071] A reference information acquisition module, used to acquire information of each reference user;

[0072] The first storage module is used to store the voiceprint features and information of each reference user in the server accordingly.

[0073] In one embodiment, the device further comprises:

[0074] A first user information acquisition module, configured to acquire the user's information when the voiceprint feature of the reference user does not match the voiceprint feature of the user;

[0075] The second storage module is used to store the user's information and voiceprint features in the server accordingly.

[0076] In one embodiment, the device further comprises:

[0077] A second user information acquisition module, used to acquire the user's information from each device;

[0078] An updating module is used to update the user information stored in the server using the user information obtained from each of the devices; the user information stored in the server is used to make personalized recommendations to the users logged in to each of the devices.

[0079] In a third aspect, an embodiment of the present application further provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program implements the steps in the above-mentioned recommended method when executed by the processor.

[0080] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps in the above-mentioned recommended method are implemented.

[0081] In a fifth aspect, the embodiments of the present application further provide a computer program product or a computer program, which includes a computer instruction stored in a computer-readable storage medium. The processor of the computer device reads the computer instruction from the computer-readable storage medium, and the processor executes the computer instruction, so that the computer device executes the method provided in various optional implementations described in the embodiments of the present application.

[0082] In summary, in the embodiments of the present application, after the initial features of the user's voice signal are obtained, the initial features can be adaptively adjusted to obtain accurate timbre features, and then accurate voiceprint features. According to the matching degree between the user's voiceprint features and the voiceprint features of each reference user, the target user can be determined from the reference users, and personalized recommendations can be made to the user based on the target user's information. Because the initial features are adaptively adjusted, the variability of the acoustic features can be corrected to obtain accurate voiceprint features, and then based on the matching degree between the user's voiceprint features and the voiceprint features of each reference user, the user can be determined as a target user, so as to achieve accurate personalized recommendations for the user based on the target user's information. BRIEF DESCRIPTION OF THE DRAWINGS

[0083] In order to more clearly illustrate the technical solutions in the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0084] Figure 1 This is a schematic diagram of the steps of a recommended method provided by an embodiment of the present application;

[0085] Figure 2is a schematic diagram of steps of another recommended method provided by an embodiment of the present application;

[0086] Figure 3 This is a schematic diagram of steps of another recommended method provided by an embodiment of the present application;

[0087] Figure 4 This is a schematic diagram of the steps of another recommended method provided by an embodiment of the present application;

[0088] Figure 5 is a structural schematic diagram of a recommended device provided in an embodiment of the present application;

[0089] Figure 6 It is a schematic diagram of the structure of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0090] The following will be combined with the drawings in this application to clearly and completely describe the technical solutions in this application. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. Based on the embodiments in the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.

[0091] In one embodiment, Figure 1 As shown, a recommendation method is provided. Although a logical order is shown in the step diagram, in some cases, the steps shown or described may be performed in an order different from that shown in the figure. Specifically, the recommendation method can be applied to a terminal or a server, wherein the terminal may include but is not limited to one or more of a smart phone, a tablet computer, a portable computer, a desktop computer, and a car computer. The server may be a physical server or a cloud server that provides various cloud services. It is worth noting that the present application does not limit the number of terminals or servers. According to the implementation needs, there may be any number of terminals or servers. For example, the server may be a single server or a server cluster consisting of multiple servers, and so on.

[0092] It should be noted that the order of description of the following embodiments is not intended to limit the priority order of the embodiments.

[0093] according to Figure 1 The recommended method shown in the figure includes at least steps S110 to S160, which are described in detail as follows:

[0094] In step S110, initial features of the user's speech signal are obtained.

[0095] The user may be any user, obtain the collected speech signal of the user when speaking, and extract initial features from the speech signal, where the initial features may at least include: a cepstral graph and / or a Mel-frequency cepstral coefficient.

[0096] Cepstrum is the Fourier transform of the spectrum of a signal. Cepstrum is a graphical representation of the signal obtained after cepstrum analysis. Cepstrum can help better analyze some hidden characteristics of the signal. By converting the speech signal into a cepstrum, the influence of the vocal tract can be removed, and key information such as the fundamental frequency of the speech can be obtained more accurately, thereby improving the accuracy of speech recognition.

[0097] Mel-Frequency Cepstral Coefficients (MFCC) are cepstral coefficients designed based on the human hearing characteristics. Mel-Frequency Cepstral Coefficients can effectively extract the characteristics of speech signals and convert speech signals into a set of coefficient vectors that can represent the essential characteristics of speech. Different people have different vocal tract shapes and vocalization habits, which will lead to different MFCC features. Therefore, different speakers can be distinguished based on MFCC features to achieve speaker recognition.

[0098] The method of extracting initial features from speech signals will be described in detail later.

[0099] In step S120, the initial features are adaptively adjusted to obtain the timbre features of the speech signal.

[0100] Because the acoustic features of human speech signals are variable, even the same person's speech may have different acoustic features at different times. Acoustic features refer to various features related to speech signals, which may include initial features, timbre features, and voiceprint features. In order to ensure that the speech of the same person in different states is not recognized as the speech of different people, the initial features can be adaptively adjusted to correct the variability of the acoustic features, obtain accurate timbre features, and then obtain accurate voiceprint features.

[0101] The initial features of the voice signal of the same user in different states, the timbre features obtained after adaptive adjustment, are less different from the real timbre features of the voice signal of the user. Optionally, the initial features of the voice signal of the same user in different states, the timbre features obtained after adaptive adjustment, are less different from the real timbre features of the voice signal of the user, than a difference threshold. The difference threshold may be preset. The real timbre features of the user's voice signal may be the timbre features of the pre-recorded voice signal.

[0102] For example, the user pre-records a voice signal A, and the timbre feature A1 can be extracted from the voice signal A; subsequently, the user's voice signal B when he has a cold and the voice signal C after he has spoken for a long time are collected, and the initial feature B1 can be extracted from the voice signal B when he has a cold, and the initial feature C1 can be extracted from the voice signal C after he has spoken for a long time. The timbre feature B2 can be obtained by adaptively adjusting the initial feature B1, and the timbre feature C2 can be obtained by adaptively adjusting the initial feature C1. Then the difference between B2 and A1 is less than the difference threshold, and the difference between C2 and A1 is also less than the difference threshold.

[0103] In one embodiment, a timbre extraction model may be trained, and the timbre extraction model may be used to adaptively adjust the initial features to obtain the timbre features of the speech signal.

[0104] In another embodiment, an adaptive filtering algorithm may be used to estimate the characteristics of background noise, remove noise features from noisy initial features, and implement adaptive adjustment of the initial features, thereby removing the influence of variation factors such as environmental interference on the timbre features, and obtaining timbre features that are less different from the user's actual timbre features.

[0105] In step S130, the voiceprint feature of the user is determined based on the timbre feature.

[0106] Voiceprint features are a collection of acoustic features in a person's voice that can reflect individual differences. They are unique and can be used to identify different individuals. Different acoustic features can reflect the information of voice signals from different angles. In order to improve accuracy, timbre features and other acoustic features can be combined to obtain accurate voiceprint features.

[0107] In step S140, a target user is determined from the reference users according to the matching degree between the voiceprint feature of the user and the voiceprint features of each reference user.

[0108] The reference user may be any user. The voiceprint features and information of the reference user are obtained in advance, and the voiceprint features and information of the reference user may be stored in the server accordingly. The voiceprint features of the reference user may be voiceprint features extracted based on the voice signals pre-recorded by each reference user. The information of the reference user may be various information of the user input by the user or obtained with the user's consent, and may include the user's attribute information and historical behavior information.

[0109] In one embodiment, a recurrent neural network (RNN) or a long short-term memory network (LSTM) can be used to process the user's voiceprint features, and finally a hidden Markov model (HMM) is used to perform state prediction and compare the prediction results with the voiceprint features of reference users to determine the degree of match between the voiceprint features of each reference user and the user's voiceprint features.

[0110] The user's voiceprint features are sorted in the time sequence of the speech to form a voiceprint feature sequence. The voiceprint feature sequence is input frame by frame into the recurrent neural network or long short-term memory network. The network learns the internal pattern of the voiceprint features by continuously updating the hidden layer state, and finally outputs a sequence. The output sequence is a representation of the voiceprint features after network learning and abstraction at different time steps, which contains high-level semantic information of the voiceprint features in the time dimension, such as the pattern information in the speech that reflects the speaker's personality characteristics and changes over time.

[0111] Recurrent neural networks can capture the dependencies of voiceprint features in the time dimension, and long short-term memory networks can effectively handle the changes in voiceprint features at different time scales, such as the timbre and pitch of speech that change over time.

[0112] Estimate the parameters of the hidden Markov model based on prior knowledge or training data, including determining the initial state probability, state transition probability, and observation probability. For example, through a large amount of annotated speaker speech data, the transition frequency between different states is counted to estimate the state transition probability and the probability of generating specific observation features in different states. The sequence output by the recurrent neural network or the long short-term memory network is input into the hidden Markov model as the observation sequence. The hidden Markov model row state prediction finds the most likely hidden state sequence given the observation sequence. This hidden state sequence can reflect the potential state changes corresponding to the voiceprint features in time, such as which parts correspond to the speaker's stable pronunciation state and which parts correspond to the pronunciation transition state.

[0113] The voiceprint features of multiple reference users are also preprocessed through the above-mentioned feature extraction, fusion, and possible recurrent neural network or long short-term memory network, and are stored and represented according to certain standards so as to be matched with the user's voiceprint features.

[0114] The voiceprint features of the user processed by the recurrent neural network or the long short-term memory network and combined with the prediction of the hidden Markov model are matched with the corresponding voiceprint features of each reference user. A variety of matching methods can be used, such as calculating the distance metric between the two, or comparing based on probability.

[0115] When using distance to measure the matching degree, taking the Euclidean distance as an example, the Euclidean distance between the user to be identified and the reference user in the corresponding voiceprint feature representation is calculated. The smaller the distance, the higher the matching degree. If the matching degree is calculated based on probability, the higher the probability of the reference user model, the higher the matching degree corresponding to it.

[0116] In step S150, the information of the target user is obtained.

[0117] The information of each reference user is acquired and stored in advance, so the information of the target user can be directly acquired from the pre-stored information of the reference user.

[0118] In step S160, personalized recommendations are made to the target user based on the target user's information.

[0119] Based on the target user's information, the target content to be recommended to the user is determined, and the target content is recommended to the user, thereby achieving personalized recommendation for the user.

[0120] By adopting the technical solution of the embodiment of the present application, after the initial features of the user's voice signal are obtained, the initial features can be adaptively adjusted to obtain accurate timbre features, and then accurate voiceprint features. According to the matching degree between the user's voiceprint features and the voiceprint features of each reference user, the target user can be determined from the reference users, and personalized recommendations can be made to the user based on the target user's information. Because the initial features are adaptively adjusted, the variability of the acoustic features can be corrected to obtain accurate voiceprint features, and then based on the matching degree between the user's voiceprint features and the voiceprint features of each reference user, the user can be determined as a target user, so as to achieve accurate personalized recommendations for the user based on the target user's information.

[0121] In one embodiment, a family has various family members (reference users), and each user has different preferences for TV programs, such as sports programs, music programs, and cartoons. Without understanding the user's preferences, it is difficult to accurately recommend programs to the user. User accounts of various reference users can be established in advance, and voice signals and information of various reference users can be recorded. Voiceprint features of reference users can be extracted based on the recorded voice signals, and voiceprint features and information of each reference user can be stored in the user account of the reference user.

[0122] The information of the reference user may include attribute information such as age, gender and / or occupation, and may also include historical behavior information. The historical behavior information may include the behavior of the reference user on the content recommended to him / her in the past, and may also include the playback history of the reference user.

[0123] During the TV broadcast process, if a voice signal of a user requesting a TV program recommendation is collected, the initial features can be extracted based on the voice signal, and the initial features can be adaptively adjusted to obtain the timbre features of the voice signal, and then the user's voiceprint features can be obtained. The matching degree between the user's voiceprint features and the pre-stored voiceprint features of each reference user is calculated, and the reference user corresponding to the voiceprint feature with the highest matching degree is determined as the target user. The stored target user information is obtained, and the target TV program to be recommended to the user is determined based on the target user information. The target TV program is displayed in a prominent position on the TV program selection interface to achieve personalized recommendations for the user.

[0124] Alternatively, when a voice signal is received from a user requesting to turn on the TV during the TV standby process, the initial features are extracted based on the voice signal, and the initial features are adaptively adjusted to obtain the timbre features of the voice signal, and then the voiceprint features of the user are obtained. The matching degree between the voiceprint features of the user and the pre-stored voiceprint features of each reference user is calculated, and the reference user corresponding to the voiceprint feature with the highest matching degree is determined as the target user. The stored target user information is obtained, and the target TV program to be recommended to the user is determined based on the target user information. The target TV program is directly used as the first TV program to be played after turning on the TV, so that the user can see the TV program of interest when the TV is initially turned on, and there is no need to look for the TV program of interest, thereby improving the user experience.

[0125] On the basis of the above technical solution, as an embodiment, the initial features include a cepstrum and a Mel-frequency cepstrum coefficient; the initial features of the user's voice signal may include: performing a discrete Fourier transform on the user's voice signal to obtain a frequency domain signal; extracting the amplitude of the voice signal from the frequency domain signal, and determining the square of the amplitude as the power spectrum; taking the logarithm of the power spectrum to obtain the logarithm of the power spectrum; performing an inverse discrete Fourier transform on the logarithm of the power spectrum to obtain a cepstrum of the user's voice signal; filtering the power spectrum through a Mel-frequency filter group to obtain a power spectrum energy on a Mel-frequency scale; taking the logarithm of the power spectrum energy on a Mel-frequency scale to obtain a power spectrum logarithm on a Mel-frequency scale; performing a discrete cosine transform on the logarithm of the power spectrum on a Mel-frequency scale to obtain a Mel-frequency cepstrum coefficient of the user's voice signal.

[0126] By performing discrete Fourier transform (DFT) on the speech signal, taking the amplitude, square, logarithm, and performing inverse discrete Fourier transform (IDFT), the cepstral coefficients can be obtained, and the cepstral graph is generated based on the cepstral coefficients.

[0127] Based on the Mel-scale power spectrum logarithm, discrete cosine transform (DCT) is applied and the first N coefficients are selected as MFCC features.

[0128] Optionally, the gas passing through the user's glottis is recorded as h(t), the user's vocal cords and vocal tract can be equivalent to a complex filter, recorded as e(t), and the user's speech signal is recorded as x(t). According to x(t) = h(t)*e(t), X(n) = H(n)·E(n), we can convert log|X(n)| 2 =2log|H(n)|+2log|E(n)|; where * represents the convolution operation, · represents the multiplication operation, X(n), H(n), and E(n) are the discrete forms of x(t), h(t), and e(t) respectively; 2log|E(n)| is the envelope signal, which extracts the fundamental frequency and formant. The fundamental frequency is the basic frequency of speech, which determines the pitch, and the formant is the resonance frequency of the vocal tract, which determines the timbre.

[0129] Performing inverse Fourier transform and discrete cosine transform on the extracted signal can obtain the cepstrum and MFCC features. mfcc C[x(n)]=F -1 [log(|F1[x(n)]| 2 )] to obtain the inverse spectrum; where x(n) represents the discretized speech signal, F1 represents the discrete Fourier transform, and F -1 represents the inverse discrete Fourier transform, C[x(n)] is the cepstral coefficient, n mfcc is the number of Mel-frequency cepstral coefficients, log(|·| 2 ) is the result of discrete Fourier transform. The amplitude of the speech signal is first taken, then the square is taken, and finally the logarithm is taken to finally obtain the result of discrete Fourier inverse transform.

[0130] Treat the logarithm of the power spectrum as a time domain signal, perform Fourier analysis on it, and then take the first n mfcc The calculated value corresponding to the frequency is used as the Mel-frequency cepstrum coefficient.

[0131] Using the technical solution of the embodiment of the present application, the cepstral graph and Mel-frequency cepstral coefficients of the speech signal are extracted as initial features. Cepstral analysis can separate the influence of the sound source and the vocal tract, and then derive the resonance characteristics and vocal cord vibration characteristics of the speech signal. Cepstral analysis can also display the essential characteristics of the signal and reduce the interference of noise on timbre analysis. The Mel-frequency cepstral coefficients are designed based on the Mel-frequency scale. The Mel-frequency scale reflects that the human ear's perception of sounds of different frequencies is a nonlinear characteristic. Therefore, the Mel-frequency cepstral coefficients can better simulate the human ear's perception of timbre and extract timbre features related to human hearing. The Mel-frequency cepstral coefficients can effectively capture the key timbre elements such as the resonance peak structure, fundamental frequency and harmonics in the speech signal. It can also reduce the complexity of data processing while retaining the main characteristics of the timbre and removing redundant information, making subsequent timbre analysis and classification tasks more efficient.

[0132] Based on the above technical solution, as an embodiment, the adaptive adjustment of the initial features to obtain the timbre features of the speech signal may include: inputting the initial features into a timbre extraction model to obtain the adaptive weights assigned by the timbre extraction model to the initial features; adjusting the initial features according to the adaptive weights to obtain the timbre features output by the timbre extraction model.

[0133] Adaptive weights are parameters assigned by the timbre extraction model to the initial features. When the initial features input to the timbre extraction model are different, the adaptive weights obtained are also different. The initial features are adjusted based on the adaptive weights, and the difference between the obtained timbre features and the real timbre features of the user's voice signal is small. The timbre extraction model uses adaptive weights to weight the initial features and obtain the timbre features of the voice signal.

[0134] By adopting the technical solution of the embodiment of the present application and using a timbre extraction model to convert initial features into timbre features, the timbre features can be quickly acquired with high efficiency.

[0135] On the basis of the above technical solution, as an embodiment, the training step of the timbre extraction model at least includes: obtaining multiple speech signal samples and timbre feature samples of each of the speech signal samples; obtaining initial feature samples of each of the speech signal samples; inputting the multiple initial feature samples into the initial timbre extraction model to obtain adaptive weight samples assigned by the initial timbre extraction model to the initial feature samples; adjusting the initial feature samples according to the adaptive weight samples to obtain predicted timbre features output by the initial timbre extraction model; training the initial timbre extraction model based on the matching degree between the timbre feature samples of each of the speech signal samples and the predicted timbre features to obtain the trained timbre extraction model.

[0136] The speech signal samples may be individual speech signal samples obtained from a timbre library, which also stores timbre feature samples of the individual speech signal samples. For each speech signal sample, an initial feature sample of each speech signal sample may be obtained. The method for obtaining the initial feature sample of each speech signal sample may refer to the method for obtaining the initial features of the speech signal described above.

[0137] The initial timbre extraction model is a timbre extraction model to be trained. The initial feature samples are input into the initial timbre extraction model. The initial timbre extraction model can predict the adaptive weight samples corresponding to the initial feature samples. The initial feature samples are weighted according to the adaptive weight samples to obtain the predicted timbre features output by the initial timbre extraction model. Because the initial timbre extraction model has not been trained, the adaptive weight samples obtained are not accurate enough, and the predicted timbre features obtained are also not accurate enough.

[0138] The training goal of the initial timbre extraction model is to make the predicted timbre features of the speech signal samples output by the model consistent with the timbre feature samples. Therefore, based on the matching degree between the timbre feature samples and the predicted timbre features of each of the speech signal samples, the initial timbre extraction model is iteratively trained until the initial timbre extraction model converges, or the matching degree between the predicted timbre features of the speech signal samples output by the initial timbre extraction model and the timbre feature samples meets the threshold condition, thereby obtaining the trained timbre extraction model.

[0139] Among them, the recursive least squares method and the method of adaptively learning model parameters by stochastic gradient descent can be used to train the timbre extraction model to achieve the best match between the predicted timbre features and the timbre feature samples. Recursive least squares is an algorithm for online system identification that updates model parameters by minimizing the sum of squares of prediction errors. For a given sequence of data points, the recursive least squares algorithm can quickly update the weight vector each time new data arrives.

[0140] The technical solution of the embodiment of the present application is adopted. Since the training goal is to make the predicted timbre features of the speech signal samples output by the model consistent with the timbre feature samples, the initial timbre extraction model is trained based on the matching degree between the timbre feature samples and the predicted timbre features of each speech signal sample, and the timbre features of the user's speech signal output by the trained timbre extraction model are less different from the actual timbre features of the user's speech signal, thereby overcoming the variability of the timbre features.

[0141] Based on the above technical solution, as an embodiment, determining the voiceprint features of the user based on the timbre features may include: extracting the sound features of the speech signal based on the timbre features; and determining the voiceprint features of the user based on the timbre features and the sound features.

[0142] The sound features of the speech signal can be extracted based on the timbre features. The sound features may include but are not limited to: pitch features and / or energy features. Optionally, the sound features of the speech signal can be extracted using a convolutional neural network (CNN) using a linear prediction coefficient (LPC) under an autoregressive model. The method for extracting sound features can also refer to related technologies.

[0143] The timbre features and sound features are integrated, and speaker verification and transfer learning are performed to obtain the user's voiceprint features.

[0144] Optionally, a subset of features that are important for speaker recognition can be selected from the timbre features and sound features, and possible redundant or highly correlated features can be removed. These selected features are then normalized, such as by using zero-mean normalization, to map the feature values ​​to the same magnitude range, to avoid the impact of dimensional differences of different features on subsequent fusion and recognition. The timbre features and sound features can be concatenated or weightedly fused to obtain the user's initial voiceprint features; the timbre features and sound features can also be fused based on deep learning, using a neural network architecture to build a simple multi-layer perceptron (MLP) or convolutional neural network, with the timbre features and sound features as different nodes of the input layer, and nonlinear transformation and feature interaction are performed through the hidden layer in the network to finally output the user's initial voiceprint features.

[0145] By inputting the user's initial voiceprint features into the trained speaker recognition model, the model outputs the corresponding speaker identity prediction results, determines whether it is the target speaker, and thus performs speaker recognition, and determines the initial voiceprint features through speaker recognition as the user's voiceprint features.

[0146] When faced with different tasks, voiceprint features suitable for the task can also be obtained through transfer learning. With the help of the general speech features learned from large-scale data by the pre-trained model, the model can converge faster on a relatively small target data set and better learn the voiceprint features under specific tasks, thereby improving the accuracy and robustness of speaker recognition. Especially when data is limited, the advantages of transfer learning are more obvious.

[0147] The technical solution of the embodiment of the present application comprehensively considers the timbre feature recognition and the sound feature to obtain accurate voiceprint features, so as to enhance the speaker recognition effect by integrating features at different levels. Based on the application of speaker recognition and transfer learning, voiceprint features can be more effectively used to accurately judge the speaker's identity and promote their application in different scenarios.

[0148] Based on the above technical solution, as an embodiment, before determining the target user from the reference users according to the matching degree between the voiceprint feature of the user and the voiceprint features of each reference user, as follows: Figure 2 As shown, the method may further include steps S210 to S230.

[0149] In step S210, the voiceprint features of each of the reference users are obtained.

[0150] In step S220, information of each of the reference users is obtained.

[0151] In step S230, the voiceprint features and information of each reference user are correspondingly stored in the server.

[0152] The method for obtaining the voiceprint features of each reference user may refer to the method for obtaining the voiceprint features of a user described above.

[0153] The information of the reference user may include attribute information such as age, gender and / or occupation, and may also include historical behavior information. The historical behavior information may include the reference user's behavior on the content recommended to him / her in the past, and may also include the reference user's playback history, etc.

[0154] A user account of a reference user may be established in advance, and the voiceprint features and information of the reference user may be stored in the server under the user account of the reference user. When the information of the reference user is updated, the information of the reference user stored in the server may be updated.

[0155] When it is necessary to determine the matching degree between the voiceprint features of the user and the voiceprint features of each reference user, the matching degree of the voiceprint features of each reference user may be obtained from the server.

[0156] By adopting the technical solution of the embodiment of the present application, the voiceprint features and information of each reference user can be generated and stored in advance. When it is necessary to determine personalized recommendations for users, the voiceprint features of each reference user can be quickly acquired to quickly match the user's voiceprint features with the voiceprint features of each reference user, thereby improving the efficiency of personalized recommendations.

[0157] Based on the above technical solution, as an embodiment, Figure 3As shown, the method may further include steps S310 to S320.

[0158] In step S310, when the voiceprint feature of the reference user does not match the voiceprint feature of the user, the information of the user is obtained.

[0159] In step S320, the user information and voiceprint features are stored in the server accordingly.

[0160] When the voiceprint features of the reference user do not match the voiceprint features of the user, it proves that the voiceprint features and information of the user are not stored in the server in advance. Therefore, the user's information can be directly obtained, and personalized recommendations can be made to the user based on the user's information. The obtained user information and the user's voiceprint features can also be used as the new reference user's information and voiceprint features, and stored in the server accordingly, so that when personalized recommendations are made to the user again based on the voice signal, the user can be directly identified from multiple reference users based on the voiceprint features of the voice signal, and the user's information can be obtained to quickly make personalized recommendations to the user.

[0161] By adopting the technical solution of the embodiment of the present application, when the voiceprint features of the reference user do not match the voiceprint features of the user, the user's information and the user's voiceprint features can be stored in the server, thereby supplementing the reference user's information and voiceprint features, so that when personalized recommendations are needed to the user again in the future, the recommendation can be completed quickly.

[0162] Based on the above technical solution, as an embodiment, Figure 4 As shown, the method may further include steps S410 to S420.

[0163] In step S410, the user's information is obtained from each device.

[0164] In step S420, the user information stored in the server is updated using the user information obtained from each of the devices.

[0165] The user information stored in the server is used to make personalized recommendations to the users logged in to each of the devices.

[0166] A unified user identity authentication mechanism can be used for each user to ensure the consistency of the user experience on different devices. Through the server, the user's data can be synchronized between different devices to maintain the consistency of the user's playback status, preference settings and other information on each device.

[0167] The user information obtained may be the information displayed by the user account of the user on each device, and the user information obtained from each device is saved to the server. Accordingly, when the same user account is logged in to each different device, the user information corresponding to the user account can be obtained from the server, so as to make personalized recommendations for the logged-in user based on the user information.

[0168] For example, the same user account shows a preference for animation videos on the logged-in mobile phone and a preference for sports videos on the logged-in tablet computer. It can be determined that the preferences of the user corresponding to the user account include animation videos and sports videos, and the preference information is stored on the server. When the user account subsequently logs in to any device (for example, the user account logs in to a smart TV), the preferences corresponding to the user account, including animation videos and sports videos, can be obtained from the server. The smart TV can then make personalized recommendations of animation videos and sports videos to the user.

[0169] By adopting the technical solution of the embodiment of the present application, cross-device synchronization can be achieved. Each device logged in to the same user account can obtain the voiceprint features and information of the user corresponding to the user account from the server, so as to make personalized recommendations based on the voiceprint features on each device.

[0170] In one embodiment, making personalized recommendations to the user may include: acquiring a recommendation strategy; determining the target content of the user using the recommendation strategy according to the information of the target user; and recommending the target content to the user.

[0171] Different recommendation strategies can be determined based on the information obtained about the target user. For example, when there is less information about the target user, the recommendation strategy adopted may be a cold start strategy, and the target content determined based on the cold start strategy may be a hot content, and the target content is recommended to the user.

[0172] In one embodiment, user feedback on the recommended target content may be collected, and the recommendation strategy may be adjusted differently according to the user feedback on the target content to improve the accuracy of personalized recommendations.

[0173] In order to better implement the recommendation method of the present application, the present application also provides a recommendation device based on the above recommendation method. The meanings of the terms are the same as those in the above recommendation method, and the specific implementation details can refer to the description in the method embodiment.

[0174] See also Figure 5 , Figure 5 : is a schematic diagram of the structure of a recommendation device provided in an embodiment of the present application, the recommendation device comprising:

[0175] A feature acquisition module 501 is used to acquire initial features of a user's speech signal;

[0176] A feature adjustment module 502, configured to adaptively adjust the initial feature to obtain a timbre feature of the speech signal;

[0177] A feature determination module 503, configured to determine a voiceprint feature of the user based on the timbre feature;

[0178] A user determination module 504, configured to determine a target user from the reference users according to a matching degree between the voiceprint feature of the user and the voiceprint features of each reference user;

[0179] The information acquisition module 505 is used to acquire the information of the target user;

[0180] The personalized recommendation module 506 is used to make personalized recommendations to the target user based on the target user's information.

[0181] In one embodiment, the feature adjustment module 502 includes:

[0182] An input unit, used for inputting the initial feature into a timbre extraction model to obtain an adaptive weight assigned by the timbre extraction model to the initial feature;

[0183] An adjustment unit is used to adjust the initial feature according to the adaptive weight to obtain the timbre feature output by the timbre extraction model.

[0184] In one embodiment, the training step of the timbre extraction model at least includes:

[0185] Acquire a plurality of speech signal samples and a timbre feature sample of each of the speech signal samples;

[0186] Acquire initial feature samples of each of the speech signal samples;

[0187] Inputting a plurality of the initial feature samples into an initial timbre extraction model to obtain adaptive weight samples assigned by the initial timbre extraction model to the initial feature samples;

[0188] Adjusting the initial feature sample according to the adaptive weight sample to obtain the predicted timbre feature output by the initial timbre extraction model;

[0189] Based on the matching degree between the timbre feature samples of each of the speech signal samples and the predicted timbre feature, the initial timbre extraction model is trained to obtain the trained timbre extraction model.

[0190] In one embodiment, the feature determination module 503 includes:

[0191] An extraction unit, configured to extract a sound feature of the speech signal based on the timbre feature;

[0192] A voiceprint feature determination unit is used to determine the voiceprint feature of the user according to the timbre feature and the sound feature.

[0193] In one embodiment, the initial features include cepstral graphs and Mel-frequency cepstral coefficients; the feature acquisition module 501 includes:

[0194] A transform unit, configured to perform a discrete Fourier transform on the user's voice signal to obtain a frequency domain signal;

[0195] an amplitude extraction unit, configured to extract the amplitude of the speech signal from the frequency domain signal, and determine the square of the amplitude as a power spectrum;

[0196] A first logarithm taking unit, used for taking the logarithm of the power spectrum to obtain the logarithm of the power spectrum;

[0197] an inverse transform unit, configured to perform an inverse discrete Fourier transform on the logarithm of the power spectrum to obtain a cepstrum of the user's speech signal;

[0198] A filtering unit, configured to filter the power spectrum through a Mel frequency filter bank to obtain power spectrum energy in a Mel frequency scale;

[0199] A second logarithm taking unit, used for taking the logarithm of the power spectrum energy of the Mel frequency scale to obtain the logarithm of the power spectrum of the Mel frequency scale;

[0200] The cosine transform unit is used to perform discrete cosine transform on the logarithm of the power spectrum of the Mel-frequency scale to obtain Mel-frequency cepstrum coefficients of the user's speech signal.

[0201] In one of the embodiments, before determining the target user from the reference users according to the matching degree between the voiceprint feature of the user and the voiceprint features of each reference user, the device further comprises:

[0202] A voiceprint feature acquisition module, used to acquire the voiceprint features of each of the reference users;

[0203] A reference information acquisition module, used to acquire information of each reference user;

[0204] The first storage module is used to store the voiceprint features and information of each reference user in the server accordingly.

[0205] In one embodiment, the device further comprises:

[0206] A first user information acquisition module, configured to acquire the user's information when the voiceprint feature of the reference user does not match the voiceprint feature of the user;

[0207] The second storage module is used to store the user's information and voiceprint features in the server accordingly.

[0208] In one embodiment, the device further comprises:

[0209] A second user information acquisition module, used to acquire the user's information from each device;

[0210] An updating module is used to update the user information stored in the server using the user information obtained from each of the devices; the user information stored in the server is used to make personalized recommendations to the users logged in to each of the devices.

[0211] By adopting the technical solution of the embodiment of the present application, after the initial features of the user's voice signal are obtained, the initial features can be adaptively adjusted to obtain accurate timbre features, and then accurate voiceprint features. According to the matching degree between the user's voiceprint features and the voiceprint features of each reference user, the target user can be determined from the reference users, and personalized recommendations can be made to the user based on the target user's information. Because the initial features are adaptively adjusted, the variability of the acoustic features can be corrected to obtain accurate voiceprint features, and then based on the matching degree between the user's voiceprint features and the voiceprint features of each reference user, the user can be determined as a target user, so as to achieve accurate personalized recommendations for the user based on the target user's information.

[0212] For the specific definition of the recommendation device, please refer to the definition of the recommendation method above, which will not be repeated here. Each module in the above recommendation device can be implemented in whole or in part by software, hardware and a combination thereof. Each of the above modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to each of the above modules.

[0213] In addition, the present application also provides an electronic device, such as Figure 6 As shown, it shows a schematic diagram of the structure of the electronic device involved in this application, specifically:

[0214] The electronic device may include one or more processors 601 of processing cores and one or more computer-readable storage media memories 602 and other components. Those skilled in the art will appreciate that Figure 6The electronic device structure shown in the figure does not constitute a limitation on the electronic device, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0215] The processor 601 is the control center of the electronic device. It uses various interfaces and lines to connect various parts of the entire electronic device. By running or executing software programs and / or modules stored in the memory 602, and calling data stored in the memory 602, it performs various functions of the electronic device and processes data, thereby monitoring the electronic device as a whole. Optionally, the processor 601 may include one or more processing cores; preferably, the processor 601 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface, and application programs, and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 601.

[0216] The memory 602 can be used to store software programs and modules. The processor 601 executes various functional applications and data processing by running the software programs and modules stored in the memory 602. The memory 602 may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 602 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices. Accordingly, the memory 602 may also include a memory controller to provide the processor 601 with access to the memory 602.

[0217] In one embodiment, the electronic device further includes a power supply 603 for supplying power to various components. Preferably, the power supply 603 can be logically connected to the processor 601 through a power management system, so as to manage charging, discharging, and power consumption through the power management system. The power supply 603 can also include any components such as one or more DC or AC power supplies, recharging systems, power supply device debugging circuits, power converters or inverters, and power status indicators.

[0218] In one embodiment, the electronic device may further include an input unit 604, which may be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal input related to user settings and function control.

[0219] Although not shown, the electronic device may further include a display unit, etc., which will not be described in detail herein. Specifically in this embodiment, the processor 601 in the electronic device will load the executable files corresponding to the processes of one or more application programs into the memory 602 according to the following instructions, and the processor 601 will run the application programs stored in the memory 602, thereby implementing the steps in any one of the recommended methods provided in the embodiments of the present application.

[0220] Those skilled in the art will understand that Figure 6 The structure shown in the figure is merely a block diagram of a partial structure related to the scheme of the present application, and does not constitute a limitation on the electronic device to which the scheme of the present application is applied. The specific electronic device may include more or fewer components than shown in the figure, or combine certain components, or have a different arrangement of components.

[0221] In one embodiment, an electronic device is provided, including a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the method described in any embodiment of the present application is implemented.

[0222] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the method described in any embodiment of the present application is implemented.

[0223] In some embodiments, a computer program product is also proposed, including a computer program or instructions, which implement the method described in any embodiment of the present application when executed by a processor.

[0224] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.

[0225] A person of ordinary skill in the art will appreciate that all or part of the steps in the various methods of the above embodiments may be completed by instructions, or by controlling related hardware through instructions. The instructions may be stored in a computer-readable storage medium and loaded and executed by a processor.

[0226] To this end, the present application provides a computer-readable storage medium, on which a computer program is stored. The computer program can be loaded by a processor to execute the steps in any one of the recommended methods provided in the present application.

[0227] The specific implementation of the above operations can be found in the previous embodiments, which will not be described in detail here.

[0228] The computer-readable storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0229] Since the instructions stored in the computer-readable storage medium can execute the steps in any of the recommendation methods provided in the present application, the beneficial effects that can be achieved by any of the recommendation methods provided in the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.

[0230] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or terminal device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or terminal device. In the absence of further restrictions, the elements defined by the sentence "including one..." do not exclude the existence of other identical elements in the process, method, article or terminal device including the elements.

[0231] The above is a detailed introduction to a recommended method, device, electronic device and computer-readable storage medium provided by the present application. Specific examples are used herein to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea. At the same time, for those skilled in the art, according to the idea of ​​the present invention, there will be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the present invention.

Claims

1. A recommendation method, characterized in that: include: Acquire initial features of a user's voice signal; Adaptively adjusting the initial features to obtain timbre features of the speech signal; Determining a voiceprint feature of the user based on the timbre feature; Determining a target user from the reference users according to a matching degree between the voiceprint feature of the user and the voiceprint features of each reference user; Acquire information of the target user; According to the information of the target user, personalized recommendations are made to the user.

2. The method according to claim 1, characterized in that The adaptively adjusting the initial features to obtain the timbre features of the speech signal includes: Inputting the initial features into a timbre extraction model to obtain an adaptive weight assigned by the timbre extraction model to the initial features; The initial feature is adjusted according to the adaptive weight to obtain the timbre feature output by the timbre extraction model.

3. The method according to claim 2, characterized in that The training step of the timbre extraction model at least includes: Acquire a plurality of speech signal samples and a timbre feature sample of each of the speech signal samples; Acquire initial feature samples of each of the speech signal samples; Inputting a plurality of the initial feature samples into an initial timbre extraction model to obtain adaptive weight samples assigned by the initial timbre extraction model to the initial feature samples; Adjusting the initial feature sample according to the adaptive weight sample to obtain the predicted timbre feature output by the initial timbre extraction model; Based on the matching degree between the timbre feature samples of each of the speech signal samples and the predicted timbre feature, the initial timbre extraction model is trained to obtain the trained timbre extraction model.

4. The method according to claim 1, characterized in that: The determining the voiceprint feature of the user based on the timbre feature includes: Extracting the sound features of the speech signal based on the timbre features; Determine the voiceprint feature of the user according to the timbre feature and the sound feature.

5. The method according to claim 1, characterized in that The initial features include cepstral graph and Mel frequency cepstral coefficients; The obtaining of initial features of the user's voice signal comprises: Performing discrete Fourier transform on the user's voice signal to obtain a frequency domain signal; Extracting the amplitude of the speech signal from the frequency domain signal, and determining the square of the amplitude as a power spectrum; Taking the logarithm of the power spectrum to obtain the logarithm of the power spectrum; Performing an inverse discrete Fourier transform on the logarithm of the power spectrum to obtain a cepstrum of the user's speech signal; Filtering the power spectrum through a Mel frequency filter bank to obtain power spectrum energy in Mel frequency scale; Taking the logarithm of the power spectrum energy of the Mel frequency scale to obtain the logarithm of the power spectrum of the Mel frequency scale; Perform discrete cosine transform on the logarithm of the power spectrum of the mel-frequency scale to obtain mel-frequency cepstrum coefficients of the user's speech signal.

6. The method according to claim 1, characterized in that Before determining the target user from the reference users according to the matching degree between the voiceprint feature of the user and the voiceprint features of each reference user, the method further includes: Acquire voiceprint features of each of the reference users; Acquire information of each of the reference users; The voiceprint features and information of each reference user are correspondingly stored in the server.

7. The method according to claim 1, characterized in that The method further comprises: When the voiceprint feature of the reference user does not match the voiceprint feature of the user, obtaining information of the user; The user's information and voiceprint features are stored in the server accordingly.

8. The method according to claim 7, characterized in that The method further comprises: Obtaining information of the user from each device; The user information stored in the server is updated using the user information obtained from each of the devices; the user information stored in the server is used to make personalized recommendations to the user logged in to each of the devices.

9. A recommendation device, characterized in that: include: A feature acquisition module, used to acquire initial features of a user's voice signal; A feature adjustment module, used for adaptively adjusting the initial feature to obtain the timbre feature of the speech signal; A feature determination module, used to determine the voiceprint feature of the user based on the timbre feature; A user determination module, configured to determine a target user from the reference users according to a matching degree between the voiceprint feature of the user and the voiceprint features of each reference user; An information acquisition module, used to acquire the information of the target user; The personalized recommendation module is used to make personalized recommendations to the user based on the information of the target user.

10. An electronic device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the recommendation method according to any one of claims 1 to 8 when executing the computer program.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the recommendation method according to any one of claims 1 to 8 are implemented.