Speech style generation method and device, electronic equipment and storage medium
By fitting style feature attributes and style feature vectors, the target speaking style is directly generated, which solves the problem of low generation efficiency caused by model retraining in existing technologies, and realizes rapid transfer of speaking style and improved efficiency.
Patent Information
- Application Number
- CN202210714001.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-22
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2042-06-22
AI Technical Summary
In existing technologies, generating new speaking styles requires retraining the model, resulting in low generation efficiency.
By fitting the target style feature attribute based on multiple style feature attributes, determining the fitting coefficient of each style feature attribute, and training the speaking style model using multiple style feature vectors, the target speaking style is generated by directly inputting the target style feature vector, thus avoiding retraining the model.
It enables rapid transfer of speaking styles and improves generation efficiency.
Smart Images

Figure CN115270922B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of computer science and natural language processing, and more particularly to a method, apparatus, electronic device, and storage medium for generating speaking styles. Background Technology
[0002] As human-computer interaction evolves from single-voice interaction to multimodal interaction, voice-driven virtual digital humans have emerged. These virtual digital humans are now entering a growth phase and have already integrated with industries such as culture and tourism, finance, live streaming, gaming, and film and entertainment. Driven by continued advancements in artificial intelligence technology, they are developing towards greater intelligence, refinement, and diversification. Different people have different speaking styles; for example, some people speak with accurate lip movements and rich facial expressions, while others speak with smaller mouth movements and a more serious expression. Therefore, it is possible to design three-dimensional virtual digital humans with different speaking styles.
[0003] However, with existing technologies, each new speaking style requires retraining the model and involves a lot of data processing, resulting in low efficiency in generating new speaking styles. Summary of the Invention
[0004] This disclosure provides a method, apparatus, electronic device, and storage medium for generating speaking styles, which can achieve rapid transfer of speaking styles and improve the efficiency of speaking style generation.
[0005] Firstly, this disclosure provides a method for generating speaking styles, including:
[0006] Fitting the target style feature attribute based on multiple style feature attributes, and determining the fitting coefficient of each style feature attribute;
[0007] Based on the fitting coefficients of each style feature attribute and multiple style feature vectors, a target style feature vector is determined, wherein each style feature vector corresponds one-to-one with the multiple style feature attributes.
[0008] The target style feature vector is input into the speaking style model, and the target speaking style parameters are output. The speaking style model is obtained based on the framework of training the speaking style model based on the multiple style feature vectors.
[0009] Based on the target speaking style parameters, a target speaking style is generated.
[0010] Secondly, this disclosure provides a speech style generation apparatus, comprising:
[0011] The determination module is used to fit a target style feature attribute based on multiple style feature attributes and determine the fitting coefficient of each style feature attribute; determine a target style feature vector based on the fitting coefficient of each style feature attribute and multiple style feature vectors, wherein the multiple style feature vectors correspond one-to-one with the multiple style feature attributes; input the target style feature vector into a speech style model and output target speech style parameters, wherein the speech style model is obtained based on the framework of training a speech style model based on the multiple style feature vectors;
[0012] The generation module is used to generate the target speaking style based on the target speaking style parameters.
[0013] Thirdly, this disclosure also provides an electronic device, including: a processor, the processor being configured to execute a computer program stored in a memory, the computer program being executed by the processor to implement the steps of the speaking style generation method described in any one of the first aspects.
[0014] Fourthly, this disclosure also provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the speech style generation method as described in any one of the first aspects.
[0015] In the technical solution of this disclosure embodiment, the fitting coefficients of each style feature attribute are determined by fitting a target style feature attribute based on multiple style feature attributes; a target style feature vector is determined based on the fitting coefficients of each style feature attribute and multiple style feature vectors, with each style feature vector corresponding to a different style feature attribute; the target style feature vector is input into a speech style model, and target speech style parameters are output, where the speech style model is obtained by training a speech style model based on multiple style feature vectors; and a target speech style is generated based on the target speech style parameters. Thus, the target style feature vector can be fitted with multiple style feature vectors. Since the speech model is trained based on multiple style feature vectors, inputting the target style feature vector fitted by multiple style feature vectors into the speech model can directly obtain a new speech style without retraining the speech style model, enabling rapid transfer of speech styles and improving the efficiency of speech style generation. Attached Figure Description
[0016] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.
[0017] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1A Schematic diagram of a three-dimensional virtual digital human provided in some embodiments of this disclosure;
[0019] Figure 1B Schematic diagram of a three-dimensional virtual digital human provided in some embodiments of this disclosure;
[0020] Figure 1C This is a schematic diagram illustrating the principle of generating new speaking styles in some embodiments of this disclosure;
[0021] Figure 2 This is a schematic diagram of a human-computer interaction scenario according to some embodiments of the present disclosure;
[0022] Figure 3 This is a flowchart illustrating a speech style generation method provided in some embodiments of this disclosure;
[0023] Figure 4 This is a schematic diagram illustrating the division of facial topological data regions according to some embodiments of this disclosure;
[0024] Figure 5 This is a flowchart illustrating a speech style generation method provided in some embodiments of this disclosure;
[0025] Figure 6 This is a flowchart illustrating a speech style generation method provided in some embodiments of this disclosure;
[0026] Figure 7 This is a flowchart illustrating a speech style generation method provided in some embodiments of this disclosure;
[0027] Figure 8 A schematic diagram illustrating the framework of a speaking style model provided in some embodiments of this disclosure;
[0028] Figure 9 This is a flowchart illustrating a speech style generation method provided in some embodiments of this disclosure;
[0029] Figure 10 A schematic diagram illustrating the framework of the speech style generation model provided in some embodiments of this disclosure;
[0030] Figure 11 This is a flowchart illustrating a speech style generation method provided in some embodiments of this disclosure;
[0031] Figure 12AA schematic diagram illustrating the framework of the speech style generation model provided in some embodiments of this disclosure;
[0032] Figure 12B A schematic diagram illustrating the framework of the speech style generation model provided in some embodiments of this disclosure;
[0033] Figure 13 This is a flowchart illustrating a speech style generation method provided in some embodiments of this disclosure;
[0034] Figure 14 This is a schematic diagram of the structure of a speech style generation apparatus provided in some embodiments of the present disclosure;
[0035] Figure 15 This is a schematic diagram of the structure of a speech style generation apparatus provided in some embodiments of the present disclosure;
[0036] Figure 16 This is a schematic diagram of the structure of a speech style generation apparatus provided in some embodiments of this disclosure. Detailed Implementation
[0037] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.
[0038] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.
[0039] The terms “first” and “second” in this disclosure are used to distinguish different objects, not to describe a specific order of objects. For example, first predicted score and second predicted score are used to distinguish different predicted scores, not to describe a specific order of predicted scores.
[0040] With the rapid development of intelligent technologies and the increasing popularity of smart terminals, voice multimodal interaction has become an increasingly important method. Traditional voice interaction involves hearing the voice but not seeing the user. The user issues voice commands to the smart device, which receives the commands, generates response information, and plays the corresponding voice response. The user can then access this response, thus enabling interaction with the smart device. In the evolution of human-computer interaction, voice-driven three-dimensional virtual digital humans have emerged. Smart devices include displays that can show these virtual digital humans, such as… Figure 1A As shown, while the smart device plays the voice response information, it simultaneously displays the facial expressions and lip movements of the three-dimensional virtual digital human as it speaks, such as... Figure 1BAs shown, users can hear the voice of the 3D virtual digital human and see its facial expressions while speaking, giving them the feeling of conversing with a person.
[0041] Typically, people exhibit different states when speaking. For example, some people speak with precise lip movements and expressive facial expressions, while others speak with smaller mouth movements and a serious expression. In other words, different people have different speaking styles. Therefore, it's possible to design 3D virtual digital avatars with different speaking styles. These avatars will have different lip movements and facial expressions, allowing users to converse with them and thus enhancing the user experience. Each time a new speaking style for a 3D virtual avatar is designed, corresponding training samples must first be obtained. The speaking style model is then retrained based on these samples, enabling the retrained model to generate new speaking style parameters. These parameters then drive the basic speaking style, such as... Figure 1C As shown, new speaking styles can be generated. However, since retraining the speaking style model requires a significant amount of time to collect training samples and process large amounts of data, generating each new speaking style takes considerable time, resulting in relatively low efficiency in speaking style generation.
[0042] To address the aforementioned issues, this disclosure involves fitting a target style feature attribute to multiple style feature attributes to determine the fitting coefficients for each style feature attribute; determining a target style feature vector based on the fitting coefficients and multiple style feature vectors, with each style feature vector corresponding one-to-one with a style feature attribute; inputting the target style feature vector into a speech style model to output target speech style parameters, where the speech style model is obtained using a framework trained with multiple style feature vectors; and generating a target speech style based on the target speech style parameters. In this way, the target style feature vector can be fitted with multiple style feature vectors. Since the speech model is trained with multiple style feature vectors, inputting the target style feature vector fitted with multiple style feature vectors into the speech model directly yields a new speech style without needing to retrain the speech style model, enabling rapid transfer of speech styles and improving the efficiency of speech style generation.
[0043] Figure 2 These are schematic diagrams illustrating human-computer interaction scenarios provided in some embodiments of this disclosure. For example... Figure 2As shown, in a voice interaction scenario between a user and a smart home, smart devices can include a smart refrigerator 110, a smart washing machine 120, and a smart display device 130, etc. When a user wants to control a smart device, they need to issue a voice command. Upon receiving the voice command, the smart device needs to perform semantic understanding to determine the corresponding semantic understanding result. Based on the semantic understanding result, it executes the corresponding control command to meet the user's needs. All smart devices in this scenario include a display screen, which can be a touchscreen or a non-touchscreen. For terminal devices with touchscreens, users can interact with the terminal device through gestures, fingers, or touch tools (e.g., a stylus). For non-touchscreen terminal devices, interaction can be achieved through external devices (e.g., a mouse or keyboard). The display screen can show a 3D virtual human, allowing users to see the 3D virtual human and its facial expressions while speaking, thus enabling dialogue interaction with the 3D virtual human.
[0044] The speech style generation method provided in this disclosure can be implemented based on a computer device, or a functional module or entity within a computer device. The computer device can be a personal computer (PC), server, mobile phone, tablet computer, laptop computer, mainframe computer, etc., and this disclosure does not specifically limit its application.
[0045] To illustrate the speech style generation scheme in more detail, the following will use an illustrative approach. Figure 3 To explain, it is understandable that Figure 3 The steps involved may include more or fewer steps in actual implementation, and the order of these steps may also be different, depending on whether the speaking style generation method provided in the embodiments of this application can be implemented.
[0046] Figure 3 This is a flowchart illustrating a speech style generation method provided in some embodiments of this disclosure, such as... Figure 3 As shown, the method specifically includes the following steps:
[0047] S101, Fit the target style feature attribute based on multiple style feature attributes, and determine the fitting coefficient of each style feature attribute.
[0048] For example, a facial topology data sequence is collected during a user's speech over a time interval Δt. In this sequence, each frame corresponds to a dynamic face topology, which includes multiple vertices. Each vertex in the dynamic face topology corresponds to a vertex coordinate (x, y, z). When the user is not speaking, a preset static face topology is used, where each vertex has coordinates (x', y', z'). Based on the difference between the vertex coordinates of the same vertex in the dynamic and static face topologies, the vertex offset (Δx, Δy, Δz) of each vertex in each dynamic face topology can be determined, i.e., Δx = x - x', Δy = y - y', Δz = z - z'. Based on the vertex offsets (Δx, Δy, Δz) of each vertex in all dynamic face topologies corresponding to the facial topology data sequence, the average vertex offset of each vertex in the dynamic face topology can be determined.
[0049] Figure 4 This is a schematic diagram illustrating the division of facial topological data regions according to an embodiment of the present disclosure, such as... Figure 4 As shown, facial topology data can be divided into multiple regions. For example, the facial topology data can be divided into three regions, namely S1, S2, and S3. S1 represents all facial regions above the lower edge of the eyes, S2 represents the facial region from the lower edge of the eyes to the upper edge of the upper lip, and S3 represents the facial region from the upper edge of the upper lip to the chin. Based on the above embodiment, the average vertex offset of all vertices of the dynamic face topology within region S1 can be determined. average Average vertex offset of all vertices of the dynamic face topology within region S2 average Average vertex offset of all vertices in the dynamic face topology within region S3 average Style characteristic attributes can be obtained, that is In summary, one style feature attribute can be obtained for a single user, and thus, multiple style feature attributes can be obtained for multiple users.
[0050] Based on the obtained multiple style feature attributes, a new style feature attribute, namely the target style feature attribute, can be fitted. For example, the target style feature attribute can be obtained by fitting based on the following formula:
[0051]
[0052] in, For target feature attributes, For user 1's style characteristics attributes, For user 2's style characteristics attributes, Let a1 be the style feature attribute of user n, a2 be the fitting coefficient of the style feature attribute of user 1, a2 be the fitting coefficient of the style feature attribute of user 2, an be the fitting coefficient of the style feature attribute of user n, n be the number of users, and a1+a2+…+an=1.
[0053] Based on the above formula, optimization methods, such as gradient descent and Gauss-Newton method, can be used to obtain the fitting coefficients for each style feature attribute.
[0054] It should be noted that this embodiment is only used as an example of dividing facial topology data into three regions, and is not intended as a specific limitation on the division of facial topology data regions.
[0055] S102, determine the target style feature vector based on the fitting coefficients of each style feature attribute and multiple style feature vectors.
[0056] The multiple style feature vectors correspond one-to-one with the multiple style feature attributes.
[0057] For example, style feature vectors, which represent style, can be based on a classification task model, using the embedding obtained from training the classification task model as style feature vectors, or one-hot feature vectors can be directly designed as style feature vectors. For example, if three users correspond to three style feature attributes as one-hot feature vectors, then the three style feature vectors can be [1; 0; 0], [0; 1; 0], and [0; 0; 1].
[0058] Based on the above embodiments, style feature attributes of n users with different speaking styles are obtained. Correspondingly, n style feature vectors for each user can be obtained. These n style feature attributes correspond one-to-one with the n style feature vectors, and the n style feature attributes and their corresponding style feature vectors form a basic style feature basis. By multiplying the fitting coefficients of each of the n style feature attributes by their corresponding style feature vectors, the target style feature vector can be represented in the form of a basic style feature basis, as shown in the following formula:
[0059] p=a1×F1+a2×F2+…+an×Fn(2)
[0060] Where F1 is the style feature vector of user 1, F2 is the style feature vector of user 2, Fn is the style feature vector of user n, and p is the target style feature vector.
[0061] For example, if the style feature vector is a one-hot feature vector, the target style feature vector p can be represented as:
[0062]
[0063] S103, input the target style feature vector into the speaking style model and output the target speaking style parameters.
[0064] The speaking style model is obtained based on the framework of training the speaking style model using the multiple style feature vectors.
[0065] For example, based on multiple style feature vectors in the basic style feature set, a framework for the speaking style model is trained, resulting in the trained framework, i.e., the speaking style model. Inputting the target style feature vector into the speaking style model can be understood as inputting the product of multiple style feature vectors and their respective fitting coefficients into the speaking style model; this is the same as the training samples used when training the framework of the speaking style model. Therefore, based on the speaking style model, using the target style feature vector as input, the target speaking style parameters can be directly output.
[0066] The target speaking style parameter can be the vertex offset between each vertex in the dynamic face topology and the corresponding vertex in the static face topology; or it can be the coefficient of the expression basis of the dynamic face topology; or it can be other parameters, which are not specifically limited in this disclosure.
[0067] S104, Generate the target speaking style based on the target speaking style parameters.
[0068] For example, the target speaking style parameter is the vertex offset of each vertex in the dynamic face topology and the corresponding vertex in the static face topology. Thus, based on the static face topology, the vertex of the static face topology is driven to move to the corresponding position according to the vertex offset, and the target speaking style can be obtained.
[0069] In this embodiment, the fitting coefficients of each style feature attribute are determined by fitting a target style feature attribute with multiple style feature attributes. Based on the fitting coefficients of each style feature attribute and multiple style feature vectors, a target style feature vector is determined, with each style feature vector corresponding one-to-one with a style feature attribute. The target style feature vector is input into a speech style model, which outputs target speech style parameters. The speech style model is obtained using a framework that trains the speech style model based on multiple style feature vectors. Based on the target speech style parameters, a target speech style is generated. Thus, the target style feature vector can be fitted with multiple style feature vectors. Since the speech model is trained based on multiple style feature vectors, inputting the target style feature vector fitted by multiple style feature vectors into the speech model directly yields a new speech style without needing to retrain the speech style model. This enables rapid transfer of speech styles and improves the efficiency of speech style generation.
[0070] Figure 5 This is a flowchart illustrating a speech style generation method provided in some embodiments of this disclosure. Figure 5 For example Figure 3 Based on the illustrated embodiment, the method further includes the following steps before executing S101:
[0071] S201 collects multi-frame facial topology data when multiple preset users read multiple segments of speech.
[0072] For example, users with different speaking styles are selected as preset users, and multiple audio segments are also selected. When each preset user reads each audio segment, multiple frames of facial topology data of that preset user are collected. For example, the duration of audio segment 1 is t1, and the frequency of collecting facial topology data is 30 frames / second. In this way, after preset user 1 finishes reading each audio segment 1, t1*30 frames of facial topology data can be collected.
[0073] S202, for each preset user: based on the speaking style parameters of the multi-frame facial topology data corresponding to the multi-segment speech and the division region of the facial topology data, determine the average value of the speaking style parameters of the multi-frame facial topology data in each division region.
[0074] For example, based on the above embodiment, for a preset user 1, after the preset user 1 reads out m segments of speech, t1*30*m frames of facial topology data can be collected. The vertex offsets (Δx, Δy, Δz) of each vertex of the dynamic face topology and each vertex of the static face topology in each frame of facial topology data can be used as the speech style parameters of each frame of facial topology data. Based on the vertex offsets (Δx, Δy, Δz) of each vertex of all dynamic face topologies corresponding to the t1*30*m frames of facial topology data of the preset user 1, the average vertex offset of each vertex of the dynamic face topology in the facial topology data of the preset user 1 can be determined.
[0075] Based on the regional division of facial topology data, for each region of the preset user 1, the average vertex offset of all vertices of the dynamic face topology in the facial topology data within the region can be obtained. The average value. For example, facial topology data is divided into three regions, where the average vertex offset of all vertices of the dynamic face topology data in region S1 is... The average vertex offset of all vertices in the dynamic face topology data within region S2 is: The average vertex offset of all vertices in the dynamic face topology data within region S3 is:
[0076] S203, the average value of the speaking style parameters of the multi-frame facial topology data in each divided region is spliced together in a preset order to obtain the style feature attributes of each preset user.
[0077] For example, the preset order can be according to... Figure 4 The order shown is from top to bottom, or it can be in the order shown below. Figure 4 The bottom-to-top order shown is not specifically limited in this disclosure. If the preset order is as follows... Figure 4 As shown in the top-to-bottom order, based on the above embodiment, the average vertex offset of all vertices in the dynamic face topology data corresponding to each region S1, S2, and S3 can be stitched together in that order. Thus, the style feature attributes of the preset user 1 can be obtained.
[0078]
[0079] In summary, style characteristic attributes can be obtained for the preset user 1. In this way, multiple style feature attributes can be obtained for multiple preset users.
[0080] Figure 6 This is a flowchart illustrating a speech style generation method provided in some embodiments of this disclosure. Figure 6 For example Figure 5 Based on the illustrated embodiment, the method further includes the following steps before executing S101:
[0081] S301, Collect multi-frame target facial topology data when the target user reads the multiple speech segments.
[0082] The target user is different from the multiple preset users.
[0083] For example, when it is necessary to generate a target speaking style that differs from the speaking styles of multiple preset users, multi-frame target facial topology data is collected when the target user reads multiple segments of speech corresponding to the target speaking style, and the content of the multiple segments of speech read by the target user is the same as the content of the multiple segments of speech read by the multiple preset users. For example, after the target user reads m segments of speech with a duration of t1, t1*30*m frames of target facial topology data can be obtained.
[0084] S302, based on the speaking style parameters of the multi-frame target facial topology data corresponding to the multi-segment speech and the division regions of the facial topology data, determine the average value of the speaking style parameters of the multi-frame target facial topology data in each division region.
[0085] The vertex offsets (Δx', Δy', Δz') of each vertex of the dynamic face topology and each vertex of the static face topology in each frame of the target face topology data can be used as speech style parameters for each frame of the target face topology data. Based on the vertex offsets (Δx', Δy', Δz') of each vertex of the dynamic face topology in the t1*30*m frames of the target user's target face topology data, the average vertex offset of each vertex of the dynamic face topology in the target user's target face topology data can be determined.
[0086] Based on the region division of the facial topology data mentioned above, for each region of the target user, the average vertex offset of all vertices of the dynamic face topology in the target facial topology data within the region can be obtained. The average value. For example, facial topology data is divided into three regions, where the average vertex offset of all vertices of the dynamic face topology data in the target facial topology data within region S1 is... The average vertex offset of all vertices in the dynamic face topology data of the target face within region S2 is: The average vertex offset of all vertices in the dynamic face topology data of the target face within region S3 is:
[0087] S303, the average value of the speech style parameters of the multi-frame target facial topology data in each divided region is concatenated in the preset order to obtain the target style feature attribute.
[0088] For example, based on the same preset order as in the above embodiments, the average vertex offset of all vertices of the dynamic face topology in the target face topology data is spliced together, for example, based on... Figure 4 The order shown from top to bottom allows us to obtain the target user's target style feature attributes by concatenating the average vertex offsets of all vertices in the dynamic face topology data corresponding to regions S1, S2, and S3 in that order.
[0089] It should be noted that you can first execute as follows: Figure 5 As shown in S201-S203, then execute as follows: Figure 6 As shown in S301-S303; or, you can first execute as follows: Figure 6 As shown in S301-S303, execute as follows: Figure 5 The present disclosure does not impose specific limitations on S201-S203 as shown.
[0090] Figure 7 This is a flowchart illustrating a speech style generation method provided in some embodiments of this disclosure. Figure 7 For example Figure 5 and Figure 4 Based on the illustrated embodiment, the method further includes the following steps before executing S103:
[0091] S401, Obtain the training sample set.
[0092] The training sample set includes an input sample set and an output sample set. The input samples include speech features and their corresponding multiple style feature vectors, and the output samples include the speech style parameters.
[0093] When a user reads aloud, the system can extract the inherent features of the speech information, primarily those that express the speech content. For example, it can extract speech features using Melp features, or it can use commonly used speech feature extraction models, or it can extract speech features based on a pre-designed deep network model. Based on the efficiency of speech feature extraction, after a user reads multiple speech segments, a speech feature sequence can be extracted. If the content of multiple speech segments read by multiple users is exactly the same, then the same speech feature sequence can be extracted for different users. Thus, for the same speech feature in the speech feature sequence, there are multiple style feature vectors corresponding to multiple users. A speech feature and its corresponding multiple style feature vectors can be used as input samples. Based on all the speech features in the speech feature sequence, multiple input samples can be obtained, resulting in an input sample set.
[0094] For example, while extracting each speech feature, corresponding facial topology data can be collected. Based on the vertex coordinates of each vertex of the dynamic face topology in the facial topology data, the vertex offsets of each vertex of the dynamic face topology in the facial topology data can be obtained. These vertex offsets are used as a set of speech style parameters, and each set of speech style parameters constitutes an output sample. Thus, based on multiple frames of facial topology data corresponding to the speech feature sequence, multiple output samples, i.e., the output sample set, can be obtained. The input sample set and the output sample set constitute the training sample set for training the speech style generation model.
[0095] S402 defines the framework of the speaking style model.
[0096] The framework of the speaking style model includes a linear combination unit and a network model. The linear combination unit is used to generate a linear combination style feature vector of the multiple style feature vectors and a linear combination output sample of the multiple output samples. The input sample and the output sample correspond one-to-one. The network model is used to generate the corresponding predicted output sample based on the linear combination style feature vector.
[0097] Figure 8 This is a schematic diagram of the framework of the speaking style model provided in some embodiments of this disclosure, such as... Figure 8 As shown, the framework of the speaking style model includes a linear combination unit 310 and a network model 320. The input of the linear combination unit 310 is used to receive training samples, and the output of the linear combination unit 310 is connected to the input of the network model 320. The output of the network model 320 is the output of the framework 300 of the speaking style model.
[0098] After the training samples are input into the linear combination unit 310, the training samples include input samples and output samples. The input samples include speech features and their corresponding multiple style feature vectors. The linear combination unit 310 can linearly combine these multiple style feature vectors to obtain a linearly combined style feature vector. It can also linearly combine the speech style parameters corresponding to each of the multiple style feature vectors to obtain a linearly combined output sample. The linear combination unit 310 can output speech features and their corresponding linearly combined style feature vectors, i.e., linearly combined input samples, and can also output corresponding linearly combined output samples. The linearly combined training samples are then input into the network model 320. The linearly combined training samples include linearly combined input samples and linearly combined output samples. Based on these linearly combined training samples, the network model 320 is trained.
[0099] S403, Based on the training sample set and loss function, train the framework of the speaking style model to obtain the speaking style model.
[0100] Based on the above embodiments, training samples from the training sample set are input into the framework of the speech style model. The framework of the speech style model can output predicted output samples. A loss function is used to determine the loss value of the predicted output sample and the output sample. Based on the direction of decreasing loss value, the model parameters of the speech style model framework are adjusted, thus completing one iteration of training. In this way, based on multiple iterations of training the framework of the speech style model, a well-trained framework of the speech style model can be obtained, i.e., the speech style model.
[0101] In this embodiment, a training sample set is obtained, which includes an input sample set and an output sample set. The input samples include speech features and their corresponding multiple style feature vectors, and the output samples include speech style parameters. A framework for a speech style model is defined, which includes a linear combination unit and a network model. The linear combination unit is used to generate a linear combination of style feature vectors of multiple style feature vectors and a linear combination of output samples of multiple output samples, with a one-to-one correspondence between input samples and output samples. The network model is used to generate corresponding predicted output samples based on the linear combination of style feature vectors. The framework for the speech style model is trained based on the training sample set and the loss function to obtain the speech style model. Thus, the speech style model is essentially obtained by training the network model based on the linear combination of style feature vectors of multiple style feature vectors, which can improve the diversity of training samples for the network model and enhance the versatility of the speech style model.
[0102] Figure 9 This is a flowchart illustrating a speech style generation method provided in some embodiments of this disclosure. Figure 9 for Figure 7Based on the illustrated embodiment, a specific description of a possible implementation of S403 is as follows:
[0103] S501, the training sample set is input into the linear combination unit, the linear combination style feature vector is generated based on the multiple style feature vectors and their respective weight values, and the linear combination output sample is generated based on the weight values of the multiple style feature vectors and the multiple output samples.
[0104] The sum of the weights of the multiple style feature vectors is 1.
[0105] For example, after training samples are input into a linear combination unit, multiple style feature vectors can be assigned weight values based on the linear combination unit, and the sum of the weight values of each style feature vector is 1. Adding the products of each style feature vector and its corresponding weight value yields a linearly combined style feature vector. Each style feature vector corresponds to an output sample. Adding the products of the weight values of each style feature vector and their corresponding output samples yields a linearly combined output sample. Thus, based on different weight values, different linearly combined style feature vectors and different linearly combined output samples can be obtained. Based on multiple speech features and their corresponding linearly combined style feature vectors, a linearly combined input sample set can be obtained; based on the output samples corresponding to each of the multiple speech features, a linearly combined output sample set can be obtained.
[0106] S502, Train the network model according to the loss function and the linear combination training sample set to obtain the speaking style model.
[0107] The linear combination training sample set includes a linear combination input sample set and a linear combination output sample set. The linear combination input sample includes the speech features and their corresponding linear combination style feature vectors.
[0108] For example, the linear combination training sample set includes a linear combination input sample set and a linear combination output sample set. The linear combination training samples are input into the network model. Based on the network model and the linear combination input samples, predicted output samples can be obtained. The model parameters of the network model are adjusted based on the direction of decrease in the loss value of the loss function, thus completing one iteration of network model training. In this way, multiple iterations of training based on the network model can yield the framework of a trained speaking style model, i.e., the speaking style model.
[0109] In this embodiment, by inputting the training sample set into the linear combination unit, a linear combination style feature vector is generated based on multiple style feature vectors and their respective weight values. A linear combination output sample is generated based on the weight values of the multiple style feature vectors and multiple output samples. The sum of the weight values of the multiple style feature vectors is 1. The network model is trained according to the loss function and the linear combination training sample set to obtain the speech style model. The linear combination training sample set includes a linear combination input sample set and a linear combination output sample set. The linear combination input sample includes speech features and their corresponding linear combination style feature vectors. The linear combination training samples can be used as training samples for the network model, which can increase the number and diversity of training samples for the network model and improve the versatility and accuracy of the speech style model.
[0110] In some embodiments of this disclosure, Figure 10 A schematic diagram of the framework of another speech style generation model provided in this disclosure embodiment is shown below. Figure 10 As shown, in Figure 8 Based on the illustrated embodiment, the speech style model framework further includes a scaling unit 330. The input of the scaling unit 330 receives training samples, and its output is connected to the input of the linear combination unit 310. The scaling unit 330 scales multiple style feature vectors and multiple output samples based on a randomly generated scaling factor, resulting in multiple scaled style feature vectors and multiple scaled output samples. It also outputs scaled training samples, which include multiple scaled style feature vectors and their corresponding scaled training samples. The scaling factor can be between 0.5 and 2, accurate to one decimal place.
[0111] The scaled training samples are input to the linear combination unit 310. Based on this unit, multiple scaled style feature vectors can be linearly combined to obtain a linearly combined style feature vector. Furthermore, the scaled output samples corresponding to each of the multiple scaled style feature vectors can be linearly combined to obtain a linearly combined output sample. The linear combination unit 310 can output speech features and their corresponding linearly combined style feature vectors (i.e., linearly combined input samples), and also output corresponding linearly combined output samples. The linearly combined training samples are input to the network model 320. These training samples include both linearly combined input samples and linearly combined output samples. The network model 320 is then trained based on these training samples.
[0112] Figure 11 This is a flowchart illustrating a speech style generation method provided in some embodiments of this disclosure. Figure 11 for Figure 7 Based on the illustrated embodiment, another possible implementation of S403 is described in detail below:
[0113] S5011, the training sample set is input to the scaling unit, and multiple scaled style feature vectors are generated based on the scaling factor and the multiple style feature vectors, and multiple scaled output samples are generated based on the scaling factor and the multiple output samples.
[0114] For example, after training samples are input into a scaling unit, the scaling unit can scale multiple style feature vectors separately using random scaling factors, resulting in multiple scaled style feature vectors. Each style feature vector corresponds to an output sample. By scaling the corresponding output samples based on the scaling factors of the multiple style feature vectors, multiple scaled output samples can be obtained. Thus, based on multiple speech features and their corresponding scaled style feature vectors, a scaled input sample set can be obtained, and based on the scaled output samples corresponding to the multiple speech features, a scaled output sample set can be obtained.
[0115] S5012, the plurality of scaling style feature vectors and the plurality of scaling output samples are input to the linear combination unit, the linear combination style feature vector is generated based on the plurality of scaling style feature vectors and their respective weight values, and the linear combination output sample is generated based on the respective weight values of the plurality of scaling style feature vectors and the plurality of scaling output samples.
[0116] The sum of the weights of the multiple scaling style feature vectors is 1.
[0117] For example, the scaled training sample set includes a scaled input sample set and a scaled output sample set. The scaled training sample set is input to a linear combination unit. Based on the linear combination unit, multiple scaled style feature vectors can be assigned weight values, and the sum of the weight values of each scaled style feature vector is 1. Adding the products of each scaled style feature vector and its corresponding weight value yields a linear combination style feature vector. Each scaled style feature vector corresponds to a scaled output sample. Adding the products of the weight values of each scaled style feature vector and its corresponding scaled output sample yields a linear combination output sample. Thus, based on different weight values, different linear combination style feature vectors and different linear combination output samples can be obtained. Based on multiple speech features and their corresponding linear combination style feature vectors, a linear combination input sample set can be obtained; based on the scaled output samples corresponding to each of the multiple speech features, a linear combination output sample set can be obtained.
[0118] S502, Train the network model according to the loss function and the linear combination training sample set to obtain the speaking style model.
[0119] The linear combination training sample set includes a linear combination input sample set and a linear combination output sample set. The linear combination input sample includes the speech features and their corresponding linear combination style feature vectors.
[0120] For example, the linear combination training sample set includes a linear combination input sample set and a linear combination output sample set. The linear combination training samples are input into the network model. Based on the network model and the linear combination input samples, predicted output samples can be obtained. The model parameters of the network model are adjusted based on the direction of decrease in the loss value of the loss function, thus completing one iteration of network model training. In this way, multiple iterations of training based on the network model can yield the framework of a trained speaking style model, i.e., the speaking style model.
[0121] In this embodiment, the framework of the speech style model also includes a scaling unit. The training sample set is input to the scaling unit, and multiple scaled style feature vectors are generated based on a scaling factor and multiple style feature vectors. Multiple scaled output samples are also generated based on a scaling factor and multiple output samples. The multiple scaled style feature vectors and multiple scaled output samples are input to a linear combination unit, and a linear combination style feature vector is generated based on the multiple scaled style feature vectors and their respective weights. A linear combination output sample is generated based on the respective weights of the multiple scaled style feature vectors and multiple scaled output samples. The sum of the weights of the multiple scaled style feature vectors is 1. The network model is trained according to the loss function and the linear combination training sample set to obtain the speech style model. The linear combination training sample set includes a linear combination input sample set and a linear combination output sample set. The linear combination input samples include speech features and their corresponding linear combination style feature vectors. Thus, using the scaled multiple style feature vectors as training samples for the network model increases the number and diversity of training samples, thereby improving the versatility and accuracy of the speech style model.
[0122] In some embodiments of this disclosure, Figure 12A This is a schematic diagram illustrating the framework of a speech style generation model provided in some embodiments of this disclosure. Figure 12B This is a schematic diagram illustrating the framework of a speech style generation model provided in some embodiments of this disclosure. Figure 12A for Figure 8 Based on the illustrated embodiment, Figure 12B for Figure 10 Based on the illustrated embodiment, the network model 320 includes a first-level network model 321, a second-level network model 322, and an overlay unit 323. The outputs of both the first-level network model 321 and the second-level network model 322 are connected to the input of the overlay unit 323. The output of the overlay unit 323 is used to output predicted output samples. The loss function includes a first loss function and a second loss function.
[0123] Linear combination training samples are input into the first-level network model 321 and the second-level network model 322, respectively. The first-level network model 321 outputs first-level predicted output samples, and the second-level network model 322 outputs second-level predicted output samples. These first-level and second-level predicted output samples are then input into the stacking unit 323, which stacks them to obtain the predicted output samples. The first-level network model 321 may include convolutional networks and fully connected networks, and its function is to extract the single-frame correspondence between speech and facial topological structure data. The second-level network model 322 may be a sequence-to-sequence (seq2seq) network model, such as a Long Short-Term Memory (LSTM) network model, a Gate Recurrent Unit (GRU) network model, or a Transformer network model, and its function is to enhance the continuity of speech features and facial expressions, as well as the subtlety of speech style.
[0124] For example, the loss function L = b1*L1 + b2*L2, where L1 is the first loss function, used to determine the loss value of the first-level predicted output sample and the linear combination output sample, L2 is the second loss function, used to determine the loss value of the second-level predicted output sample and the linear combination output sample, b1 is the weight of the first loss function, and b2 is the weight of the second loss function. b1 and b2 are adjustable. By setting b2 close to 0, the first-level network model 321 can be trained, and by setting b1 close to 0, the second-level network model 322 can be trained. In this way, the first-level network model and the second-level network model can be trained separately in stages, which can improve the convergence speed of network model training, save network model training time, and thus improve the efficiency of speech style generation.
[0125] Figure 13 This is a flowchart illustrating a speech style generation method provided in some embodiments of this disclosure. Figure 13 for Figure 9 or Figure 11 Based on the illustrated embodiment, a specific description of a possible implementation of S502 is as follows:
[0126] S5021, Based on the linear combination training sample set and the first loss function, train the first-level network model to obtain the middle speaking style model.
[0127] The intermediate speaking style model includes the second-level network model and the trained first-level network model.
[0128] For example, based on the above embodiments, in the first stage, the weight b2 of the second loss function is set to approach 0. The loss function of the current network model can be understood as the first loss function. The linear combination training samples are input into the first-level network model and the second-level network model respectively. Based on the predicted output samples of the overlay unit, the first loss function, and the corresponding linear combination output samples, the first loss value can be obtained. The model parameters of the first-level network model are adjusted according to the direction of decreasing the first loss value until the first loss value converges, thus obtaining the trained first-level network model. The framework of the speaking style model trained in the first stage is the intermediate speaking style model.
[0129] S5022, Fix the model parameters of the trained first-level network model.
[0130] For example, after training the first-level network model, the second stage begins by fixing the model parameters of the trained first-level network model.
[0131] S5023, based on the linear combination training sample set and the second loss function, train the second-level network model in the intermediate speaking style model to obtain the speaking style model.
[0132] The speaking style model includes the trained first-level network and the trained second-level network.
[0133] Secondly, the weight b1 of the first loss function is set to approach 0. The loss function of the current network model can be understood as the second loss function. The linear combination training samples are input into the second-level network model and the trained first-level network model. Based on the predicted output samples of the stacking unit, the second loss function, and the corresponding linear combination output samples, the second loss value can be obtained. The model parameters of the second-level network model are adjusted according to the direction of decreasing the second loss value until the second loss value converges, resulting in the trained second-level network model. The framework of the first-stage trained speaking style model is the speaking style model.
[0134] In this embodiment, the network model includes a first-level network model, a second-level network model, and an overlay unit. The outputs of both the first-level and second-level network models are connected to the input of the overlay unit, which outputs the predicted output samples. The loss functions include a first loss function and a second loss function. The first-level network model is trained using a linear combination of the training sample set and the first loss function to obtain an intermediate speaking style model, which includes the second-level network model and the trained first-level network model. The model parameters of the trained first-level network model are fixed. The second-level network model in the intermediate speaking style model is trained using a linear combination of the training sample set and the second loss function to obtain a speaking style model, which includes the trained first-level network and the trained second-level network. In this way, the network model can be trained in stages, which can improve the convergence speed of the network model, i.e., shorten the training time of the network model, thereby improving the efficiency of speaking style generation.
[0135] Figure 14 This is a schematic diagram of the structure of a speech style generation apparatus provided in some embodiments of this disclosure. The apparatus is configured in a computer device and can implement the speech style generation method described in any embodiment of this application. The apparatus specifically includes the following:
[0136] The determination module 410 is used to fit a target style feature attribute based on multiple style feature attributes and determine the fitting coefficient of each style feature attribute; determine a target style feature vector based on the fitting coefficient of each style feature attribute and multiple style feature vectors, wherein the multiple style feature vectors correspond one-to-one with the multiple style feature attributes; input the target style feature vector into the speaking style model and output the target speaking style parameters, wherein the speaking style model is obtained based on the framework of training the speaking style model based on the multiple style feature vectors.
[0137] The generation module 420 is used to generate a target speaking style based on the target speaking style parameters.
[0138] As an optional implementation of this disclosure, Figure 15 This is a schematic diagram of the structure of a speech style generation apparatus provided in some embodiments of this disclosure. Figure 15 for Figure 14 Based on the illustrated embodiment, the speech style generation device further includes:
[0139] The acquisition module 430 is used to acquire multi-frame facial topology data when multiple preset users read out multiple segments of speech.
[0140] The determining module 410 is further configured to, for each preset user, determine the average value of the speech style parameters of the multi-frame facial topology data in each divided region based on the speech style parameters of the multi-segment facial topology data corresponding to the multi-segment speech and the division region of the facial topology data; and concatenate the average values of the speech style parameters of the multi-frame facial topology data in each divided region in a preset order to obtain the style feature attributes of each preset user.
[0141] As an optional implementation of this disclosure, based on the above embodiments, the acquisition module 430 is further used to acquire multi-frame target facial topology data when the target user reads the multiple speech segments, wherein the target user and the multiple preset users are different users.
[0142] The determining module 410 is further configured to determine the average value of the speech style parameters of the multi-frame target facial topology data in each of the division regions based on the speech style parameters of each of the multi-segment speech corresponding to the multi-frame target facial topology data and the division regions of the facial topology data; and to concatenate the average values of the speech style parameters of the multi-frame target facial topology data in each of the division regions according to the preset order to obtain the target style feature attribute.
[0143] As an optional implementation of this disclosure, Figure 16 This is a schematic diagram of the structure of a speech style generation apparatus provided in some embodiments of this disclosure. Figure 16 for Figure 15 Based on the illustrated embodiment, the speech style generation device further includes:
[0144] The acquisition module 440 is used to acquire a training sample set, which includes an input sample set and an output sample set. The input samples include speech features and their corresponding multiple style feature vectors, and the output samples include the speech style parameters.
[0145] The framework definition module 450 is used to define the framework of the speaking style model. The framework of the speaking style model includes a linear combination unit and a network model. The linear combination unit is used to generate a linear combination style feature vector of the multiple style feature vectors and a linear combination output sample of the multiple output samples. The input sample and the output sample correspond one-to-one. The network model is used to generate the corresponding predicted output sample based on the linear combination style feature vector.
[0146] Training module 460 is used to train the framework of the speaking style model based on the training sample set and the loss function, so as to obtain the speaking style model.
[0147] As an optional implementation of this disclosure, based on the above embodiments, the training module 460 is further configured to input the training sample set into the linear combination unit, generate the linear combination style feature vector based on the plurality of style feature vectors and their respective weight values, generate the linear combination output sample based on the respective weight values of the plurality of style feature vectors and the plurality of output samples, wherein the sum of the respective weight values of the plurality of style feature vectors is 1; train the network model according to the loss function and the linear combination training sample set to obtain the speech style model, wherein the linear combination training sample set includes the linear combination input sample set and the linear combination output sample set, and the linear combination input sample includes the speech features and their corresponding linear combination style feature vectors.
[0148] As an optional implementation of this disclosure, based on the above embodiments, the framework of the speaking style model further includes a scaling unit.
[0149] The training module 460 is further configured to input the training sample set into the scaling unit, generate multiple scaled style feature vectors based on the scaling factor and the multiple style feature vectors, and generate multiple scaled output samples based on the scaling factor and the multiple output samples; input the multiple scaled style feature vectors and the multiple scaled output samples into the linear combination unit, generate the linear combination style feature vector based on the multiple scaled style feature vectors and their respective weight values, and generate the linear combination output samples based on the respective weight values of the multiple scaled style feature vectors and the multiple scaled output samples, wherein the sum of the respective weight values of the multiple scaled style feature vectors is 1; train the network model according to the loss function and the linear combination training sample set to obtain the speech style model, wherein the linear combination training sample set includes a linear combination input sample set and a linear combination output sample set, and the linear combination input sample includes the speech features and their corresponding linear combination style feature vectors.
[0150] As an optional implementation of this disclosure, based on the above embodiments, the network model includes a first-level network model, a second-level network model, and an overlay unit. The output terminals of the first-level network model and the second-level network model are both connected to the input terminal of the overlay unit. The output terminal of the overlay unit is used to output the predicted output samples. The loss function includes a first loss function and a second loss function.
[0151] The training module 460 is further configured to train the first-level network model according to the linear combination training sample set and the first loss function to obtain an intermediate speaking style model, wherein the intermediate speaking style model includes the second-level network model and the trained first-level network model; fix the model parameters of the trained first-level network model; and train the second-level network model in the intermediate speaking style model according to the linear combination training sample set and the second loss function to obtain the speaking style model, wherein the speaking style model includes the trained first-level network and the trained second-level network.
[0152] The speaking style generation apparatus provided in this disclosure can execute the speaking style generation method provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects of executing the method.
[0153] This disclosure provides an electronic device, including: a processor, the processor being configured to execute a computer program stored in a memory, the computer program being executed by the processor to implement the steps of any method embodiment of this disclosure.
[0154] This disclosure also provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of any method embodiment of this disclosure.
Claims
1. A speaking style generation method characterized by, The method comprises the following steps: fitting target style feature attributes based on a plurality of style feature attributes, determining the fitting coefficients of each style feature attribute; the plurality of style feature attributes are style feature attributes of a plurality of preset users; determining a target style feature vector according to the fitting coefficients of each style feature attribute and a plurality of style feature vectors, the plurality of style feature vectors correspond one-to-one to the plurality of style feature attributes; inputting the target style feature vector into a speaking style model to output a target speaking style parameter, the speaking style model being obtained based on a product of the plurality of style feature vectors and the respective fitting coefficients; generating a target speaking style based on the target speaking style parameter; Before the step of fitting target style feature attributes based on a plurality of style feature attributes, determining the fitting coefficients of each style feature attribute, the method further comprises the following step: collecting a plurality of frames of face topology structure data when a plurality of preset users read a plurality of segments of speech; 2. The method of claim 1, wherein, for each preset user: determining the average value of the speaking style parameters of the plurality of frames of face topology structure data in each divided region according to the speaking style parameters and the divided regions of the plurality of frames of face topology structure data corresponding to the plurality of segments of speech; and splicing the average values of the speaking style parameters of the plurality of frames of face topology structure data in each divided region in a preset order to obtain the style feature attribute of each preset user. The method further comprises the following steps: collecting a plurality of frames of target face topology structure data when a target user reads the plurality of segments of speech, the target user being a different user from the plurality of preset users; determining the average value of the speaking style parameters of the plurality of frames of target face topology structure data in each divided region according to the speaking style parameters and the divided regions of the plurality of frames of target face topology structure data corresponding to the plurality of segments of speech; 3. The method of claim 2, wherein, splicing the average values of the speaking style parameters of the plurality of frames of target face topology structure data in each divided region in the preset order to obtain the target style feature attribute. Before the step of inputting the target style feature vector into a speaking style model to output a target speaking style parameter, the method further comprises the following steps: obtaining a training sample set, the training sample set comprising an input sample set and an output sample set, the input sample comprising a voice feature and a plurality of style feature vectors corresponding thereto, and the output sample comprising a speaking style parameter; defining a framework of the speaking style model, the framework of the speaking style model comprising a linear combination unit and a network model, the linear combination unit being configured to generate a linear combination style feature vector of the plurality of style feature vectors, and to generate a linear combination output sample of a plurality of output samples, the input sample and the output sample corresponding one-to-one; and the network model being configured to generate a corresponding predicted output sample according to the linear combination style feature vector; 4. The method of claim 3, wherein, training the framework of the speaking style model according to the training sample set and a loss function to obtain the speaking style model. The step of training the framework of the speaking style model according to the training sample set and a loss function to obtain the speaking style model comprises the following steps: inputting the training sample set into the linear combination unit, generating the linear combination style feature vector based on the plurality of style feature vectors and respective weight values of the plurality of style feature vectors, and generating the linear combination output sample based on respective weight values of the plurality of style feature vectors and the plurality of output samples, wherein a sum of the respective weight values of the plurality of style feature vectors is 1; training the network model according to the loss function and the linear combination training sample set to obtain the speaking style model, wherein the linear combination training sample set includes a linear combination input sample set and a linear combination output sample set, and a linear combination input sample includes the speech feature and a corresponding linear combination style feature vector.
5. The method of claim 3, wherein, The framework of the speaking style model further includes a scaling unit. The framework of training the speaking style model according to the training sample set and the loss function to obtain the speaking style model includes: inputting the training sample set into the scaling unit, generating a plurality of scaled style feature vectors based on a scaling factor and the plurality of style feature vectors, and generating a plurality of scaled output samples based on the scaling factor and the plurality of output samples; inputting the plurality of scaled style feature vectors and the plurality of scaled output samples into the linear combination unit, generating the linear combination style feature vector based on the plurality of scaled style feature vectors and respective weight values of the plurality of scaled style feature vectors, and generating the linear combination output sample based on respective weight values of the plurality of scaled style feature vectors and the plurality of scaled output samples, wherein a sum of the respective weight values of the plurality of scaled style feature vectors is 1; training the network model according to the loss function and the linear combination training sample set to obtain the speaking style model, wherein the linear combination training sample set includes a linear combination input sample set and a linear combination output sample set, and a linear combination input sample includes the speech feature and a corresponding linear combination style feature vector.
6. The method according to claim 4 or 5, characterized in that, The network model includes a primary network model, a secondary network model, and a superposition unit, an output end of the primary network model and an output end of the secondary network model are both connected to an input end of the superposition unit, and an output end of the superposition unit is used to output the predicted output sample; the loss function includes a first loss function and a second loss function; The framework of training the network model according to the loss function and the linear combination training sample set to obtain the speaking style model includes: training the primary network model according to the linear combination training sample set and the first loss function to obtain an intermediate speaking style model, wherein the intermediate speaking style model includes the secondary network model and the trained primary network model; fixing model parameters of the trained primary network model; training the secondary network model in the intermediate speaking style model according to the linear combination training sample set and the second loss function to obtain the speaking style model, wherein the speaking style model includes the trained primary network and the trained secondary network.
7. A speaking style generation apparatus characterized by comprising: The determining module is configured to fit target style feature attributes based on a plurality of style feature attributes, determine fitting coefficients of the style feature attributes, and determine a target style feature vector based on the fitting coefficients of the style feature attributes and a plurality of style feature vectors corresponding to the plurality of style feature attributes; input the target style feature vector into a speaking style model, and output target speaking style parameters, wherein the speaking style model is obtained based on a product of the plurality of style feature vectors and the fitting coefficients. The generating module is configured to generate a target speaking style based on the target speaking style parameters. Before the determining module determines the fitting coefficients of the style feature attributes based on the plurality of style feature attributes, the determining module is further configured to collect a plurality of frames of face topology structure data when a plurality of preset users read a plurality of segments of speech; for each preset user, determine average values of speaking style parameters of the plurality of frames of face topology structure data in each divided region based on the divided regions of the face topology structure data and the speaking style parameters corresponding to the plurality of segments of speech; and splice the average values of the speaking style parameters of the plurality of frames of face topology structure data in the divided regions in a preset order to obtain a style feature attribute of each preset user.
8. An electronic device, comprising: The processor is configured to execute a computer program stored in the memory, and the computer program, when executed by the processor, implements the steps of the method of any one of claims 1-6. The program, when executed by the processor, implements the method of any one of claims 1-6.
9. A computer readable storage medium having stored thereon a computer program, characterized in that,
Citation Information
Patent Citations
Many-to-many voice conversion method and system based on speaker style feature modeling
CN111816156A
Information processing method and device and computer readable storage medium
CN114328806A