Audio synthesis method, apparatus, device, and computer storage medium
By extracting voiceprints and emotions from video data during VR viewing and combining this with the emotions of the characters and the environment, audiobooks are generated, solving the problem of seamless switching between VR viewing and audiobooks and improving the user experience.
Patent Information
- Application Number
- CN202310001688.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-03
- Publication Date
- 2025-11-25
- Estimated Expiration
- 2043-01-03
AI Technical Summary
In existing technologies, VR movie viewing and audiobook listening cannot be seamlessly switched, resulting in a choppy user experience.
By acquiring video data and novel text content during the viewing process, voiceprint and emotion extraction are performed. Combined with character voiceprint information, video emotion classification, text emotion classification, and external environmental emotions, audio synthesis is performed to generate audiobooks.
It achieves a seamless switch from VR movie viewing to audiobook listening, reducing the jarring feeling during the switch and enhancing the user's immersive experience.
Smart Images

Figure CN116229937B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The embodiment of the present application relates to the technical field of audio-visual combination, and in particular to an audio synthesis method, device, equipment and computer storage medium. BACKGROUND
[0002] Currently, VR videos can be normally watched in virtual devices, and audio books are more used in traditional devices such as mobile phones. However, the inventors of the present application found in the process of implementing the present application that the switching between VR watching and audio book reading can only be independent and parallel at present, and seamless switching between VR watching and audio book reading cannot be achieved. SUMMARY
[0003] In view of the above problems, the embodiment of the present application provides an audio synthesis method for solving the problem that VR watching cannot be seamlessly switched with audio book reading in the prior art.
[0004] According to an aspect of the embodiment of the present application, an audio synthesis method is provided, and the method comprises:
[0005] Obtaining video data when switching from a watching state to an audio book reading state and novel text content corresponding to the video data;
[0006] Performing voiceprint and emotion extraction processing on the video data to obtain voiceprint information of each character and video emotion classification of each character;
[0007] Performing dialogue and emotion extraction processing on the novel text content to obtain dialogue text of each character and text emotion classification of each character;
[0008] According to the voiceprint information of each character, the video emotion classification of each character, the dialogue text of each character, the text emotion classification of each character and the external environment emotion classification, performing audio synthesis processing on the dialogue text of each character to obtain audio book audio corresponding to the novel text content.
[0009] In an optional manner, the obtaining of the video data when switching from the watching state to the audio book reading state and the novel text content corresponding to the video data comprises: obtaining the video data in the watching state and the novel text in the audio book reading state; inputting the video data into a preset key dialogue text discrimination model to obtain target key dialogue text; and matching the target key dialogue text with novel dialogue in the novel text to obtain novel text content corresponding to the video data.
[0010] In an optional manner, the audio synthesis processing of the dialogue text of each character according to the voiceprint information of each character, the video emotion classification of each character, the text emotion classification of each character, and the external environment emotion classification, to obtain the audiobook audio corresponding to the novel text content, comprises: determining the target voiceprint information corresponding to each character in the novel text content according to the voiceprint information of each character and each character in the novel text content; inputting the video emotion classification of each character, the text emotion classification of each character, and the external environment emotion classification into a preset emotion classification fusion model to obtain an overall emotion fusion classification; and synthesizing the audiobook audio corresponding to the novel text content according to the target voiceprint information and the overall emotion fusion classification.
[0011] In an optional manner, before the video emotion classification of each character, the text emotion classification of each character, and the external environment emotion classification are input into a preset emotion classification fusion model to obtain an overall emotion fusion classification, the method comprises: collecting environment parameters corresponding to each external environment factor; calculating the emotion score of each external environment factor according to each environment parameter and a corresponding segmentation function; and determining the external environment emotion classification according to each emotion score.
[0012] In an optional manner, the audio synthesis processing of the dialogue text of each character according to the voiceprint information of each character, the video emotion classification of each character, the text emotion classification of each character, and the external environment emotion classification, to obtain the audiobook audio corresponding to the novel text content, comprises: determining whether each character in the video data is consistent with each character in the novel text content; if the characters are consistent, extracting the voiceprint information of each character in the video data as the voiceprint information of each character in the novel text content in the audio synthesis processing; if the characters are inconsistent, determining a similar character according to the character information of a target character in the novel text content, and taking the voiceprint information of the similar character as the voiceprint information of the target character in the novel text content in the audio synthesis processing; the target character is a character in the novel text content that is inconsistent with the character in the video data.
[0013] In an optional mode, the audio synthesis processing of the dialogue text of each character according to the voiceprint information of each character, the video emotion classification of each character, the text emotion classification of each character and the external environment emotion classification comprises: determining whether the video emotion classification of each character is consistent with the text emotion classification of each character; if the emotion classifications are consistent, the emotion of each character in the audio synthesis processing of the novel text content is obtained according to the video emotion classification of each character and the external environment emotion classification; if the emotion classifications are inconsistent, the emotion of each character in the audio synthesis processing of the novel text content is obtained according to the text emotion classification of each character and the external environment emotion classification.
[0014] In an optional mode, before the audio synthesis processing of the dialogue text of each character according to the voiceprint information of each character, the video emotion classification of each character, the text emotion classification of each character and the external environment emotion classification, the method further comprises: obtaining the voice content of each character from the video data; determining whether the voice content of each character in the video data is consistent with the dialogue text of each character in the novel text content; if consistent, the current audio of each character in the video data is extracted, and the current audio of each character is taken as the audiobook audio corresponding to the novel text content; if inconsistent, the audio synthesis processing of the dialogue text of each character is performed according to the voiceprint information of each character, the video emotion classification of each character, the text emotion classification of each character and the external environment emotion classification, and the audiobook audio corresponding to the novel text content is obtained.
[0015] According to another aspect of the embodiment of the present application, an audio synthesis device is provided, comprising:
[0016] The acquisition module is configured to acquire video data when switching from a viewing state to an audiobook state and novel text content corresponding to the video data;
[0017] The first processing module is configured to perform voiceprint and emotion extraction processing on the video data to obtain voiceprint information of each character and video emotion classification of each character;
[0018] The second processing module is configured to perform dialogue and emotion extraction processing on the novel text content to obtain dialogue text of each character and text emotion classification of each character;
[0019] The synthesis module is configured to perform audio synthesis processing of the dialogue text of each character according to the voiceprint information of each character, the video emotion classification of each character, the dialogue text of each character, the text emotion classification of each character and external environment emotion classification, and obtain audiobook audio corresponding to the novel text content.
[0020] According to another aspect of the embodiments of the present application, an audio synthesis device is provided, comprising:
[0021] a processor, a memory, a communication interface and a communication bus, the processor, the memory and the communication interface accomplish communication with each other through the communication bus;
[0022] The memory is used to store at least one executable instruction, and the executable instruction makes the processor execute the steps of the above-mentioned audio synthesis method.
[0023] According to still another aspect of the embodiments of the present application, a computer readable storage medium is provided, and the storage medium stores at least one executable instruction, and the executable instruction makes the processor execute the steps of the above-mentioned audio synthesis method.
[0024] The embodiments of the present application can find the corresponding novel when watching the video according to the watching video, and acquire the video data and the novel text content corresponding to the video data when switching from the watching state to the audiobook state; the voiceprint and emotion extraction processing is performed on the video data to acquire the voiceprint information of each role and the video emotion classification of each role; the dialogue and emotion extraction processing is performed on the novel text content to acquire the dialogue text of each role and the text emotion classification of each role; the audio synthesis processing is performed on the dialogue text of each role according to the voiceprint information of each role, the video emotion classification of each role, the dialogue text of each role, the text emotion classification of each role and the external environment emotion classification, and the audiobook audio corresponding to the novel text content is acquired, so that the user can be more smooth when switching from the watching state to the audiobook state, the sense of discomfort generated when switching states is reduced, the user can also have the experience in the watching state when in the audiobook state, and the immersion experience of the user is increased.
[0025] The above description is only a summary of the technical solutions of the embodiments of the present application, in order to more clearly understand the technical means of the embodiments of the present application, the embodiments of the present application can be implemented according to the content of the specification, and in order to make the above and other purposes, characteristics and advantages of the embodiments of the present application more obvious and easy to understand, the specific embodiments of the present application are described below. BRIEF DESCRIPTION OF DRAWINGS
[0026] The accompanying drawings are only used to show the embodiments, and are not considered as limitations of the present application. Moreover, the same reference signs are used to represent the same parts throughout the drawings. In the drawings:
[0027] Figure 1 A flowchart of an audio synthesis method provided by the embodiments of the present application is shown;
[0028] Figure 2 A schematic diagram of a segmented function corresponding to a temperature factor in the audio synthesis method provided by the embodiments of the present application is shown;
[0029] Figure 3 A schematic diagram of emotion fusion in the audio synthesis method provided by the embodiment of the present application is shown.
[0030] Figure 4 A structural schematic diagram of the audio synthesis device provided by the embodiment of the present application is shown.
[0031] Figure 5 A structural schematic diagram of the audio synthesis device provided by the embodiment of the present application is shown. DETAILED DESCRIPTION
[0032] Exemplary embodiments of the present application will be described herein below with reference to the accompanying drawings. Although exemplary embodiments of the present application are shown in the drawings, it should be understood that the present application can be implemented in various forms and should not be limited by the embodiments set forth herein.
[0033] With the development of virtual-real combination technology, the price of portable VR devices is also becoming more affordable. Existing devices can set a safe area, and automatically start the front camera to collect image information of the physical space when leaving the safe area to ensure the safety of the device movement. Virtual devices can normally watch VR videos and brush video dramas, while audiobooks are more commonly used with traditional devices such as mobile phones. Existing proposals indicate that mobile terminals can switch to audio or audiobooks when watching videos, and watching dramas and listening to books can be freely switched, such as setting an audio-only entry, but the current switching between VR watching dramas and mobile phone audiobooks can only be independent and parallel, and seamless connection cannot be achieved.
[0034] The existing technology cannot achieve seamless switching of audio and video in VR immersive viewing, and watching dramas and listening to books can be freely switched, such as setting an audio-only entry, but the current switching between VR watching dramas and mobile phone audiobooks can only be independent and parallel, and cannot achieve seamless switching. Therefore, in order to make the VR viewing and audiobook switching more smooth and the perception more natural during switching, the present application synthesizes the audio of the audiobook by combining external environmental factors, so that the audiobook in the VR experience has an immersive effect.
[0035] Figure 1 A flowchart of the audio synthesis method provided by the embodiment of the present application is shown. The audio synthesis method of the embodiment of the present application is applied to a computer device, which can be a server device, a desktop computer, a tablet computer, a smart terminal device, etc. In one specific embodiment of the present application, it can be a wearable device, such as a VR device, etc. The embodiment of the present application does not make specific limitations. As shown in the figure, the method comprises the following steps: Figure 1
[0036] Step 110: acquiring video data and novel text content corresponding to the video data when switching from a viewing state to an audiobook state.
[0037] Currently, there are VR devices that provide a safety area setting, within which the user can move freely, and if the user goes beyond the safety area, the external environment is imaged into the user's vision using a front camera to avoid safety accidents. In the present application, if the user goes beyond the safety area during VR viewing, the immersive audiobook mode is automatically switched, rather than simply switching to the ordinary audio mode. The immersive audiobook mode refers to a multi-person broadcasting scene, and the ordinary audio mode refers to audio of dialogues between characters in the video. The immersive audiobook mode. The novel text of the embodiment of the present application can be a drama-related audio book, a drama-related e-book, and a drama-related e-book without a corresponding drama.
[0038] In the embodiment of the present application, the video data when switching from the viewing state to the audiobook state is the video data corresponding to the currently playing video segment. It can be the current episode corresponding to the video (such as a drama) being watched by the user through the VR, and / or the current playing frame corresponding to the video being watched, and / or the dialogue scene or dialogue subtitles of the current playing frame and the next n sentences of the video being watched. In an embodiment of the present application, the video data can include video plot, character lines, video episode, video duration, audio information, target character voiceprint information, etc. For example, the video data can be the video segment data of the viewing video for 5 minutes before and after the switching from the viewing video state to the audiobook state, including the video plot, timestamp, character lines, video episode, video duration, audio information, etc. of the video segment of the viewing video for 5 minutes before and after the switching from the viewing video state to the audiobook state.
[0039] In an embodiment of the present application, the novel text content is novel text content at a corresponding position in the novel text corresponding to the video data. In an embodiment of the present application, the video data and the novel text content corresponding to the video data when switching from a viewing state to an audiobook state are obtained in the following manner: obtaining the video data in the viewing state and the novel text in the audiobook state; inputting the video data or dialogue text in the video data into a preset key dialogue text discrimination model to obtain target key dialogue text; and matching the target key dialogue text with novel dialogue in the novel text to obtain the novel text content corresponding to the video data. The key dialogue text discrimination model is obtained by training a sentence frequency algorithm according to video dialogue samples in advance. The specific training process is as follows: obtaining a video dialogue library from dialogue subtitles or audio text conversion of sample videos (such as films and television series); obtaining common dialogue text and uncommon dialogue text according to text sorting and key information sorting; obtaining video dialogue samples according to a preset proportion and threshold; marking the video dialogue samples with labels (the labels represent whether there are or not) and inputting them into the sentence frequency algorithm for training; calculating the loss of the sentence frequency algorithm according to the output result of the training and the labels; adjusting the parameters of the sentence frequency algorithm according to the loss; continuing to iterate the training until a preset iteration number or a preset accuracy threshold is reached to obtain the key dialogue text discrimination model. After obtaining the key dialogue text of the video data, the embodiment of the present application matches the target key dialogue text with novel dialogue in the novel text to obtain the novel text content corresponding to the video data. Specifically, a correspondence between the timestamp of the video data and the novel text content can be established in advance, the novel text corresponding position can be quickly located according to the timestamp, and text search and matching are performed near the corresponding position to determine the most matched text content as the novel text content corresponding to the current video data.
[0040] In another embodiment of the present application, the novel text can also be marked with a plot in advance, and the video scene in the video data is identified to match the plot marking of the novel text to determine the novel text content corresponding to the current video data.
[0041] Step 120: performing voiceprint and emotion extraction processing on the video data to obtain voiceprint information of each character and video emotion classification of each character.
[0042] In the embodiment of the present application, after obtaining the novel text content corresponding to the video data, the characters in the video data and the dialogue text of each character are determined. Specifically, the dialogue text and the character voice can be marked with a character in advance in the viewing video. Therefore, according to the dialogue text and the character voice of each character, the characters in the video data can be determined. After determining the characters in the video data, the voiceprint corresponding to each character is extracted to obtain the voiceprint information of each character in the video data. The voiceprint information of each character can be determined by the actor ID and the actor voiceprint information corresponding to each character in the video data stored in the database in advance. The actor ID and the actor voiceprint information are stored in the database in advance, and the voiceprint information of each character is obtained by searching the database according to the actor ID corresponding to each character in the video data. The actor ID and the actor voiceprint information stored in the database in advance can be the voiceprint information of each actor obtained from the Internet or the voiceprint information of the actors in the video data extracted in advance.
[0043] In the embodiment of the present application, the emotions of each character in the video data are classified to obtain the video emotion classification of each character. In one embodiment, the video emotion classification of each character can be determined according to the audio features of each character in the video data. Specifically, the audio features of each character in the video data can be input into a first emotion recognition model to obtain the video emotion classification of each character. The first emotion recognition model can recognize the emotion of each character according to the audio features such as voiceprint change features, speech rate, pause, volume, tone, etc., to obtain the video emotion classification of each character. The first emotion recognition model can be obtained by training a neural network model according to audio feature samples in advance.
[0044] In another embodiment, the video data of each character in the video data can be input into a second emotion recognition model to obtain the video emotion classification of each character. The video data of each character includes dialogue text data, speech feature data, micro-expression data, body movement data, etc. The second emotion recognition model can recognize the relationship between characters in the video data and determine the emotion classification of each character in the video data according to the relationship between characters and the video features, which can better classify the emotions of each character in the video.
[0045] The video emotion classification of each character can be positive emotion, neutral emotion and neutral emotion, or more detailed emotion classification such as stable emotion, low emotion, anxious emotion, angry emotion, worried emotion, etc. The emotion classification can be set according to the specific scene, and the embodiment of the present application does not make specific limitation.
[0046] Step 130: dialogue and emotion extraction processing is performed on the novel text content to obtain dialogue text of each character and text emotion classification of each character.
[0047] In the embodiment of the present application, for the novel text content, the character names and the context pronouns can be distinguished through text recognition, so as to determine each character in the novel text content. For the context pronouns, the person corresponding to the pronoun can be located through text resolution of natural language processing, and the pronoun is replaced by the name of the person, so that the character corresponding to the dialogue text is recognized. After determining the character, the dialogue text corresponding to each character can be determined through text recognition, so as to obtain the dialogue text of each character.
[0048] The dialogue text of each character in the novel content text is identified by the text emotion recognition model, and the text emotion classification of each character is obtained. The text emotion recognition model can be trained by inputting a text emotion sample into a preset neural network. The text emotion sample includes a sample text and a corresponding emotion classification label.
[0049] Step 140: according to the voiceprint information of each character, the video emotion classification of each character, the dialogue text of each character, the text emotion classification of each character, and the external environment emotion classification, audio synthesis processing is performed on the dialogue text of each character to obtain the audiobook audio corresponding to the novel text content.
[0050] After obtaining the characters of the video data, the voiceprint information of each character, the video emotion classification of each character, the characters of the novel text content, and the text emotion classification of each character, the dialogue text of each character can be audio synthesized according to the voiceprint information of each character, the video emotion classification of each character, the dialogue text of each character, the text emotion classification of each character, and the external environment emotion classification, so as to obtain the audiobook audio corresponding to the novel text content.
[0051] In one embodiment of the present application, the characters in the video data are matched with the characters in the novel content. Specifically, it is determined whether the characters in the video data are consistent with the characters in the novel text content. If the characters are consistent, the voiceprint information of each character in the video data corresponding to each character is extracted as the voiceprint information of each character in the novel text content in the audio synthesis processing. If the characters are not consistent, a similar character is determined according to the character information of a target character in the novel text content, and the voiceprint information corresponding to the similar character is taken as the voiceprint information of the target character in the novel text content in the audio synthesis processing. The target character is the character in the novel text content that is not consistent with the character in the video data. After obtaining the voiceprint information of each character in the novel text content, audio synthesis processing can be performed according to the voiceprint information of each character and the dialogue text of each character to obtain the audiobook audio corresponding to the novel text content. The basic information of the voiceprint includes basic time domain and frequency domain information, which represents the frequency and loudness of the sound. Through the basic information of the voiceprint, the user can feel the audio consistent with the voiceprint of the character in the video data, and to some extent, the emotion of the character in the video data can be reflected.
[0052] In another embodiment of the present application, after the characters in the video data are matched with the characters in the novel content, if the characters are consistent, it is determined whether the video emotion classification of each character is consistent with the text emotion classification of each character. If the emotion classifications are consistent, the emotion of each character in the novel text content in the audio synthesis processing is obtained according to the video emotion classification of each character and the external environment emotion classification. If the emotion classifications are not consistent, the emotion of each character in the novel text content in the audio synthesis processing is obtained according to the text emotion classification of each character and the external environment emotion classification. The audiobook audio corresponding to the novel text content can be obtained by performing audio synthesis processing according to the obtained emotion of each character in the novel text content and the dialogue text of each character. In this way, the user can feel the emotion consistent with the character in the video data.
[0053] In yet another embodiment of the present application, it is determined whether each character in the video data is consistent with each character in the novel text content. If the characters are consistent, the voiceprint information of each character in the video data corresponding to each character is extracted as the voiceprint information of each character in the novel text content. When the characters are consistent, it is determined whether the video emotion classification of each character is consistent with the text emotion classification of each character. If the emotion classifications are consistent, the emotion of each character in the video data is used as the emotion of each character in the novel text content. According to the voiceprint information of each character and the emotion of each character in the video data, the dialogue text of each character is subjected to audio synthesis processing to obtain the audiobook audio corresponding to the novel text content. In addition to the voiceprint information change characteristics in the time domain and the frequency domain, the synthesized audio has a basic emotion expression, but it is relatively flat and the user perception is not enough. The embodiment of the present application considers the video emotion classification of the character performance in different speech speeds, pauses, volumes (loudness), and pitches, which can enrich the emotional content expression of the novel, increase the audio dynamic effect of the actual scene, and make the audiobook realize the feeling of being in the scene, thereby enhancing the user perception. The two features of volume (loudness) and pitch can be reflected by two-dimensional spectrum information. Because the pitch and the loudness have a complementary relationship, for example, the same loudness can be felt, and the frequency can be supplemented by the sound intensity to make people feel the same pitch. When the user terminal volume interface does not support adaptive adjustment, the sound intensity is supplemented by the frequency, and when the user terminal open volume interface supports adaptive adjustment, the frequency is supplemented by the sound intensity or both are adjusted to satisfy the audio features corresponding to the emotion. Therefore, the embodiment of the present application combines the video emotion classification of the character to increase the emotional content of the audiobook audio corresponding to the novel text content in different speech speeds, pauses, and the like. For example, considering the same late scene, two characters, and three different dialogue examples in different contexts, different text widths represent different audio speeds, and [stop] represents a pause marker.
[0054] The following is the mother's prompt:
[0055] "Quick, quick, you will be late. [stop] Yes, take your bag, and do your own things!"
[0056] The following is the child's monologue:
[0057] "Damn it, there are still ten minutes, I will be late."
[0058] The following is the child's confession to the classmates:
[0059] "Oh, [stop] I was late again today, and I was criticized by the teacher. I am so sad."
[0060] The corresponding role video emotion classification of the above content is anxious + angry, anxious + worried, and sad, so the different speech speed and pause features corresponding to these emotion tags can be extracted as a supplement to the role voiceprint information of each role. In this way, both the original voice of the role in the video data and the emotion of the role in the video data are combined, so that when a user watching a TV series switches to an audiobook scene, the user can have a seamless transition, and the user can have a more immersive audiobook experience.
[0061] In another embodiment of the application, the overall emotion fusion classification of the dialogue text of each role is also determined in combination with the external environment emotion classification. When the VR device switches to the front camera, the external environment emotion classification in the current shooting scene frame, such as the weather, character action, and environment sound of the current environment of the user, can be input as a parameter for the emotion classification of the audiobook audio as an emotion supplement for the audiobook audio. Specifically, the target voiceprint information corresponding to each role in the novel text content is determined according to the voiceprint information of each role and the novel text content; the role video emotion classification, the role text emotion classification, and the external environment emotion classification are input into a preset emotion classification fusion model to obtain an overall emotion fusion classification; and the audiobook audio corresponding to the novel text content is synthesized according to the target voiceprint information and the overall emotion fusion classification. The external environment emotion classification is obtained by classifying the emotions of external environment factors. The external environment factors include weather emotion factors, environment sound emotion factors, and character action emotion factors. The weather emotion factors can include at least one external environment factor, such as temperature, humidity, air pressure, wind speed, and sunshine, etc. Before the role video emotion classification, the role text emotion classification, and the external environment emotion classification are input into the preset emotion classification fusion model to obtain the overall emotion fusion classification, the environment parameters corresponding to each external environment factor are collected; the emotion scores of each external environment factor are calculated according to each environment parameter and a corresponding piecewise function; and the external environment emotion classification is determined according to the emotion scores. Taking the weather emotion factor as an example, each external environment factor in the weather emotion factor has an impact on the user's emotions. The impact of each external environment factor on emotions can be defined as a piecewise function (i.e., a piecewise function). The functions at both ends can be linear or nonlinear, and in real life, they are mostly nonlinear functions. In the embodiment of the application, the piecewise function can set a linear function according to the experience range of the impact of each external environment factor in the weather emotion factor on emotions, or it can be realized according to the experience model of the change of the weather emotion factor and emotions. In the embodiment of the application, a linear function is set according to the experience range of the impact of each external environment factor in the weather emotion factor on emotions, such as Figure 2As shown, assuming x is the weather element factor value, y is the mood score, the piecewise linear step function is set as: y=ax+b x<left critical value of mood stability; y=y*left critical value of mood stability<=x<=right critical value of mood stability; y=cx+d x>right critical value of mood stability. Among them, in order to simplify the processing, the weather emotional factor corresponding to the weather emotion classification can be divided into neutral, positive and negative three kinds, and for the mood score parameter, 0 is neutral, greater than 0 is positive, and less than 0 is negative, and the x interval is the upper and lower limit value range of the weather element factor. According to the experience range of the influence of weather emotional factors on emotion, the coefficients of the piecewise function can be set according to the environmental parameters of one of the external environment factors, and the environmental parameters of all other external environment factors are normalized. The environmental parameter range of the external environment factor and the mood score are extended and mapped to obtain the corresponding weather emotional classification. Taking temperature as an example, when the temperature is 11-25℃, people are most likely to keep a happy and stable mood. When the temperature exceeds 34℃, people are prone to feel irritable and are prone to overreaction. At the same time, too low temperature also has a negative impact. When the indoor temperature is lower than 10℃, people will feel depressed and depressed. When it is lower than 4℃, the thinking efficiency of people is seriously affected, the work quality is decreased, and the error is prone to occur. Therefore, 4-34℃ is set as the temperature factor interval of the weather emotional factor, 11-25℃ is set as the highest positive mood score interval, 25-34℃ is set as y=cx+d, that is, the interval of positive mood subsiding and turning to negative mood, and 4-11℃ is set as y=ax+b, that is, the interval of negative mood subsiding and turning to positive mood, so that the piecewise function coefficients can be calculated according to the above interval critical value. For the humidity factor, according to different seasons, the mood is best, that is, the mood score is highest, which is usually between 65% and 90%. The humidity factor and the temperature factor are normalized, the interval scale value of the humidity factor is calculated as 1.8 by fixing the interval scale value of the temperature factor: (25-11) / 1=(90-65) / humidity factor interval scale value, and the piecewise function coefficients of the humidity factor are obtained according to the corresponding relationship with the environmental emotion classification. Through the above method, the piecewise function of each external environment factor can be obtained. Through the environmental parameters of the external environment factors and the corresponding piecewise function, the mood score of each external environment factor can be determined. Correspondingly, the environmental sound emotional factor and the character action emotional factor can also be determined according to the corresponding relationship between the environmental sound and the character action and the preset emotional classification.
[0062] As Figure 3As shown, after obtaining the weather emotion classification corresponding to the weather emotion factor, the environmental sound emotion classification corresponding to the environmental sound emotion factor, and the character action emotion classification corresponding to the character action emotion factor, the external environment emotion classification, the character video emotion classification, and the character text emotion classification are input into a preset emotion classification fusion model to obtain the overall emotion fusion classification. The preset emotion classification fusion model is a deep learning network such as a feature fusion learning network, a bidirectional Transformer representation encoder, and an attention encoder.
[0063] In another specific embodiment of the application, before the audio synthesis processing of the character dialogue text according to the character voiceprint information, the character video emotion classification, the character text emotion classification, and the external environment emotion classification, the character voice content in the video data is obtained, and it is determined whether the character voice content in the video data is consistent with the character dialogue text in the novel text content. If consistent, the current audio of each character in the video data is extracted and used as the audiobook audio corresponding to the novel text content. If the text is inconsistent, the audio synthesis processing of the character dialogue text according to the character voiceprint information, the character video emotion classification, the character text emotion classification, and the external environment emotion classification is performed to obtain the audiobook audio corresponding to the novel text content.
[0064] In the embodiment of the application, when watching a video through VR, the corresponding novel is found according to the watching video, the video data when switching from a watching state to an audiobook state and the novel text content corresponding to the video data are obtained, the voiceprint and emotion extraction processing of the video data is performed to obtain character voiceprint information and character video emotion classification, the dialogue and emotion extraction processing of the novel text content is performed to obtain character dialogue text and character text emotion classification, and the audio synthesis processing of the character dialogue text according to the character voiceprint information, the character video emotion classification, the character dialogue text, the character text emotion classification, and the external environment emotion classification is performed to obtain the audiobook audio corresponding to the novel text content. This can make the user switch from a watching state to an audiobook state more smoothly, reduce the sense of discomfort caused by state switching, and enable the user to have the experience of the watching state when in the audiobook state.
[0065] Figure 4 A structural schematic diagram of an audio synthesis device provided by an embodiment of the present application is shown. As shown in the figure, the device 200 comprises: Figure 4
[0066] An acquisition module 210, configured to acquire video data when switching from a viewing state to an audiobook state and novel text content corresponding to the video data;
[0067] A first processing module 220, configured to perform voiceprint and emotion extraction processing on the video data to obtain voiceprint information of each character and video emotion classification of each character;
[0068] A second processing module 230, configured to perform dialogue and emotion extraction processing on the novel text content to obtain dialogue text of each character and text emotion classification of each character;
[0069] A synthesis module 240, configured to perform audio synthesis processing on the dialogue text of each character according to the voiceprint information of each character, the video emotion classification of each character, the dialogue text of each character, the text emotion classification of each character, and external environment emotion classification, to obtain audiobook audio corresponding to the novel text content.
[0070] In an optional manner, the acquisition module 210 is further configured to:
[0071] acquire the video data in the viewing state and novel text in the audiobook state;
[0072] input the video data into a preset key dialogue text discrimination model to obtain target key dialogue text;
[0073] match the target key dialogue text with novel dialogue in the novel text to obtain novel text content corresponding to the video data.
[0074] In an optional manner, the audio synthesis processing on the dialogue text of each character according to the voiceprint information of each character, the video emotion classification of each character, the text emotion classification of each character, and external environment emotion classification, to obtain audiobook audio corresponding to the novel text content, comprises:
[0075] determining target voiceprint information corresponding to each character in the novel text content according to the voiceprint information of each character and each character in the novel text content;
[0076] inputting the video emotion classification of each character, the text emotion classification of each character, and the external environment emotion classification into a preset emotion classification fusion model to obtain overall emotion fusion classification;
[0077] According to the target voiceprint information and the overall emotion fusion classification, the audiobook audio corresponding to the novel text content is synthesized.
[0078] In an optional manner, before the role video emotion classification, the role text emotion classification, and the external environment emotion classification are input into a preset emotion classification fusion model to obtain an overall emotion fusion classification, the method comprises the following steps:
[0079] An environment parameter corresponding to each external environment factor is collected.
[0080] According to each environment parameter and a corresponding segmentation function, an emotional score of each external environment factor is calculated.
[0081] According to each emotional score, the external environment emotion classification is determined.
[0082] In an optional manner, the audio synthesis processing of the role dialogue text according to the role voiceprint information, the role video emotion classification, the role text emotion classification, and the external environment emotion classification to obtain the audiobook audio corresponding to the novel text content comprises the following steps: determining whether each role in the video data is consistent with each role in the novel text content; if the roles are consistent, extracting each role voiceprint information corresponding to each role in the video data as the role voiceprint information of each role in the novel text content in the audio synthesis processing; if the roles are inconsistent, determining a similar role according to the role information of a target role in the novel text content, and taking the voiceprint information corresponding to the similar role as the voiceprint information of the target role in the novel text content in the audio synthesis processing; the target role is a role in the novel text content that is inconsistent with the role in the video data.
[0083] In an optional manner, the audio synthesis processing of the role dialogue text according to the role voiceprint information, the role video emotion classification, the role text emotion classification, and the external environment emotion classification to obtain the audiobook audio corresponding to the novel text content comprises the following steps: determining whether the role video emotion classification is consistent with the role text emotion classification; if the emotion classifications are consistent, obtaining the emotion of each role in the novel text content in the audio synthesis processing according to the role video emotion classification and the external environment emotion classification; if the emotion classifications are inconsistent, obtaining the emotion of each role in the novel text content in the audio synthesis processing according to the role text emotion classification and the external environment emotion classification.
[0084] In an optional manner, before the audio synthesis processing of the dialogue text of each character according to the voiceprint information of each character, the video emotion classification of each character, the text emotion classification of each character, and the external environment emotion classification, the method further comprises: obtaining the voice content of each character from the video data; determining whether the voice content of each character in the video data is consistent with the dialogue text of each character in the novel text content; if consistent, extracting the current audio of each character in the video data, and taking the current audio of each character as the audiobook audio corresponding to the novel text content; if inconsistent, performing the audio synthesis processing of the dialogue text of each character according to the voiceprint information of each character, the video emotion classification of each character, the text emotion classification of each character, and the external environment emotion classification, to obtain the audiobook audio corresponding to the novel text content.
[0085] In the embodiment of the application, when watching a video, the corresponding novel is searched according to the video, the video data when switching from a video watching state to an audiobook state and the novel text content corresponding to the video data are obtained, the voiceprint and emotion extraction processing is performed on the video data to obtain voiceprint information of each character and video emotion classification of each character, the dialogue and emotion extraction processing is performed on the novel text content to obtain dialogue text of each character and text emotion classification of each character, and the audio synthesis processing of the dialogue text of each character is performed according to the voiceprint information of each character, the video emotion classification of each character, the dialogue text of each character, the text emotion classification of each character, and the external environment emotion classification, to obtain the audiobook audio corresponding to the novel text content, so that the user can be more smoothly switched from the video watching state to the audiobook state, the sense of discomfort caused by the state switching is reduced, and the user can also have the experience in the video watching state when in the audiobook state.
[0086] Figure 5 The structure of the audio synthesis device provided by the embodiment of the application is shown, and the embodiment of the application does not limit the specific implementation of the audio synthesis device.
[0087] As shown in Figure 5 The audio-video synchronization test device can include a processor 302, a communications interface 304, a memory 306, and a communications bus 308.
[0088] The processor 302, the communications interface 304, and the memory 306 can communicate with each other through the communications bus 308. The communications interface 304 is configured to communicate with network elements such as clients or other servers. The processor 302 is configured to execute the program 310, and can execute the related steps in the above-mentioned audio-video synchronization test method embodiments.
[0089] In particular, the program 310 can include program code comprising computer-executable instructions.
[0090] The processor 302 can be a central processing unit (CPU), or an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement one or more embodiments of the present application. The one or more processors included in the audio-video synchronization test device can be the same type of processor, such as one or more CPUs, or different types of processors, such as one or more CPUs and one or more ASICs.
[0091] The memory 306 is configured to store the program 310. The memory 306 can include a high-speed RAM memory, and can also include a non-volatile memory, such as at least one disk memory.
[0092] The program 310 can be specifically invoked by the processor 302 to cause the audio-video synchronization test device to perform the following operations:
[0093] Obtain video data when switching from a viewing state to an audiobook state and novel text content corresponding to the video data;
[0094] Perform voiceprint and emotion extraction processing on the video data to obtain voiceprint information of each character and video emotion classification of each character;
[0095] Perform dialogue and emotion extraction processing on the novel text content to obtain dialogue text of each character and text emotion classification of each character;
[0096] Perform audio synthesis processing on the dialogue text of each character according to the voiceprint information of each character, the video emotion classification of each character, the dialogue text of each character, the text emotion classification of each character, and the external environment emotion classification, to obtain audiobook audio corresponding to the novel text content.
[0097] In an optional manner, the obtaining of the video data when switching from a viewing state to an audiobook state and novel text content corresponding to the video data includes: obtaining the video data in the viewing state and novel text in the audiobook state; inputting the video data into a preset key dialogue text discrimination model to obtain target key dialogue text; and matching the target key dialogue text with novel dialogue in the novel text to obtain novel text content corresponding to the video data.
[0098] In an optional manner, the audio synthesis processing of the dialogue text of each character according to the voiceprint information of each character, the video emotion classification of each character, the text emotion classification of each character, and the external environment emotion classification, to obtain the audiobook audio corresponding to the novel text content, comprises: determining the target voiceprint information corresponding to each character in the novel text content according to the voiceprint information of each character and each character in the novel text content; inputting the video emotion classification of each character, the text emotion classification of each character, and the external environment emotion classification into a preset emotion classification fusion model to obtain an overall emotion fusion classification; and synthesizing the audiobook audio corresponding to the novel text content according to the target voiceprint information and the overall emotion fusion classification.
[0099] In an optional manner, before the video emotion classification of each character, the text emotion classification of each character, and the external environment emotion classification are input into a preset emotion classification fusion model to obtain an overall emotion fusion classification, the method comprises: collecting environment parameters corresponding to each external environment factor; calculating the emotion score of each external environment factor according to each environment parameter and a corresponding segmentation function; and determining the external environment emotion classification according to each emotion score.
[0100] In an optional manner, the audio synthesis processing of the dialogue text of each character according to the voiceprint information of each character, the video emotion classification of each character, the text emotion classification of each character, and the external environment emotion classification, to obtain the audiobook audio corresponding to the novel text content, comprises: determining whether each character in the video data is consistent with each character in the novel text content; if the characters are consistent, extracting the voiceprint information of each character in the video data as the voiceprint information of each character in the novel text content in the audio synthesis processing; if the characters are inconsistent, determining a similar character according to the character information of a target character in the novel text content, and taking the voiceprint information of the similar character as the voiceprint information of the target character in the novel text content in the audio synthesis processing; the target character is a character in the novel text content that is inconsistent with the character in the video data.
[0101] In an alternative manner, the audio synthesis processing of the dialogue text of each character according to the voiceprint information of each character, the video emotion classification of each character, the text emotion classification of each character and the external environment emotion classification to obtain the audiobook audio corresponding to the novel text content comprises: determining whether the video emotion classification of each character is consistent with the text emotion classification of each character; if the emotion classifications are consistent, obtaining the emotion of each character in the novel text content in the audio synthesis processing according to the video emotion classification of each character and the external environment emotion classification; and if the emotion classifications are inconsistent, obtaining the emotion of each character in the novel text content in the audio synthesis processing according to the text emotion classification of each character and the external environment emotion classification.
[0102] In an alternative manner, before the audio synthesis processing of the dialogue text of each character according to the voiceprint information of each character, the video emotion classification of each character, the text emotion classification of each character and the external environment emotion classification to obtain the audiobook audio corresponding to the novel text content, the method further comprises: obtaining the voice content of each character from the video data; determining whether the voice content of each character in the video data is consistent with the dialogue text of each character in the novel text content; if consistent, extracting the current audio of each character in the video data and taking the current audio of each character as the audiobook audio corresponding to the novel text content; and if inconsistent, performing the audio synthesis processing of the dialogue text of each character according to the voiceprint information of each character, the video emotion classification of each character, the text emotion classification of each character and the external environment emotion classification to obtain the audiobook audio corresponding to the novel text content.
[0103] According to still another aspect of the embodiment of the present application, a computer readable storage medium is provided, wherein the storage medium stores at least one executable instruction, and the executable instruction causes the processor to execute the steps of the above-mentioned audio synthesis method. The executable instruction causes the processor to execute the following steps:
[0104] obtaining the video data when switching from a viewing state to an audiobook state and the novel text content corresponding to the video data;
[0105] performing voiceprint and emotion extraction processing on the video data to obtain voiceprint information of each character and video emotion classification of each character;
[0106] performing dialogue and emotion extraction processing on the novel text content to obtain dialogue text of each character and text emotion classification of each character;
[0107] According to the role voiceprint information, the role video emotion classification, the role dialogue text, the role text emotion classification, and the external environment emotion classification, the role dialogue text is subjected to audio synthesis processing to obtain the audiobook audio corresponding to the novel text content.
[0108] In an optional manner, the video data in the viewing state and the novel text content in the audiobook state are obtained, including: obtaining the video data in the viewing state and the novel text content in the audiobook state; inputting the video data into a preset key dialogue text discrimination model to obtain target key dialogue text; matching the target key dialogue text with novel dialogue in the novel text to obtain the novel text content corresponding to the video data.
[0109] In an optional manner, the role voiceprint information, the role video emotion classification, the role text emotion classification, and the external environment emotion classification are used to perform audio synthesis processing on the role dialogue text to obtain audiobook audio corresponding to the novel text content, including: determining target voiceprint information corresponding to each role in the novel text content according to the role voiceprint information and each role in the novel text content; inputting the role video emotion classification, the role text emotion classification, and the external environment emotion classification into a preset emotion classification fusion model to obtain overall emotion fusion classification; and synthesizing audiobook audio corresponding to the novel text content according to the target voiceprint information and the overall emotion fusion classification.
[0110] In an optional manner, before the role video emotion classification, the role text emotion classification, and the external environment emotion classification are input into a preset emotion classification fusion model to obtain overall emotion fusion classification, the method includes: collecting environment parameters corresponding to each external environment factor; calculating emotion scores of each external environment factor according to each environment parameter and a corresponding segmentation function; and determining the external environment emotion classification according to the emotion scores.
[0111] In an optional manner, the audio synthesis processing of the dialogue text of each character according to the character voiceprint information, the character video emotion classification, the character text emotion classification, and the external environment emotion classification to obtain the audiobook audio corresponding to the novel text content comprises: determining whether the characters in the video data are consistent with the characters in the novel text content; if the characters are consistent, extracting the character voiceprint information of each character in the video data as the character voiceprint information of each character in the novel text content in the audio synthesis processing; if the characters are inconsistent, determining a similar character according to the character information of a target character in the novel text content, and taking the voiceprint information of the similar character as the voiceprint information of the target character in the novel text content in the audio synthesis processing; the target character is a character in the novel text content that is inconsistent with the character in the video data.
[0112] In an optional manner, the audio synthesis processing of the dialogue text of each character according to the character voiceprint information, the character video emotion classification, the character text emotion classification, and the external environment emotion classification to obtain the audiobook audio corresponding to the novel text content comprises: determining whether the character video emotion classification is consistent with the character text emotion classification; if the emotion classifications are consistent, obtaining the emotion of each character in the novel text content in the audio synthesis processing according to the character video emotion classification and the external environment emotion classification; if the emotion classifications are inconsistent, obtaining the emotion of each character in the novel text content in the audio synthesis processing according to the character text emotion classification and the external environment emotion classification.
[0113] In an optional manner, before the audio synthesis processing of the dialogue text of each character according to the character voiceprint information, the character video emotion classification, the character text emotion classification, and the external environment emotion classification to obtain the audiobook audio corresponding to the novel text content, the method further comprises: obtaining the voice content of each character from the video data; determining whether the voice content of each character in the video data is consistent with the dialogue text of each character in the novel text content; if consistent, extracting the current audio of each character in the video data and taking the current audio of each character as the audiobook audio corresponding to the novel text content; if inconsistent, performing the audio synthesis processing of the dialogue text of each character according to the character voiceprint information, the character video emotion classification, the character text emotion classification, and the external environment emotion classification to obtain the audiobook audio corresponding to the novel text content.
[0114] The embodiment of the present application provides a computer program product, the computer program product comprises a computer program stored on a computer readable storage medium, the computer program comprises program instructions, when the program instructions run on a computer, the computer executes the audio synthesis method in any method embodiment described above.
[0115] The algorithms or displays provided herein are not inherently related to any particular computer, virtual system, or other apparatus. Various general purpose systems can be used with programs in accordance with the teachings herein, or it can prove convenient to construct more specialized apparatus to perform the required method steps. The required structure for a variety of these systems will be apparent from the description above. In addition, the present embodiments are not described with reference to any particular programming language. It will be appreciated that a variety of programming languages can be used to implement the teachings of the application as described herein, and any references below to specific languages are provided for disclosure of enablement only.
[0116] In the description provided herein, numerous specific details are set forth. However, it is understood that embodiments of the application can be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been described in detail in order to not obscure the understanding of this description.
[0117] Similarly, it is to be understood that the mechanical details of the application sometimes are grouped into single embodiments, figures, or descriptions for the purpose of streamlining the disclosure and aiding in the apprehension of one or more of the various aspects of the application. In some instances, reference has been made to acts or symbolic representations of operations that are related to one or more of such aspects. However, it is to be understood that the disclosed subject matter is not limited by the filing of these acts or symbolic representations— the specifics with respect to individual features of the application can later be modified in a manner that still falls within the scope of the present application. Certain features and subcombinations are thus, to the maximum extent possible, described in terms of
[0118] Those skilled in the art will appreciate that the modules in the device in the embodiments can be adaptively changed and arranged in one or more devices different from the embodiments. The modules or units or components in the embodiments can be combined into one module or unit or component, and can be divided into multiple sub-modules or sub-units or sub-components. All the features disclosed in the specification (including the claims, abstract and drawings) and all the processes or units of any method or apparatus disclosed can be combined in any combination, except that at least some of such features and / or processes or units are mutually exclusive. Unless specifically stated otherwise, each feature disclosed in the specification (including the claims, abstract and drawings) can be replaced by alternative features providing the same, equivalent or similar functionality.
[0119] It should be noted that the above-mentioned embodiments illustrate rather than limit the application, and that those skilled in the art will be able to design many alternative embodiments without departing from the scope of the appended claims. In the claims, any reference signs placed between parentheses shall not be construed as limiting the claim. The word 'comprising' does not exclude the presence of elements or steps other than those listed in a claim. The word 'a' or 'an' preceding an element does not exclude the presence of a plurality of such elements. The application can be implemented by means of both hardware and software, and any combination thereof. In a unitary claim, several devices, apparatuses or means can be listed, comprising means for carrying out a certain task. The use of the term'means' in a claim is intended to refer to a combination of devices, apparatuses or means for carrying out a task. The word 'first','second', 'third', etc. do not imply any order. The use of these terms is to be construed as an indication of particular embodiments. Steps in the above-described embodiments, unless otherwise specified, are not to be construed as necessarily limiting the order in which the steps are performed.
Claims
1. An audio synthesis method, characterized by, The method comprises: acquiring video data when switching from a viewing state to an audiobook state and novel text content corresponding to the video data; performing voiceprint and emotion extraction processing on the video data to obtain voiceprint information of each character and video emotion classification of each character; performing dialogue and emotion extraction processing on the novel text content to obtain dialogue text of each character and text emotion classification of each character; performing audio synthesis processing on the dialogue text of each character according to the voiceprint information of each character, the video emotion classification of each character, the dialogue text of each character, the text emotion classification of each character, and external environment emotion classification, to obtain audiobook audio corresponding to the novel text content.
2. The method of claim 1, wherein, The acquisition of video data when switching from a viewing state to an audiobook state and novel text content corresponding to the video data comprises: acquiring the video data in the viewing state and novel text in the audiobook state; inputting the video data into a preset key dialogue text discrimination model to obtain target key dialogue text; matching the target key dialogue text with novel dialogue in the novel text to obtain novel text content corresponding to the video data.
3. The method of claim 1, wherein, The audio synthesis processing on the dialogue text of each character according to the voiceprint information of each character, the video emotion classification of each character, the dialogue text of each character, the text emotion classification of each character, and external environment emotion classification to obtain audiobook audio corresponding to the novel text content comprises: determining target voiceprint information corresponding to each character in the novel text content according to the voiceprint information of each character and each character in the novel text content; inputting the video emotion classification of each character, the text emotion classification of each character, and the external environment emotion classification into a preset emotion classification fusion model to obtain overall emotion fusion classification; synthesizing audiobook audio corresponding to the novel text content according to the target voiceprint information and the overall emotion fusion classification.
4. The method of claim 3, wherein, Before the inputting of the video emotion classification of each character, the text emotion classification of each character, and the external environment emotion classification into a preset emotion classification fusion model to obtain overall emotion fusion classification, the method comprises: collecting environmental parameters corresponding to each external environmental factor; calculating emotion scores of each external environmental factor according to each environmental parameter and a corresponding segmentation function; determining the external environment emotion classification according to the emotion scores.
5. The method according to any one of claims 1-3, characterized in that, The audio synthesis processing on the dialogue text of each character according to the voiceprint information of each character, the video emotion classification of each character, the dialogue text of each character, the text emotion classification of each character, and external environment emotion classification to obtain audiobook audio corresponding to the novel text content comprises: determining whether each character in the video data is consistent with each character in the novel text content; if the characters are consistent, extracting the voiceprint information of each character in the video data as the voiceprint information of each character in the novel text content in the audio synthesis processing; If the characters are inconsistent, a similar character is determined based on the character information of the target character in the novel's text content, and the voiceprint information corresponding to the similar character is used as the voiceprint information of the target character in the novel's text content in the audio synthesis process; the target character is the character in the novel's text content that is inconsistent with the character in the video data.
6. The method according to any one of claims 1-3, characterized in that, The step of performing audio synthesis processing on the dialogue text of each character based on the voiceprint information of each character, the emotional classification of each character's video, the dialogue text of each character, the emotional classification of the text of each character, and the emotional classification of the external environment, to obtain the audiobook corresponding to the novel text content, includes: Determine whether the video emotion classification of each character is consistent with the text emotion classification of each character. If the emotion classifications are consistent, then the emotions of each character in the novel text content during the audio synthesis process are obtained based on the emotion classifications of each character's video and the emotion classifications of the external environment. If the emotion classifications are inconsistent, the emotions of each character in the novel text content during the audio synthesis process are obtained based on the character text emotion classification and the external environment emotion classification.
7. The method according to any one of claims 1-3, characterized in that, Before obtaining the audiobook audio corresponding to the novel text content by performing audio synthesis processing on the dialogue text of each character based on the voiceprint information of each character, the emotional classification of each character's video, the dialogue text of each character, the emotional classification of the character's text, and the emotional classification of the external environment, the method further includes: The voice content of each character is obtained from the video data; Determine whether the voice content of each character in the video data is consistent with the dialogue text of each character in the novel text content; If they match, the current audio of each character in the video data is extracted, and the current audio of each character is used as the audiobook corresponding to the novel text content. If there is a discrepancy, then based on the voiceprint information of each character, the emotional classification of each character's video, the emotional classification of each character's text, and the emotional classification of the external environment, audio synthesis processing is performed on the dialogue text of each character to obtain the audiobook corresponding to the novel's text content.
8. An audio synthesizing apparatus characterized by comprising: The device includes: The acquisition module is used to acquire video data and novel text content corresponding to the video data when switching from movie viewing mode to audiobook mode; The first processing module is used to perform voiceprint and emotion extraction processing on the video data to obtain voiceprint information of each character and video emotion classification of each character; The second processing module is used to perform dialogue and emotion extraction processing on the novel text content to obtain the dialogue text of each character and the emotion classification of each character's text. The synthesis module is used to perform audio synthesis processing on the dialogue text of each character based on the voiceprint information of each character, the emotional classification of the video of each character, the dialogue text of each character, the emotional classification of the text of each character, and the emotional classification of the external environment, so as to obtain the audiobook corresponding to the novel text content.
9. An audio synthesizing apparatus characterized by comprising: include: The processor, memory, communication interface, and communication bus are provided, wherein the processor, memory, and communication interface communicate with each other via the communication bus. The memory is used to store at least one executable instruction that causes the processor to perform the operation of the audio synthesis method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores at least one executable instruction, which, when executed on the audio synthesis device, causes the audio synthesis device to perform the operation of the audio synthesis method as described in any one of claims 1-7.
Citation Information
Patent Citations
Speech synthesis model training method and related device
CN112820265A
Voice generation method and device, equipment and storage medium
CN115472185A