Method, system and electronic device for generating real-time audio based on conversation content

By acquiring the audio correlation parameters of the speech signal and using algorithms to generate audio and sound effects that match the human-computer interaction scenario, the problem of insufficient sound effects in the existing technology is solved, and the immersive experience of users in scenarios such as amusement parks is improved.

CN119889310BActive Publication Date: 2025-11-18BEIJING APAILANG CREATIVITY TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411980781.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-11-18
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

Existing human-computer dialogue scenarios lack the ability to generate sound effects that match the dialogue scenario and atmosphere, resulting in insufficient user immersion.

Method used

By acquiring audio correlation parameters of the speech signal, including topic type, emotional tendency, interaction location, interaction time, and interaction event type, the system generates and plays matching audio using a preset algorithm model, and combines sound effects to adjust audio intensity and switch task audio in real time.

Benefits of technology

It enables the generation of sound effects that match the atmosphere based on the human-computer interaction scenario, enhancing the user's immersive experience. It is especially suitable for scenarios with strong interactive experiences, such as amusement parks and theme parks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119889310B_ABST
    Figure CN119889310B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a method, system and electronic device for generating real-time audio based on conversation content, wherein the method obtains an input voice signal, determines audio-related parameters corresponding to the voice signal, such as a subject type, an emotional tendency, an interaction position, an interaction time, an interaction event type, and the like, and then generates and plays audio matching the audio based on the determined audio-related parameters. By selecting embodiments of the present application, real-time audio can be generated and played based on specific audio-related parameters of a voice signal in human-computer interaction, such as a subject type, an emotional tendency, an interaction position, an interaction time, an interaction event type, and the like, and the generated audio can match the current conversation scene and atmosphere. For some scenes that emphasize interaction experience, such as amusement parks, theme parks, exhibition halls, and the like, selecting embodiments of the present application can provide users with a more immersive use experience and interaction experience.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to a method and system for generating real-time audio based on dialogue content and an electronic device. BACKGROUND

[0002] With the development of artificial intelligence technology, the application scenarios of human-computer dialogue have appeared in various aspects of daily life. Existing human-computer dialogue scenarios mainly focus on generating and displaying text content according to human-computer dialogue content, and lack the ability to generate sound effects that conform to the dialogue scene and atmosphere in combination with human-computer dialogue content, and cannot provide users with more immersive use experience and interaction experience. SUMMARY

[0003] Therefore, the embodiments of the present application provide a method and system for generating real-time audio based on dialogue content and an electronic device to generate adaptive sound effects according to human-computer interaction scenarios and interaction content in real time, so as to provide users with more immersive use experience and interaction experience.

[0004] In a first aspect, the embodiments of the present application provide a method for generating real-time audio based on dialogue content, wherein the method comprises:

[0005] obtaining an input voice signal, determining audio association parameters corresponding to the voice signal, wherein the audio association parameters include: theme type, emotional tendency, interaction position, interaction time, and interaction event type; wherein the interaction event type is an interaction task set in advance according to the interaction position;

[0006] generating and playing audio matching the audio association parameters based on the audio association parameters.

[0007] In some possible embodiments, the generating and playing audio matching the audio association parameters based on the audio association parameters comprises:

[0008] inputting the audio association parameters into a preset background music generation algorithm model to generate a plurality of initial music segments by the preset background music generation algorithm model;

[0009] determining an initial music segment corresponding to the identity label information of the user as the initial music based on the identity label information of the user.

[0010] In some possible embodiments, the method further comprises:

[0011] obtaining and analyzing the input voice signal at a preset sampling interval to determine audio association parameters corresponding to each sampling interval;

[0012] And based on the audio correlation coefficient corresponding to each sampling interval, the process music segment corresponding to different sampling intervals is generated by combining the preset process music generation algorithm model.

[0013] According to the sampling interval sequence, the process music is obtained by connecting each process music segment in sequence.

[0014] In some possible embodiments, the method further comprises:

[0015] Based on the pre-trained sound effect matching model, a matched sound effect sound combination is determined, wherein the sound effect sound combination includes multiple sound effects, and each sound effect in the sound effect sound combination is associated with the existence of the interaction event type;

[0016] In the process of continuously playing the initial music and the process music without pause, the target sound effect matched with the interaction event type is embedded and played in the gap between different process music segments.

[0017] In some possible embodiments, the method further comprises:

[0018] The sound intensity of the input voice signal is acquired in real time, and if the sound intensity is lower than a preset sound intensity threshold, the sound intensity of the played audio is reduced according to a preset sound intensity reduction ratio.

[0019] In some possible embodiments, the method further comprises:

[0020] If the interactive task is changed, the audio associated with the current interactive task is terminated, and the audio corresponding to the changed new interactive task is switched and played.

[0021] In some possible embodiments, the theme type is determined in the following manner:

[0022] According to a preset voice-to-text function, the input voice signal is converted into target text content;

[0023] Based on the target text content, semantic analysis is performed to extract associated words related to a preset theme type contained in the target text content;

[0024] According to the associated words, a theme type with the highest association degree with the associated words is determined as the theme type corresponding to the voice signal;

[0025] The emotional tendency is determined in the following manner:

[0026] The audio features of the voice signal are acquired, and the audio features include: speech rate, volume, and pitch;

[0027] normalize the speech features based on a reference speech speed, a reference volume, and a reference pitch to obtain normalized audio coefficients;

[0028] According to the audio coefficients and preset emotional tendency label information, a target emotional tendency score is calculated, and an emotional tendency corresponding to the speech signal is determined based on the target emotional tendency score.

[0029] In a second aspect, the present application provides a system for generating real-time audio based on conversation content, wherein the system comprises:

[0030] An acquisition module is configured to acquire an input speech signal and determine audio-related parameters corresponding to the speech signal, wherein the audio-related parameters comprise a theme type, an emotional tendency, an interaction position, an interaction time, and an interaction event type; and the interaction event type is an interaction task set in advance according to the interaction position.

[0031] An audio generation and playing module is configured to generate and play audio matching the audio-related parameters based on the audio-related parameters.

[0032] In a third aspect, the present application provides an electronic device, wherein the electronic device comprises a processor and a memory storing a program; the program comprises instructions which, when executed by the processor, cause the processor to execute the method for generating real-time audio based on conversation content according to the first aspect.

[0033] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are configured to cause a computer to execute the method for generating real-time audio based on conversation content according to the first aspect.

[0034] The present application has the following beneficial effects:

[0035] The present application provides a method, system, and electronic device for generating real-time audio based on conversation content, wherein the method acquires an input speech signal and determines audio-related parameters such as a theme type, an emotional tendency, an interaction position, an interaction time, and an interaction event type corresponding to the speech signal, and then generates and plays audio matching the audio based on the determined audio-related parameters. By using the present application, real-time audio matching the theme type, emotional tendency, interaction position, interaction time, and interaction event type of the speech signal of human-computer interaction can be generated and played, and audio matching the current conversation scene and atmosphere can be generated, thereby providing users with more immersive use experience and interaction experience in scenes such as amusement parks, theme parks, and exhibition halls which emphasize interaction experience. BRIEF DESCRIPTION OF DRAWINGS

[0036] More details, features and advantages of the present application will be disclosed in the following description of exemplary embodiments in conjunction with the accompanying drawings, in which:

[0037] Figure 1 Fig. 1 shows a flow diagram of a method for generating real-time audio based on conversation content according to an embodiment of the present application;

[0038] Figure 2 Fig. 2 shows a system architecture diagram of a system for generating real-time audio based on conversation content according to an embodiment of the present application;

[0039] Figure 3 Fig. 3 shows another system architecture diagram of a system for generating real-time audio based on conversation content according to an embodiment of the present application;

[0040] Figure 4 Fig. 4 shows a structural block diagram of an exemplary electronic device that can be used to implement embodiments of the present application. DETAILED DESCRIPTION

[0041] Embodiments of the present application will be described in more detail below with reference to the accompanying drawings. While certain embodiments of the present application are shown in the drawings, it is understood that the present application can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein; rather, these embodiments are provided so that the present application will be more thoroughly and completely understood. It should be understood that the drawings and embodiments of the present application are only for illustrative purposes and should not be construed as limiting the scope of protection of the present application.

[0042] It should be understood that the various steps in the method embodiments of the present application can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present application is not limited in this respect.

[0043] The term "comprising" and variations thereof as used herein are open-ended, and mean "including but not limited to". The term "based on" means "based, at least in part, on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Related terms are defined as follows: the term "another", as used in "another embodiment", means "at least one additional embodiment"; the term "some", as used in "some embodiments", means "at least some embodiments". It is noted that the terms "first", "second", and the like used in the description and the claims are used to differentiate between similar elements, and are not used to denote a sequence or a relative importance of the elements.

[0044] It should be noted that the modification of "one", "a plurality of" mentioned in the present application is illustrative but not restrictive, and those skilled in the art should understand that unless otherwise explicitly indicated in the context, it should be understood as "one or more".

[0045] In order to realize real-time generation of adaptive sound effects according to human-computer conversation content, to provide users with more immersive use experience and interaction experience. The present application provides a method, system and electronic device for generating real-time audio based on conversation content, wherein the method for generating real-time audio based on conversation content provided by the present application can be applied to any electronic device with the ability to generate real-time audio based on conversation content. Similarly, the system for generating real-time audio based on conversation content provided by the present application can be deployed in the electronic device with the ability to generate real-time audio based on conversation content. The type of the electronic device can be different according to the application scenario, including but not limited to: personal mobile terminal device, computer or server, etc.

[0046] As an example, the electronic device can be a trigger device in a theme park for providing interactive services to tourists. The trigger device can be a portable interactive device or a fixed interactive device arranged in the theme park.

[0047] In the first aspect, the present application provides a method for generating real-time audio based on conversation content, which can be as shown in the figure Figure 1 The method comprises the following steps:

[0048] S11, obtaining an input voice signal, and determining audio association parameters corresponding to the voice signal, wherein the audio association parameters include: theme type, emotional tendency, interaction position, interaction time, and interaction event type; wherein the interaction event type is a pre-set interactive task according to the interaction position;

[0049] S12, generating and playing audio matching the audio association parameters based on the audio association parameters.

[0050] The embodiments of the present application obtain an input voice signal, and determine the theme type, emotional tendency, interaction position, interaction time, and interaction event type corresponding to the voice signal, and then generate and play audio matching the audio based on the determined audio association parameters.

[0051] By selecting the embodiments of the present application, the theme type, emotional tendency, interaction position, interaction time, interaction event type and other specific audio-related parameters of the voice signal of human-computer interaction can be generated in real time according to the theme type, emotional tendency, interaction position, interaction time, interaction event type and other specific audio-related parameters of the voice signal of human-computer interaction, and the matching audio can be played, and the sound effect matching the current dialogue scene and atmosphere can be generated. For some scenes that need to focus on interactive experience, such as amusement parks, theme parks, exhibition halls and the like, the embodiments of the present application can provide users with more immersive use experience and interactive experience.

[0052] In order to facilitate understanding of the beneficial effects of the embodiments of the present application, the method provided by the embodiments of the present application is taken as an example applied to an amusement park, it is assumed that there are several play sites in a certain amusement park, different play sites form a complete play route, if a tourist arrives at a play site on the play route, the tourist can interact with the interactive device in the play site, complete the specially designed interactive game of the play site, and there is continuity between the plots between the interactive games of each play site. The interactive game is an interaction event, the interaction event types of different play sites are different, and the interaction events of each play site constitute a complete story. The interactive device serves as the execution subject of the embodiments of the present application, by interacting with the tourist of the play site arriving at the interaction position of the interactive device, it can be determined how many interaction events have been completed by the current tourist, and the interaction event to be completed at present, the theme corresponding to the current play site, the emotional tendency of the tourist at present and the reply of the tourist based on the interaction event type, and the audio matching the current play site, the current theme, the emotional tendency of the current tourist and the interaction event type is generated, so as to improve the play experience of the tourist.

[0053] The steps S11-S12 will be described in detail below in conjunction with specific examples:

[0054] In step S11, the input voice signal is acquired, which can be the voice signal collected by the microphone module inside the device, for example, the execution subject of the present application is a portable interactive device, and the input voice signal is the voice signal collected by the microphone of the portable interactive device. As another embodiment, the input voice signal can be the voice signal collected by an external microphone device and sent to the execution subject, wherein the external microphone device and the execution subject of the method of the embodiments of the present application are not one device, and the two devices are connected through near field communication NFC, radio frequency identification RFID, Bluetooth and the like.

[0055] As an example, the guide robot is the execution subject of the method of the embodiment of the application, and the microphone on the visitor is the specific sound pickup device, which is connected with the fixed interactive device in the amusement park through Bluetooth. In the embodiment of the application, the device with the ability to collect voice signals is referred to as a sound pickup device. Based on this, in some embodiments, since the application scenario of the method provided by the embodiment of the application may be a large public place, there will be other sounds in the place in addition to the dialogue user, that is, there will be noise interference. In order to ensure the accuracy of the generated sound effects, in the process of executing the above step S11, the method for generating real-time audio based on dialogue content provided by the embodiment of the application can further include the following steps:

[0056] S11-1, acquiring a voice signal collected by a target sound pickup device, and adding identity tag information to the voice signal.

[0057] S11-2, based on a preset noise reduction algorithm, performing noise reduction processing on the voice signal of the target sound pickup device to obtain a noise-reduced to-be-analyzed voice signal.

[0058] In the execution of step S11-1, the source device of the voice signal, that is, the sound pickup device inputting the voice signal (hereinafter referred to as the target sound pickup device), can be determined, and then the input voice signal can be encoded according to the input time of the voice signal and the device number of the voice signal collection device, and the identification information that can distinguish different voice signals is added. The identification information is the identity tag information. As another implementation, the identity tag of different users can be determined by interacting with other identity recognition devices in the application scenario, so as to generate matching audio according to the identity tag subsequently. As an example, the identity tag of different users can be determined by a face recognition camera. In the embodiment of the application, the identity tag information can be set as I, and specific descriptions of using the identity tag I will be given below, which will not be described here.

[0059] As another possible implementation, when the user uses the execution subject of the method of the embodiment of the application, the user can register to use the service, the execution subject can generate user tag information according to the registration information of the user, and then when the user determines to use the function of automatically generating sound effects, the identity tag information of the voice signal issued by the user is automatically added according to the user tag information of the user. The user here can be the owner of the device, or the user of the device, which can also be referred to as the participant herein. As another possible implementation, when the user restarts the function of automatically generating sound effects, the participant manually inputs the audio tag, and the participant manually inputs the audio tag as the identity tag information of the voice signal.

[0060] In step S11-2, the AI noise reduction algorithm in the preset execution subject is executed to filter out other irrelevant sounds in the noisy environment and ensure the voice of a participant is recognized. For example, assuming the input voice signal is S voice (t), t represents time. By the preset AI noise reduction algorithm, S voice (t) can be converted into the noise-reduced voice signal S’ voice (t).

[0061] Further, the voice signal S’ voice (t) is analyzed to determine the subject type and sentiment tendency corresponding to the voice signal S’ voice (t). As an implementation, the subject type corresponding to the voice signal can be determined by the following steps:

[0062] S11-3, converting the input voice signal into target text content according to a preset voice-to-text function;

[0063] S11-4, performing semantic analysis based on the target text content to extract associated words related to the preset subject type contained in the target text content;

[0064] S11-5, determining the subject type with the highest correlation degree with the associated words as the subject type corresponding to the voice signal according to the associated words.

[0065] In step S11-3, the preset voice-to-text function can be any algorithm function with voice-to-text function. For example, a trained neural network model for voice-to-text conversion can be used to convert the input voice signal into target text content or convert the input voice signal to be analyzed into target text content. Assuming the preset voice-to-text function is T trans (), the voice signal S’ voice (t) to be analyzed can be converted into the corresponding target text content W, where W=T trans (S’ voice (t)).

[0066] Further, step S11-4 is performed. As an implementation form, the target text content W is subjected to semantic analysis by using a preset natural language processing tool or model, and each noun and mood-related auxiliary word contained in the target text content W is extracted, so that the corresponding theme type C is determined based on the extracted nouns and mood-related auxiliary words. As another implementation form, the theme types that can exist are preset based on the applicable scenarios of the method provided in the embodiments of the present application. The theme types preset in this way are preset theme types. For example, assuming that the application scenario of the amusement park is divided into technology, forest, farm, ocean, etc., the associated keywords contained in the target text content are extracted by matching the extracted nouns with the preset keywords, and it is further determined which theme the current target text content belongs to.

[0067] As another implementation form, a language text recognition and classification model, such as a text classification model based on word vectors and deep learning, is trained in advance based on a specific application scenario, and the theme type with the highest association degree with the dialog content is determined. Then, step S11-5 is performed, and the theme type with the highest association degree with the associated keywords is determined as the theme type corresponding to the voice signal according to different associated keywords. For example, assuming that the amusement park is divided into different plot areas, each plot area corresponds to a theme type, and the plot area corresponding to the keywords contained in the current target text content is determined according to the keywords associated with the plot of the amusement park, and the plot area with the highest association degree is determined as the theme type corresponding to the current voice signal.

[0068] The specific algorithm implementation of steps S11-4 and S11-5 can be understood as follows: assuming that a theme classification function F topic () is trained based on an actual application scenario, when the target text content is input into the function F topic (), the corresponding theme type C is obtained, that is, C=F topic (W). trans (S voice (t)).

[0069] As an implementation form, the target text content is subjected to semantic analysis, and the mood auxiliary word contained in the target text content is extracted, to further assist in determining the current mood and emotional tendency of the participant. As another implementation form, the emotional tendency corresponding to the voice signal is determined by the following steps:

[0070] S11-6, audio features of the voice signal are obtained, and the audio features include: speech speed v, volume l, and pitch p.

[0071] S11-7, based on the reference speech speed, the reference volume, and the reference pitch, performing normalization processing on the speech features to obtain normalized audio coefficients;

[0072] S11-8, calculating a target sentiment tendency score according to the audio coefficients and the preprocessed sentiment tendency label information, and determining the sentiment tendency corresponding to the speech signal based on the target sentiment tendency score.

[0073] In step S11-6, as an implementation form, the speech speed v of the speech signal can be calculated by counting the number of words or the number of syllables per unit time. Specifically, the word or syllable boundaries in the speech signal are recognized, the time interval between the word or syllable boundaries is measured, and then the number of words or syllables per unit time is calculated.

[0074] As an implementation form, the amplitude of the speech signal can be extracted as the volume l of the speech signal by performing audio feature extraction on the speech signal. Specifically, the volume can be calculated by calculating the sum of the amplitudes of all sampling points in an audio frame, or by calculating the sum of the squares of the amplitudes of all sampling points and then taking the logarithm as the volume of the speech signal.

[0075] As an implementation form, the fundamental frequency of the speech signal can be extracted by performing frequency spectrum analysis on the speech signal, such as short-time Fourier transform, and then the fundamental frequency is determined as the pitch p of the speech signal.

[0076] In step S11-7, the reference speech speed v0, the reference volume l0, and the reference pitch p0 are standards for measuring the speed, volume, and pitch of the speech signal. By comparing the speech speed v of the speech signal with the reference speech speed v0, the speed of the speech of the participant can be determined. Similarly, by comparing with the reference volume, the volume of the speech of the participant can be determined, and by comparing with the reference pitch, the pitch of the speech of the participant can be determined. Since the speech speed, the volume, and the pitch are data of different dimensions, in order to facilitate unified calculation, in the embodiments of the present application, the actual speech features of the participant are normalized based on the reference values to obtain normalized audio coefficients. Corresponding to the audio features of the speech, the audio coefficients also include three types: speech speed audio coefficient, volume audio coefficient, and pitch audio coefficient.

[0077] In the embodiments of the present application, the normalization processing can be performed by taking the ratio between the actual audio features and the reference values as the audio coefficients. Specifically, taking the speech speed as an example, the speech speed audio coefficient k v = the actual speech speed v of the speech signal / the reference speech speed v0, and the volume audio coefficient k l = the actual volume l of the speech signal / the reference volume l0, and the pitch audio coefficient k p= actual pitch p / reference pitch p0 of the voice signal. As an implementation, an integrated audio coefficient K = k v *k l *k p i.e. K = actual speed v / reference speed v0 of the voice signal * actual volume l / reference volume l0 of the voice signal * actual pitch p / reference pitch p0 of the voice signal.

[0078] Further, step S11-8 is performed, first, a real-time emotion analysis algorithm (for example, a support vector machine SVM algorithm) is used to perform real-time emotion analysis on the text of each piece of human-computer dialogue, to obtain the emotion tendency label information E0 of the dialogue text. Then, in combination with the integrated audio coefficient K representing the voice audio features of the dialogue, the actual emotion tendency score E = E0*K of the dialogue is calculated.

[0079] Wherein, the emotion tendency label corresponding to the dialogue text can be simply negative, neutral, positive three levels, or pessimistic, relatively pessimistic, negative, relatively negative, neutral, relatively positive, positive, relatively optimistic, optimistic, a total of 9 levels. Different levels correspond to different emotion tendency label values, i.e. different E0 values. As an implementation, when the participant first uses the service of generating real-time audio based on dialogue content provided by the embodiments of the present application, the emotion tendency score of the participant can be calculated by performing the above steps S11-6 to S11-8.

[0080] As another implementation, the participant's speech speed v', volume l' and pitch p' can be compared with different times of the participant's own speech speed, volume and pitch to learn the participant's voice characteristics and gradually perform normalization processing. This can make the calculation of the above-mentioned integrated audio coefficient K more in line with the participant's regular voice characteristics, avoid the integrated audio coefficient obtained by single calculation from excessively affecting the results of emotion tendency analysis, and finally update the emotion tendency score E and the corresponding emotion tendency category in the above-mentioned manner.

[0081] Specifically, the values of the volume, speed and pitch used to determine the emotion tendency can be updated by using statistical methods such as mean and standard deviation. For example, the mean value of the volume, the mean value of the speed and the mean value of the pitch can be calculated, and then the final integrated audio coefficient K is calculated by means of the mean value of the volume, the mean value of the speed and the mean value of the pitch, so that the result of K is more reasonable.

[0082] Based on the execution result of the above step S11, the topic theme type of the participant and the emotional tendency of the participant can be basically determined, and further, step S12 can be executed to generate and play audio matched with the topic type and the emotional tendency based on the topic type and the emotional tendency. For example, if the current game theme is a knowledge quiz, it can be determined whether the participant answers correctly according to the voice signal of the participant. If the answer is wrong, a relatively sad audio is generated and played, and if the voice signal of the participant is correct, a relatively happy audio can be generated and played. In the embodiment of the application, the audio is divided into music and sound effects, wherein the music is a continuously played audio, and the sound effects are non-continuous audio embedded in the process of playing music. As an implementation, step S12 can be implemented by the following steps:

[0083] S12-1, input the audio association parameter into a preset background music generation algorithm model to generate a plurality of initial music segments by the preset background music generation algorithm model;

[0084] S12-2, based on the identity tag information of the user, determine the initial music segment corresponding to the identity tag information of the user as the initial music.

[0085] In step S12-1, the preset background music generation algorithm model can be any type of algorithm model that can generate corresponding audio based on input text content. The preset process music generation algorithm in the following can also be any type of algorithm model that can generate corresponding audio based on input text content. The two music generation algorithm models can be the same or different. As an example, the preset background music generation algorithm model or the preset process music generation algorithm model can be MusicGen, MusicLM, etc. algorithm model to generate music segments matched with the audio association parameter.

[0086] Specifically, the interactive position P of the participant, the interactive role R of the participant, the interactive time T, the basic event O, etc. can be used as input parameters of the music generation algorithm model, wherein the basic event is an interactive event, that is, the interactive task set by the party providing the audio according to the application scenario at the corresponding interactive position P. The theme type C and the emotional tendency E are used as input parameters of the music generation algorithm model to generate initial music segments M initial .

[0087] The interactive role R can be flexibly set according to the scene plot, for example, in the application scene of the amusement park, the interactive role can be a magician or a fairy, and the specific type of the interactive role can be flexibly set. The interactive position P can be set according to the position where the interactive behavior occurs, and can be a position obtained by GPS positioning in the execution subject, or a position of a device interacting therewith. The interactive time can be determined according to the device system time, and the basic event can also be flexibly set according to the scene plot, which is mainly set by the scene sponsor according to actual needs. Taking the amusement park as an example, the basic event can be an interactive task that needs to be completed by the visitor at the current play site.

[0088] As an implementation manner, a plurality of corresponding initial music segments can be generated in advance for different user identity tag information I, and then in the actual interaction process, step S12-2 can be performed to determine the initial music segment corresponding to the user's identity tag information as the background music based on the user's identity tag information. For example, taking the amusement park as an example again, because the identity tags of different visitors are different, or the progress of different visitors in completing the script is different, the background music corresponding to different visitors at the same play site should be different. The initial music conforming to the visitor can be determined according to the identity tag of the visitor to provide an immersive play experience for the visitor.

[0089] Specifically, it is assumed that the preset music generation algorithm model includes a background music generation algorithm, and the background music generation algorithm is G music , a set of generation parameter prompt phrases B = {P, R, T, O, E, K, C} is obtained, and an initial music segment M initial = G music (B) is obtained. i Different identity tag information I initial corresponds to different initial music set M i (I play ), and the actually played background music M initial = M i (I ).

[0090] Further, in the process of performing the above step S12, the sound effect can also be generated in the following manner:

[0091] S12-3, the input voice signal is obtained and analyzed according to a preset sampling interval to determine the audio correlation coefficient corresponding to each sampling interval; that is, the set of generation parameter prompt phrases B = {P, R, T, O, E, K, C} under different sampling intervals is determined.

[0092] S12-4. Based on the audio correlation coefficients corresponding to each sampling interval, and combined with a preset process music generation algorithm model, process music segments corresponding to different sampling intervals are generated. These process music segments are then joined together according to the sampling interval order to obtain the final process music. Based on the set of generation parameter prompts B = {P, R, T, O, E, K, C} for different sampling intervals, and combined with the preset process music generation algorithm model, process music for different sampling intervals is generated. Finally, the process music segments corresponding to each sampling interval are joined together according to the sampling interval order, which can also be understood as splicing, to obtain the final process music.

[0093] S12-5. Based on the pre-trained sound effect matching model, determine the matching sound effect combination, wherein the sound effect combination includes multiple sound effects, and each sound effect in the sound effect combination has a correlation with the interaction event type;

[0094] S12-6. During the continuous and uninterrupted playback of the initial music and the process music, a target sound effect matching the type of the interactive event is embedded and played in the gaps between different process music segments.

[0095] Specifically, in steps S12-3 and S12-4, the preset sampling interval is determined based on the sampling time interval set by the device. The sampling time interval can be flexibly set according to the sound effect accuracy requirements of the actual application scenario, and this application does not impose strict limitations on it. As an example, the preset music generation algorithm may also include a process music generation algorithm G. process (), the sampling time interval can be set to t interval =15s, sampling and analysis of human-computer dialogue is performed at 15-second intervals to obtain the topic type C of the dialogue within the current sampling period. current and emotional tendency E current The theme type and emotional tone are passed as input parameters to the process music generation algorithm G. process (), generate the corresponding background music segment M segment By connecting the process music segments corresponding to different sampling time intervals, background music can be generated in real time in an accompanying manner.

[0096] The specific algorithm implementation process can be understood as follows: at each sampling time t n (n represents the sample number), generated by the process music generation algorithm G process () Based on the topic type C determined within the current sampling interval current Emotional Tendency E current Output the corresponding music segments M within different sampling intervals segment (t n ) = G process (Ccurrent (t n ),E current (t n Then, the various music segments are connected according to the sampling order: M process =∑ n M segment (t n ).

[0097] Further, steps S12-5 are executed to determine the matching sound effect combination based on the pre-trained sound effect matching model, wherein the sound effect combination includes multiple sound effects. Specifically, a sound effect library S including various types of sounds can be created based on the interactive content and the environment corresponding to the current location. These various types of sounds can refer to animal sounds, machine sounds, natural sounds, human sounds, etc., or they can be sound effects specifically designed according to the actual scene. For example, in a war-themed scene in an amusement park, different shooting sound effects can be set for that theme.

[0098] Then, using the sound effects in the sound effect library S as training sample data, the sound effect matching model is trained to obtain a sound effect matching model M that can generate matching effects based on interactive content and the overall environment. match With the help of this sound effect matching model M match Construct a many-to-many matching relationship between each sound effect and the dialogue topic type C and sentiment tendency E. Then, according to step S11, obtain the current dialogue topic type C corresponding to the current input speech signal. current Emotional Tendency E current Input to sound effect matching model M match In the process, a sound effect combination that matches the current dialogue topic type and emotional tone is determined, and this sound effect combination is denoted as S. combination The sound effect combination is a collection of multiple sound effects, namely S. combination ={s1,s2,...,s m}, where m is the number of matched sound effects.

[0099] Furthermore, in steps S12-6, during the playback of background music and background music, since the background music is composed of multiple music segments, the aforementioned sound effect combination S is embedded in the gaps between the different music segments. combination ={s1,s2,...,s m This is to achieve the effect of uninterrupted background music playback and the superposition of sound effects in the sound effect combination with the background music and the background music.

[0100] In some embodiments, the method for generating real-time audio based on dialogue content provided in this application further includes the following steps:

[0101] acquiring the sound intensity of the input voice signal in real time, and if the sound intensity is lower than a preset sound intensity threshold, reducing the sound intensity of the played audio according to a preset sound intensity reduction ratio.

[0102] Specifically, the sound intensity L of the voice of the human-computer dialogue of the participant can be detected in real time by the above method for determining the volume. voice Then, the dialogue volume of the current participant is judged according to a predefined sound intensity threshold L. threshold The predefined sound intensity threshold is the preset sound intensity threshold described above, which can be flexibly set according to different actual application scenarios, and is not strictly limited in the present application.

[0103] If L voice < L threshold The volume of the played audio can be reduced to reduce the sound intensity of the audio. Specifically, the sound intensity reduction ratio can be determined according to the range of L voice < L threshold For example, if it is lower than 10%, the volume of the played audio is reduced by 10%, and if it is lower than 20%, the volume of the played audio is reduced by 20%. The specific sound intensity reduction ratio can also be flexibly set according to actual application scenarios, and is not strictly limited in the present application.

[0104] In this way, the sound intensity of the played music and sound effects can be automatically adjusted to determine that the played sound effects will not interfere with the clarity of the dialogue.

[0105] In some possible embodiments, the method for generating real-time audio based on dialogue content provided by the embodiments of the present application can further include:

[0106] If the interactive task is changed, the audio associated with the current interactive task is terminated, and the audio corresponding to the changed new interactive task is switched and played.

[0107] Specifically, the task switching instruction can be listened for, and if the listening result is that the task exists switching, the audio matched with the label information of the new task, the new topic type and the new emotional tendency is generated and played based on the label information of the switched task and in combination with the topic type and emotional tendency of the input new voice signal.

[0108] Or, when it is received that the user triggers the task switching button or interface, it can be determined that there is a task switching instruction and the task exists switching. At this time, the initial music, the process music and the target sound effect corresponding to the new interactive task can be generated and played based on the label information of the switched interactive task and in combination with the input new voice signal in steps S11 and S12.

[0109] The above process can be specifically implemented by the following algorithm:

[0110] Assume that the mixed audio signal is A(t), at time t, if t is in the background music playing stage and not in the gap of embedded sound effects, let A(t) = M(t); if t is in the gap between music segments, let:

[0111]

[0112] Where M(t) is the background music corresponding to time t, m is the number of sound effects, s i (t) is the signal of the i-th sound effect in the sound effect combination S combination at time t.

[0113] Then, based on the sound intensity adjustment function F adjust , the adjusted sound intensity of the playing music L music = F adjust (L voice , L threshold )*L original_music , where L original_music is the sound intensity of the currently playing music, the adjusted sound intensity of the playing sound effect L effect = F adjust (L voice , L threshold )*L original_effect , where L original_effect is the sound intensity of the currently playing sound effect.

[0114] When the task switches, assume that the background music and sound effect termination function associated with the current task is F terminate (M current , S current ), and the start function of the background music and sound effect corresponding to the new task is F start (M new , S new ), then execute F terminate (M current , S current ) first, and then execute F start (M new , S new ), where M current is the currently playing music, S current is the currently playing sound effect, M new is the music corresponding to the new task, and S new is the sound effect corresponding to the new task.

[0115] By selecting the embodiments of the present application, the overall experience of the background music and sound effects of the human-computer dialogue can be intelligently created, starting from personalized music combined with the environment and the task, generating matching music and sound effects around the human-computer dialogue between the participant and the AI in the process. In addition, sentiment analysis is performed on the participant, and the speech features of the dialogue are analyzed and applied, which can make the generated audio have higher fault tolerance, achieve good user experience, and not easily cause obvious errors and serious problems.

[0116] In the embodiments of the present application, the playing of audio can be played by the execution subject itself, or can be controlled by the execution subject to play by external sound equipment. As an example, after the execution subject determines the corresponding audio in step S12, the audio is played through the sound equipment in the current interactive position in the amusement park.

[0117] Based on the method of the first aspect, the second aspect, the embodiments of the present application provide a system for generating real-time audio based on dialogue content, wherein, as shown in the figure, Figure 2 The system 20 includes:

[0118] The acquisition module 201 is configured to acquire an input voice signal and determine an audio-related parameter corresponding to the voice signal, wherein the audio-related parameter includes a theme type, a sentiment tendency, an interactive position, an interactive time, and an interactive event type; wherein the interactive event type is an interactive task set in advance according to the interactive position.

[0119] The audio generation and playing module 202 is configured to generate and play audio matching the audio-related parameter based on the audio-related parameter. As an embodiment, the system can be subdivided into the following modules as shown in the figure. Figure 3

[0120] The user identity information recognition module 21 is configured to recognize the identity of the participant by performing step S11-1 in the above method.

[0121] The voice-to-text and target participant voice recognition module 22 is configured to convert the interactive voice signal of the participant into text content and perform noise reduction processing by performing steps S11-2 to S11-5 in the above method.

[0122] The dialogue theme analysis and dialogue sentiment tendency analysis module 23 is configured to analyze the dialogue theme and sentiment tendency of the participant's dialogue voice by performing steps S11-6 to S11-8.

[0123] The initial music generation and process music generation module 24 is configured to generate corresponding initial music and process music by performing steps S12-1 to S12-4.

[0124] ​The sound effect synthesis and audio mixing and playing module 25 is configured to perform the steps S12-5 to S12-6, synthesize the sound effect, and perform embedding processing on the sound effect to obtain mixed audio, and finally output and play the mixed audio.

[0125] By selecting the embodiments of the present application, the overall experience of background music and sound effects of human-computer dialogue can be intelligently created, starting from personalized music combined with environment and tasks, generating matching music and sound effects around the human-computer dialogue between participants and AI in the process. In addition, sentiment tendency analysis is performed on the participants, and the speech features of their dialogue are analyzed and applied, which can make the generated audio have higher fault tolerance, achieve good user experience, and not easily cause obvious errors and serious problems. The music and sound effects can be generated around the interactive position, interactive role, interactive time, basic event, and other comprehensive information, and the sound effect generation model is trained in combination with the interactive content and the overall environment.

[0126] In the present application, the collection, storage, use, processing, transmission, provision and disclosure of user personal information comply with relevant laws and regulations and do not violate public order and good customs.

[0127] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present application are only used for illustrative purposes, and are not used to limit the scope of the messages or information.

[0128] In a third aspect, the example embodiments of the present application further provide an electronic device, comprising: at least one processor; and a memory connected with the at least one processor in communication. The memory stores a computer program capable of being executed by the at least one processor, and the computer program is used to make the electronic device execute the method according to the embodiments of the present application when executed by the at least one processor.

[0129] The example embodiments of the present application further provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is used to make the computer execute the method according to the embodiments of the present application when executed by the processor of the computer.

[0130] The example embodiments of the present application further provide a computer program product comprising a computer program, wherein the computer program is used to make the computer execute the method according to the embodiments of the present application when executed by the processor of the computer.

[0131] Reference Figure 4An architectural block diagram of an electronic device 400 that can be a server or a client of the present application, which is an example of a hardware device that can be applied to aspects of the present application, will now be described. The electronic device is intended to represent a wide variety of digital electronic computer devices, such as laptops, desktops, workstations, personal digital assistants, servers, blade servers, mainframes, and other appropriate computers. The electronic device can also represent a wide variety of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices, and other similar computing devices. The components, their connections, and their functions, as described herein, are meant to be examples only, and are not intended to limit the implementations of the present application described and / or claimed herein.

[0132] As shown in FIG. 4, Figure 4 The electronic device 400 includes a computing unit 401 that can perform various appropriate actions and processes in accordance with a computer program stored in a read only memory (ROM 402) or a computer program loaded into a random access memory (RAM 403) from a storage unit 408. In the RAM 403, various programs and data required for the operation of the electronic device 400 can also be stored. The computing unit 401, the ROM 402, and the RAM 403 are connected to each other through a bus 404. An input / output interface (I / O interface 405) is also connected to the bus 404.

[0133] A plurality of components in the electronic device 400 are connected to the I / O interface 405, including an input unit 406, an output unit 407, the storage unit 408, and a communication unit 409. The input unit 406 can be any type of device that can input information to the electronic device 400, and can receive inputted digital or character information, and generate key signal inputs related to user settings and / or function controls of the electronic device. The output unit 407 can be any type of device that can present information, and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 408 can include, but is not limited to, a magnetic disk, an optical disk. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks, and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth™ device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.

[0134] The computing unit 401 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 401 performs various methods and processes described above. For example, in some embodiments, the aforementioned method of generating real-time audio based on conversation content can be implemented as a computer software program tangibly embodied in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 400 via the ROM 402 and / or the communication unit 409. In some embodiments, the computing unit 401 can be configured to perform the aforementioned method of generating real-time audio based on conversation content by any other appropriate means, such as by means of firmware.

[0135] Program code for carrying out methods of the present application can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces a means for implementing the functions / acts specified in the flowcharts and / or block diagrams. The program code can be retrieved from a machine-readable medium or device and executed by a processor or controller to produce a machine for implementing the functions / acts specified in the flowcharts and / or block diagrams.

[0136] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. Machine-readable storage media can include, without limitation, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media can include one or more lines of electrical connections, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical storage devices, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0137] As used in this application, the terms "machine-readable medium" and "computer- readable medium" refer to any computer program product, apparatus and / or device (e.g., magnetic discs, optical disks, memory, Programmable Logic Devices (PLDs)) used to provide machine instructions and / or data to a programmable processor, including a machine-readable medium that receives machine instructions as a machine-readable signal. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0138] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0139] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.

[0140] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other.

Claims

1. A method for generating real-time audio based on dialogue content, characterized in that, The method is applied to triggering devices in amusement theme parks to provide interactive services to visitors. These triggering devices include user-carried interactive devices and fixed interactive devices located at preset interactive positions within the amusement park. The interactive devices are used to interact with users to trigger interactive events. The method includes: Acquire the voice signal input by the user during a human-computer dialogue process collected by a sound pickup device, wherein the human-computer dialogue process is used to perform interactive events pre-set according to the interaction position, and the interactive events include interactive tasks; Determine the audio association parameters corresponding to the speech signal, wherein the audio association parameters include: the topic type and emotional tendency of the dialogue content in the speech signal, and the interaction position, interaction time, and interaction event type of the human-computer dialogue; Real-time audio matching the audio association parameters is generated based on the audio association parameters, and the real-time audio is played to the user during the interactive task. The system monitors for changes in the interactive task. If the interactive task changes, it terminates the real-time audio associated with the current interactive task and returns to the step of acquiring the voice signal input by the user during the human-computer dialogue collected by the sound pickup device, so as to generate and play the real-time audio corresponding to the new interactive task after the change.

2. The method according to claim 1, characterized in that, The real-time audio includes initial music, and the generation of real-time audio that matches the audio association parameters based on the audio association parameters includes: The audio association parameters are input into a preset background music generation algorithm model, which then generates multiple initial music segments. The initial music segment corresponding to the user's identity tag information is determined as the initial music.

3. The method according to claim 2, characterized in that, The real-time audio includes process music; The acquisition of the voice signal input by the user during the human-computer dialogue includes: acquiring the voice signal input by the user during the human-computer dialogue according to a preset sampling interval; The step of determining the audio correlation parameters corresponding to the speech signal includes: determining the audio correlation coefficient corresponding to the speech signal at each sampling interval; The method further includes: generating process music segments corresponding to each sampling interval based on the audio correlation coefficients corresponding to each sampling interval and combining them with a preset process music generation algorithm model; and connecting the process music segments corresponding to each sampling interval in the order of each sampling interval to obtain process music.

4. The method according to claim 3, characterized in that, The real-time audio includes sound effects, and the method further includes: Based on the topic type and sentiment of the dialogue content, a pre-trained sound effect matching model is invoked to determine the matching sound effect combination. The sound effect combination includes multiple sound effects, and each sound effect in the sound effect combination has a correlation with the type of the interaction event. During the continuous, uninterrupted playback of the process music, target sound effects matching the type of the interactive event are embedded and played in the gaps between different process music segments.

5. The method according to claim 1, characterized in that, The method further includes: The sound intensity of the input voice signal is acquired in real time. If the sound intensity is lower than a preset sound intensity threshold, the sound intensity of the real-time audio being played is reduced according to a preset sound intensity reduction ratio.

6. The method according to claim 1, characterized in that, The topic type is determined in the following way: The input speech signal is converted into target text content according to a preset speech-to-text function; Semantic analysis is performed on the target text content to extract related words that are associated with a preset topic type contained in the target text content; Based on the associated words, the topic type with the highest degree of association with the associated words is determined as the topic type corresponding to the speech signal; The emotional tendency is determined in the following ways: The audio features of the speech signal are obtained, including speech rate, volume, and pitch. Based on the reference speech rate, reference volume, and reference pitch, the speech features are normalized to obtain normalized audio coefficients. Based on the audio coefficients and preset sentiment tendency label information, a target sentiment tendency score is calculated, and the sentiment tendency corresponding to the speech signal is determined based on the target sentiment tendency score.

7. A system for generating real-time audio based on dialogue content, characterized in that, The system is used to implement the method according to any one of claims 1 to 6, the system comprising: The acquisition module is used to acquire the voice signal input by the user during the human-computer dialogue process collected by the sound pickup device, wherein the human-computer dialogue process is used to perform interactive events pre-set according to the interaction position, and the interactive events include interactive tasks; and to determine the audio association parameters corresponding to the voice signal, wherein the audio association parameters include: the topic type and emotional tendency of the dialogue content in the voice signal, and the interaction position, interaction time and interaction event type of the human-computer dialogue; An audio generation and playback module is used to generate real-time audio that matches the audio association parameters based on the audio association parameters, and to play the real-time audio to the user during the interactive task.

8. An electronic device, characterized in that, The electronic device includes: a processor and a memory storing a program; wherein the program includes instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-6.

9. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Adding background sound to speech-containing audio data

    CN107464555A

  • Method and system for generating dynamic image in real time based on dialogue content and electronic equipment

    CN119815135A

  • Adding background sound to speech-containing audio data

    US20170352361A1