Broadcast television program production method and system based on artificial intelligence

By using voiceprint embedding vectors in audio processing technology to divide the audio bands and obtain enhanced audio, the problems of inaccurate and incomplete subtitles caused by multiple people speaking at the same time are solved, and more accurate and complete subtitles are achieved, enhancing the audience's impression.

CN119996597AActive Publication Date: 2025-05-13SHANGRAO DAWAN NETWORK TECHNOLOGY CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510214684.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-05-13
Estimated Expiration
2045-02-26

AI Technical Summary

Technical Problem

In variety shows or interview programs, due to problems such as multiple people speaking at the same time, different speaking habits of different people, and omitting speech, the subtitles display is inaccurate and incomplete, which affects the audience's understanding.

Method used

By obtaining the audio bands in the audio file, dividing them into single-sound sub-segment and multi-sound sub-segment based on the voiceprint embedding vector, obtaining enhanced audio for each character in each multi-sound sub-segment, and accurately obtaining the display subtitles of the TV program by identifying and analyzing the structure and narration of each audio band.

Benefits of technology

It improves the processing efficiency of each character's audio in multiple sub-segments, accurately identify the speech content of each character, ensures the accuracy and completeness of subtitles display, and thus enhances the audience's viewing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119996597A_ABST
    Figure CN119996597A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of audio processing, in particular to a radio and television program making method and system based on artificial intelligence. The method comprises: acquiring an audio segment; dividing the audio segment into a single-voice sub-segment and a multi-voice sub-segment based on the voiceprint embedding vector corresponding to the human voice at each moment in the audio segment; based on the intensity corresponding to the voice of each person at each moment in the single-voice sub-segment corresponding to each person in the multi-voice sub-segment under each frequency, acquiring an enhanced audio of the voice of each person in the multi-voice sub-segment, and further acquiring each sentence corresponding to each person; and obtaining display subtitles of the television program according to the structure and the lecture condition of each sentence corresponding to each character. According to the subtitle display method and the subtitle display device, the voice frequency of each person in the multi-voice sub-segments is enhanced, so that the accurate recognition of the speaking content of each person is improved, each sentence corresponding to each person is accurately obtained, then each sentence is analyzed, and the subtitle display accuracy and integrity are effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of audio processing, and in particular to a radio and television program production method and system based on artificial intelligence. Background Art

[0002] The production of radio and television programs is a systematic project, including four stages: pre-planning, production execution, post-production, and broadcast feedback. After recording according to the script, there will be a large amount of original materials. Post-production is to edit these original materials according to the characteristics of the program, retain effective materials and playback logic, and synchronize the sound and picture. Whether it is a variety show, news or interview program, it contains a lot of character speeches. The video needs to edit subtitles according to the content of the character's speech to make the film more complete. Therefore, it is necessary to ensure that the subtitles in the TV program are displayed accurately to improve the audience's viewing experience.

[0003] With the rise of artificial intelligence, computers are able to analyze and process audio and identify the speech content of each person in the audio. However, in actual situations, there are multiple people speaking at the same time in variety shows or interviews, different people have different speaking habits, and speech is omitted. As a result, the subtitles displayed in TV programs by identifying audio are inaccurate and incomplete, affecting the audience's accurate understanding of the speech content of the characters in the TV program. Summary of the invention

[0004] In order to solve the technical problem that the subtitles displayed in a television program are inaccurate and incomplete due to the existence of multiple people speaking at the same time in variety shows or interviews, different people's speaking habits, and omissions of speech, the purpose of the present invention is to provide a radio and television program production method and system based on artificial intelligence, and the technical solution adopted is as follows:

[0005] In a first aspect, an embodiment of the present invention provides a method for producing a radio and television program based on artificial intelligence, the method comprising the following steps:

[0006] Get each audio segment in the audio file;

[0007] Based on the voiceprint embedding vector corresponding to the human voice at each moment in each audio segment, the audio segment is divided into single-sound sub-segments and multi-sound sub-segments;

[0008] Based on the intensities of the human voice at each frequency at each moment in the single sound subsegment corresponding to each character in each multi-sound subsegment, obtaining the reference intensities of the human voice at each frequency at each moment in each multi-sound subsegment, and then obtaining the enhanced audio of the human voice of each character in each multi-sound subsegment;

[0009] Each single sound sub-segment and the enhanced audio are identified to obtain each sentence corresponding to each character; and display subtitles of the television program are obtained according to the structure and narration of each sentence corresponding to each character.

[0010] Furthermore, the method for obtaining the single sound sub-segment and the multi-sound sub-segment is:

[0011] Get the voiceprint embedding vector corresponding to each person appearing in the audio file, and use them as the reference voiceprint embedding vector;

[0012] For any moment in any audio segment, obtain the cosine similarity between the voiceprint embedding vector corresponding to the human voice at that moment and each reference voiceprint embedding vector, and use them as the degree of character matching corresponding to the human voice at that moment;

[0013] When the character matching degree is greater than the character matching degree threshold, the character corresponding to the reference voiceprint embedding vector is used as the matching character of the human voice at that moment;

[0014] When the number of matching characters is greater than 1, this moment is a multi-voice overlapping moment;

[0015] When the number of matching characters is equal to 1, the moment is a single voice moment;

[0016] The local audio segment formed by the continuous overlapping moments of multiple voices in the audio segment is used as a multi-voice sub-segment;

[0017] A local audio segment formed by continuous single voice moments in the audio segment is taken as a single voice sub-segment.

[0018] Furthermore, the method for obtaining the character matching degree threshold is:

[0019] Obtain the cosine similarity of any two reference voiceprint embedding vectors as the first similarity value;

[0020] The largest first similarity value is used as the character matching degree threshold.

[0021] Furthermore, the reference intensity is obtained by:

[0022] For the a-th polyphonic sub-segment, the local polyphonic sub-segments corresponding to the continuous human voices corresponding to the same character in the a-th polyphonic sub-segment are all used as reference sound segments; wherein the characters corresponding to the same reference sound segment at each moment are the same;

[0023] For the t-th moment in the i-th reference sound segment of the a-th multi-sound sub-segment, the intensity of the human voice at the k-th frequency at the t-th moment is used as the target intensity;

[0024] For the p-th person in the ith reference sound segment, the intensity of the human voice at the k-th frequency at each moment in all the single sound sub-segments corresponding to the p-th person is used as the reference intensity;

[0025] Taking the difference between the target intensity and each of the reference intensities as a first difference;

[0026] The result of negatively correlating and normalizing the mean of the first difference is used as the intensity adjustment weight of the voice of the p-th character at the k-th frequency at the t-th moment;

[0027] The product of the target intensity and the intensity adjustment weight is used as the reference intensity corresponding to the voice of the p-th character at the k-th frequency at the t-th moment in the i-th reference sound segment of the a-th multi-sound sub-segment.

[0028] Furthermore, the method for obtaining the display subtitles of a television program is:

[0029] For any character in the audio file, the degree of component deficiency of each sentence of the character is obtained according to the difference in balance factor and the difference in the number of leaf nodes of the syntactic tree corresponding to each sentence of the character and each other sentence of the character;

[0030] When the degree of component deficiency is greater than a preset component deficiency threshold, the corresponding sentence is a component deficiency sentence of the character; wherein a sentence is a complete sentence;

[0031] For any component constituting a sentence and any component-missing sentence, the possibility degree of the lack of the component in the component-missing sentence is obtained according to the lack of the component in the component-missing sentence and the proportion of the sentences containing the component in all the sentences of the character;

[0032] When the probability is greater than a preset probability threshold, the content corresponding to the component in the adjacent sentence before the sentence where the component is missing is used as the content of the component in the sentence where the component is missing, and is supplemented and displayed in brackets in the subtitles;

[0033] For any sentence, the ratio of the corresponding duration of the sentence to the number of words in the sentence is used as the reference time interval between any two adjacent words in the sentence;

[0034] When the time interval between any two adjacent words in the sentence is greater than a preset multiple of the reference time interval, the sentence is displayed separately in the subtitles.

[0035] Furthermore, the method for obtaining the degree of deficiency of the component is:

[0036] For any sentence of the character, obtain the balance factor difference between the sentence and the syntax tree corresponding to each other sentence of the character, and use them as the second difference;

[0037] Obtain the difference in the number of leaf nodes of the syntactic tree corresponding to the sentence and each other sentence of the character, and use them as the third difference;

[0038] The result of adding the accumulated value of the second difference and the accumulated value of the third difference and then normalizing them is used as the degree of component deficiency of the sentence.

[0039] Furthermore, the method for obtaining the possibility degree is:

[0040] The result of negative correlation and normalization of the number of occurrences of the component in the sentence lacking the component is used as the reference lack degree of the component in the sentence lacking the component;

[0041] The ratio of the number of sentences containing the component in all sentences of the character to the number of all sentences of the character is used as the reference weight of the component in the corresponding sentences of the character;

[0042] The product of the reference weight and the reference lack degree is taken as the possible degree of lack of the component in the component-lacking sentence.

[0043] Furthermore, the method for obtaining the audio segment is:

[0044] The audio in the audio file with continuous moments where the sound is greater than the preset decibel level is divided into audio segments.

[0045] Furthermore, the artificial intelligence-based radio and television program production method also includes:

[0046] Each moment in each audio segment is consistent with the moment corresponding to each video frame.

[0047] In the second aspect, another embodiment of the present invention provides a radio and television program production system based on artificial intelligence, the system comprising: a memory, a processor, and a computer program stored in the memory and running on the processor, when the processor executes the computer program, it implements the steps of any one of the above methods.

[0048] The present invention has the following beneficial effects:

[0049] The present invention first divides the audio segment into a single-sound sub-segment and a multi-sound sub-segment based on the voiceprint embedding vector corresponding to the human voice at each moment in each audio segment, thereby improving the efficiency of audio processing; then, based on the intensity of the human voice at each moment in the single-sound sub-segment corresponding to each character in each multi-sound sub-segment at each frequency, obtains the reference intensity of the human voice of each character at each moment in each multi-sound sub-segment, accurately reflects the sound intensity corresponding to the human voice of each character at each moment in the multi-sound sub-segment, effectively reduces the interference between the human voices of different characters in the multi-sound sub-segment, and improves the accuracy of obtaining the content spoken by each character in the multi-sound sub-segment; then, the enhanced audio of the human voice of each character in each multi-sound sub-segment is obtained, so that each sentence corresponding to each character in the audio file is accurately identified; in order to ensure that the subtitles of a TV program are displayed more accurately and completely, the structure and narration of each sentence corresponding to each character are analyzed, so as to avoid the situation where the subtitles are incompletely displayed and the sound and picture are not synchronized due to the missing of sentence components or sentence pauses, so that the displayed subtitles of the TV program are accurately obtained, and the audience's viewing experience is effectively improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] In order to more clearly illustrate the technical solutions and advantages in the embodiments of the present invention or the prior art, the drawings required for use in the embodiments or the prior art descriptions are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0051] Figure 1 A schematic flow chart of a radio and television program production method based on artificial intelligence provided by one embodiment of the present invention;

[0052] Figure 2 A flow chart of a method for obtaining subtitles for a television program provided by an embodiment of the present invention;

[0053] Figure 3 A structural diagram of a radio and television program production system based on artificial intelligence provided by one embodiment of the present invention;

[0054] Figure 4 A schematic diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0055] In order to further explain the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the following is a detailed description of the specific implementation method, structure, features and effects of a radio and television program production method and system based on artificial intelligence proposed by the present invention in combination with the accompanying drawings and preferred embodiments. In the following description, different "one embodiment" or "another embodiment" does not necessarily refer to the same embodiment. In addition, specific features, structures or characteristics in one or more embodiments may be combined in any suitable form.

[0056] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs.

[0057] The following is a detailed description of a specific solution of a radio and television program production method and system based on artificial intelligence provided by the present invention in conjunction with the accompanying drawings.

[0058] Embodiment 1:

[0059] The present invention proposes a method for producing radio and television programs based on artificial intelligence. Figure 1 , which shows a schematic flow chart of a radio and television program production method based on artificial intelligence provided by an embodiment of the present invention, the method comprising the following steps:

[0060] Step S1: Acquire each audio segment in the audio file.

[0061] Specifically, this embodiment takes any TV program as an example for analysis. It should be noted that all subsequent TV programs are this TV program. In the post-production process, after the editor completes the editing of the original material, the video of the TV program is filled with subtitles, and then the audio file corresponding to the TV program is obtained. It is known that the minimum sound that the human ear can hear is 20 decibels, and then this embodiment sets the preset decibel to 20 decibels. The implementer can set the size of the preset decibel according to the actual situation, and it is not limited here. The audio at the continuous moments when the sound in the audio file is greater than the preset decibel is divided into audio segments. Among them, each audio segment contains the human voice of the character.

[0062] It should be noted that, in order to ensure the consistency of the audio and video time axes, this embodiment keeps each moment in each audio segment consistent with the moment corresponding to each video frame.

[0063] Step S2: Based on the voiceprint embedding vector corresponding to the human voice at each moment in each audio segment, the audio segment is divided into single-sound sub-segments and multi-sound sub-segments.

[0064] Specifically, it is known that the speech between characters in the audio file is overlapping and carried out sequentially, and there are situations where multiple people speak at the same time. When the existing technology is used to identify the content of each person speaking when multiple people speak at the same time, it is easy to identify inaccurately because the voices of different characters interfere with each other. It is very easy to accurately identify the audio segment in which only one person speaks by using the existing technology. Therefore, for the local audio segment in which multiple people speak at the same time, it is necessary to enhance the voice of each character to ensure that the subtitles of subsequent TV programs are more accurate. Therefore, this embodiment first determines whether the voice at each moment in each audio segment is a single character or multiple characters based on the voiceprint embedding vector corresponding to the voice at each moment in each audio segment, and then divides the audio segment into a single sound sub-segment and a multi-sound sub-segment. Among them, the single sound sub-segment is an audio segment corresponding to the voice of only one character, and the multi-sound sub-segment is an audio segment corresponding to multiple characters speaking at the same time.

[0065] Preferably, in a method that can be implemented in this embodiment, the method for obtaining single sound sub-segments and multiple sound sub-segments is as follows: It is known that the voiceprint features of each character are different, and then this embodiment first manually obtains the exclusive audio segments corresponding to a preset number of sentences spoken by each character appearing in the TV program. This embodiment sets the preset number to 10, that is, each character speaks 10 complete sentences alone. The implementer can set the size of the preset number according to actual conditions, and it is not limited here; then a pre-trained voiceprint encoder is used to extract the voiceprint embedding vector of the exclusive audio segment of each character appearing in the audio file, and all of them are used as reference voiceprint embedding vectors. At this point, the reference voiceprint embedding vector corresponding to each character is obtained, accurately representing the voiceprint features corresponding to each character. Among them, using a pre-trained voiceprint encoder to extract voiceprint embedding vectors is a well-known technology and will not be repeated;

[0066] For any moment in any audio segment, the pre-trained voiceprint encoder is also used to extract the voiceprint embedding vector corresponding to the human voice at that moment, and then the cosine similarity between the voiceprint embedding vector corresponding to the human voice at that moment and each reference voiceprint embedding vector is obtained, and all of them are used as the degree of character matching corresponding to the human voice at that moment; wherein, the greater the degree of character matching, the more likely the character corresponding to the corresponding reference voiceprint embedding vector is to be the character corresponding to the human voice at that moment, therefore, when the character matching degree is greater than the character matching degree threshold, the character corresponding to the corresponding reference voiceprint embedding vector is used as the matching character of the human voice at that moment. Among them, the method for obtaining the character matching degree threshold is: considering that there is a certain similarity in the voiceprint features between different characters, this embodiment obtains the cosine similarity of any two reference voiceprint embedding vectors as the first similarity value; then the largest first similarity value is used as the character matching degree threshold;

[0067] When the number of matching characters is greater than 1, this moment is a moment of overlapping multi-voices; when the number of matching characters is equal to 1, this moment is a moment of single voice; then the local audio segment formed by the continuous overlapping moments of multi-voices in the audio segment is taken as a multi-voice sub-segment; the local audio segment formed by the continuous single voice moments in the audio segment is taken as a single voice sub-segment.

[0068] At this point, the single-sound sub-segments and multi-sound sub-segments in each audio segment are accurately obtained.

[0069] Step S3: Based on the corresponding intensities of the human voice at each frequency in the single sound sub-segment corresponding to each character in each multi-sound sub-segment at each moment, obtain the corresponding reference intensities of the human voice of each character at each moment in each multi-sound sub-segment, and then obtain the enhanced audio of the human voice of each character in each multi-sound sub-segment.

[0070] Specifically, for a single sound sub-segment, speech recognition can be performed directly, and for a multi-sound sub-segment, when the speech content of each character contained in the multi-sound sub-segment is to be identified, the human voices of other characters are relatively noise, and then the human voices of other characters should be weakened, and the spectrum part of the desired character is retained, so as to obtain the enhanced audio of the desired character. When the difference between the intensity corresponding to the human voice at a certain moment in a multi-sound sub-segment at each frequency and the intensity corresponding to the human voice at each moment in the single sound sub-segment corresponding to a certain character at that moment in the multi-sound sub-segment is greater, it means that the human voice of the character at that moment in the multi-sound sub-segment is more disturbed by the human voices of other characters, and the intensity corresponding to the human voice at that moment at each frequency needs to be adjusted to a greater extent, so as to accurately obtain the intensity corresponding to the human voice of the character at that moment at each frequency, and then accurately obtain the speech content of the character at that moment. Therefore, this embodiment obtains the reference intensity of the human voice of each character at each moment in each multi-sound subsegment at each frequency based on the intensity of the human voice at each frequency in the single sound subsegment corresponding to each character in each multi-sound subsegment, and then obtains the enhanced audio of the human voice of each character in each multi-sound subsegment.

[0071] Preferably, in a method that can be implemented in this embodiment, the method for obtaining the reference intensity is: taking into account that not all of the multiple characters in a multi-sound sub-segment keep speaking continuously, in order to more accurately enhance the audio of each character in the multi-sound sub-segment, and then accurately obtain the speech content of each character in each multi-sound sub-segment, it is necessary to further divide each multi-sound sub-segment, and take the audio segments corresponding to the continuous speech of the same character in each multi-sound sub-segment as new independent audio segments for subsequent analysis. This embodiment takes the ath polyphonic subsegment as an example for analysis. For the ath polyphonic subsegment, the local polyphonic subsegments corresponding to the continuous human voice corresponding to the same character in the ath polyphonic subsegment are all used as reference sound segments; wherein the characters corresponding to each moment in the same reference sound segment are the same; for the tth moment in the ith reference sound segment of the ath polyphonic subsegment, the intensity corresponding to the human voice at the tth moment at the kth frequency is used as the target intensity; for the pth character in the ith reference sound segment, the intensity corresponding to the human voice at each moment in all the single sound subsegments corresponding to the pth character at the kth frequency is used as the reference intensity; it should be noted that if there is no corresponding single sound subsegment for the pth character, the intensity corresponding to the human voice at each moment in the exclusive audio segment corresponding to the pth character at the kth frequency is used as the reference intensity;

[0072] The absolute value of the difference between the target intensity and each reference intensity is taken as the first difference; the larger the first difference is, the greater the interference of the voice of the p-th character at the k-th frequency by the voices of other characters is, and the denoising degree needs to be increased. Therefore, in this embodiment, the mean of the first difference is negatively correlated and normalized as the intensity adjustment weight of the voice of the p-th character at the t-th moment at the k-th frequency; then the product of the target intensity and the intensity adjustment weight is taken as the reference intensity corresponding to the voice of the p-th character at the k-th frequency at the t-th moment in the i-th reference sound segment of the a-th multi-sound sub-segment. It should be noted that in this embodiment, the mean of the first difference is negatively correlated and normalized by exp(-x), where exp is an exponential function with a natural constant as the base, and x represents the mean of the first difference.

[0073] At this point, the reference intensity corresponding to the human voice of each character at each moment in the i-th reference sound segment at each frequency is obtained, and then the enhanced audio corresponding to the human voice of each character in the i-th reference sound segment is obtained; according to the method of obtaining the enhanced audio corresponding to the human voice of each character in the i-th reference sound segment, the enhanced audio corresponding to the human voice of each character in each reference sound segment in the a-th multi-sound sub-segment is obtained, the enhanced audio corresponding to the human voice of each character in the a-th multi-sound sub-segment is determined, and then the enhanced audio corresponding to the human voice of each character in each multi-sound sub-segment is obtained.

[0074] Step S4: Identify each single sound sub-segment and the enhanced audio to obtain each sentence corresponding to each character; obtain the display subtitles of the TV program according to the structure and narration of each sentence corresponding to each character.

[0075] Specifically, each single sound sub-segment and enhanced audio are identified by the existing technology to accurately obtain each sentence corresponding to each character. It is known that subtitles are to synchronously display the speaker's speech content. Due to the different speaking habits of different characters, some characters use inverted sentences, which leads to inaccurate punctuation of the character's speech content in the subtitles. At the same time, some characters omit the said content, resulting in incomplete content of some subtitles, which seriously affects the display effect of subtitles. For example, taking the two sentences "I like this program very much" and "Me too" as examples, the former sentence has the object front, and the correct word order should be "I like this program very much", and the latter sentence omits "I like this program very much", resulting in insufficient sentence integrity of "Me too". In order to make the subtitles more intelligent and barrier-free, sentences with missing components should be supplemented, and inverted sentences can be judged by analyzing the structure of each sentence to determine the integrity of each sentence, and by analyzing the structure of each sentence, it can also be judged whether each sentence has missing components. Therefore, this embodiment analyzes the structure of each sentence of each character to obtain the speaking habits of each character, which is conducive to more reasonable and accurate acquisition of subtitles corresponding to the audio of each character.

[0076] On the other hand, when a character is speaking a certain sentence, in order to create a sense of suspense, there will be a pause in the sentence. If the subtitles are presented during the pause, it will lead to spoilers. Therefore, not every sentence of every character needs to be presented in full. For sentences that create a sense of suspense, it is necessary to first show the part of the sentence before the speaker pauses, and show the complete sentence after the pause rather than the part after that, so as to ensure the complete effect of the subtitles and the effect of the TV program.

[0077] Therefore, this embodiment accurately obtains the display subtitles of the television program according to the structure and narration of each sentence corresponding to each character.

[0078] Preferably, in one possible implementation of this embodiment, the method for obtaining the display subtitles of a television program can be found in Figure 2 , which shows a flow chart of a method for obtaining subtitles of a television program provided by this embodiment, the method comprising the following steps:

[0079] Step S201: For any character in the audio file, the degree of component deficiency of each sentence of the character is obtained according to the difference in balance factor and the difference in the number of leaf nodes of the syntax tree corresponding to each sentence of the character and each other sentence of the character.

[0080] For the sake of clear description, this embodiment takes any character in the audio file as an example for analysis. First, each sentence of the character is converted into text using the speech recognition model Whisper, and then the text corresponding to each sentence is analyzed using the jsSyntaxTree software to draw the syntax tree of each sentence; wherein, using the speech recognition model Whisper to convert each sentence into text and using the jsSyntaxTree software to draw the syntax tree are both well-known technologies and will not be repeated. The balance factor of the syntax tree of each sentence corresponding to the character should be similar, because the speaking habits of the character will definitely not change quickly; considering that some of the words spoken by the character may have some components omitted, there will be obvious differences between the syntax trees corresponding to the omitted sentences and all the sentences spoken by the character, therefore, this embodiment obtains the degree of lack of components of each sentence of the character according to the difference in the balance factor and the difference in the number of leaf nodes of the syntax tree corresponding to each sentence of the character and other sentences of the character. It should be noted that if the number of sentences spoken by the character in the audio file is less than the first preset number, all sentences in the exclusive audio segment of the character are taken as the words spoken by the character in the audio file, wherein each sentence in the exclusive audio segment is a complete sentence. In this embodiment, the first preset number is set to 10, and the implementer can set the size of the first preset number according to actual conditions, which is not limited here. Among them, the greater the degree of component missing, the more likely the corresponding sentence is to lack sentence components.

[0081] Preferably, in a method that can be implemented in this embodiment, the method for obtaining the degree of component lack is: for any sentence of the character, the absolute value of the difference between the balance factor of the syntactic tree corresponding to the sentence and each other sentence of the character is obtained, and all are used as the second difference; the larger the second difference, the more abnormal the sentence is, and the more likely the sentence is incomplete; wherein, the method for obtaining the balance factor of the syntactic tree is a well-known technology and will not be repeated; further obtain the absolute value of the difference between the number of leaf nodes of the syntactic tree corresponding to the sentence and each other sentence of the character, and all are used as the third difference; the larger the third difference, the greater the possibility that the sentence is abnormal, which indirectly indicates that the sentence is more likely to be incomplete; therefore, in this embodiment, the cumulative value of the second difference is added to the cumulative value of the third difference and the result of normalization is used as the degree of component lack of the sentence. It should be noted that in this embodiment, the cumulative value of the second difference and the cumulative value of the third difference are normalized by the norm normalization function. It should be noted that among all the sentences said by the character, sentences with missing components must account for a minority.

[0082] At this point, the degree of missing components in each sentence of each character is obtained.

[0083] Step S202: When the component lacking degree is greater than a preset component lacking degree threshold, the corresponding sentence is a component lacking sentence of the character; wherein a sentence is a complete sentence.

[0084] It is known that the greater the degree of missing components, the more likely the corresponding sentence is to have missing components. Therefore, this embodiment sets the preset component missing degree threshold to 0.7. The implementer can set the value of the preset component missing degree threshold according to the actual situation, which is not limited here. When the degree of missing components is greater than the preset component missing degree threshold, the corresponding sentence is a component missing sentence of the character; wherein a sentence is a complete sentence.

[0085] At this point, the components of each character are accurately obtained.

[0086] Step S203: For any component constituting a sentence and any component-missing sentence, the possibility of the component-missing sentence missing the component is obtained according to the situation of the component-missing sentence missing the component and the proportion of sentences containing the component in all sentences of the character.

[0087] In order to determine the missing components of each component-deficient sentence of the character, and then for any component constituting a sentence and any component-deficient sentence of the character, the more sentences containing the component in the sentences corresponding to the character, the more likely the component is to be a component that should exist in every sentence of the character. When the component is missing in the component-deficient sentence, the more the component needs to be supplemented in the component-deficient sentence. Therefore, this embodiment obtains the possibility of the component-deficient sentence lacking the component based on the situation of the component-deficient sentence lacking the component and the proportion of the sentences containing the component in all the sentences of the character. The greater the possibility, the more likely the component-deficient sentence is to lack the content corresponding to the component.

[0088] Preferably, in a method that can be implemented in this embodiment, the method for obtaining the degree of possibility is: the number of times the component appears in the sentence lacking the component is negatively correlated and normalized, as the reference degree of lack of the component in the sentence lacking the component; wherein, the number of times the component appears in the sentence lacking the component must be 0 or 1, therefore, the greater the reference degree of lack, the less likely the component is to exist in the sentence lacking the component. It should be noted that in this embodiment, exp(-X) is negatively correlated and normalized with the number of times the component appears in the sentence lacking the component, wherein exp is an exponential function with a natural constant as the base, and X represents the number of times the component appears in the sentence lacking the component. In order to analyze whether the content of the component needs to be supplemented when the component is missing in the sentence lacking the component, the ratio of the number of sentences containing the component in all the sentences of the character to the number of all the sentences of the character is obtained as the reference weight of the component in the corresponding sentence of the character; wherein, the greater the reference weight, the more important the component is in the sentence of the character, and if the component is missing in the sentence lacking the component, the more the content of the component needs to be supplemented. In order to accurately obtain the possibility that the content of the component missing in the component-missing sentence needs to be supplemented, the product of the reference weight and the reference lack degree is taken as the possible degree of the lack of the component in the component-missing sentence.

[0089] Step S204: When the probability is greater than a preset probability threshold, the content corresponding to the component in the previous adjacent sentence of the sentence where the component is missing is used as the content of the component in the sentence where the component is missing, and is supplemented and displayed in brackets in the subtitles.

[0090] It is known that the greater the degree of possibility, the more the component in the sentence where the component is missing needs to be supplemented, and thus this embodiment sets the preset degree of possibility threshold to 0.6. The implementer can set the size of the preset degree of possibility threshold according to the actual situation, and it is not limited here. When the degree of possibility is greater than the preset degree of possibility threshold, the component is the missing component of the sentence where the component is missing and the content of the component in the sentence where the component is missing needs to be supplemented; it is known that the conversations between the characters on the timeline are closely connected, and thus this embodiment uses the content corresponding to the component in the previous adjacent sentence of the sentence where the component is missing as the content of the component in the sentence where the component is missing, and displays it in brackets in the subtitles. For example, the subtitles are displayed as "I too (like this show)". It should be noted that if the number of adjacent sentences before the sentence where the component is missing is not unique, the content corresponding to the component in the sentence with the smallest degree of component missing will be used as the corresponding supplementary content of the component in the sentence where the component is missing.

[0091] At this point, the components of each character are accurately obtained and the content that needs to be supplemented in the missing components in the missing sentence is missing.

[0092] Step S205: For any sentence, the ratio of the corresponding duration of the sentence to the number of words contained in the sentence is used as the reference time interval between any two adjacent words in the sentence; when the time interval between any two adjacent words in the sentence is greater than a preset multiple of the reference time interval, the sentence is displayed separately in the subtitles.

[0093] After obtaining all the complete sentences of each character, it is also necessary to adjust the filling of the subtitles according to the pause of the speaker. For any sentence, the ratio of the corresponding duration of the sentence to the number of words contained in the sentence is used as the reference time interval between any two adjacent words in the sentence; when the sentence is a sentence that needs to have a sense of suspense, there will be a pause in the process of telling the sentence. Therefore, when the time interval between any two adjacent words in the sentence is greater than the preset multiple of the reference time interval, for the effect of the subtitles, the subtitles of the sentence should be filled separately, that is, the half sentence before the pause is first filled in the video of the corresponding time, and for the sake of completeness, when filling the second half of the sentence, the complete sentence should be filled instead of only the content of the second half of the sentence after the sentence is broken. For example, taking "My favorite program is XXX" as an example, the half sentence before the pause is displayed in the subtitles as "My favorite program is" and the second half of the sentence is filled as "My favorite program is XXX". Among them, the preset multiple is set to 2 times in this embodiment, and the implementer can set the size of the preset multiple according to the actual situation, which is not limited here.

[0094] At this point, the subtitles corresponding to each sentence of each character are accurately obtained, and then all the subtitle sentences are filled into the preset position of the corresponding video through the editing software. Usually, the preset position is set to the center position below the video, which is not limited here, and the subtitle production is accurately completed, which effectively improves the accuracy and completeness of the subtitle display, and finally accurately obtains the finished TV program.

[0095] In summary, the present embodiment obtains an audio segment; based on the voiceprint embedding vector corresponding to the human voice at each moment in the audio segment, the audio segment is divided into a single-sound sub-segment and a multi-sound sub-segment; based on the corresponding intensity of the human voice at each moment in the single-sound sub-segment corresponding to each character in the multi-sound sub-segment at each frequency, the enhanced audio of the human voice of each character in the multi-sound sub-segment is obtained, and then each sentence corresponding to each character is obtained; according to the structure and narration of each sentence corresponding to each character, the display subtitles of the TV program are obtained. The present invention enhances the audio of each character in the multi-sound sub-segment, improves the accuracy of identifying the speech content of each character, and then accurately obtains each sentence corresponding to each character, and then analyzes each sentence, effectively improving the accuracy and completeness of the subtitle display.

[0096] Embodiment 2:

[0097] The present invention also proposes a radio and television program production system based on artificial intelligence, see Figure 3 , which shows a structural diagram of a radio and television program production system based on artificial intelligence provided by an embodiment of the present invention, the system includes: an audio segment acquisition module 10, an audio division module 20, an enhanced audio acquisition module 30 and a display subtitle acquisition module 40.

[0098] The audio segment acquisition module 10 is used to acquire each audio segment in the audio file.

[0099] The audio segmentation module 20 is used to divide the audio segment into single-sound sub-segments and multi-sound sub-segments based on the voiceprint embedding vector corresponding to the human voice at each moment in each audio segment.

[0100] The enhanced audio acquisition module 30 is used to obtain the reference intensity of the human voice of each character at each moment in each multi-sound subsegment at each frequency based on the intensity of the human voice at each frequency in the single sound subsegment corresponding to each character in each multi-sound subsegment, and then obtain the enhanced audio of the human voice of each character in each multi-sound subsegment.

[0101] The display subtitle acquisition module 40 is used to identify each single sound sub-segment and the enhanced audio to obtain each sentence corresponding to each character; and obtain the display subtitles of the TV program according to the structure and narration of each sentence corresponding to each character.

[0102] It should be noted that: the system provided in the above embodiment is only illustrated by the division of the above functional modules. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the above embodiment provides a radio and television program production system based on artificial intelligence and an embodiment of a radio and television program production method based on artificial intelligence, which belongs to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0103] Embodiment 3:

[0104] The present invention also proposes a radio and television program production device based on artificial intelligence, the device includes a memory and a processor, wherein the memory stores an executable program code, and the processor is used to call and execute the executable program code to execute a radio and television program production method based on artificial intelligence provided in an embodiment of the present application. The device can be a chip, a component or a module, and the chip may include a connected processor and a memory; wherein the memory is used to store instructions, and when the processor calls and executes the instructions, the chip can execute a radio and television program production method based on artificial intelligence provided in the above embodiment.

[0105] In addition, the present application embodiment also protects a computer device, see Figure 4 The computer device includes a memory 401, a processor 402, and a computer program 403 stored in the memory 401 and running on the processor 402, wherein when the processor 402 executes the computer program 403, the computer device can execute any one of the artificial intelligence-based radio and television program production methods introduced above.

[0106] Embodiment 4:

[0107] This embodiment also provides a computer-readable storage medium, in which a computer program code is stored. When the computer program code is run on a computer, the computer executes the above-mentioned related method steps to implement an artificial intelligence-based radio and television program production method provided in the above embodiment.

[0108] Embodiment 5:

[0109] This embodiment also provides a computer program product. When the computer program product runs on a computer, it enables the computer to execute the above-mentioned related steps to implement a radio and television program production method based on artificial intelligence provided by the above embodiment.

[0110] Among them, the device, computer-readable storage medium, computer program product or chip provided in this embodiment is used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be repeated here.

[0111] It should be noted that the sequence of the above embodiments of the present invention is only for description and does not represent the advantages and disadvantages of the embodiments. The processes depicted in the accompanying drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0112] The various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referenced to each other, and each embodiment focuses on the differences from other embodiments.

Claims

1. A method for producing radio and television programs based on artificial intelligence, characterized in that: The method comprises the following steps: Get each audio segment in the audio file; Based on the voiceprint embedding vector corresponding to the human voice at each moment in each audio segment, the audio segment is divided into single-sound sub-segments and multi-sound sub-segments; Based on the intensities of the human voice at each frequency at each moment in the single sound subsegment corresponding to each character in each multi-sound subsegment, obtaining the reference intensities of the human voice at each frequency at each moment in each multi-sound subsegment, and then obtaining the enhanced audio of the human voice of each character in each multi-sound subsegment; Each single sound sub-segment and the enhanced audio are identified to obtain each sentence corresponding to each character; and display subtitles of the television program are obtained according to the structure and narration of each sentence corresponding to each character.

2. A method for producing radio and television programs based on artificial intelligence as claimed in claim 1, characterized in that: The method for obtaining the single sound sub-segment and the multi-sound sub-segment is: Get the voiceprint embedding vector corresponding to each person appearing in the audio file, and use them as the reference voiceprint embedding vector; For any moment in any audio segment, obtain the cosine similarity between the voiceprint embedding vector corresponding to the human voice at that moment and each reference voiceprint embedding vector, and use them as the degree of character matching corresponding to the human voice at that moment; When the character matching degree is greater than the character matching degree threshold, the character corresponding to the reference voiceprint embedding vector is used as the matching character of the human voice at that moment; When the number of matching characters is greater than 1, this moment is a multi-voice overlapping moment; When the number of matching characters is equal to 1, the moment is a single voice moment; The local audio segment formed by the continuous overlapping moments of multiple voices in the audio segment is used as a multi-voice sub-segment; A local audio segment formed by continuous single voice moments in the audio segment is taken as a single voice sub-segment.

3. A method for producing radio and television programs based on artificial intelligence as claimed in claim 2, characterized in that: The method for obtaining the character matching degree threshold is as follows: Obtain the cosine similarity of any two reference voiceprint embedding vectors as the first similarity value; The largest first similarity value is used as the character matching degree threshold.

4. The method for producing a radio and television program based on artificial intelligence according to claim 1, characterized in that: The method for obtaining the reference intensity is: For the a-th polyphonic sub-segment, the local polyphonic sub-segments corresponding to the continuous human voices corresponding to the same character in the a-th polyphonic sub-segment are all used as reference sound segments; wherein the characters corresponding to the same reference sound segment at each moment are the same; For the t-th moment in the i-th reference sound segment of the a-th multi-sound sub-segment, the intensity of the human voice at the k-th frequency at the t-th moment is used as the target intensity; For the p-th person in the ith reference sound segment, the intensity of the human voice at the k-th frequency at each moment in all the single sound sub-segments corresponding to the p-th person is used as the reference intensity; Taking the difference between the target intensity and each of the reference intensities as a first difference; The result of negatively correlating and normalizing the mean of the first difference is used as the intensity adjustment weight of the voice of the p-th character at the k-th frequency at the t-th moment; The product of the target intensity and the intensity adjustment weight is used as the reference intensity corresponding to the voice of the p-th character at the k-th frequency at the t-th moment in the i-th reference sound segment of the a-th multi-sound sub-segment.

5. The method for producing a radio and television program based on artificial intelligence according to claim 1, characterized in that: The method for obtaining the display subtitles of a television program is: For any character in the audio file, the degree of component deficiency of each sentence of the character is obtained according to the difference in balance factor and the difference in the number of leaf nodes of the syntactic tree corresponding to each sentence of the character and each other sentence of the character; When the degree of component deficiency is greater than a preset component deficiency threshold, the corresponding sentence is a component deficiency sentence of the character; wherein a sentence is a complete sentence; For any component constituting a sentence and any component-missing sentence, the possibility degree of the lack of the component in the component-missing sentence is obtained according to the lack of the component in the component-missing sentence and the proportion of the sentences containing the component in all the sentences of the character; When the probability is greater than a preset probability threshold, the content corresponding to the component in the adjacent sentence before the sentence where the component is missing is used as the content of the component in the sentence where the component is missing, and is supplemented and displayed in brackets in the subtitles; For any sentence, the ratio of the corresponding duration of the sentence to the number of words in the sentence is used as the reference time interval between any two adjacent words in the sentence; When the time interval between any two adjacent words in the sentence is greater than a preset multiple of the reference time interval, the sentence is displayed separately in the subtitles.

6. A method for producing radio and television programs based on artificial intelligence as claimed in claim 5, characterized in that: The method for obtaining the degree of deficiency of the component is as follows: For any sentence of the character, obtain the balance factor difference between the sentence and the syntax tree corresponding to each other sentence of the character, and use them as the second difference; Obtain the difference in the number of leaf nodes of the syntax tree corresponding to the sentence and each other sentence of the character, and use them as the third difference; The cumulative value of the second difference and the cumulative value of the third difference are added together and then normalized, and the result is used as the degree of component deficiency of the sentence.

7. A method for producing radio and television programs based on artificial intelligence as claimed in claim 5, characterized in that: The method for obtaining the possible degree is: The result of negative correlation and normalization of the number of occurrences of the component in the sentence lacking the component is used as the reference lack degree of the component in the sentence lacking the component; The ratio of the number of sentences containing the component in all sentences of the character to the number of all sentences of the character is used as the reference weight of the component in the corresponding sentence of the character; The product of the reference weight and the reference lack degree is taken as the possible degree of lack of the component in the component-lacking sentence.

8. The method for producing a radio and television program based on artificial intelligence according to claim 1, characterized in that: The method for obtaining the audio segment is: The audio in the audio file with continuous moments where the sound is greater than the preset decibel level is divided into audio segments.

9. The method for producing a radio and television program based on artificial intelligence according to claim 1, characterized in that: The artificial intelligence-based radio and television program production method also includes: Each moment in each audio segment is consistent with the moment corresponding to each video frame.

10. A radio and television program production system based on artificial intelligence, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When executing the computer program, the processor implements the steps of the radio and television program production method based on artificial intelligence as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Sentence processing method and device and computer readable storage medium

    CN109325234A

  • Voice translation method and device, computer readable medium and electronic equipment

    CN113299309A

  • Audio data enhancement method and related equipment

    CN113611318A

  • Voice segmentation method and device, computer equipment and medium

    CN118155648A

  • Diversity controllable text rewriting method and device based on sub-tree bank

    CN118211574A