Text-to-video conversion method, device, electronic device, and readable storage medium
By generating pictures and audio data for text and converting them into video files, the problem of time-consuming and laborious reading of users is solved, and fast and efficient information acquisition is achieved.
Patent Information
- Application Number
- CN202210572962.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-24
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2042-05-24
AI Technical Summary
In the prior art, users can read text in a single way and time-consuming manner, and cannot quickly and efficiently obtain information.
Split the pending text into multiple sentences, and generate corresponding picture data and audio data for each sentence, display it in video form, and generate the target video file.
It increases the way users obtain text information, reduces the difficulty of reading, and improves the user experience.
Smart Images

Figure CN115034181B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video technology, and in particular to a method, device, electronic device, and readable storage medium for converting text to video. Background Art
[0002] Reading is an important way for people to obtain information. As China's economy rapidly develops, the demand for reading is growing across all social classes. The results of the 14th National Reading Survey, conducted by the China Institute of Publishing Research, were released on April 18th. Data show that digital reading rates among Chinese adults have increased significantly over the past year. The average daily mobile phone usage time for adults reached 74.4 minutes, a year-on-year increase of 19.6%. The report shows that in 2016, online and mobile phone reading rates among Chinese adults increased, while other digital reading methods decreased. Since most articles are presented in text format, and users have limited and time-consuming methods for reading text, there is a pressing need to develop methods that reduce the difficulty of accessing text information. Summary of the Invention
[0003] The purpose of the embodiments of the present application is to provide a method, device, electronic device, and readable storage medium for converting text to video, so as to solve the problem in the prior art that users have only a single way to read text and that reading text is time-consuming and laborious. The specific technical solution is as follows:
[0004] In the first aspect of the implementation of the present application, a method for converting text to video is first provided, including: generating corresponding image data and audio data for the text to be processed, the image data being used to describe the text to be processed through images, and the audio data being used to describe the text to be processed through audio; generating a target video file of the text to be processed based on the image data and the audio data.
[0005] In the second aspect of the implementation of the present application, a device for converting text to video is also provided, including: a generation module for generating corresponding image data and audio data for the text to be processed, the image data being used to describe the text to be processed through images, and the audio data being used to describe the text to be processed through audio; a writing module for generating a target video file of the text to be processed based on the image data and the audio data.
[0006] In the third aspect of the implementation of the present application, an electronic device is also provided, including: a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus, and the memory is used to store computer programs; the processor is used to implement the steps of the text-to-video method as described above when executing the program stored in the memory.
[0007] In another aspect of the present application, a computer-readable storage medium is provided, wherein instructions are stored in the computer-readable storage medium. When the computer-readable storage medium is executed on a computer, the computer executes the steps of any of the above-mentioned text-to-video conversion methods.
[0008] In an embodiment of the present application, the text to be processed is split into multiple sentences, and corresponding image data and audio data are generated for each sentence, and the audio data is converted into audio frames and written into a video file, and the image data is converted into video frames and written into a video file according to the audio frames in the video file, so as to convert the text into a corresponding video file, and allow users to obtain information in the form of video, thereby increasing the ways for users to obtain text information. At the same time, the video form reduces the difficulty of reading for users, solves the problem that users are time-consuming and labor-intensive in reading texts and cannot quickly and efficiently obtain text information, and improves the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art.
[0010] Figure 1 This is a basic flow chart of the method for converting text to video in an embodiment of the present application;
[0011] Figure 2 This is a basic flow chart of an optional text-to-video method in an embodiment of the present application;
[0012] Figure 3 This is a basic flow chart of an optional text-to-video method in an embodiment of the present application;
[0013] Figure 4 This is a basic structural diagram of an optional text-to-video device in an embodiment of the present application;
[0014] Figure 5 This is a schematic diagram of the structure of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0015] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0016] In order to solve the problems existing in the prior art, the present invention provides a method for converting text to video. Figure 1 As shown, the text-to-video method includes but is not limited to:
[0017] S101, generating corresponding image data and audio data for a text to be processed, wherein the image data is used to describe the text to be processed through images, and the audio data is used to describe the text to be processed through audio;
[0018] S102, generating a target video file of the text to be processed according to the image data and the audio data;
[0019] The text-to-video method provided in this embodiment generates corresponding image data and audio data for the text to be processed, wherein the image data is used to describe the text to be processed through images, and the audio data is used to describe the text to be processed through audio; a target video file of the text to be processed is generated based on the image data and the audio data; wherein the text is converted into a corresponding target video file by generating audio data and image data for the text to be processed, and then generating a target video file based on the audio data and image data, allowing users to obtain information in the form of video, thereby increasing the ways in which users obtain text information. At the same time, the video form reduces the difficulty of reading for users, solves the problem that users are time-consuming and labor-intensive in reading text and cannot quickly and efficiently obtain text information, and improves the user experience.
[0020] It should be understood that the text-to-video method provided in this embodiment can be applied to terminals, including but not limited to mobile terminals such as mobile phones, tablet computers, laptop computers, PDAs, portable media players (PMPs), and navigation devices, as well as fixed terminals such as digital TVs and desktop computers. The subsequent description will use mobile terminals as an example, but the embodiments of the present invention can also be applied to fixed terminals.
[0021] It should be understood that the text to be processed can be data obtained by the mobile terminal itself, for example, the text to be processed is text input by the user through the acquisition interface of the mobile terminal, or the text to be processed is text obtained by the mobile terminal from other locally installed applications according to instructions; the text to be processed can also be text obtained by the mobile terminal from an external device, for example, the text to be processed is text obtained by the mobile terminal from a server.
[0022] In some examples of this embodiment, generating a video file of the text to be processed based on the image data and the audio data includes: writing the audio data into a video file and obtaining the duration of the audio data; generating video frames of the text to be processed based on the duration of the audio data in the video file and the image data, and writing the video frames of the text to be processed into the video file to obtain the target video file; in some examples, the above-mentioned video file can be a video file created before writing the audio data, and the video file does not carry audio data or video frames. It should be understood that generating corresponding audio data for the text to be processed can be achieved by using speech synthesis technology; wherein the speech synthesis technology includes but is not limited to at least one of the following: waveform splicing technology, pitch synchronization superposition technology, LMA vocal tract model technology, or other speech synthesis models, etc., obtaining the audio of the text to be processed through speech synthesis technology can generate audio with multiple timbres, thereby improving the diversity of audio data, and the parameters of the generated audio data are not restricted, for example, the sampling rate, bit rate and other parameters of the audio data can be flexibly set by relevant personnel.
[0023] It should be understood that, in some examples, obtaining the duration of the audio data includes: directly obtaining the duration of the audio data generated by the above-mentioned speech synthesis technology. For example, if the duration of the audio data generated by the speech synthesis technology is n, then the obtained audio duration is n; in some examples, obtaining the duration of the audio data includes: obtaining the duration of the audio data in the video file; it can be understood that after the generated audio data is written into the video file, the duration of the audio data may increase or decrease due to adaptation to the encoding of the video file, and at this time it is necessary to obtain the duration of the audio data in the video file.
[0024] Continuing with the above example, after generating audio data for the text to be processed and writing the audio data into a video file, the duration of the audio data is obtained, and then based on the duration of the audio data and the image data, the video frame of the text to be processed is generated. The total duration of the generated video frame is equal to the duration of the audio data in the video file, and then the video frame of the text to be processed is written into the video file to obtain the target video file.
[0025] In some examples of this embodiment, based on the duration of the audio data and the image data, the video frames of the text to be processed are generated, including: based on the acquired duration of the audio data and the duration of each video frame in the video file, calculating the number of video frames corresponding to the text to be processed; based on the number of video frames, converting the image data into video frames of the text to be processed; wherein, the duration of the video frame is related to the frame rate of the video file, if the frame rate of the video file is 60HZ, the duration of a video frame is one-sixtieth of a second, if the frame rate of the video file is 30HZ, the duration of a video frame is one-thirtieth of a second, according to the acquired duration of the audio data in the video file and the duration of a video frame, the required number of video frames can be calculated, and the image data is converted into the video frames of the text to be processed according to the number of video frames; it should be understood that the frame rate of the video file is related to the video encoder, which will be explained later and will not be repeated here.
[0026] Continuing with the above example, for example, the frame rate of the video file is 60HZ, and the audio data of the text to be processed in the video file is 2 minutes long, then it can be calculated that the required number of video frames is 120, and the image data is converted into a video frame of 120 frames.
[0027] In some examples of this embodiment, generating corresponding image data for the text to be processed includes: setting a background view for the text to be processed, creating a text control on the background view, the text control being used to display the text to be processed; and generating image data corresponding to the text to be processed based on the background view and the text control. The background view is provided by a view control, and the image data corresponding to the text to be processed is generated by adding a view control A and a text control B. Specifically, for example, B is added to A. The font, font size, font color, and position of control B are set, and the text content of control B is set to the text to be processed.
[0028] In some examples of this embodiment, before generating corresponding image data and audio data for the text to be processed, the method further includes: obtaining punctuation marks of the text to be processed; segmenting the text to be processed according to the punctuation marks to generate multiple sentences; generating corresponding image data and audio data for the text to be processed, including: generating corresponding image data and audio data for each sentence; specifically, obtaining segmentation points based on punctuation marks such as commas, periods, exclamation marks, and ellipsis marks contained in the text to be processed, and segmenting the text to be processed based on the segmentation points to obtain multiple sentences. It should be understood that in some examples, dividing the text to be processed into multiple sentences may also include: segmenting the text to be processed according to the semantics of the text to be processed to generate multiple sentences; specifically, segmenting the text to be processed according to the segmentation points obtained based on the semantics of the text to be processed using a semantic segmentation model to obtain multiple sentences. It should be understood that in some examples, the text to be processed can be segmented based on both the punctuation marks and the semantics contained in the text to be processed to obtain multiple sentences.
[0029] In some examples of this embodiment, after the text to be processed is segmented according to the punctuation marks and multiple sentences are generated, the following further includes: when there are sentences with empty content among the multiple sentences, the sentences with empty content are deleted; it should be understood that if there are continuous punctuation marks in the text to be processed, when the text to be processed is segmented and multiple sentences are generated, there may be sentences with empty content. Therefore, it is necessary to filter out the sentences with empty content, that is, delete the sentences with empty content, so that the video converted according to the text to be generated is more accurate, avoiding the problem of converting empty sentences into videos, and also making the image data and audio data generated using such sentences clean and tidy.
[0030] Continuing with the above example, when the text to be processed is divided into multiple sentences, generating corresponding image data and audio data for the text to be processed includes: directly generating the corresponding image data and audio data for each of the sentences; specifically, after generating multiple sentences, performing speech synthesis on each sentence at the same time through speech synthesis technology, synthesizing the audio data of each sentence, so as to achieve the effect of generating audio data for the text to be processed. The steps of generating audio data for each sentence are consistent with the above steps of generating audio data for the text to be processed, and will not be repeated here. Through the image generation technology, an image is generated for each sentence, and the image data of each sentence is generated, that is, after the image data and audio data of all sentences are generated at the same time, the subsequent steps are executed.
[0031] It should be understood that when the mobile terminal generates audio data and picture data for a sentence, it will occupy the memory of the mobile terminal, and generating picture data and audio data for each sentence at the same time will occupy too much memory of the mobile terminal, causing the mobile terminal to freeze. Therefore, in some examples, generating corresponding picture data and audio data for each sentence can be: after generating multiple sentences, according to the order of sentence generation, each sentence is voice synthesized in turn through speech synthesis technology, and the audio data of each sentence is synthesized; according to the order of sentence generation, each sentence is image generated and the picture data of each sentence is generated, that is, after the picture data and audio data of all sentences are generated in sequence, the subsequent steps are executed; it should be understood that the picture data and audio data for each sentence in sequence may also occupy too much memory of the mobile terminal, causing the mobile terminal to freeze.
[0032] Therefore, in order to avoid occupying too much memory of the mobile terminal, in some examples of this embodiment, generating corresponding image data and audio data for each of the sentences includes: obtaining the first generated sentence according to the order in which the sentences are generated in the text to be processed, generating the corresponding image data and audio data for the first generated sentence; and after the image data and audio data of the first sentence are written into the video file, generating the image data and audio data corresponding to the next sentence of the first generated sentence. For example, after the text to be processed is divided into multiple sentences, image data and audio data are generated for the first sentence according to the order in which the sentences are generated, and written into the video file. After the audio data and image data of the sentence are written into the video file and the encoding is completed, audio data and image data are generated for the next sentence, and then the above steps are repeated until the image data and audio data of all sentences are written into the video file.
[0033] It should be understood that in some examples, after the text to be processed is divided into multiple sentences, image data is generated for all sentences, and audio data is generated for the first sentence, and the image data and audio data of the first sentence are written to the video file. When the audio data of the first sentence is written to the video file and the encoding is completed, audio data is generated for the next sentence, and then the above steps are repeated until the image data and audio data of all sentences are written to the video file. For example, an audio write semaphore and a video write semaphore are created, the initial value of the audio write semaphore is 1, representing an unlocked state, and the initial value of the video write semaphore is 0, representing a locked state. When the first sentence is obtained, the value of the audio write semaphore is subtracted by 1, so that the value of the audio write semaphore is changed to 0, and it becomes a locked state. Then, audio data is generated for the sentence, and the audio data of the sentence is written into the video file, and the value of the audio write semaphore is changed to 1, and it becomes an unlocked state. The next sentence is obtained, and audio data is generated for the next sentence. The above steps are repeated until the audio data of all sentences are written into the video file. After the audio data of the first sentence is written into the video file, the value of the video write semaphore is added by 1, so that the value of the video write semaphore is changed to 1, and it becomes an unlocked state. At this time, the picture data of the sentence is obtained, and the value of the video write semaphore is changed to 0, and it becomes a locked state. Then, the picture data of the sentence is written into the video file accordingly. When the audio data of the next sentence is written into the video file, the above steps are repeated until the audio data of all sentences are written into the video file.
[0034] In some examples of this embodiment, before generating the target video file of the text to be processed based on the image data and the audio data, the method further includes: obtaining a video encoder for the video file, the video encoder including: a video frame input source for controlling the video frame input, and an audio frame input source for controlling the audio frame input; the video frame input source is used to write the video frames of the text to be processed into the video file; the audio frame input source is used to write the audio data into the video file. It should be understood that the video encoder includes: a video frame input source for controlling the video frame input, and an audio frame input source for controlling the audio data input. The parameters of the video encoder include bit rate, frame rate, encoding level, encoding format, and video width and height.
[0035] In some examples of this embodiment, writing the audio data into the video file includes: converting the audio data into the audio frames according to the parameters of the video encoder; and writing the audio frames into the video file through the audio frame input source. Specifically, after obtaining the audio data corresponding to the sentence, converting the audio data into audio frames according to the audio frame encoding parameters of the video encoder, and then writing the audio frames corresponding to the sentence into the video file through the audio frame input source.
[0036] In some examples of this embodiment, converting the image data into video frames corresponding to the number of video frames, and writing video frames matching the audio frames of the sentence into the video file includes: converting the image data into initial video frames according to the parameters of the video encoder, and converting the initial video frames into the video frames corresponding to the audio data according to the number of video frames; and writing the video frames into the video file through the video frame input source. Specifically, taking the example of dividing the text to be processed into sentences and generating image data for each sentence, after obtaining the image data corresponding to the sentence, converting the image data into initial video frames according to the parameters of the video encoder, obtaining the duration of the audio data of the sentence in the video file, and calculating the number of video frames required for the sentence based on the duration and the frame rate of the video frame encoding parameters, thereby writing the video frames corresponding to the sentence into the video file through the video frame input source, the video frames being the initial video frames of the required number of video frames; wherein, using the audio duration of the sentence to calculate the number of video frames corresponding to the sentence efficiently and conveniently controls the display problem of the video frames, while solving the problem of video duration, and improving the efficiency of text-to-video conversion. For example, after obtaining the image data corresponding to the sentence and converting the image data into initial video frames according to the parameters of the video encoder, the duration of the audio frame of the sentence in the video file is obtained to be 2.238 seconds, and the frame rate of the video frame encoding parameters of the video file is 30 frames per second. It is calculated that the number of video frames required for the sentence is 68 frames. Therefore, 68 frames of initial video frames corresponding to the sentence are written into the video file through the video frame input source, thereby completing the writing of the video frames of the sentence.
[0037] In order to better understand the present invention, the present invention is described below with reference to the specific implementation of the embodiment of the present invention; a method for converting text to video is provided in the specific implementation, such as Figure 2 As shown, the text-to-video method includes:
[0038] S201, dividing the text to be processed into multiple sentences;
[0039] In some examples of the present embodiment, punctuation marks are first used to separate the text into several sentences. Among the several sentences, sentences with empty content are filtered out (if there are continuous punctuation marks in the text, there will be a situation where the sentence content is empty).
[0040] S202, generating corresponding image data for each sentence;
[0041] In some examples of this embodiment, image generation technology is used to generate corresponding image data for each sentence. For example, image data corresponding to a sentence is generated by adding a view control A and a text control B. Specifically, B is added to A. The font, size, color, and position of control B are set. The text content of control B is set to a single sentence after the text to be processed is segmented. The content of control A is then written to an image and saved to memory. This process is repeated until corresponding image data is generated for each segmented sentence.
[0042] S203: Create a video file and create a video encoder for the video file;
[0043] In some examples of this embodiment, the mobile terminal creates an empty video file and the video encoder assetWriter corresponding to the video file, and sets the video encoder parameters. The video encoder is divided into two input sources: the video frame input source assetVideoWriterInput and the audio data input source assetAudioWriterInput. For example, create an empty video file 123.mp4, and create a corresponding video encoder based on the video file 123.mp4. Among them, the video encoder includes two input sources: the video frame input source and the audio frame data input source; when creating the video file and the corresponding video encoder, it is necessary to add the video input source assetVideoWriterInput and the audio input source assetAudioWriterInput to the same group group using dispatch_group_enter. In this way, when both assetVideoWriterInput and assetAudioWriterInput have completed their work, we can get the callback through the group and close the video encoder.
[0044] S204, creating a video write semaphore and an audio write semaphore;
[0045] In some examples of this embodiment, two semaphores are created: a video write semaphore and an audio write semaphore, which are used to lock the video and audio input sources. The initial value of the video frame write semaphore is 0, indicating a locked state. The initial value of the audio write semaphore is 1, indicating an unlocked state.
[0046] S205, generating corresponding audio data for each of the sentences, writing the audio data of each sentence into a video file, and converting the picture data into corresponding video frames according to the written audio data and writing the frames into the video file;
[0047] In some examples of this embodiment, the audio signal value is first decremented by 1, changing the state from unlocked to locked. Then, based on the order in which the sentences were generated, the previously generated sentence is retrieved, and the speech synthesis module is called to generate audio data for that sentence using speech synthesis technology. After obtaining the audio data (in PCM format by default), an audio frame buffer is created based on the audio data. The audio frame buffer is then written to the created video file using the audio input source assetAudioWriterInput.
[0048] Continuing with the previous example, after writing the sentence's audio frame to the video file, calculate the sentence's audio duration (in seconds) and store it in the array audioBufferDurationArr. audioDuration = totalSize * 8 / mSampleRate / mBitsPerChannel / mChannelsPerFrame. Then, increment the video semaphore by 1, changing it from the default locked state to unlocked, and begin writing the sentence's video frame data to the video input source. At this point, when writing the sentence's video frame data to the video input source, the audio semaphore is also incremented by 1, changing it from locked to unlocked. Determine whether the current sentence is the last one. If not, retrieve the next sentence and write the audio data. If the current sentence is the last one, end writing the audio input source and call dispatch_group_leave to leave the group.
[0049] Continuing with the previous example, before writing the image data to the video file, the default value of the video signal quantity value is 0. Therefore, the system will wait until the video signal quantity value is increased by 1 in the above step. Only then will the image data of the sentence be converted into video frames corresponding to the number of video frames, and the video frames that match the audio frames will be written to the video file. When the image data of the sentence is obtained, the video signal quantity value is reduced by 1, changing the state from unlocked to locked, and the system will wait until the next time the video signal quantity value is 1, and then the image data of the next sentence will be obtained.
[0050] In some examples of this embodiment, after the video frame input source extracts the image data of a sentence, it uses the image data to create an initial video frame, and then calculates the number of video frames to be written based on the audio duration audioDuration of the sentence, with 30 frames displayed per second by default. Through the video frame input source assetVideoWriterInput, the initial video frames of the corresponding number of frames are written, and then the video frames corresponding to the audio frames are written to the video file. It is then determined whether the current sentence is the last sentence. If the current sentence is not the last sentence, the subsequent steps are continued until the Value of the video signal quantity is modified to 1; if the current sentence is the last sentence, the writing of the video input source is terminated, and dispatch_group_leave is called to leave the group.
[0051] Continuing with the previous example, the mobile terminal monitors the status of the group. After all the audio data and image data of the sentences are written into the video file, the video encoder is closed, the video encoding is completed, and the video file is output.
[0052] The following is an example of the present application, which is described again with reference to the specific implementation of the embodiment of the present application; a method for converting text to video is provided in the specific implementation, such as Figure 3 As shown, the text-to-video method includes:
[0053] S301, prepare text to be processed;
[0054] The pending text reads: Cui Liulang smiled and said, "Haven't you had breakfast yet? I brought you a cake." He then handed over a steaming hot sesame cake, the front dotted with shiny, large sesame seeds, its aroma tangy. The old official pinched it and discovered a small, straight silver ingot pressed deeply into the back of the cake. He secretly weighed it and estimated it to be at least two taels. While not enough for cash, it would be enough to make a nice hairpin for his daughter.
[0055] S302, segmenting the text to be processed to obtain multiple sentences;
[0056] Continuing with the above example, we segment the text to be processed according to punctuation marks and obtain 12 sentences, which are:
[0057] Cui Liulang said with a smile,
[0058] Haven't had breakfast yet?
[0059] I brought you a cake,
[0060] Then he handed over a steaming hot sesame cake.
[0061] The front is decorated with shiny big sesame seeds.
[0062] The aroma is fragrant,
[0063] The old official pinched it.
[0064] I found a small straight silver stick pressed deep into the back of the dough.
[0065] He weighed it up secretly.
[0066] I'm afraid it's two taels.
[0067] Although it cannot be used as cash,
[0068] But I can also make a nice hairpin for my daughter.
[0069] S303, generating picture data according to the generated sentence;
[0070] Continuing with the above example, use the above 12 sentences to create 12 sentences of image data images. Specifically, use the video control to create an image background view for each sentence, create a text control to display the sentence, then set the text style in the text, generate a screenshot from the view, and then get 12 pictures corresponding to the 12 sentences. For example, for the sentence "Haven't you had breakfast yet?", generate image data images. First, use the video control to create an image background view for each sentence. After creating the text control, set the text control to the view. Add it to the middle of the background view, then add the sentences "Cui Liulang said with a smile", "Haven't you had breakfast yet", and "I brought you a cake" to the text of the text control, segment the three added sentences, at this time, the sentence "Haven't you had breakfast yet" is in the middle of the segment, and set the sentences "Cui Liulang said with a smile" and "I brought you a cake" to black, and the sentence "Haven't you had breakfast yet" to red for highlighting. Finally, generate a screenshot from the view, and get the image data images of the sentence "Haven't you had breakfast yet";
[0071] S304, creating a local video file and a video encoder;
[0072] In some examples of this embodiment, an empty video file and a corresponding video encoder (assetWriter) are created, and the video encoder parameters are set. The video encoder has two input sources: the video frame input source (assetVideoWriterInput) and the audio data input source (assetAudioWriterInput). For example, a mobile terminal creates an empty video file (123.mp4) and creates a corresponding video encoder based on the video file (123.mp4). It should be understood that the video encoder includes two input sources: the video frame input source (assetVideoWriterInput) and the audio data input source (assetAudioWriterInput). For example, a mobile terminal creates an empty video file (123.mp4) and creates a corresponding video encoder based on the video file (123.mp4). Two input sources also need to be created for the video encoder: a video frame input source and an audio frame data input source. When creating the video file (123.mp4) and the corresponding video encoder, the video input source (assetVideoWriterInput) and the audio input source (assetAudioWriterInput) need to be added to the same group (group) using dispatch_group_enter. In this way, when both assetVideoWriterInput and assetAudioWriterInput have completed their work, we can get the callback through the group and close the video encoder.
[0073] S305: Create a speech synthesis module handle;
[0074] It should be understood that after creating the speech synthesis module handle synthesizerSpeaker, speech synthesis technology can be used through the speech synthesis module handle synthesizerSpeaker to generate audio data corresponding to the sentence.
[0075] S306: Generate audio data for each sentence, convert the audio data of each sentence into audio frames and write them into a video file, and convert the picture data into corresponding video frames according to the written audio frames and write them into the video file;
[0076] Continuing with the previous example, after creating the speech synthesis module handle, we create a group using dispatch_group_create() and then add the video frame input source assetVideoWriterInput and the audio input source assetAudioWriterInput to the group using dispatch_group_enter(). This requires two group enter operations. We then use dispatch_semaphore_create() to create the video write semaphore videoWriteLock and the audio write semaphore audioWriteLock, setting the value of videoWriteLock to 0 and the value of audioWriteLock to 1.
[0077] Continuing with the previous example, we traverse the 12 sentences generated in step 2 in the audio input source assetAudioWriterInput, generate audio data for each sentence, generate audio frames based on the audio data, and write the audio frames to the video file 123.mp4. Specifically, we retrieve the first sentence generated in the order in which the sentences were generated, and decrement the value of the audio write semaphore audioWriteLock by 1, changing it from unlocked to locked. We then call dispatch_semaphore_wait(audioWriteLock, DISPATCH_TIME_FOREVER). We then use the speech synthesis handle synthesizerSpeaker to retrieve the audio data for the current sentence using audioBuffer[weakSelf.synthesizerSpeakeraudioBufferWithString:titles[i]bufferCallback:^(AVAudioBuffer*_NonnullaudioBuffer){}] . After obtaining the audio data audioBuffer corresponding to the sentence, we need to create an audio frame buffer (CMSampleBufferRef type) based on the audioBuffer. The audio frame buffer is then written to the video file 123.mp4 through the audio frame input source assetAudioWriterInput [weakSelf.assetAudioWriterInputappendSampleBuffer:buffer]; the audio duration of the audio data audioBuffer, audioDuration, is then calculated and saved to the array audioBufferDurationAr. Dispatch_semaphore_signal(videoWriteLock) is called; dispatch_semaphore_signal(audioWriteLock), and then the audio semaphore value is incremented by 1 to retrieve the next sentence, changing the state from locked to unlocked. Finally, the program determines whether the current sentence is the last one. If so, it ends the audio data input and exits the group. If not, after retrieval of the next sentence, the audio write semaphore audioWriteLock value is decremented by 1 to change the state from unlocked to locked. The audio data generation steps continue until audio data for all sentences has been generated.
[0078] Continuing with the previous example, in the video input source assetVideoWriterInput, the 12 image data generated from the 12 sentences are iterated over, converted into video frames, and written to the video file 123.mp4. Specifically, the program waits for the video write semaphore to become unlocked. After obtaining the duration of the audio frame corresponding to the sentence, it retrieves a semaphore value and obtains the image data corresponding to that sentence. Then, the value of the video write semaphore videoWriteLock is decremented by 1, changing it from unlocked to locked, and stopping the acquisition of image data for the remaining sentences. Dispatch_semaphore_wait(videoWriteLock, DISPATCH_TIME_FOREVER) is called. Based on the sentence image data, CVPixelBufferCreate() is used to create the corresponding initial video frame buffer (of type CVPixelBufferRef). The audio duration of the current sentence is then retrieved from audioBufferDurationArr. Here, the audio duration of the first sentence, "Cui Liulang smiled and said," is 2.238 seconds. The calculated number of video frames to be displayed is 68, or 30 frames per second. The video input source writerAdapter (the adapter of assetVideoWriterInput) then writes 68 initial video frames into the video, completing the conversion of the sentence's image data into video frames and writing them to the video file 123.mp4. The process then checks whether the current image is the last. If so, it ends the video data input and exits the group. If not, it waits for the video write semaphore to become unlocked before continuing to convert the next sentence's image data into video frames and writing them to the video file 123.mp4. This continues until all sentence image data has been converted into video frames and written to the video file 123.mp4.
[0079] Continuing with the previous example, we continue monitoring the group status. When all sentences have been converted to videos, we close the video encoder. The video encoding is complete and the output video file is 123.mp4. The total video length of the text is 34 seconds, with a width of 1366 and a height of 768.
[0080] The present application also provides a device for converting text to video. Figure 4 As shown, it includes but is not limited to:
[0081] Generating module 1, for generating corresponding picture data and audio data for the text to be processed, wherein the picture data is used to describe the text to be processed through pictures, and the audio data is used to describe the text to be processed through audio;
[0082] The writing module 2 is used to generate a target video file of the text to be processed according to the image data and the audio data.
[0083] It can be understood that the text-to-video device provided in this embodiment can implement each step of the above-mentioned text-to-video method and achieve the same technical effects as the various steps of the above-mentioned text-to-video method. Therefore, it will not be described in detail here.
[0084] The present application also provides an electronic device, such as Figure 5 As shown, it includes a processor 501, a communication interface 502, a memory 503 and a communication bus 504, wherein the processor 501, the communication interface 502, and the memory 503 communicate with each other through the communication bus 504.
[0085] Memory 503, used for storing computer programs;
[0086] The processor 501 is configured to implement the steps of the text-to-video method when executing the program stored in the memory 503 .
[0087] It should be noted that the role played by the processor 501 when executing the program stored in the memory 503 is similar to the steps of the above-mentioned text-to-video method, and will not be repeated here.
[0088] The communication bus mentioned in the terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus.
[0089] The communication interface is used for communication between the above terminal and other devices.
[0090] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.
[0091] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.
[0092] In another embodiment provided in the present application, a computer-readable storage medium is also provided, which stores instructions. When the computer-readable storage medium is run on a computer, the computer executes the steps of the text-to-video method described in any of the above embodiments.
[0093] In another embodiment provided by the present application, a computer program product comprising instructions is also provided, which, when executed on a computer, enables the computer to execute the steps of the text-to-video method described in any one of the above embodiments.
[0094] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0095] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0096] Each embodiment in this specification is described in a related manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiment is generally similar to the method embodiment, so the description is relatively simple. For related parts, refer to the description of the method embodiment.
[0097] The above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application are included in the scope of protection of the present application.
Claims
1. A method for converting text to video, characterized in that: include: Get the punctuation marks of the text to be processed; Segmenting the text to be processed according to the punctuation marks to generate multiple sentences; Generate corresponding image data and audio data for each sentence in the text to be processed, wherein the image data is used to describe the text to be processed through images, and the audio data is used to describe the text to be processed through audio; Generate a target video file of the text to be processed according to the image data and the audio data; The step of generating a video file of the text to be processed based on the image data and the audio data includes: Writing the audio data into a video file and obtaining the duration of the audio data in the video file; Calculate the number of video frames corresponding to the sentence based on the acquired duration of the audio data and the duration of each video frame in the video file; Based on the number of video frames, converting the image data into video frames of the sentence; Writing the video frame of each sentence into the video file to obtain the target video file; Generating corresponding image data and audio data for each sentence in the text to be processed includes: Create an audio write semaphore and a video write semaphore, where the initial value of the audio write semaphore is 1, representing an unlocked state, and the initial value of the video write semaphore is 0, representing a locked state; When the first sentence is obtained, the value of the audio write semaphore is reduced by 1, and the audio write semaphore becomes locked; After generating audio data for the sentence, the audio data of the sentence is written into the video file, and the value of the audio write semaphore is changed to 1, and the audio write semaphore becomes unlocked to obtain the next sentence and generate audio data for the next sentence; repeat the above steps until the audio data of all sentences are written into the video file; After the audio data of the first sentence is written into the video file, the value of the video write semaphore is increased by 1, the video write semaphore becomes unlocked, the picture data of the sentence is obtained, the value of the video write semaphore is changed to 0, and it becomes locked, and the picture data of the sentence is written into the video file accordingly. When the audio data of the next sentence is written into the video file, the above steps are repeated until the audio data of all sentences are written into the video file.
2. The method for converting text to video according to claim 1, wherein: Generate corresponding image data for the text to be processed, including: Setting a background view for the text to be processed, and creating a text control on the background view, wherein the text control is used to display the text to be processed; Image data corresponding to the text to be processed is generated according to the background view and the text control.
3. The method for converting text to video according to claim 1, wherein: Before generating corresponding image data and audio data for the text to be processed, the method further includes: When there is a sentence with empty content among the multiple sentences generated by the text to be processed, the sentence with empty content is deleted.
4. The method for converting text to video according to claim 1, wherein: Before generating the target video file of the text to be processed according to the image data and the audio data, the method further includes: Obtain a video encoder for the video file, the video encoder comprising: a video frame input source for controlling video frame input, and an audio frame input source for controlling audio data input; the video frame input source is used to write the video frames of the text to be processed into the video file; the audio frame input source is used to write the audio data into the video file.
5. A device for converting text to video, characterized in that: include: A generation module for obtaining punctuation marks of the text to be processed; Segmenting the text to be processed according to the punctuation marks to generate multiple sentences; Generate corresponding image data and audio data for each sentence in the text to be processed, wherein the image data is used to describe the text to be processed through images, and the audio data is used to describe the text to be processed through audio; A writing module, configured to generate a target video file of the text to be processed according to the image data and the audio data; Wherein, the writing module is used for: Writing the audio data into a video file and obtaining the duration of the audio data in the video file; Calculate the number of video frames corresponding to the sentence based on the acquired duration of the audio data and the duration of each video frame in the video file; Based on the number of video frames, converting the image data into video frames of the sentence; Writing the video frame of each sentence into the video file to obtain the target video file; Wherein, the generating module is used for: Create an audio write semaphore and a video write semaphore, where the initial value of the audio write semaphore is 1, representing an unlocked state, and the initial value of the video write semaphore is 0, representing a locked state; When the first sentence is obtained, the value of the audio write semaphore is reduced by 1, and the audio write semaphore becomes locked; After generating audio data for the sentence, the audio data of the sentence is written into the video file, and the value of the audio write semaphore is changed to 1, and the audio write semaphore becomes unlocked to obtain the next sentence and generate audio data for the next sentence; repeat the above steps until the audio data of all sentences are written into the video file; After the audio data of the first sentence is written into the video file, the value of the video write semaphore is increased by 1, the video write semaphore becomes unlocked, the picture data of the sentence is obtained, the value of the video write semaphore is changed to 0, and it becomes locked, and the picture data of the sentence is written into the video file accordingly. When the audio data of the next sentence is written into the video file, the above steps are repeated until the audio data of all sentences are written into the video file.
6. An electronic device, characterized in that: include: Processor, communication interface, memory and communication bus, wherein the processor, communication interface and memory communicate with each other through the communication bus. Memory for storing computer programs; The processor is configured to implement the steps of the text-to-video method according to any one of claims 1 to 4 when executing the program stored in the memory.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the text-to-video method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Information sharing method, information sharing device and storage medium
CN107517323A