Method and device for multimedia content generation, equipment and storage medium
By transforming and time-aligning audio and image data based on delay compensation, the method synchronizes audio and video in real-time multimedia content generation, resolving desynchronization issues.
Patent Information
- Application Number
- CN202410059722.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-15
- Publication Date
- 2025-07-15
AI Technical Summary
In real-time shooting scenes, it is difficult for the prior art to synchronize the audio data and image data, especially when complex processing such as tone conversion, the audio and video are out of synchronization due to delay problems.
By receiving concurrently captured image data and input sound data, performing a target conversion operation to generate converted sound data, and timely aligning the audio data with the image data according to the time delay to generate multimedia content.
The synchronization of audio data and image data is realized, and the problem of audio and video out-synchronization in real-time shooting scenes is solved, ensuring that the audio and image in the generated multimedia content are consistent.
Smart Images

Figure CN120321466A_ABST
Abstract
Description
Technical Field
[0001] Example embodiments of the present disclosure generally relate to the field of computers, and particularly to methods, apparatuses, devices, and computer-readable storage media for multimedia content generation. Background Art
[0002] Currently, more and more applications are designed to provide various services to users. For example, users can browse, comment on, and forward various types of content in the application, including, for example, videos, images, image sets, audio, etc. In addition, content sharing applications also support shooting multimedia content, such as videos, image sets with sound, etc. Summary of the Invention
[0003] In a first aspect of the present disclosure, a method for multimedia content generation is provided. The method includes: receiving concurrently captured image data and input sound data from a target object; generating audio data based at least on transformed sound data corresponding to the input sound data, the transformed sound data being obtained by performing a target transformation operation on at least a part of the input sound data; aligning the audio data with the image data in time according to a time delay associated with the transformed sound data; and generating multimedia content associated with the target object based on the aligned audio data and image data.
[0004] In a second aspect of the present disclosure, an apparatus for multimedia content generation is provided. The apparatus includes: a data receiving module configured to receive concurrently captured image data and input sound data from a target object; an audio processing module configured to generate audio data based at least on transformed sound data corresponding to the input sound data, the transformed sound data being obtained by performing a target transformation operation on at least a part of the input sound data; a first alignment module configured to align the audio data with the image data in time according to a time delay associated with the transformed sound data; and a first merging module configured to generate multimedia content associated with the target object based on the aligned audio data and image data.
[0005] In a third aspect of the present disclosure, an electronic device is provided. The device includes at least one processing unit; and at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit. The instructions, when executed by the at least one processing unit, cause the device to perform the method of the first aspect.
[0006] In a fourth aspect of the present disclosure, a computer-readable storage medium is provided. A computer program is stored on the computer-readable storage medium and can be executed by a processor to implement the method of the first aspect.
[0007] It should be understood that the content described in this content part is not intended to define the key features or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0008] In combination with the accompanying drawings and with reference to the following detailed description, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. In the drawings, the same or similar reference numerals denote the same or similar elements, where:
[0009] Figure 1 FIG. shows a schematic diagram of an example environment in which embodiments of the present disclosure can be implemented;
[0010] Figure 2 FIG. shows a schematic diagram of an example processing architecture for audio data according to some embodiments of the present disclosure;
[0011] Figure 3A FIG. shows a schematic diagram of an example user interface for initiating content capture according to some embodiments of the present disclosure;
[0012] Figure 3B FIG. shows a schematic diagram of an example user interface for content capture according to some embodiments of the present disclosure;
[0013] Figure 4 FIG. shows a schematic diagram of an example data stream for latency compensation according to some embodiments of the present disclosure;
[0014] Figure 5 FIG. shows an example signaling diagram for audio processing according to some embodiments of the present disclosure;
[0015] Figure 6 FIG. shows a flowchart of a process for multimedia content generation according to some embodiments of the present disclosure;
[0016] Figure 7 FIG. shows a block diagram of a device for multimedia content generation according to some embodiments of the present disclosure; and
[0017] Figure 8 FIG. shows a block diagram of a device capable of implementing multiple embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0019] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the present disclosure technical solution based on the prompt message.
[0020] As an optional but non-limiting implementation manner, when responding to receiving an active request from a user, the manner of sending a prompt message to the user may, for example, be in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0021] It can be understood that the above process of notifying and obtaining user authorization is only illustrative and does not limit the implementation manners of the present disclosure. Other manners that comply with relevant laws and regulations can also be applied to the implementation manners of the present disclosure.
[0022] It can be understood that the data involved in the present technical solution (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of corresponding laws, regulations and related provisions.
[0023] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0024] It should be noted that the title of any section / subsection provided herein is not restrictive. Various embodiments are described throughout this document, and any type of embodiment can be included under any section / subsection. In addition, the embodiments described in any section / subsection can be combined with any other embodiments described in the same section / subsection and / or different section / subsections in any manner.
[0025] In this document, unless explicitly stated, performing a step "in response to A" does not mean that the step is immediately performed after "A", but may include one or more intermediate steps.
[0026] In the description of the embodiments of the present disclosure, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "an embodiment" or "the embodiment" should be understood as "at least one embodiment". The term "some embodiments" should be understood as "at least some embodiments". There may also be other explicit and implicit definitions hereinafter. The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.
[0027] As used herein, the term "model" can learn the association between corresponding inputs and outputs from training data, so that after training, for a given input, the corresponding output can be generated. The generation of the model can be based on machine learning techniques. Deep learning is a machine learning algorithm that processes inputs and provides corresponding outputs by using multiple layers of processing units. In this document, "model" can also be referred to as "machine learning model", "machine learning network" or "network", and these terms can be used interchangeably herein. A model can further include different types of processing units or networks.
[0028] Figure 1 A schematic diagram of an exemplary environment 100 in which the embodiments of the present disclosure can be implemented is shown. In this exemplary environment 100, an application 120 is installed in the terminal device 110. The user 140 can interact with the application 120 via the terminal device 110 and / or an attached device of the terminal device 110. Exemplarily, the application 120 can be a content generation application, a content sharing application or a social application, which can provide services related to media content to the user 140, including content browsing, commenting, forwarding, creation (e.g., shooting and / or editing), publishing, etc. "Media content" can include one or more types of content, such as videos, images, animated graphics, image sets, audio, text, etc. The application 120 can support the user 140 to create multimedia content. Such multimedia content can include image data and audio data.
[0029] In Figure 1 the environment 100, the terminal device 110 can present the user interface 150 of the application 120. The user interface 150 can include various interfaces that the application 120 can provide, such as a content presentation interface, a content creation interface, a content publishing interface, a message interface, a personal homepage, etc. The application 120 can provide a content browsing function to browse various types of content published in the application 120. The application 120 can also provide a content creation function, including shooting, uploading, editing and / or publishing media content.
[0030] In some embodiments, the terminal device 110 communicates with the server 130 to implement the supply of services for the application 120. The terminal device 110 can be any type of mobile terminal, fixed terminal or portable terminal, including mobile phones, desktop computers, laptop computers, notebook computers, netbook computers, tablet computers, media computers, multimedia tablets, personal communication system (PCS) devices, personal navigation devices, personal digital assistants (PDAs), audio / video players, digital cameras / video cameras, positioning devices, television receivers, radio broadcast receivers, e-book devices, gaming devices, or any combination of the foregoing, including accessories and peripherals of these devices or any combination thereof. In some embodiments, the terminal device 110 is also capable of supporting any type of user interface (such as a "wearable" circuit, etc.). The server 130 can be various types of computing systems / servers capable of providing computing capabilities, including but not limited to mainframes, edge computing nodes, computing devices in a cloud environment, and so on.
[0031] It should be understood that the structures and functions of the various components in the environment 100 are described only for exemplary purposes, without implying any limitation on the scope of the present disclosure.
[0032] As mentioned above, the user can create multimedia content through an application. In the creation of multimedia content, additional processing of the captured sound data may be required, such as conversion operations. For example, voice conversion (VC) refers to converting the input user voice into a voice with a specified timbre. The specified timbre can be an existing timbre in the timbre library or a timbre that has been authorized for use. Using voice conversion can achieve voice fun voice change and enrich the voice interaction experience.
[0033] However, this additional processing may consume relatively large resources, thus causing a time delay. In some cases, this additional processing may be relatively complex (for example, in the case of VC), and thus it is necessary to rely on the server. The communication with the server will cause further delays, such as network delays and server-side processing delays, etc. These delays will cause problems of out-of-sync audio and video in a real-time shooting scenario, thus unable to meet the real-time shooting scenario.
[0034] To this end, embodiments of the present disclosure provide an improved solution for multimedia content generation. In this solution, concurrently captured image data and input sound data from a target object are received. Transformed sound data corresponding to the input sound data is obtained by performing a target transformation operation on at least a part of the input sound data. Audio data is generated at least based on the transformed sound data. According to the time delay associated with the transformed sound data, the audio data is time-aligned with the image data to generate multimedia content associated with the target object.
[0035] In an embodiment of the present disclosure, the audio data and the image data in the multimedia content are aligned according to the time delay associated with the converted sound data, thereby achieving audio-visual synchronization. In other words, the audio track is temporally aligned with the video track according to the delay of the audio link. This is a delay compensation scheme based on the audio link. In this way, the problem of audio-visual asynchronization in the real-time content generation scenario of streaming sound conversion (e.g., VC) can be solved, enabling the streaming sound conversion to be used in the real-time content generation scenario.
[0036] Some example embodiments of the present disclosure will be further described below with reference to the accompanying drawings.
[0037] Figure 2 A schematic diagram of an example processing architecture 200 for audio data according to some embodiments of the present disclosure is shown. It will be described with reference to Figure 1 the Figure 2 example processing architecture 200. As Figure 2 shown, the user 140 interacts with the application 120 through the prop user interface (UI) element 210 presented on the terminal device 110. The prop UI element 210 receives a user selection of a sound conversion operation to be performed. The sound conversion operation to be performed may also be referred to as a target conversion operation. Alternatively, in some embodiments, the target conversion operation may be default or randomly set without user selection.
[0038] The target conversion operation may include any suitable conversion adapted to be performed on the sound data. For example, the target conversion operation may include timbre conversion, that is, converting the input speech into a user-specified, randomly selected, or default timbre. Another example is that the target conversion operation may include pitch conversion, that is, raising or lowering the pitch of the input sound. It should be understood that the sound conversion operations listed here are only exemplary and are not intended to limit the scope of the present disclosure.
[0039] In response to a user operation on the target conversion operation, the prop UI element 210 may send the target conversion operation information to the prop audio processing unit 220. Thus, the prop audio processing unit 220 can learn the sound conversion operation to be performed. If the user 140 triggers content shooting, the prop audio processing unit 220 may send a message or notification to the audio-video encoding unit 230 to trigger the capture of sound data. For example, the audio-video encoding unit 230 may trigger the capture of sound data via the microphone of the terminal device 110 or other suitable devices. The captured sound data may come from a target object. The target object may be any suitable object in the environment where the microphone is located, such as the user 140 or other people, animals, physical sound sources, etc.
[0040] In addition to the capture of input audio data, if user 140 triggers content capture, the capture of image data is also triggered. For example, the image data can be captured via the camera of terminal device 110. In some embodiments, the image data may include a face image of the target object, such as the face image of user 140. Note that the image data and the input audio data are captured concurrently, e.g., captured simultaneously.
[0041] Refer to Figure 3A and Figure 3B to describe an example scenario. Figure 3A FIG. shows a schematic diagram of an example user interface 300A for initiating content capture according to some embodiments of the present disclosure. As Figure 3A shown, a selection entry 310 for a voice conversion operation is displayed in user interface 300A. Through this selection entry 310, user 140 can select the conversion operation desired to be performed on the input audio data, such as selecting or specifying a specific timbre in the case of VC. If the capture control 312 is triggered, terminal device 110 can start capturing image data and input audio data concurrently. Accordingly, terminal device 110 can present Figure 3B the example user interface 300B for content capture as shown. In this example, an image of user 140 captured in real time is displayed in user interface 300B.
[0042] It should be understood that the example scenarios and user interfaces described with reference to Figure 3A and Figure 3B are for illustrative purposes only and are not intended to be limiting. The latency compensation scheme according to the embodiments of the present disclosure can be applied to any concurrently captured image data and audio data.
[0043] Continuing to refer to Figure 2 . The captured input audio data can be returned to the prop audio processing unit 220 through the audio-video encoding unit 230. The prop audio processing unit 220 can then send the input audio data to the audio rendering unit 240, and can request the audio rendering unit 240 to create a conversion operation algorithm for the target conversion operation, or send the conversion operation algorithm created by itself to the audio rendering unit 240. That is, the target conversion is performed while capturing the audio data. Streaming conversion of the input audio data is implemented, such as streaming timbre conversion.
[0044] Depending on the specific implementation or the capabilities of the terminal device 110, the target conversion operation may be performed by any suitable entity. For example, the audio rendering unit 240 may perform the target conversion operation on at least a part of the input sound data, such as timbre conversion, to obtain the converted sound data. Alternatively, the audio rendering unit 240 may send the input sound data to the server 130 for the server 130 to perform the target conversion operation. Correspondingly, the audio rendering unit 240 may receive the converted sound data from the server 130. For example, an interface for streaming sound conversion (e.g., streaming timbre conversion) may be configured. The interface may include various parameters for sound conversion for exchanging data and information required for sound conversion between the server 130 and the terminal device 110.
[0045] Note that the target conversion operation may be performed on all of the input sound data or on a part of the input sound data. For example, if the input sound data includes human speech and background noise, the target conversion operation may be performed only on the speech part. Embodiments of the present disclosure are not limited in terms of how the target conversion operation is performed.
[0046] After obtaining the converted sound data, the audio rendering unit 240 may generate audio data based at least on the converted sound data. In some embodiments, the terminal device 110 may play background sound data, such as background music, concurrently with the capture of the input sound data and image data. Exemplarily, if the user 140 selects or designates background music through the background music entry 301 in the user interface 300A, then after the shooting control 312 is triggered, the terminal device 110 may play the selected background music simultaneously.
[0047] The playback of the background sound may cause a problem that the sound of the target object is out of sync with the background sound, i.e., the audio-audio out-of-sync problem. For this reason, in such embodiments, the audio rendering unit 240 may align the converted sound data and the background sound data in time, and generate audio data based on the aligned converted sound data and background sound data. Example embodiments of time alignment will be described below with reference to Figure 4 Describe example embodiments of time alignment.
[0048] In the case of successful conversion, the audio rendering unit 240 sends the obtained audio data to the prop audio processing unit 220. If the conversion fails, the audio rendering unit 240 may send a status code indicating failure to the prop audio processing unit 220.
[0049] The prop audio processing unit 220 may forward the audio data to the audio and video encoding unit 230, or the audio rendering unit 240 may send the audio data directly to the audio and video encoding unit 230. The audio and video encoding unit 230 may time-align the audio data with the concurrently captured image data based on the time delay associated with the converted sound data. The audio and video encoding unit 230 may then generate multimedia content associated with the target object based on the aligned audio data and image data. For example, the audio and video encoding unit 230 may align the audio track with the video track based on the time delay to generate a corresponding multimedia file.
[0050] The above-mentioned time delay may include any delay in time related to the capture, processing, etc. of the converted sound data. Such a time delay may be understood, for example, as the actual difference between the moment when the sound is emitted and the moment when the corresponding audio data is used to generate multimedia content. In some embodiments, the time delay may include a conversion time delay caused by obtaining the converted sound data based on the input sound data. The conversion time delay at least includes the time consumed by performing the target conversion operation. If the target conversion operation is at least partially performed by the server 130 or other remote devices, the conversion time delay may also include the delay caused by the communication between the terminal device 110 and the server 130 or other remote devices. In this embodiment, the terminal device 110 may send the input sound data and a request to perform the target conversion operation on the input sound data to the remote device (e.g., the server 130), and may receive the converted sound data from the remote device. The terminal device 110 may obtain the conversion time delay based on the sending of the input sound data and the receiving of the converted sound data. For example, the terminal device 110 may calculate the time experienced from the sending of the input sound data to the receiving of the converted sound data. For another example, the interface may include parameters for obtaining the conversion time delay. The terminal device 110 may obtain the conversion time delay through the interface.
[0051] Alternatively or additionally, the time delay may include a capture time delay caused by obtaining input sound data using a microphone. For example, the sound arrives at the microphone from the sound source and is converted into sound data by the microphone, and the sound data then arrives at the audio rendering unit 240 via the audio and video encoding unit 230 and the prop audio processing unit 220. This sound data link may cause a capture time delay.
[0052] By performing time delay compensation during audio and video encoding, the audio data and the image data can be aligned in time. In this way, the audio and video alignment of the generated multimedia content can be achieved. For example, the voice and mouth shape of the user 140 are consistent.
[0053] The above describes an example processing architecture for delay compensation. It should be understood thatFigure 2 The functional descriptions and divisions of the respective units shown are merely exemplary and are not intended to limit the scope of the present disclosure. The specific functions of the respective units can be divided in any suitable manner. In addition, the prop UI element 210, the prop audio processing unit 220, the audio-video encoding unit 230, the audio rendering unit 240, etc. can be at least partially implemented in the terminal device 110, or at least partially implemented in the server 130.
[0054] Reference is made below Figure 4 to a schematic diagram of an example data stream 400 for describing latency compensation. In the data stream 400, except after the audio-video encoding unit 230, the remaining nodes can be at least a part of an audio graph, which can be, for example, executed by the audio rendering unit 240. In addition, in this example, a case where background sound data (e.g., background music) is played simultaneously is shown.
[0055] As Figure 4 shown, at the file node 401, a file storing background sound data is read. After that, the processing of the background sound data is divided into two branches, a branch for concurrent playback and a branch for writing multimedia content. In the branch for concurrent playback, the read background sound data is passed to the playback node 402 and reaches the external speaker node 404 via the aggregation node 403. At the external speaker node 404, the background sound data, such as background music, can be played through the sound output device (e.g., speaker) of the terminal device 110 or other attached sound output devices. The played background sound data can be captured by a microphone and returned to the microphone node 406 in the audio graph. At the same time, as described with reference to Figure 2 , the audio-video encoding unit 230 can trigger the microphone to capture the input sound data of the target object, and the input sound data can also enter the processing flow of the audio graph through the microphone node 406.
[0056] At the latency calculation node 410, the capture time latency caused by capturing sound data using the microphone (which can include input sound data and background sound data) can be calculated. The capture latency time can be reported to the audio-video encoding unit 230 and the latency setting node 405.
[0057] Continuing the description of the data flow in the audio graph. The played background sound can enter the processing flow of the audio graph through the microphone node 406 together with the sound of the captured target object. At the conversion node 407, at least a part of the input sound data is subjected to a target conversion operation to obtain the converted sound data. For example, the audio rendering unit 240 can perform the target conversion operation, or the converted sound data can be obtained from the server 130, as described above with reference to Figure 2As described. At the conversion node 407, the conversion time delay can also be obtained. For example, a parameter for obtaining the conversion time delay can be set in the interface of the streaming timbre conversion, so as to obtain the conversion time delay via the interface.
[0058] The conversion time delay can be reported to the audio - video encoding unit 230 and the delay setting node 405. The delay setting node 405 is located in the branch for writing multimedia content of the background sound data. The delay setting node 405 can obtain the conversion time delay and the capture time delay. In this way, the delay setting node 405 can set the total time delay, for example, add the conversion time delay and the capture time delay, and send the set time delay to the writing node 408.
[0059] The writing node 408 can receive two - channel sound data, that is, the converted sound data from the conversion node 407 and the background sound data from the delay setting node 405. At the writing node 408, according to the set time delay, the converted sound data is aligned with the background sound data in time, and based on the aligned converted sound data and background sound data, audio data is generated.
[0060] The converted sound data can be regarded as coming from the recording link, while the background sound data from the file node 401 can be regarded as coming from the non - recording link. Any suitable alignment method can be adopted to align the two links. In some embodiments, according to the time delay, the starting position of the background sound data can be moved backward in time. That is, in the non - recording link, the background sound data is delayed by the delay amount in the recording link in time.
[0061] Thus, a delay scheme for the recording link and the non - recording link can be implemented in terms of audio. In this way, the final generated multimedia content can achieve audio - audio synchronization.
[0062] Continuing to describe the data stream 400. The audio data generated at the writing node 408 has achieved the asynchronization of the sounds of different links. Such audio data can reach the audio - video encoding unit 230 via the aggregation node 403. Although not shown, it should be understood that the audio - video encoding unit 230 also receives image data captured concurrently with the input sound data. In addition, the audio - video encoding unit 230 also obtains the total time delay, including the conversion time delay and the capture time delay. The audio - video encoding unit 230 can then align the audio data and the image data in time according to the time delay, and generate multimedia content 409, such as a video, based on the aligned audio data and image data.
[0063] The alignment of the audio track and the video track can be achieved in any suitable alignment manner. In some embodiments, the audio-video encoding unit 230 may determine a target data volume based on the time delay and the target bit rate, and may remove the target data volume of audio data from the starting position of the audio data to align the audio data with the image data in time. Exemplarily, the audio-video encoding unit 230 may calculate the sum of the conversion time delay and the capture time delay, and calculate the number of samples to be discarded based on the delay and the bit rate. For example, the delay value multiplied by the bit rate is used as the number of samples. Then, the first number of samples in the audio data may be discarded. Thus, audio-visual delay compensation can be achieved, thereby achieving audio-visual consistency.
[0064] Figure 5 FIG. 500 shows an example signaling diagram for audio processing according to some embodiments of the present disclosure. The signaling diagram 500 relates to a prop UI element 210, an audio-video encoding unit 230, a special effect system 501, an audio rendering unit 240, and multimedia content 409. Figure 2 The shown prop audio processing unit 230 may be regarded as a part of the special effect system 501.
[0065] In the initialization stage 502, in response to a target conversion operation being selected or specified, at 510, the prop UI element 210 may instruct the audio-video encoding unit 230 to load the prop. At 512, the audio-video encoding unit 230 may instruct the special effect system 501 to create a corresponding handle. At 514, the special effect system 501 may instruct the audio rendering unit 240 to create a corresponding backend.
[0066] At 516, the prop UI element 210 may instruct the special effect system 501 to interact with a message indicating the creation of an audio graph, binding the audio graph and the backend, and enabling the audio graph. At 518, the special effect system 501 may forward the message to the audio rendering unit 240. Accordingly, the audio rendering unit 240 may create an audio graph, bind the audio graph and the backend, and enable the audio graph.
[0067] Next, in the data stream processing stage 503, at 520, the audio-video encoding unit 230 may push and pull concurrently captured sound data and image data to the special effect system 501. At 522, the special effect system 501 may send the sound data to the audio rendering unit 240 to perform the target conversion operation.
[0068] Next, enter the delay reporting stage 504. At 524, the special effects system 501 can send a message to the audio rendering unit 240 to request obtaining the time delay, such as the conversion time delay and the capture time delay described above. At 526, the audio rendering unit 240 can return the time delay to the special effects system 501. At 528, the special effects system 501 can report the obtained time delay to the audio-video encoding unit 230. Next, enter the delay compensation stage 505. At 530, the audio-video encoding unit 230 can generate the multimedia content 409 through delay compensation as described above.
[0069] After the streaming generation of the multimedia content 409 ends, enter the prop unloading stage 506. At 532, the prop UI element 210 can interact with the special effects system 501 to instruct it to destroy the handle and the audio graph. The special effects system 501 can destroy the handle, and at 534, can instruct the audio rendering unit 240 to unbind the audio graph and the backend, stop the audio graph, and destroy the backend and the audio graph. Accordingly, the audio rendering unit 240 can unbind the audio graph and the backend, stop the audio graph, and destroy the backend and the audio graph.
[0070] The above reference Figure 5 The interactions between the various units described above are merely exemplary and are not intended to impose any limitations. In the embodiments of the present disclosure, any suitable number and functions of units can be set to implement delay compensation.
[0071] Example processes, devices, and apparatuses
[0072] Figure 6 The flowchart of a process 600 for multimedia content generation according to some embodiments of the present disclosure is shown. The process 600 can be implemented at the terminal device 110 or at the terminal device 110 and the server 130.
[0073] At block 610, the terminal device 110 receives concurrently captured image data and input sound data from a target object. At block 620, the terminal device 110 generates audio data based at least on the converted sound data corresponding to the input sound data. The converted sound data is obtained by performing a target conversion operation on at least a part of the input sound data. At block 630, the terminal device 110 aligns the audio data and the image data in time according to the time delay associated with the converted sound data. At block 640, the terminal device 110 generates multimedia content associated with the target object based on the aligned audio data and the image data.
[0074] In some embodiments, generating the audio data includes: obtaining background sound data that is played concurrently with capturing the input sound data and the image data; aligning the converted sound data with the background sound data in time according to the time delay; and generating the audio data based on the aligned converted sound data and the background sound data.
[0075] In some embodiments, aligning the converted sound data with the background sound data in time includes: moving the start position of the background sound data backward in time according to the time delay.
[0076] In some embodiments, aligning the audio data with the image data in time includes: determining a target data amount based on the time delay and a target bit rate; and aligning the audio data with the image data in time by removing the target data amount of the audio data starting from the start position of the audio data.
[0077] In some embodiments, the time delay includes at least one of the following: a conversion time delay caused by obtaining the converted sound data based on the input sound data, or a capture time delay caused by obtaining sound data using a microphone.
[0078] In some embodiments, process 600 further includes: sending the input sound data and a request to perform the target conversion operation on the input sound data to a remote device; receiving the converted sound data from the remote device; and obtaining the conversion time delay based on the sending of the input sound data and the receiving of the converted sound data.
[0079] In some embodiments, the input sound data includes the speech of the target object, and the target conversion operation includes converting the speech into a tone specified by the user.
[0080] In some embodiments, the image data includes a face image of the target object.
[0081] In some embodiments, process 600 further includes: receiving a user input indicating content shooting; and in response to the user input, triggering concurrent capture of the image data and the input sound data.
[0082] Figure 7 A schematic structural block diagram of a device 700 for multimedia content generation according to certain embodiments of the present disclosure is shown. Device 700 may be implemented as or included in a terminal device 110. Each module / component in device 700 may be implemented by hardware, software, firmware, or any combination thereof.
[0083] As shown in the figure, device 700 includes a data receiving module 710 configured to receive concurrently captured image data and input sound data from a target object. Device 700 further includes an audio processing module 720 configured to generate audio data based at least on the converted sound data corresponding to the input sound data, where the converted sound data is obtained by performing a target conversion operation on at least a part of the input sound data. Device 700 further includes a first alignment module 730 configured to align the audio data and the image data in time according to a time delay associated with the converted sound data. Device 700 further includes a first merging module 740 configured to generate multimedia content associated with the target object based on the aligned audio data and the image data.
[0084] In some embodiments, the first merging module includes: a sound data acquisition module configured to obtain background sound data played concurrently with the capture of the input sound data and the image data; a second alignment module configured to align the converted sound data and the background sound data in time according to the time delay; and a second merging module configured to generate the audio data based on the aligned converted sound data and the background sound data.
[0085] In some embodiments, the second alignment module is further configured to: move the start position of the background sound data backward in time according to the time delay.
[0086] In some embodiments, the first alignment module 730 is further configured to: determine a target data volume based on the time delay and a target bit rate; and align the audio data and the image data in time by removing the target data volume of the audio data starting from the start position of the audio data.
[0087] In some embodiments, the time delay includes at least one of the following: a conversion time delay caused by obtaining the converted sound data based on the input sound data, or a capture time delay caused by obtaining sound data using a microphone.
[0088] In some embodiments, device 700 further includes: a data sending module configured to send the input sound data and a request for performing the target conversion operation on the input sound data to a remote device; a data receiving module configured to receive the converted sound data from the remote device; and a delay acquisition module configured to acquire the conversion time delay based on the sending of the input sound data and the receiving of the converted sound data.
[0089] In some embodiments, the input voice data includes the voice of the target object, and the target conversion operation includes converting the voice into a tone specified by the user.
[0090] In some embodiments, the image data includes a face image of the target object.
[0091] In some embodiments, the apparatus 700 further includes: a user input receiving module configured to receive a user input indicating content capture; and a triggering module configured to trigger concurrent capture of the image data and the input voice data in response to the user input.
[0092] Figure 8 A block diagram of an electronic device 800 in which one or more embodiments of the present disclosure may be implemented is shown. It should be understood that Figure 8 The illustrated electronic device 800 is merely exemplary and should not constitute any limitation to the functions and scope of the embodiments described herein. Figure 8 The illustrated electronic device 800 may be used to implement Figure 1 the terminal device 110.
[0093] As Figure 8 shown, the electronic device 800 is in the form of a general-purpose electronic device. The components of the electronic device 800 may include, but are not limited to, one or more processors or processing units 810, a memory 820, a storage device 830, one or more communication units 840, one or more input devices 850, and one or more output devices 860. The processing unit 810 may be an actual or virtual processor and be capable of performing various processes according to the programs stored in the memory 820. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing ability of the electronic device 800.
[0094] The electronic device 800 generally includes multiple computer storage media. Such media may be any accessible media that can be obtained by the electronic device 800, including but not limited to volatile and non-volatile media, removable and non-removable media. The memory 820 may be a volatile memory (e.g., registers, caches, random access memory (RAM)), a non-volatile memory (e.g., read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory), or some combination thereof. The storage device 830 may be a removable or non-removable medium and may include a machine-readable medium, such as a flash drive, a magnetic disk, or any other medium that can be used to store information and / or data and can be accessed within the electronic device 800.
[0095] The electronic device 800 may further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in Figure 8 , a disk drive for reading from and writing to a removable, non-volatile disk (e.g., a "floppy disk") and an optical disk drive for reading from and writing to a removable, non-volatile optical disk may be provided. In these cases, each drive may be connected to a bus (not shown) by one or more data media interfaces. The memory 820 may include a computer program product 825 having one or more program modules configured to perform the various methods or actions of the various embodiments of the present disclosure.
[0096] The communication unit 840 enables communication with other electronic devices via a communication medium. Additionally, the functions of the components of the electronic device 800 may be implemented by a single computing cluster or multiple computer machines capable of communicating via a communication connection. Thus, the electronic device 800 may operate in a networked environment using a logical connection to one or more other servers, network personal computers (PCs), or another network node.
[0097] The input device 850 may be one or more input devices such as a mouse, keyboard, trackball, etc. The output device 860 may be one or more output devices such as a display, speaker, printer, etc. The electronic device 800 may also communicate with one or more external devices (not shown) as needed via the communication unit 840, such as a storage device, a display device, etc., communicate with one or more devices that enable a user to interact with the electronic device 800, or communicate with any device that enables the electronic device 800 to communicate with one or more other electronic devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface (not shown).
[0098] According to an exemplary implementation of the present disclosure, there is provided a computer-readable storage medium having computer-executable instructions stored thereon, where the computer-executable instructions are executed by a processor to implement the method described above. According to an exemplary implementation of the present disclosure, there is also provided a computer program product tangibly stored on a non-transitory computer-readable medium and including computer-executable instructions, and the computer-executable instructions are executed by a processor to implement the method described above.
[0099] Aspects of the present disclosure are described herein with reference to flowcharts and / or block diagrams of methods, apparatuses, devices, and computer program products according to the present disclosure. It should be understood that each block of the flowcharts and / or block diagrams, and combinations of blocks in the flowcharts and / or block diagrams, can be implemented by computer-readable program instructions.
[0100] These computer-readable program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus to produce a machine such that, when the instructions are executed by the processing unit of the computer or other programmable data processing apparatus, an apparatus is created that implements the functions / acts specified in one or more boxes of the flowchart and / or block diagram. These computer-readable program instructions can also be stored in a computer-readable storage medium that causes a computer, a programmable data processing apparatus, and / or other devices to operate in a particular manner, so that the computer-readable medium storing the instructions comprises a manufacture, the instructions therein implement various aspects of the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0101] The computer-readable program instructions may be loaded onto a computer, other programmable data processing apparatus, or other devices, such that a series of operational steps are performed on the computer, other programmable data processing apparatus, or other devices to produce a computer-implemented process, so that the instructions executed on the computer, other programmable data processing apparatus, or other devices implement the functions / acts specified in one or more boxes of the flowchart and / or block diagram.
[0102] The flowcharts and block diagrams in the figures illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various implementations of the present disclosure. In this regard, each box in the flowchart or block diagram may represent a module, a segment of a program, or a portion of an instruction, which contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions noted in the boxes may occur out of the order noted in the figures. For example, two consecutive boxes may, in fact, be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each box of the block diagrams and / or flowcharts, and combinations of boxes in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or acts, or by a combination of dedicated hardware and computer instructions.
[0103] The various implementations of the present disclosure have been described above. The above description is exemplary, not exhaustive, and is not limited to the disclosed implementations. Many modifications and variations will be apparent to those of ordinary skill in the art in the field without departing from the scope and spirit of the described implementations. The choice of terms used herein is intended to best explain the principles of the implementations, the practical application, or the improvement of the technology in the market, or to enable other ordinary skilled persons in the art in the field to understand the various implementations disclosed herein.
Claims
1. A method for generating multimedia content, comprising: Receiving concurrently captured image data and input sound data from a target object; Generating audio data based at least on transformed sound data corresponding to the input sound data, the transformed sound data being obtained by performing a target transformation operation on at least a part of the input sound data; Temporally aligning the audio data with the image data according to a time delay associated with the transformed sound data; And Generating multimedia content associated with the target object based on the temporally aligned audio data and image data.
2. The method according to claim 1, wherein generating the audio data comprises: Obtaining background sound data that is played concurrently with the capturing of the input sound data and the image data; Temporally aligning the transformed sound data with the background sound data according to the time delay; And Generating the audio data based on the temporally aligned transformed sound data and background sound data.
3. The method according to claim 2, wherein temporally aligning the transformed sound data with the background sound data comprises: Moving the start position of the background sound data backward in time according to the time delay.
4. The method according to claim 1, wherein temporally aligning the audio data with the image data comprises: Determining a target data amount based on the time delay and a target bit rate; And Temporally aligning the audio data with the image data by removing the target data amount of the audio data starting from the start position of the audio data.
5. The method according to claim 1, wherein the time delay comprises at least one of the following: A conversion time delay caused by obtaining the transformed sound data based on the input sound data, or A capture time delay caused by obtaining sound data using a microphone.
6. The method according to claim 5, further comprising: Sending the input sound data and a request for performing the target transformation operation on the input sound data to a remote device; Receiving the transformed sound data from the remote device; And Obtaining the conversion time delay based on the sending of the input sound data and the receiving of the transformed sound data.
7. The method according to claim 1, wherein the input sound data comprises speech of the target object, and the target transformation operation comprises converting the speech into a tone specified by a user.
8. The method according to claim 7, wherein the image data comprises a face image of the target object.
9. The method according to claim 1, further comprising: Receiving a user input indicating content shooting; And Triggering the concurrent capture of the image data and the input sound data in response to the user input.
10. An apparatus for generating multimedia content, comprising: A data receiving module configured to receive concurrently captured image data and input sound data from a target object; An audio processing module, configured to generate audio data based at least on transformed sound data corresponding to the input sound data, the transformed sound data being obtained by performing a target transformation operation on at least a part of the input sound data; A first alignment module, configured to align the audio data and the image data in time according to a time delay associated with the transformed sound data; And A first merging module, configured to generate multimedia content associated with the target object based on the aligned audio data and the image data.
11. An electronic device, comprising: At least one processing unit; And At least one memory, the at least one memory being coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions, when executed by the at least one processing unit, causing the electronic device to perform the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, having stored thereon a computer program, the computer program being executable by a processor to implement the method according to any one of claims 1 to 9.
Citation Information
Cited By
Multimedia data synchronization method and multimedia synchronization system
CN122138027A