A video processing method, apparatus, computer device, and storage medium
By selecting and encapsulating multiple target data sources from a data source set for video directing or editing, the problem of cost waste and effect impact caused by data switching in virtual object live streaming is solved, achieving flexible and efficient video generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-10-11
- Publication Date
- 2026-04-03
AI Technical Summary
In virtual object live streaming scenarios, switching data in the live video requires re-recording, resulting in a waste of human and material resources and affecting the display effect.
Select multiple target data sources from the data source set, obtain their output source data, encapsulate it according to a unified data protocol format, and perform video directing or editing processing, including data alignment, editing, and switching operations.
It improves the flexibility and versatility of video processing, saves time and manpower costs in video generation, enhances the stability and accuracy of data transmission, and improves the efficiency of video processing.
Smart Images

Figure CN115604413B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer technology, and more specifically, to a video processing method, apparatus, computer device, and storage medium. Background Technology
[0002] In the context of virtual object live streaming, live videos (or recorded videos) can be generated by uniformly recording various types of data (such as voice data, motion data, and facial data).
[0003] When it's necessary to switch certain data in a live stream (or recorded video), it's usually necessary to re-record after switching the data to generate an updated live stream (or recorded video). This approach not only wastes manpower and resources but also easily affects the display quality of the live stream (or recorded video). Summary of the Invention
[0004] This disclosure provides at least one video processing method, apparatus, computer device, and storage medium.
[0005] In a first aspect, embodiments of this disclosure provide a video processing method, the method comprising: in response to selecting a plurality of target data sources from a data source set, acquiring source data output by the plurality of target data sources respectively; each data source in the data source set is used to provide source data corresponding to at least one video element in a video to be generated; the video element includes at least one of scene elements, character elements, and sound elements; encapsulating the source data output by the plurality of target data sources according to a set unified data protocol format to obtain a first encapsulation file; the first encapsulation file includes a plurality of data blocks, each data block corresponding to one type of source data; and performing video directing processing or video editing processing based on the first encapsulation file.
[0006] In one optional implementation, the video directing processing based on the first encapsulated file includes: inputting the first encapsulated file into directing software, performing directing processing on the first encapsulated file through the directing software, and then outputting the live video; wherein the directing processing includes data alignment of the source data of each data block in the first encapsulated file.
[0007] In an optional implementation, the method further includes: recording the live video using the broadcasting software to obtain a second encapsulated file under a unified data protocol format; and storing the second encapsulated file in the data source set as data in one of the data sources in the data source set.
[0008] In one optional implementation, the video editing process based on the first encapsulated file includes: inputting the first encapsulated file into editing software, editing the first encapsulated file using the editing software, recording the edited video as a second encapsulated file under a unified data protocol format, and storing the second encapsulated file into the data source set as data in one of the data sources in the data source set.
[0009] In one optional implementation, storing the second encapsulated file into the data source set includes: splitting the second encapsulated file into source data corresponding to each of the data sources, storing each of the split source data into the replay service module in the data source set, and using the replay service module as a new data source.
[0010] In one optional implementation, the data sources in the data source set include at least one of the following target service devices, which are used to provide source data for real-time output: motion capture device, facial capture device, artificial intelligence (AI) service device, and audio recording device.
[0011] In one optional implementation, during video directing or video editing, the method further includes: switching the selected target data source in response to a data source switching request.
[0012] Secondly, embodiments of this disclosure also provide a video processing apparatus, comprising: an acquisition unit, configured to acquire source data output by the plurality of target data sources respectively in response to selecting a plurality of target data sources from a data source set; each data source in the data source set is configured to provide source data corresponding to at least one video element in a video to be generated; the video element includes at least one of scene elements, character elements, and sound elements; an encapsulation unit, configured to encapsulate the source data output by the plurality of target data sources according to a set unified data protocol format to obtain a first encapsulation file; the first encapsulation file includes a plurality of data blocks, each data block corresponding to one type of source data; and a processing unit, configured to perform video directing processing or video editing processing based on the first encapsulation file.
[0013] Thirdly, embodiments of this disclosure also provide a computer device, including: a processor, a memory, and a bus, wherein the memory stores machine-readable instructions executable by the processor, and when the computer device is running, the processor communicates with the memory via the bus, and when the machine-readable instructions are executed by the processor, the steps of the first aspect above, or any possible implementation of the first aspect, are performed.
[0014] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing a computer program that, when executed by a processor, performs the steps of the first aspect or any possible implementation of the first aspect.
[0015] As described above, multiple target data sources can be selected from a data source set, and the source data output by each of these target data sources can be obtained. This allows for free selection of the source data used to generate the video, improving the flexibility and diversity of video processing. Next, the source data output by the multiple target data sources can be encapsulated according to a unified data protocol format to obtain a first encapsulated file. This enables the transmission of multiple source data according to a unified data protocol format, improving the stability and accuracy of data transmission. Subsequently, based on this first encapsulated file, video directing or editing processing can be performed on the multiple source data, thereby saving time and manpower costs in video generation and improving the efficiency of video processing.
[0016] To make the above-mentioned objects, features and advantages of this disclosure more apparent and understandable, preferred embodiments are described below in detail with reference to the accompanying drawings. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the embodiments will be briefly described below. These drawings are incorporated in and constitute a part of this specification. They illustrate embodiments conforming to this disclosure and, together with the specification, serve to explain the technical solutions of this disclosure. It should be understood that the following drawings only show some embodiments of this disclosure and should not be considered as limiting the scope. Those skilled in the art can obtain other related drawings based on these drawings without creative effort.
[0018] Figure 1 A flowchart of a video processing method provided by an embodiment of this disclosure is shown;
[0019] Figure 2 A schematic diagram of a unified data protocol format provided by an embodiment of this disclosure is shown;
[0020] Figure 3 A flowchart illustrating a video processing method provided in an embodiment of this disclosure is shown.
[0021] Figure 4 A schematic diagram of a video processing apparatus provided in an embodiment of the present disclosure is shown;
[0022] Figure 5 A schematic diagram of a computer device provided in an embodiment of this disclosure is shown. Detailed Implementation
[0023] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. The components of the embodiments of this disclosure described and shown in the accompanying drawings can generally be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of this disclosure provided in the accompanying drawings is not intended to limit the scope of the claimed disclosure, but merely represents selected embodiments of this disclosure. All other embodiments obtained by those skilled in the art based on the embodiments of this disclosure without inventive effort are within the scope of protection of this disclosure.
[0024] It should be noted that similar labels and letters in the following figures indicate similar items. Therefore, once an item is defined in one figure, it does not need to be further defined and explained in subsequent figures.
[0025] In this document, the term "and / or" merely describes a relationship, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Furthermore, the term "at least one" in this document means any combination of at least two of any one or more elements. For example, including at least one of A, B, and C can mean including any one or more elements selected from the set consisting of A, B, and C.
[0026] Research has found that in virtual object live streaming scenarios, live videos (or recorded videos) can be generated by uniformly recording various types of data (such as voice data, motion data, and facial data).
[0027] When it's necessary to switch certain data in a live stream (or recorded video), it's usually necessary to re-record after switching the data to generate an updated live stream (or recorded video). This approach not only wastes manpower and resources but also easily affects the display quality of the live stream (or recorded video).
[0028] Based on the above research, this disclosure provides a video processing method, apparatus, computer device, and storage medium. As described above, multiple target data sources can be selected from a data source set, and the source data output by each of these target data sources can be obtained. This allows for free selection of source data for video generation, improving the flexibility and diversity of video processing. Next, the source data output by the multiple target data sources can be encapsulated according to a unified data protocol format to obtain a first encapsulated file. This enables the transmission of multiple source data according to a unified data protocol format, improving the stability and accuracy of data transmission. Subsequently, based on this first encapsulated file, video directing or editing processing can be performed on the multiple source data, thereby saving time and manpower costs in video generation and improving the efficiency of video processing.
[0029] To facilitate understanding of this embodiment, a video processing method disclosed in this disclosure will first be described in detail. The execution subject of the video processing method provided in this disclosure is generally a computer device with certain computing power.
[0030] See Figure 1 The diagram shows a flowchart of a video processing method provided in an embodiment of this disclosure. The method includes steps S101 to S105, wherein:
[0031] S101: In response to selecting multiple target data sources from the data source set, obtain the source data output by the multiple target data sources respectively.
[0032] Each data source in the data source set provides source data corresponding to at least one video element in the video to be generated. Therefore, the data sources in the data source set can be data acquisition devices providing source data and / or AI service devices providing source data.
[0033] In this embodiment of the disclosure, video elements may include at least one of scene elements, character elements, and sound elements.
[0034] When the video element is a scene element, the source data for generating that scene element can be provided by a data acquisition device. In this case, the data acquisition device can be an image acquisition device, such as a mobile phone or a network camera.
[0035] When the video element is a character element, the source data for generating that character element may include at least one of the following: character model data, character motion data, character expression data, etc. Here, the character can be a virtual character or a real person character, and this disclosure does not specifically limit it.
[0036] In the case of a virtual character, the source data for that character element can be provided through artificial intelligence (AI) service devices.
[0037] When the character is a real person, the character elements can be provided through data acquisition devices. In this case, image acquisition devices can be used to capture image information of the real person and determine the character model data based on that image information; motion capture devices can be used to acquire motion data of the real person, obtaining character motion data; facial capture devices can be used to capture facial data of the real person, obtaining character expression data; and sound recording devices can be used to record the voice data of the real person, obtaining character voice data.
[0038] When a video element is an audio element, the source data for generating that audio element can be provided by a data acquisition device. In this case, the data acquisition device can be an audio recording device, such as a recorder or microphone.
[0039] For example, when the data acquisition device includes the aforementioned image acquisition device, AI service device, motion capture device, facial capture device, and audio recording device, then the data sources in the data source set can include: the image acquisition device, the AI service device, the motion capture device, the facial capture device, and the audio recording device. In this case, in response to a selection operation for multiple target data sources in the data source set, the source data output by those multiple target data sources can be obtained.
[0040] S103: The source data output by the multiple target data sources are encapsulated according to the set unified data protocol format to obtain the first encapsulated file.
[0041] The first encapsulation file includes multiple data blocks, each corresponding to one type of source data.
[0042] In this embodiment of the disclosure, the unified data protocol format can be understood as a format for simultaneously transmitting source data corresponding to multiple data blocks. In this case, the unified data protocol format may contain multiple pre-defined index information. Based on this, when encapsulating source data output from multiple target data sources according to the set unified data protocol format, the index information matching each data block can be determined within the unified data protocol format, and the source data corresponding to that data block can be placed at the position corresponding to the index information matching that data block.
[0043] For example, the unified data protocol format provided in this disclosure can be as follows: Figure 2 As shown. At this point, the index information included in this unified data protocol format can be Meta data, that is, as... Figure 2The metadata shown, for example, can be "motion metadata (i.e., motion metadata)," "facial metadata (i.e., facial metadata)," ..., "xx metadata (i.e., xx metadata)." In this case, when the source data output from multiple target data sources includes both motion data and facial data, "motion data" can be placed in the corresponding position of "motion metadata (i.e., motion metadata)" (e.g., ...). Figure 2 As shown, "motion data" is stored below "motion metadata," and "facial data" is placed in the corresponding position in "facial metadata (i.e., facial meta data)" (e.g., ...). Figure 2 As shown, the "facial data" is stored below the "facial metadata". At this point, the first encapsulated file can be obtained.
[0044] S105: Based on the first encapsulated file, perform video directing or video editing.
[0045] In this embodiment of the disclosure, after obtaining the first packaged file, the first packaged file can be transmitted to data receiving software for video directing or video editing. The data receiving software can be either video directing software or video editing software; this disclosure does not limit the application to either.
[0046] In practice, after the first packaged file is transmitted to the data receiving software, it can be split into multiple data blocks, each corresponding to a type of source data. Then, video directing or editing can be performed on this source data.
[0047] As described above, multiple target data sources can be selected from a data source set, and the source data output by each of these target data sources can be obtained. This allows for free selection of the source data used to generate the video, improving the flexibility and diversity of video processing. Next, the source data output by the multiple target data sources can be encapsulated according to a unified data protocol format to obtain a first encapsulated file. This enables the transmission of multiple source data according to a unified data protocol format, improving the stability and accuracy of data transmission. Subsequently, based on this first encapsulated file, video directing or editing processing can be performed on the multiple source data, thereby saving time and manpower costs in video generation and improving the efficiency of video processing.
[0048] In an optional implementation, for S105: based on the first encapsulated file, video directing processing is performed, specifically including the following process: inputting the first encapsulated file into the directing software, performing directing processing on the first encapsulated file through the directing software, and then outputting the live video; wherein, the directing processing includes the operation of data alignment of the source data of each data block in the first encapsulated file.
[0049] In this embodiment of the disclosure, after the first encapsulated file is input into the broadcast control software, the first encapsulated file can be split to obtain the source data corresponding to each data source.
[0050] Afterwards, the source data corresponding to each data source can be processed for broadcasting. This broadcasting process can be understood as a data alignment operation on the source data corresponding to each data source, so that the source data corresponding to each data source corresponds to the same time point.
[0051] For example, if the source data in the first encapsulation file includes motion data and voice data, assuming that motion data 1 and voice data 1 are acquired simultaneously during the time period from the 1st second to the 3rd second, then the start time of motion data 1 and voice data 1 can be aligned, and the end time of motion data 1 and voice data 1 can be aligned.
[0052] In addition, the directing processing performed on the first packaged file by the directing software may also include at least one of the following operations: adding, deleting, or adjusting video elements.
[0053] For example, when the video element is a scene element, the director's processing can be at least one of the following: adding a scene element, deleting a scene element, adjusting the position, size, orientation, etc. of the scene element.
[0054] In this embodiment of the disclosure, the live video can be used in live streaming scenarios or in recorded streaming scenarios. This disclosure does not make specific limitations in this regard, but rather follows the actual needs.
[0055] In the above embodiments, the first encapsulated file can be input into the broadcast software, and the broadcast software can be used to process the first encapsulated file for broadcasting, thereby enabling free editing of the first encapsulated file to obtain a higher quality and more compliant video.
[0056] In an optional implementation, after outputting the live video, this disclosure further includes the following steps:
[0057] Step S21: Record the live video using the broadcasting software to obtain a second encapsulated file under the unified data protocol format;
[0058] Step S22: Store the second encapsulated file into the data source set, as data in one of the data sources in the data source set.
[0059] In this embodiment, live video can be recorded in multiple tracks using broadcast control software, resulting in multiple track data points. Each track data point can correspond to a data source. It should be noted that the track data can correspond to a specific data source. For example, if the data source is an audio recording device, then the track data can be the source data provided by that audio recording device. These multiple track data points are then encapsulated according to a unified data protocol format to obtain a second encapsulated file. At this point, recording is complete (i.e., recording audio, motion, and other scene data, and then encapsulating it to obtain the recorded data in the unified data protocol format). This second encapsulated file can then be stored in a data source set, serving as data from one of the data sources in the set (e.g., a replay service within the data source set).
[0060] In the above embodiments, live video can be recorded to obtain a second encapsulated file under a unified data protocol format. This second encapsulated file is then stored in a data source set as data from one of the data sources in the set. This broadens the sources of source data for generating videos and increases the diversity of generated videos. Furthermore, the above embodiments also enable the reuse of live video, thereby reducing video recording costs.
[0061] In an optional implementation, for S105: based on the first encapsulated file, video editing processing is performed, specifically including the following steps:
[0062] Step S31: Input the first packaged file into the editing software, and after the editing software processes the first packaged file, record the edited video as a second packaged file under the unified data protocol format;
[0063] Step S32: Store the second encapsulated file into the data source set, as data in one of the data sources in the data source set.
[0064] In this embodiment of the disclosure, after detecting an editing operation on the first packaged file, the first packaged file is split into source data corresponding to each data source, and then the source data corresponding to each data source can be edited.
[0065] In this embodiment of the disclosure, the editing process can be any one of adding, deleting, or adjusting the source data corresponding to each data block in the first encapsulation file.
[0066] The operation of adjusting the source data can be either adjusting the order of the source data or adjusting the content of the source data.
[0067] For example, the motion data in the first packaged file can be adjusted using editing software. Let's assume the motion data corresponds to three actions: "Action 1, raising hand," "Action 2, shooting," and "Action 3, drinking water." The order of these three actions can be adjusted using editing software; for instance, the sequence can be changed from "Action 1 -> Action 2 -> Action 3" to "Action 3 -> Action 1 -> Action 2." Alternatively, the content of the actions can be adjusted using editing software (which can also be understood as adjusting the posture corresponding to the actions). For example, the height of the raised hand in Action 1 can be adjusted.
[0068] In this embodiment of the disclosure, when performing an add operation on the source data corresponding to a data block, the data source corresponding to the data block can first be determined from the data source set. Then, the source data output by the corresponding data source can be obtained. Afterward, the add operation on the source data corresponding to the data block can be performed based on the output source data.
[0069] In the above embodiments, the first packaged file can be input into editing software, and the first packaged file can be edited by the editing software, thereby enabling free editing of the source data in the first packaged file to obtain a recorded video that better meets the requirements.
[0070] In this embodiment of the disclosure, after editing the first encapsulated file, the edited video can be recorded as a second encapsulated file in a unified data protocol format. The unified data protocol format is the same as described in S103 above, and will not be described in detail here.
[0071] In an optional implementation, the process of storing the second encapsulated file into the data source set, as described in steps S22 and S32 above, specifically includes the following steps:
[0072] The second encapsulated file is split into source data corresponding to each of the data sources, and the split source data is stored in the replay service module in the data source set, and the replay service module is used as a new data source.
[0073] In the above embodiments, the second encapsulated file can be split into source data corresponding to each data source, thereby enabling the splitting of the video after editing or directing, resulting in multiple separate source data. Subsequently, the split source data can be stored in the replay service module of the data source set, and this replay service module can be used as a new data source. This allows for the reuse of at least some source data in the video after editing or directing, thus saving video recording costs.
[0074] In one optional implementation, the data sources in the data source set provided in this disclosure include at least one of the following target service devices: motion capture device, facial capture device, artificial intelligence (AI) service device, and audio recording device. The aforementioned target service device is used to provide source data for real-time output.
[0075] In this embodiment of the disclosure, the artificial intelligence (AI) service device may instruct a neural network model for generating at least one of motion data, voice data, and facial data. For example, the neural network model may be a GAN (Generative Adversarial Networks) model. This disclosure does not specifically limit the neural network model, but only to the extent that it can be implemented.
[0076] In an optional implementation, during the video directing or video editing process, the present disclosure embodiments may also switch the selected target data source in response to a data source switching request.
[0077] As described above, before performing video directing or editing, the first encapsulated file can be split into source data corresponding to each data source.
[0078] Next, you can select the source data to switch from the source data corresponding to each data source, and upon receiving a data source switching request, display a list of data sources corresponding to the data source set. Then, based on the selection operation of this data source list, you can determine the target data source.
[0079] In the above implementation, the target data source can be switched in response to a data source switching request, which not only simplifies the operation of switching target data sources and improves the user experience, but also allows for the free splicing of various source data by freely switching target data sources to obtain a wide variety of videos, increasing the diversity and interest of the generated videos.
[0080] In one optional implementation, the video processing method provided in this disclosure can be applied to scenarios involving live streaming or recorded broadcasts of digital humans. This will be described in detail below with reference to flowcharts.
[0081] like Figure 3 As shown, assuming the data source set includes: motion capture devices, facial capture devices, AI service devices (i.e., the aforementioned artificial intelligence AI service devices), audio recording devices, and replay service modules, then, in response to multiple target data sources selected from the data sources, the source data output by these multiple target data sources can be obtained. Subsequently, the obtained source data can be encapsulated according to a unified data protocol format, resulting in a first encapsulated file.
[0082] In one possible implementation, the first encapsulated file can be input into the digital human broadcasting software. Then, the first encapsulated file can be processed for broadcasting, resulting in a live video. This live video can then be broadcast live using the digital human broadcasting software. Alternatively, the live video can be recorded using the recording function of the digital human broadcasting software, resulting in a second encapsulated file.
[0083] In another possible implementation, the first encapsulated file can be input into digital human editing software. In this case, the first encapsulated file can be edited, and the edited video can be recorded to obtain a second encapsulated file.
[0084] Next, the second encapsulated file can be split to obtain the source data corresponding to each data source, and the source data can be stored in the replay service module.
[0085] In this scenario, in the context of digital human live streaming, the system can respond to data source switching requests and switch the target data source to the replay service module. This enables seamless switching between live streaming and recorded broadcasting for digital humans, thereby reducing the pressure on digital human live streaming and lowering labor costs.
[0086] Those skilled in the art will understand that, in the above-described method of the specific implementation, the order in which each step is written does not imply a strict execution order and does not constitute any limitation on the implementation process. The specific execution order of each step should be determined by its function and possible internal logic.
[0087] Based on the same inventive concept, this disclosure also provides a video processing apparatus corresponding to the video processing method. Since the principle of the apparatus in this disclosure for solving the problem is similar to that of the video processing method described above, the implementation of the apparatus can refer to the implementation of the method, and repeated details will not be repeated.
[0088] Reference Figure 4 The diagram shown is a schematic representation of a video processing apparatus provided in an embodiment of this disclosure. The apparatus includes: an acquisition unit 41, an encapsulation unit 42, and a processing unit 43; wherein,
[0089] The acquisition unit 41 is configured to, in response to selecting multiple target data sources from a data source set, acquire the source data output by the multiple target data sources respectively; each data source in the data source set is used to provide source data corresponding to at least one video element in the video to be generated; the video element includes at least one of scene elements, character elements, and sound elements;
[0090] The encapsulation unit 42 is used to encapsulate the source data output by the multiple target data sources according to a set unified data protocol format to obtain a first encapsulation file; the first encapsulation file includes multiple data blocks, each data block corresponding to one type of source data.
[0091] The processing unit 43 is used to perform video directing processing or video editing processing based on the first encapsulated file.
[0092] As described above, multiple target data sources can be selected from a data source set, and the source data output by each of these target data sources can be obtained. This allows for free selection of the source data used to generate the video, improving the flexibility and diversity of video processing. Next, the source data output by the multiple target data sources can be encapsulated according to a unified data protocol format to obtain a first encapsulated file. This enables the transmission of multiple source data according to a unified data protocol format, improving the stability and accuracy of data transmission. Subsequently, based on this first encapsulated file, video directing or editing processing can be performed on the multiple source data, thereby saving time and manpower costs in video generation and improving the efficiency of video processing.
[0093] In one possible implementation, the processing unit 43 is further configured to: input the first encapsulated file into the broadcast software, perform broadcast processing on the first encapsulated file through the broadcast software, and output the live video; wherein the broadcast processing includes the operation of data alignment of the source data of each data block in the first encapsulated file.
[0094] In one possible implementation, the processing unit 43 is further configured to: record the live video using the broadcasting software to obtain a second encapsulated file in a unified data protocol format; and store the second encapsulated file in the data source set as data in one of the data sources in the data source set.
[0095] In one possible implementation, the processing unit 43 is further configured to: input the first encapsulated file into editing software, edit the first encapsulated file using the editing software, record the edited video as a second encapsulated file under a unified data protocol format, and store the second encapsulated file into the data source set as data in one of the data sources in the data source set.
[0096] In one possible implementation, the device is further configured to: split the second encapsulated file into source data corresponding to each of the data sources, and store the split source data into a replay service module in the data source set, and use the replay service module as a new data source.
[0097] In one possible implementation, the data sources in the data source set include at least one of the following target service devices, which are used to provide source data for real-time output: motion capture device, facial capture device, artificial intelligence (AI) service device, and audio recording device.
[0098] In one possible implementation, the processing unit 43 is further configured to: switch the selected target data source in response to a data source switching request.
[0099] The processing flow of each module in the device and the interaction flow between each module can be referred to the relevant descriptions in the above method embodiments, and will not be detailed here.
[0100] Corresponding to Figure 1 In addition to the video processing methods described in this disclosure, this embodiment also provides a computer device 500, such as... Figure 5 The diagram shown is a structural schematic of a computer device 500 provided in an embodiment of this disclosure, including:
[0101] The system includes a processor 51, a memory 52, and a bus 53. The memory 52 stores execution instructions and includes main memory 521 and external memory 522. The main memory 521, also called internal memory, temporarily stores the processing data in the processor 51, as well as data exchanged with external memory such as a hard disk. The processor 51 exchanges data with the external memory 522 through the main memory 521. When the computer device 500 is running, the processor 51 communicates with the memory 52 through the bus 53, causing the processor 51 to execute the following instructions:
[0102] In response to selecting multiple target data sources from a data source set, the source data output by the multiple target data sources is obtained respectively; each data source in the data source set is used to provide source data corresponding to at least one video element in the video to be generated; the video element includes at least one of scene elements, character elements, and sound elements;
[0103] The source data output from the multiple target data sources are encapsulated according to a set unified data protocol format to obtain a first encapsulation file; the first encapsulation file includes multiple data blocks, each data block corresponding to one type of source data;
[0104] Based on the first encapsulated file, video directing or video editing is performed.
[0105] This disclosure also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, performs the steps of the video processing method described in the above-described method embodiments. The storage medium may be a volatile or non-volatile computer-readable storage medium.
[0106] This disclosure also provides a computer program product carrying program code. The program code includes instructions that can be used to execute the steps of the video processing method described in the above method embodiments. For details, please refer to the above method embodiments, which will not be repeated here.
[0107] The aforementioned computer program product can be implemented through hardware, software, or a combination thereof. In one optional embodiment, the computer program product is specifically embodied in a computer storage medium; in another optional embodiment, the computer program product is specifically embodied in a software product, such as a software development kit (SDK), etc.
[0108] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems and devices described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. In the several embodiments provided in this disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division; in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection may be through some communication interfaces; the indirect coupling or communication connection of devices or units may be electrical, mechanical, or other forms.
[0109] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0110] In addition, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0111] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0112] Finally, it should be noted that the above-described embodiments are merely specific implementations of this disclosure, used to illustrate the technical solutions of this disclosure, and not to limit it. The protection scope of this disclosure is not limited thereto. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features, within the scope of the technology disclosed in this disclosure. Such modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure, and should all be covered within the protection scope of this disclosure. Therefore, the protection scope of this disclosure should be determined by the protection scope of the claims.
Claims
1. A video processing method, characterized in that, include: In response to selecting multiple target data sources from a data source set, the system acquires the source data output by each of the multiple target data sources, wherein the data sources in the data source set are data acquisition devices that provide source data and / or artificial intelligence (AI) service devices that provide source data; each data source in the data source set is used to provide source data corresponding to at least one video element in the video to be generated; the video element includes at least one of scene elements, character elements, and sound elements; The source data output from the multiple target data sources are encapsulated according to a set unified data protocol format to obtain a first encapsulation file; the first encapsulation file includes multiple data blocks, each data block corresponding to one type of source data; Based on the first encapsulated file, the source data of each data block in the first encapsulated file is subjected to video directing processing or video editing processing to obtain the second encapsulated file. The video processing method further includes, during the video directing or editing process: The second encapsulated file is split to switch the selected target data source in response to a data source switching request.
2. The method according to claim 1, characterized in that, The step of performing video directing processing on the source data of each data block in the first encapsulation file based on the first encapsulation file includes: The first encapsulated file is input into the broadcast control software, and after the broadcast control software performs broadcast control processing on the first encapsulated file, the live video is output; wherein, the broadcast control processing includes the operation of data alignment of the source data of each data block in the first encapsulated file.
3. The method according to claim 2, characterized in that, The method further includes: The live video is recorded using the broadcast control software to obtain the second encapsulated file under the unified data protocol format; The second encapsulated file is stored in the data source set as data in one of the data sources in the data source set.
4. The method according to claim 1, characterized in that, The step of performing video editing processing on the source data of each data block in the first encapsulated file to obtain a second encapsulated file includes: The first packaged file is input into the editing software, and after the first packaged file is edited by the editing software, the edited video is recorded as the second packaged file under the unified data protocol format; The second encapsulated file is stored in the data source set as data in one of the data sources in the data source set.
5. The method according to claim 3 or 4, characterized in that, The step of storing the second encapsulated file into the data source set includes: The second encapsulated file is split into source data corresponding to each of the data sources, and the split source data is stored in the replay service module in the data source set, and the replay service module is used as a new data source.
6. The method according to claim 1, characterized in that, The data sources in the data source set include at least one of the following target service devices, which are used to provide source data for real-time output: Motion capture equipment, facial capture equipment, artificial intelligence (AI) service equipment, and audio recording equipment.
7. A video processing apparatus, characterized in that, include: The acquisition unit is configured to, in response to selecting multiple target data sources from a data source set, acquire the source data output by the multiple target data sources respectively, wherein the data sources in the data source set are data acquisition devices that provide source data and / or artificial intelligence (AI) service devices that provide source data; each data source in the data source set is used to provide source data corresponding to at least one video element in the video to be generated; the video element includes at least one of scene elements, character elements, and sound elements; The encapsulation unit is used to encapsulate the source data output by the multiple target data sources according to a set unified data protocol format to obtain a first encapsulation file; the first encapsulation file includes multiple data blocks, each data block corresponding to one type of source data. The processing unit is configured to perform video directing or video editing processing on the source data of each data block in the first encapsulated file, based on the first encapsulated file, to obtain a second encapsulated file. In the process of video directing or video editing, the processing unit is also used for: The second encapsulated file is split to switch the selected target data source in response to a data source switching request.
8. A computer device, characterized in that, include: The computer device includes a processor, a memory, and a bus. The memory stores machine-readable instructions executable by the processor. When the computer device is running, the processor communicates with the memory via the bus. When the machine-readable instructions are executed by the processor, they perform the steps of the video processing method as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, performs the steps of the video processing method as described in any one of claims 1 to 6.
Citation Information
Patent Citations
Encapsulation method and device of interactive video, and electronic equipment
CN113254393A
Content generation device and content generation program
JP2015060539A