A video processing method, electronic device and medium

By forming a recording group with multiple electronic devices, a collection of material clips is generated and preprocessed, which solves the problems of single-device video recording having limited shot variety and complex operation, enabling ordinary users to synthesize high-quality videos from multiple angles on their own.

CN116980765BActive Publication Date: 2025-12-19HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210803438.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2022-04-14
Filing Date
2022-07-07
Publication Date
2025-12-19
Estimated Expiration
2042-07-07

AI Technical Summary

Technical Problem

In existing technologies, when recording video using a single electronic device, the lens is limited, the sound pickup effect is poor, it is difficult to obtain ideal video recording data, and the operation is complicated, making it difficult to apply to the daily life scenarios of ordinary users.

Method used

By forming a recording group using multiple electronic devices, a collection of material clips is generated, which are then preprocessed and synthesized into a video. This simplifies the operation process and allows ordinary users to select material clips to synthesize videos.

Benefits of technology

There is no need to transfer the recorded data to a dedicated compositing device, which simplifies the operation process and allows ordinary users to compose videos themselves, improving the multi-angle video recording and audio quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116980765B_ABST
    Figure CN116980765B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of video processing, and discloses a video processing method, an electronic device and a medium. The video processing method comprises the following steps: a first processing device acquires multiple sets of recording data from multiple recording devices, wherein the multiple recording devices belong to a same recording group; the first processing device belongs to a device in the recording group; a first interface of the first processing device displays the multiple sets of recording data; in response to a first operation of a user, the first processing device pre-processes the multiple sets of recording data to generate a material segment set; a second interface of the first processing device displays each material segment in the material segment set; and in response to a second operation of the user, a third interface of the first processing device displays a synthesized video. Based on the above scheme, the recording data does not need to be transferred to an additional special synthesis device, thereby saving the operation process. Moreover, based on the material segments, the user's later synthesis steps can be greatly simplified, so that ordinary non-professional users can also select material segments to synthesize videos.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video processing, in particular to a video processing method, an electronic device and a medium. BACKGROUND

[0002] With the popularization of electronic devices with camera functions, users often use mobile phones, cameras and other electronic devices to record videos in various scenes, such as recording wedding, celebration, concert and other scenes. However, a single electronic device can usually only record at a single fixed angle, and the lens is single. If the photographer is too far away from the target, the sound pickup effect will be poor, and if the distance is too close, it will be difficult to obtain a suitable angle. Therefore, it is difficult to obtain ideal video recording data through a single electronic device.

[0003] To obtain better video recording data, some commonly used methods are to use two or more cameras and other electronic devices to record the same scene from multiple angles and multiple directions at the same time to obtain more comprehensive sound and picture materials. For example, as shown in Figure 1 Camera A is needed to shoot close-up scenes, camera B is needed to shoot long shots, and microphone C is needed to record sound. In post-production, different shooting pictures and recorded sounds are selected according to actual needs to synthesize videos.

[0004] However, the above scheme generally needs to arrange photographers in advance and divide the shooting positions and scenes to complete material collection, and needs to collect all the shooting materials to a professional video processing device to perform video editing and other operations to synthesize videos. The whole process is complex, and the above method can only be applied to professional camera teams and is difficult to be applied to ordinary users for shooting of daily life scenes. SUMMARY

[0005] To solve the technical problem that the whole process of the video processing and synthesis method is complex, and can only be applied to professional camera teams and is difficult to be applied to ordinary users for shooting of daily life scenes, the embodiments of the present application provide a video processing method, an electronic device and a medium.

[0006] In a first aspect, the present application provides a video processing method, which comprises:

[0007] The first processing device obtains a plurality of sets of recording data from a plurality of recording devices, wherein the plurality of recording devices belong to a same recording group; the first processing device belongs to the devices in the recording group; a first interface of the first processing device displays the plurality of sets of recording data; in response to a first operation of a user, the first processing device pre-processes the plurality of sets of recording data to generate a set of material clips, a second interface of the first processing device displays each material clip in the set of material clips; and in response to a second operation of the user, a third interface of the first processing device displays a synthesized video.

[0008] It can be understood that in the embodiments of the present application, the recording data generated by each electronic device in the recording group can be aggregated on any one or several electronic devices in the recording group for video synthesis processing, so that the recording data does not need to be transferred to an additional dedicated synthesis device, saving the operation process and facilitating operation.

[0009] In addition, the electronic device as the receiving device can pre-process all video data to generate material clips of various scenes corresponding to different targets and material tags corresponding to the material clips, so that the user's post-synthesis step in the video synthesis application can be greatly simplified, and ordinary non-professional users can also select material clips to synthesize videos by themselves. In the embodiments of the present application, there can be multiple receiving devices in the plurality of electronic devices, and then multiple users can synthesize videos according to their own needs.

[0010] In a possible implementation, the first operation is an operation of triggering a pre-processing control of the first interface.

[0011] It can be understood that the first interface in the embodiments of the present application can be the material display interface mentioned in the embodiments of the present application.

[0012] In a possible implementation, when the pre-processing control is clicked, a selection submenu can be displayed in the material display interface, and the selection submenu can display a clip generation control, a subtitle generation control and a barrage generation control. When the clip generation control, the subtitle generation control and the barrage generation control are selected and the pre-processing control is clicked, the electronic device can pre-process all recording data in the recording group material folder to generate different scenes of material clips corresponding to each target person, subtitle clips corresponding to each material clip, and barrage clips corresponding to non-target persons.

[0013] In a possible implementation, the second interface includes a track area, and the second operation is an operation of selecting at least one material clip in each material clip to a corresponding position of the track area and triggering a preview control.

[0014] It can be understood that the first interface in the embodiments of the present application can be the interface corresponding to the timeline tab in the video processing interface in the embodiments of the present application.

[0015] It can be understood that when the user clicks the preview control, the electronic device can perform video synthesis according to the material segments of the track area and the positions corresponding to the material segments.

[0016] In a possible implementation, the track area includes a video track, an audio track, a subtitle track, and a barrage track.

[0017] The video track includes a plurality of video tracks of different shots.

[0018] It can be understood that the plurality of video tracks can include tracks of various shots, such as a long shot track, a close-up shot track, and a close-up shot track, etc. It is convenient to better perform video synthesis and superimpose videos of multiple shots, for example, a long shot video and a close-up shot video can be superimposed in the same time period to present the superimposed video effect.

[0019] In a possible implementation, the preview area of the third interface displays the synthesized video.

[0020] It can be understood that the third interface in the embodiments of the present application can be an interface corresponding to a preview tab in the video processing interface mentioned in the embodiments of the present application.

[0021] The interface corresponding to the preview tab includes a preview area, and the preview area can be used to display and play the synthesized video.

[0022] In a possible implementation, the preview area includes one or more display areas, and each display area displays a corresponding synthesized video according to the material segments of the corresponding track area.

[0023] It can be understood that in some embodiments, the preview area of the electronic device can also display a plurality of display areas in split screen, and each display area can select different pictures according to user demand.

[0024] In some embodiments, the preview area can include a first display area and a second display area. When the user clicks the first display area, the user can edit the preview content of the first display area, that is, make any adjustment such as increase, decrease, and move to the material segments of the track area corresponding to the first display area. When the user clicks the second display area, the user can edit the preview content of the second display area, that is, make any adjustment such as increase, decrease, and move to the material segments of the track area corresponding to the second display area. In this way, the user can compare multiple preview video effects to select a suitable synthesized video.

[0025] In a possible implementation, the preview area includes a plurality of display areas, and a part of the plurality of display areas displays the synthesized video, and another part of the plurality of display areas displays the material segments selected by the user.

[0026] It can be understood that in some embodiments, the preview area of the electronic device can also display multiple display areas in split screen mode, and some display areas can be used to preview the synthesized video, and other display areas can be used to magnify the user-selected material segment for the user to watch.

[0027] In a possible implementation, the video of the same time point in the synthesized video includes video of at least one view.

[0028] It can be understood that in some embodiments, the synthesized video displayed in the preview area can be superimposition of videos of multiple views, for example, superimposition of long shot video and close-up video, etc. In this way, the video content can be enriched.

[0029] In a possible implementation, in response to a third operation of the user, the synthesized video is switched from the first synthesized video to the second synthesized video.

[0030] It can be understood that in some embodiments, the panoramic video can be automatically displayed in the preview area or each display area, and a selection switching function is added to the video synthesis interface, for example, the panoramic video of a certain time period can be switched to a close-up video, etc.

[0031] In some embodiments, the first synthesized video can be a panoramic video, that is, the preview area can automatically display the panoramic video. The video synthesis interface is provided with a switching control, and when the user clicks the switching control, the video synthesis interface can pop up other view videos existing in the current corresponding time period, for example, there are long shot target X1 video and close-up target X1 video, and when the user clicks the corresponding other view video, the other view video of the time period, that is, the second synthesized video, will replace the panoramic video.

[0032] It can be understood that in some embodiments, the first synthesized video can be a current synthesized video, and when the user clicks the switching control, the video synthesis interface can pop up each view video existing in the current corresponding time period, and when the user clicks the corresponding view video, the other view video of the time period, that is, the second synthesized video, will replace the current synthesized video.

[0033] In a possible implementation, the third operation is an operation of clicking the switching control on the third interface.

[0034] In a possible implementation, the first processing device, in response to a fourth operation of the user on the target material segment, prompts the region position corresponding to the target material segment in the track area.

[0035] It can be understood that in the embodiments of the present application, the first processor device can display the corresponding position of the target material segment in the track area according to the time mark information of the target material segment. In this way, it is convenient for the user to drag the material label to the corresponding track area.

[0036] In a possible implementation, the fourth operation is a click operation.

[0037] In a possible implementation, the region position corresponding to the target material segment is prompted in the track region, including:

[0038] The region position corresponding to the target material segment is highlighted in the track region.

[0039] It can be understood that, in the embodiments of the present application, the prompt manner of the region position corresponding to the target material segment in the track region can be any manner such as highlighting or color highlighting.

[0040] In a possible implementation, the first processing device receives time mark information and recording content labels corresponding to the recording data sent by the plurality of recording devices.

[0041] It can be understood that, in the embodiments of the present application, the recording content labels corresponding to the recording data can be labels obtained by each electronic device identifying the data recorded by itself, or can be labels freely selected by the user through the group interface operation of the electronic device. The time mark information corresponding to each recording data can be obtained by each electronic device recording the time of the data recorded by itself.

[0042] In a possible implementation, the first processing device identifies the plurality of groups of recording data to obtain the recording content labels of each group of recording data.

[0043] It can be understood that, in some embodiments, the recording content labels corresponding to the recording data can be labels obtained by the receiving device uniformly identifying the recording data after obtaining the recording data.

[0044] In a possible implementation, the first processing device identifies the plurality of groups of recording data to obtain the time mark information corresponding to the plurality of groups of recording data.

[0045] It can be understood that, in some embodiments, the electronic devices in the recording group can also not obtain the time mark information, but the receiving device can align the time points by identifying the audio and video content of different recording data after receiving each recording data, so as to obtain uniform time mark information.

[0046] In a possible implementation, the time mark information includes the start time and / or end time and / or special mark time corresponding to each recording data.

[0047] It can be understood that during the recording process, the electronic device can automatically record the start time and end time of the recording. The user can mark a special time point during the recording of the video when recording a key part or a moment that the user wants to record, to obtain a special marked time. The special time point marking method can be various, for example, through a voice instruction method, the user can send a "mark" voice instruction, and the electronic device can mark the current time point when the voice instruction is recognized. It can also be any other implementable method.

[0048] In a possible implementation, the recording content label includes a scene label and a recording category label.

[0049] In a possible implementation, the scene label includes a long shot label, a medium shot label, a close-up label, and a close-up label.

[0050] The recording category label includes a video label and an audio label.

[0051] In a possible implementation, the material segment set includes video material segments of different scenes corresponding to each target respectively, audio material segments corresponding to each target respectively, material labels corresponding to the video material segments and the audio material segments respectively, time mark information, and subtitle information included in the plurality of sets of recording data.

[0052] In a possible implementation, the material segment set further includes barrage information corresponding to audio of non-targets included in the plurality of sets of recording data, and time mark information corresponding to the barrage information.

[0053] In a possible implementation, the plurality of sets of recording data are preprocessed to generate the material segment set, including:

[0054] The video data in each set of recording data is processed to obtain video material segments of different scenes corresponding to each target respectively, and material labels and time mark information corresponding to the video material segments.

[0055] The audio data in each set of recording data, and the audio stream data in the video data of each set of recording data are audio processed to obtain audio segments corresponding to each target respectively, and material labels, time mark information, and subtitle information corresponding to the audio segments.

[0056] It can be understood that in the embodiments of the present application, the manner of audio processing of the audio data can include pre-processing the data using a multi-channel joint processing algorithm, wherein the multi-channel joint processing algorithm can include a multi-channel noise reduction algorithm, a speech separation algorithm, and a speech recognition algorithm, etc. In some embodiments, the manner of audio processing of each data can be to first perform noise reduction processing on the data by a multi-channel noise reduction algorithm, and then perform separation processing on the data obtained after noise reduction processing by a speech separation algorithm, to obtain audio clips corresponding to each target respectively, and obtain the material tags, time mark information and subtitle information corresponding to the audio clips.

[0057] In the embodiments of the present application, the manner of processing the video data can include processing each recording data using a pedestrian re-identification algorithm to obtain video material clips of different scenes corresponding to each target respectively. Then the material tags and time mark information corresponding to the video material clips are obtained.

[0058] It can be understood that in some embodiments, when the synthesized video is played, the corresponding audio can be played when the picture switches to the picture corresponding to the scene.

[0059] It can be understood that the synthesized video mentioned in the embodiments of the present application includes video stream data and audio stream data, that is, the corresponding audio can be played at the same time as the video picture is played.

[0060] In some embodiments, after previewing the video, the electronic device can play the synthesized video in response to the user's play operation. Wherein, the electronic device can play in any implementable form based on the user's selection, for example, it can be played at any play speed, etc.

[0061] In a possible implementation, the plurality of recording devices includes a first processing device.

[0062] It can be understood that in the embodiments of the present application, any device in the recording group, for example, the first processing device, can only receive video data sent by other devices for processing without recording, or can both record and receive video data sent by other devices for processing.

[0063] In a possible implementation, the devices in the recording group include first type devices, or first type devices and second type devices;

[0064] The first type device is a device with wireless network function, the first type devices are connected to the same wireless network, and the first processing device is a first type device;

[0065] The second type device is a device without wireless network function but with Bluetooth function, and the second type device is connected to the first type device through the Bluetooth function.

[0066] In a possible implementation, the first processing device sends the recording data acquired by the first processing device to a second processing device in the recording group.

[0067] It can be understood that in the embodiments of the present application, the first processing device can also send the acquired recording data to other receiving devices in the recording group.

[0068] In a second aspect, the embodiments of the present application provide an electronic device, comprising: a data transceiver module, configured to acquire a plurality of sets of recording data from a plurality of recording devices, wherein the plurality of recording devices belong to a same recording group; a first processing device belongs to a device in the recording group; a data processing module, configured to, in response to a first operation of a user, pre-process the plurality of sets of recording data to generate a set of material segments; and the data processing module is configured to, in response to a second operation of the user, generate a synthesized video based on a material segment selected by the user.

[0069] In a possible implementation, the electronic device further comprises a data acquisition module, configured to acquire the recording data.

[0070] In a possible implementation, the electronic device further comprises a center control module, configured to, in response to a group creation operation of a user, create a recording group;

[0071] In a possible implementation, the center control module is further configured to, in response to a join existing group operation of the user, control the electronic device to join an existing recording group.

[0072] In a third aspect, the present application provides an electronic device, comprising: a memory, configured to store instructions executed by one or more processors of the electronic device, and a processor, which is one of the one or more processors of the electronic device, and is configured to execute the video processing method mentioned in the present application.

[0073] In a fourth aspect, the present application provides a readable storage medium, characterized in that the readable medium stores instructions, and the instructions, when executed on an electronic device, cause the electronic device to execute the video processing method mentioned in the present application.

[0074] In a fifth aspect, the present application provides a computer program product, characterized in that the computer program product comprises instructions, and the instructions, when executed on an electronic device, cause the electronic device to execute the video processing method mentioned in the present application. BRIEF DESCRIPTION OF DRAWINGS

[0075] Figure 1 According to some embodiments of the present application, a recording scene schematic diagram is shown;

[0076] Figure 2 According to some embodiments of the present application, a schematic diagram of a video shooting method is shown;

[0077] Figure 3 According to some embodiments of the present application, a recording scene schematic diagram is shown;

[0078] Figure 4a According to some embodiments of the present application, a video project interface schematic diagram of an electronic device is shown;

[0079] Figure 4b According to some embodiments of the present application, a group selection interface schematic diagram of an electronic device is shown;

[0080] Figure 4c According to some embodiments of the present application, a new group interface schematic diagram of an electronic device is shown;

[0081] Figure 4d According to some embodiments of the present application, a WIFI setting interface schematic diagram of an electronic device is shown;

[0082] Figure 4e According to some embodiments of the present application, a group interface schematic diagram of an electronic device is shown;

[0083] Figure 4f According to some embodiments of the present application, a group interface schematic diagram of an electronic device is shown;

[0084] Figure 4g According to some embodiments of the present application, a video project interface schematic diagram of an electronic device is shown;

[0085] Figure 4h According to some embodiments of the present application, a material display interface schematic diagram of an electronic device is shown;

[0086] Figure 4i According to some embodiments of the present application, a video processing interface schematic diagram of an electronic device is shown;

[0087] Figure 4j According to some embodiments of the present application, a video processing interface schematic diagram of an electronic device is shown;

[0088] Figure 4k According to some embodiments of the present application, a video processing interface schematic diagram of an electronic device is shown;

[0089] Figure 4l According to some embodiments of the present application, a group interface schematic diagram of an electronic device is shown;

[0090] Figure 4m According to some embodiments of the present application, a video processing interface schematic diagram of an electronic device is shown;

[0091] Figure 4n According to some embodiments of the present application, a schematic diagram of a video processing interface of an electronic device is shown;

[0092] Figure 4o According to some embodiments of the present application, a schematic diagram of a video processing interface of an electronic device is shown;

[0093] Figure 4p According to some embodiments of the present application, a schematic diagram of a video processing interface of an electronic device is shown;

[0094] Figure 4q According to some embodiments of the present application, a schematic diagram of a video processing interface of an electronic device is shown;

[0095] Figure 4r According to some embodiments of the present application, a schematic diagram of a video processing interface of an electronic device is shown;

[0096] Figure 5 According to some embodiments of the present application, a schematic diagram of a structure of an electronic device is shown;

[0097] Figure 6 According to some embodiments of the present application, a schematic diagram of a video processing method is shown;

[0098] Figure 7 According to some embodiments of the present application, a schematic diagram of connection of a plurality of electronic devices is shown;

[0099] Figure 8 According to some embodiments of the present application, a schematic diagram of an existing group interface of an electronic device is shown;

[0100] Figure 9 According to some embodiments of the present application, a schematic diagram of a recording interface of an electronic device is shown;

[0101] Figure 10 According to some embodiments of the present application, a schematic diagram of data transmission between electronic device A1, electronic device B1 and electronic device C1 in a recording group is shown;

[0102] Figure 11 According to some embodiments of the present application, a schematic diagram of a video processing interface of an electronic device is shown;

[0103] Figure 12 According to some embodiments of the present application, a schematic diagram of a structure of an electronic device is shown. DETAILED DESCRIPTION

[0104] Illustrative embodiments of the present application include, but are not limited to, a video processing method, an electronic device, and a medium.

[0105] As described above, in the prior art, the video shooting scheme is complex to operate, and can only be applied to professional camera teams, and is difficult to be applied to ordinary users for shooting daily life scenes.

[0106] To solve the above problems, in some embodiments, a video shooting method is provided, as shown in the figure. Figure 2 The shooting method includes: forming a local area network by a plurality of electronic devices, wherein the plurality of electronic devices can include a control terminal for sending a control instruction, a teleprompter for playing subtitles during shooting of video files by all other shooting terminals, and a plurality of shooting terminals for shooting videos, such as a first shooting terminal, a second shooting terminal, a third shooting terminal, and a fourth shooting terminal. The plurality of electronic devices can form a local area network by being connected under the same WiFi or by being connected to each other through Bluetooth, and the control terminal can send a clock instruction for unified shooting time to other shooting terminals through built-in online management software, so as to ensure that all shooting terminals shoot at the same time, and the teleprompter plays subtitles during shooting of video files by all other shooting terminals. Finally, the videos shot by all shooting terminals are manually copied and summarized on a professional video editing device for synthesis processing.

[0107] The above scheme can realize simultaneous shooting of video files by the control terminal and other shooting terminals through online connection of the control terminal and other shooting terminals. In this way, manual arrangement is reduced to a certain extent, and the operation process is simplified.

[0108] However, in the above shooting process, if any shooting terminal delays the start of shooting due to failure to receive a shooting instruction or other reasons, the shooting data cannot be aligned with the data shot by other shooting terminals, affecting subsequent video editing and synthesis.

[0109] In addition, when synthesizing the video, all shot video files still need to be summarized on another video editing device through manual copying, and all audio and video materials collected are preprocessed, edited, spliced, and other complex operations by professional operators. In addition, the specific content of the video data can be known only after all video data are watched during video processing. Therefore, a lot of time is also needed in the process of selecting materials for video synthesis. In summary, the above scheme still has the problems of complex scheme and difficulty in being applied to ordinary users.

[0110] To solve the above problems, the embodiment of the present application provides a video processing method, comprising: a plurality of electronic devices establish a recording group by connecting in the same wireless network or by Bluetooth connection and the like; after the recording group is established, any electronic device in the plurality of electronic devices can perform audio and / or video recording to obtain recording data, and each electronic device can obtain time mark information corresponding to the recording data and a recording content label; then the recording data and the corresponding time mark information and the recording content label are stored correspondingly, and the recording data and the corresponding time mark information and the recording content label are sent to a receiving device in the plurality of devices. It can be understood that each electronic device in the recording group can become a receiving device.

[0111] The receiving device can obtain the recording data and the corresponding time mark information and the recording content label sent by other devices in the recording group; and pre-process each recording data based on the time mark information and the recording content label of each recording data obtained, to generate a material segment set, wherein the material segment set can include material segments of different scenes corresponding to each target in the recording data respectively, and a material label and time mark information corresponding to the material segment. And the receiving device can send the material segment set to a video application program in the receiving device. And display the material segment in the video application program in a video synthesis interface, and display the material segment on a unified time axis based on the time mark information of each material segment. Then the receiving device synthesizes a corresponding video based on the material segment selected by the user.

[0112] It can be understood that in some embodiments, the plurality of devices can also record respectively to obtain recording data, and then establish a recording group. The non-receiving devices of the plurality of devices send the recording data to the receiving device in the plurality of devices, the receiving device receives the recording data sent by other non-receiving devices, and pre-processes each recording data received and the recording data obtained by the receiving device itself to obtain a material segment set.

[0113] It can be understood that the electronic devices in the recording group can be divided into two categories. The first category of devices are mobile phones, tablet computers and the like with wireless network function, wherein the devices with wireless network function can refer to devices with WiFi function or mobile communication and the like wireless network function. The first category of devices can join or create a recording group by connecting in the same wireless network. The second category of devices are microphones, Bluetooth earphones and the like without wireless network function but with Bluetooth function. Among them, the second category of devices can join the recording group through the first category of devices as relay nodes.

[0114] It can be understood that, since the receiving device in the recording group in the embodiment of the present application needs to perform video synthesis through the video application program, in some embodiments, only the electronic device that can install the video application program can become the receiving device. The way to become the receiving device can be that each electronic device selects to become the receiving device in the recording group, or each electronic device sends a pop-up window to ask the user during recording, and selects based on the user's determination.

[0115] It can be understood that, in some embodiments, the electronic devices in the recording group can also transmit the recording data to the cloud, so that any electronic device in the recording group can download each recording data from the cloud to perform video synthesis.

[0116] It can be understood that, in the embodiment of the present application, the recording data generated by each electronic device in the recording group can be aggregated on any one or several electronic devices in the recording group for video synthesis processing, so that it is not necessary to transfer the recording data to an additional dedicated synthesis device, saving the operation process and facilitating the operation.

[0117] It can be understood that the time mark information can include the recording start time and / or the recording end time recorded by the electronic device. The time mark information is used to align the time of each video data or audio data for facilitating post-editing and synthesis. It can be understood that, in some embodiments, the recording start time and the recording end time in the time mark information can be standard network time.

[0118] In this way, if the shooting of any electronic device is interrupted or not unified with other electronic devices, so that a non-full-time period video is obtained, the non-full-time period video can still be aligned with the video in the corresponding time period of other complete videos according to the time mark information of the non-full-time period video, i.e., the recording start time and the recording end time, during video processing.

[0119] In some embodiments, the time mark information can also include a special time point mark. For example, the user can perform a special time point mark when recording a key part or an important time point that the user wants to record during recording the video. The way of the special time point mark can be various, for example, through a voice instruction mode, the user can send a voice instruction of "mark", and the electronic device can perform a special mark on the current time point when recognizing the voice instruction. It can be understood that the above-mentioned way of performing a special time point mark through a voice instruction is only an example, and the way of performing a special time point mark in the embodiment of the present application can also be any implementable way such as a key mode.

[0120] In some embodiments, the electronic devices in the recording group can also not acquire the time mark information, but receive the devices to acquire the unified time mark information by identifying the audio and video content of different recording data to realize the alignment of time points after receiving each recording data. For example, the starting time of the longest time recording data can be set as 00:00, and then the ending time is determined according to the time length of the longest time recording data. In this way, by identifying the audio or video of the recording data, the starting time or ending time of other recording data is determined, for example, the 00:06 time point of the longest time recording data of the starting time point of the recording data E is aligned, then the starting time of the recording data E can be set as 00:06, and then the ending time is determined according to the time length of the recording data E.

[0121] It can be understood that the recording content label described above can include a category label reflecting the recording category and a scene label reflecting the main shooting scene in the video data, etc. Among them, the category label can include a video or audio label; the scene label can include one or several of medium shot, close-up and long shot, close-up label, etc. Among them, different scenes can be understood as different content or different scenes. Different scenes can be different corresponding lens pictures, for example, close-up can refer to a picture in which the target object occupies more than half of the picture frame when the target object is represented. Medium shot can refer to a picture that can describe most of the area of the target object, for example, a picture of the upper part of the knee of a person, and the viewing distance is slightly farther than close-up, which is beneficial to display the shape and action of the person. Long shot can refer to a picture that includes distant scenery and target objects, which can be used to show the spatial background or environmental atmosphere of the target object activity. Close-up can refer to a detail picture of the target object. In some embodiments, in different scenes, the proportion of the same object in the picture is different. For example, the proportion of the same person in close-up, close-up, medium shot and long shot decreases in turn.

[0122] It can be understood that the recording content label described above can be a label freely selected by the user through the relevant interface operation of the electronic device, so that the user can record according to the label.

[0123] In some embodiments, the recording content label described above can also be used to identify the data recorded by each electronic device, or the label obtained by unified identification of each recording data after the recording data is obtained by the receiving device. For example, each electronic device can determine whether the recording data is audio data or video data according to the content of the recording data to obtain the category label of the recording data. Each electronic device can determine the scene label of the video data according to the size of the target element in the recorded video data, for example, when the size of the target person in the video data is greater than a set value, it can be determined that the video data is close-up data, and the corresponding scene label is close-up label; when the size of the target object in the video data is less than a set value, it can be determined that the video data is long shot data, and the corresponding scene label is long shot label, wherein the target object can be a target person, or a target object, etc.

[0124] In some embodiments, the recording content label described above can also be used to identify the data recorded by each electronic device, or the label obtained by unified identification of each recording data after the recording data is obtained by the receiving device. For example, each electronic device can determine whether the recording data is audio data or video data according to the content of the recording data to obtain the category label of the recording data. Each electronic device can determine the scene label of the video data according to the size of the target element in the recorded video data, for example, when the size of the target person in the video data is greater than a set value, it can be determined that the video data is close-up data, and the corresponding scene label is close-up label; when the size of the target object in the video data is less than a set value, it can be determined that the video data is long shot data, and the corresponding scene label is long shot label, wherein the target object can be a target person, or a target object, etc.

[0125] It can be understood that through the recording content label, the recording content of each recording data can be directly obtained, without the need to watch all video data to know the specific content of the video data, which facilitates subsequent video processing.

[0126] It can be understood that the above-mentioned pre-processing of all video data based on the time mark information and the recording content label of each recording data can generate material segments including segments corresponding to different scenes and different targets, and the material segment label is a label marking the content of the material segment. For example, the material segment can include a medium shot segment corresponding to target X1 and a long shot segment corresponding to target X2, and the corresponding material segment label can be medium shot target X1 and long shot target X2.

[0127] In some embodiments, the electronic device can pre-process all video data based on the time mark information and the video recording label of each video data to generate material segments and the labels corresponding to the material segments; which can include:

[0128] The electronic device can preliminarily classify the recording data based on the recording content label, for example, the recording data can be divided into an audio set and a video set according to the category label;

[0129] The audio processing is performed on each data in the audio set to obtain corresponding sound material segments of different target persons, and labels and time mark information corresponding to the sound material segments are generated. For example, the material segments in the audio set can include sound material segments corresponding to target X1 and sound material segments corresponding to target X2, and the corresponding material labels can be target X1 audio and target X2 audio, respectively.

[0130] The manner of audio processing on each data in the audio set can include pre-processing the data using a multi-channel joint processing algorithm, wherein the multi-channel joint processing algorithm can include a multi-channel noise reduction algorithm, a speech separation algorithm, and a speech recognition algorithm, etc. In some embodiments, the manner of audio processing on each data can be to first perform noise reduction processing on the data by a multi-channel noise reduction algorithm, and then perform separation processing on the data obtained after the noise reduction processing by a speech separation algorithm to obtain the speech corresponding to each target person or non-target person.

[0131] In some embodiments, when the recorded video is played, the corresponding audio can be played when the screen switches to the screen corresponding to the scene. For example, when the playing screen of the video is a close-up screen of target X1, the sound can be automatically switched to the sound of target X1, and the sounds of other targets can be considered as noise and reduced or removed, for example, the sound of target X2 and / or the environmental sound can be reduced or removed.

[0132] Then, according to the scene labels of each data in the video data set, each data in the video set can be divided into a long shot sub-set, a close-up sub-set, and a close-up sub-set, etc. Then, each recording data in each sub-set is identified and extracted to obtain material segments corresponding to different target objects in each sub-set and material labels. For example, the long shot sub-set in the video set can include material segments corresponding to target X1 and material segments corresponding to target X2, etc., and the material labels can be long shot target X1 video and long shot target X2 video.

[0133] In the embodiments of the present application, the manner of processing each recording data in the video set can include processing each recording data using a pedestrian re-identification algorithm to obtain material segments corresponding to different target persons in each sub-set. It can be understood that in the embodiments of the present application, if a same material segment includes multiple target persons, the material segment can belong to the material segments corresponding to multiple target persons, respectively. For example, a material segment includes target X1 and target X2, and the material segment belongs to the material segment corresponding to target X1 and the material segment corresponding to target X2.

[0134] It can be understood that the video synthesis interface in the video application program can display the above-mentioned material segments so as to facilitate the user to select the material segments for video synthesis.

[0135] In the embodiment of the present application, the electronic device as the receiving device can preprocess all the video data according to the time mark information of each video data and the recording content label, generate the material segments of various scenes corresponding to different targets and the material labels corresponding to the material segments, so that the post-synthesis step of the user in the video synthesis application program can be greatly simplified, and even the ordinary non-professional user can also select the material segments for video synthesis. In the embodiment of the present application, there can be multiple receiving devices in multiple electronic devices, and then multiple users can synthesize videos according to their own needs.

[0136] The above-mentioned video processing method in the present application can be applied in any scene where the video recording using the electronic device is needed, such as the wedding, the celebration, the concert, the competition site and the like. The following will take the recording of the speech site by the electronic device A, the electronic device B and the electronic device C as an example to describe the video processing scheme of the present application.

[0137] As shown in Figure 3 , the electronic device A1, the electronic device B1, the electronic device C1 and the electronic device D1 record the speech site, the speech person includes the target X1 and the target X2, the electronic device A1, the electronic device B1 and the electronic device C1 can all be mobile phones, and the electronic device D1 can be a microphone. The electronic device A1, the electronic device B1 and the electronic device C1 are simultaneously connected to the wireless network with the account name of "PH888". In some embodiments, the electronic device D1 can not have the wireless network function, but have the Bluetooth function, and the electronic device D1 can be connected to the electronic device B1 through the Bluetooth, so as to transmit the recorded audio data to other devices in the wireless network.

[0138] At this time, the electronic device A1, the electronic device B1, the electronic device C1 and the electronic device D1 first perform networking. For example, in some embodiments, as shown in Figure 4a , the user can open the video synthesis application program through the electronic device A1, and when the new video project control 4011 in the video project interface 401 of the video synthesis application program is clicked, the electronic device A1 will jump to the group selection interface 402 as shown in Figure 4b , the group selection interface 402 can display the new group control 4021 and the search existing group control 4022. When the new group control is clicked, it will jump to the new group interface 403 as shown in Figure 4c . When the search existing group control 4022 is clicked, the existing groups will be displayed.

[0139] In some embodiments, the way of creating a new recording group can also be any implementable way of creating a group, such as scanning a two-dimensional code, creating a group face to face, inputting the same password, and the like.

[0140] The new group creation interface 403 can display a WIFI control 4031 and the name of the current wireless network to which the electronic device A1 is connected. When the WIFI control 4031 is clicked, the electronic device A1 will jump to the WIFI setting interface 002 as shown in Figure 4d The new group creation interface 403 can also display the icons and names of the electronic devices connected to the wireless network with the account name "PH888", for example, the icons 4032, 4033 and 4034 and the corresponding names PHA, PHB and PHC of the electronic devices A1, B1 and C1 respectively. The new group creation interface 403 can also display the icons and names of other electronic devices connected to the electronic devices through Bluetooth, for example, the icon 4035 and the corresponding name PHD of the electronic device D1 connected to the electronic device B1. The new group creation interface 403 can also display a group name control 4036 and a new group control 4037. When the group name control 4036 is clicked, the group name can be changed and set, for example, the group name is set to "recording group 1". When the new group control 4037 is clicked, a recording group can be created.

[0141] When the user clicks the new group control 4037, the electronic device A1 can display the group interface 404 as shown in Figure 4e The group interface can display the icons of the electronic devices A1, B1, C1 and D1, which are the icons 4032, 4032, 4034 and 4035 respectively. The group interface 404 can also display a label row, such as a close-up label row 4044, a medium shot label row 4045, a long shot label row 4046, an audio label row 4047, and the like. In addition, the group interface 404 can display a receiving device row 4048 and a start recording control 4049. Figure 4fAs shown, each user can select the recording content label by dragging the icon corresponding to his or her electronic device or other electronic device to the corresponding label row, and each user can also record according to the recording content label of each electronic device. In addition, the receiving device can also be determined by dragging the electronic device to the receiving device row 4048. For example, user A is the owner of electronic device A1, and user A can drag the icon 4041 corresponding to electronic device A1 to the close-up label row 4044 and the receiving device row 4048, respectively, drag the icon 4042 corresponding to electronic device B1 to the long shot label row 4046, drag the icon 4042 corresponding to electronic device C1 to the audio label row 4047, and drag the icon 4042 corresponding to electronic device D1 to the audio label row 4047.

[0142] It can be understood that in some embodiments, electronic device A1, electronic device B1, and electronic device C1 can all display any interface shown in FIG. 4B. Electronic device D1 cannot display the interface shown in FIG. 4B because it has no display, and therefore, electronic device D1 can control the recording process through the connected relay electronic device electronic device B1. For example, recording can be started synchronously with electronic device B1. Figures 4a-4r Figures 4a-4q It can be understood that in some embodiments, electronic device A1, electronic device B1, and electronic device C1 can all display any interface shown in FIG. 4B. Electronic device D1 cannot display the interface shown in FIG. 4B because it has no display, and therefore, electronic device D1 can control the recording process through the connected relay electronic device electronic device B1. For example, recording can be started synchronously with electronic device B1.

[0143] When each user clicks the start recording control on the corresponding electronic device, the camera or microphone of the electronic device itself can be controlled to start recording according to the selected recording category label. For example, user A clicks the start recording control 4049 of electronic device A1, and electronic device A1 can control the camera or microphone of electronic device A1 to start recording according to the selected recording category label. For example, when the recording content label selected by electronic device A1 is the audio label, electronic device A1 can control the microphone of electronic device A1 to start recording, and when the recording content label selected by electronic device A1 is the medium shot label, the long shot label, or other video category label, electronic device A1 can control the camera of electronic device A1 to start recording. The recording start time and recording end time are recorded, and the recorded content can be stored in correspondence with the recording content label corresponding to the label row where electronic device A1 is located.

[0144] In some embodiments, when user A clicks the start recording control 4049 of electronic device A1, electronic device A1 can not record, but as a receiving device, electronic device A1 can receive the recording data and corresponding recording content label and time mark information sent by electronic device B1, electronic device C1, and electronic device D1. In some embodiments, when user A clicks the start recording control 4049 of electronic device A1, electronic device A1 can also control electronic device B1, electronic device C1, and electronic device D1 to start recording.​

[0145] It is understandable that in some embodiments, a user may not shoot based on the selected scene type for some time. For example, the user may have selected a distant view as the scene type, but during the shooting process, the user may shoot close-ups for some time. In this case, when the receiving device A1 identifies and extracts the recorded data from each subset, it will also extract footage clips with different scene types corresponding to that subset. For example, when the receiving device A1 identifies and extracts the recorded data from the distant view subset, it will also extract close-up footage clips corresponding to different target people. The tag for this footage clip can be close-up target X1 or close-up target X2.

[0146] In some embodiments, an electronic device can record more than two contents or scenes. For example, electronic device A1 can simultaneously record video of a distant view and video of a close-up view, and so on.

[0147] Furthermore, in some embodiments, such as Figure 4g In the example shown, the video project interface 401 of electronic device A1 displays the recording group 1 material folder 4012 corresponding to the current recording group. When the recording group 1 material folder 4012 is clicked, the electronic device can display as follows: Figure 4h The material display interface 405 shown can display the recording data recorded by the electronic device itself and the recording data received from other electronic devices. For example, it can display the close-up 1 recording data recorded by electronic device A1, the far-view 1 recording data recorded by electronic device B1, the audio 1 recording data recorded by electronic device C1, and the audio 2 recording data recorded by electronic device D1.

[0148] It is understood that in some embodiments, if electronic device A1 also receives recording data from other recording groups, the video project interface 401 of electronic device A1 may also display the material folders corresponding to other recording groups.

[0149] The pre-processing control 4051 can be displayed in the material display interface 405. When the pre-processing control 4051 is clicked, a selection submenu 4052 can be displayed in the material display interface 405. The selection submenu 4052 can display a clip generation control 4053, a subtitle generation control 4054, and a bullet screen generation control 4055. The clip generation control 4053 is used to generate different scene material clips corresponding to each target person based on the video recording data. The subtitle generation control 4054 is used to generate a subtitle clip corresponding to each material clip based on the audio and the audio of the target person in the video recording data. The bullet screen generation control 4055 is used to generate a corresponding subtitle clip based on the audio of a non-target person in the audio and the video recording data. The subtitle clip corresponding to the audio of the non-target person can be displayed as a bullet screen in the video.

[0150] When the clip generation control 4053, the subtitle generation control 4054, and the bullet screen generation control 4055 are selected, and the pre-processing control 4051 is clicked, the electronic device can pre-process all recording data in the recording group 1 material folder 4012 to generate different scene material clips corresponding to each target person, subtitle clips corresponding to each material clip, and bullet screen clips corresponding to non-target persons.

[0151] When the pre-processing control 4051 is clicked, the electronic device A1 can display a video processing interface 406 as shown in Figure 4i The video processing interface can include a material tab 4061, a timeline tab 4062, and a preview tab 4063. When the material tab 4061 is clicked, the electronic device A1 will display each material clip pre-processed by the electronic device A1 as shown in Figure 4i When the timeline tab 4062 is clicked, the electronic device A1 will display a material clip arrangement area 4064, a timeline 4065, and a track area as shown in Figure 4j The track area can include a video track 4066, an audio track 4067, a subtitle track 4068, and a bullet screen track 4069. The materials displayed in the material clip arrangement area 4062 can be directly dragged to the corresponding track area to complete video editing and synthesis.

[0152] When the preview tab 4063 is clicked, the synthesized video can be previewed in the preview area 4060 as shown in Figure 4k The video can include corresponding subtitles and bullet screens.

[0153] It can be understood that in some embodiments, when the synthesized video is played, when the picture switches to the picture corresponding to the scene, the corresponding audio can be played. For example, when the playing picture of the video is the close-up picture of target X1, the sound can be automatically switched to the sound of target X1, and the sounds of other targets can be considered as noise and reduced or removed, for example, the sound of target X2 and / or the environmental sound can be reduced or removed.

[0154] It can be understood that in some embodiments, each video-type material segment in the material segment set has a corresponding audio segment and a subtitle segment. When the user drags the video-type material segment to the video track 4066, the electronic device can automatically identify the audio segment and the subtitle segment corresponding to the video-type material segment and automatically place the audio segment and the subtitle segment corresponding to the video-type material segment in the corresponding positions of the corresponding track area.

[0155] For example, as shown in Figure 4k , when the user drags the close-up material segment of target X2 to the video track 4066, the electronic device places the audio segment of target X2 in the same time period of the close-up material segment of target X2 in the corresponding area of the audio track 4067, and places the subtitle segment of target X2 in the same time period in the corresponding area of the subtitle track 4068.

[0156] It can be understood that in some embodiments, in order to better synthesize the video, there can be multiple video tracks in the track area, which can include various types of tracks, such as long shot tracks, close-up tracks, and close-up tracks, etc. For example, as shown in Figure 4l , the video track in the track area includes a long shot track 4071 and a close-up track 4072. In this way, the superposition of videos of multiple types can be realized, for example, the long shot video 4081 and the close-up video 4082 in Figure 4l can be superimposed in the same time period, presenting the superimposed video effect as shown in Figure 4l .

[0157] It can be understood that in some embodiments, the preview area of the electronic device can also display multiple display areas in split screen, and each display area can select different pictures according to user needs. For example, as shown in Figure 4m , the preview area 4060 can include a first display area 4091 and a second display area 4092. When the user clicks the first display area 4091, the user can edit the preview content of the first display area 4091, for example, as shown in Figure 4m , the first display area displays the barrage information "good" by dragging the barrage information "good" to the barrage track; as shown in Figure 4nAs shown, when a user clicks on the second display area 4092, they can edit the preview content in the second display area 4092. For example, dragging the bullet comment "666" to the bullet comment track will then display the bullet comment "666" in the preview video of the second display area 4092. This allows users to easily compare various preview video effects and select the appropriate composite video.

[0158] It is understood that in some embodiments, a save selection control 4093 and a delete selection control 4094 may be provided below the first display area 4091, and a save selection control 4095 and a delete selection control 4096 may be provided below the second display area 4091. Clicking the save selection control will save the video in the corresponding display area, and clicking the delete selection control will delete the video in the current display area, allowing the preview content of the display area to be edited again.

[0159] It is understood that in some embodiments, the first display area 4091 can also be used to preview the composite video, and the second display area 4092 can be used to zoom in on the material clips selected by the user for easier viewing.

[0160] It is understood that in some embodiments, such as Figure 4o As shown, the second display area 4092 can be superimposed on the first display area 4091. Specifically, the first display area 4091 can display a preview of the composite video, while the second display area 4092 can be used to display close-up videos of the target person, such as displaying a close-up video of target X1.

[0161] In some embodiments, panoramic video can be automatically displayed in the preview area or each display area, and a selection and switching function is added to the video compositing interface, such as switching panoramic video of a certain time period to close-up video. For example, Figure 4p As shown, the preview area 4060 automatically displays the panoramic video. A switching control 4097 is provided on the video compositing interface 406. When the user clicks the switching control 4097, other videos of different scene types existing in the corresponding time period will pop up on the video compositing interface 406. For example, there may be a close-up target X1 video and a extreme close-up target X1 video. When the user clicks on a video of another scene type, that video will replace the panoramic video for that time period. For example, when the user clicks on the extreme close-up target X1 video, the panoramic video for that time period will appear as follows: Figure 4q As shown, it will switch to a close-up of the target X1 video.

[0162] It can be understood that in some embodiments, the current synthesized video can also be displayed in the preview area 4060, and when the user clicks the switching control 4097, the videos of each scene existing in the current time period can be popped up on the video synthesis interface 406, and when the user clicks the video of the corresponding scene, the videos of other scenes in the time period will replace the current synthesized video.

[0163] It can be understood that the setting mode of the video synthesis interface 606 in the embodiments of the present application is only an example, and the setting mode of the video synthesis interface 606 in the embodiments of the present application can be any combination of the above-mentioned multiple schemes. For example, as shown in FIG. 6c, in the embodiments of the present application, the video synthesis interface can include a first display area 4091, a second display area 4092, a save selection control 4093 corresponding to the first display area 4091, a delete selection control 4094, and a switching control 4097. The lower part of the second display area 4091 can be provided with a corresponding save selection control 4095, a delete selection control 4096, and a corresponding switching control 4098. When the save selection control is clicked, the video of the corresponding display area can be saved, and when the delete selection control is clicked, the video of the corresponding display area can be deleted, and the preview content of the display area can be edited again. When the user clicks the switching control, the display content of the display area corresponding to a certain time period can be replaced with other material segments. Figure 4r

[0164] Before the video processing method provided in the embodiments of the present application is described in detail, the electronic device provided in the embodiments of the present application is first briefly introduced. As shown in FIG. 7a, the electronic device can include a data acquisition module, a data transceiver module, a central control module, an audio processing module, a video processing module, a user interaction module, and an audio / video synthesis module. Figure 5

[0165] The user interaction module is configured to detect user operations and send execution instructions corresponding to the user operations to the central control module.

[0166] The central control module can be configured to execute the received instructions. For example, the user interaction module can detect some user operations of creating a recording group and send a group creation instruction to the central control module. The central control module can create a group. The user interaction module can detect some user operations of joining a recording group and send a group joining instruction to the central control module. The central control module can control the electronic device to join the group.

[0167] ​​The data collection module can be used to control the microphone and the camera to record audio or video respectively, to obtain video data collected by the camera of the electronic device and audio data collected by the microphone, and to obtain recording content tags and time mark information corresponding to the recording data. The audio data, the corresponding recording content tags and the time mark information are sent to the audio processing module for audio processing, and the video data, the corresponding recording content tags and the time mark information are sent to the video processing module for video processing.

[0168] The data transceiving module can be used to obtain video data and audio data from other electronic devices in the recording group and recording content tags and time mark information corresponding to the data when the electronic device is a receiving device, and to send the video data and the audio data and the recording content tags and the time mark information corresponding to the data to other electronic devices in the recording group through a WIFI or other wireless communication module when the electronic device is a sending device.

[0169] The data processing module is configured to preprocess all the recording data based on the time mark information and the recording content tags of the recording data, and to generate a material set including material segments and material tags corresponding to the material segments.

[0170] The data processing module can include a classification module, an audio processing module and a video processing module.

[0171] The classification module is configured to preliminarily classify the recording data based on the recording content tags. For example, the recording data can be divided into an audio set and a video set according to the category tags.

[0172] The audio processing module is configured to perform audio processing on the data in the audio set to obtain corresponding voice material segments of different target persons, and to generate tags corresponding to the voice material segments.

[0173] The video processing module is configured to divide the data in the video set into a long shot sub-set, a close-up sub-set and a close-up sub-set according to the scene tags of the recording data in the video data set. Then, the recording data in each sub-set is identified and extracted to obtain material segments corresponding to different target objects in each sub-set and material tags.

[0174] The sending module obtains the material segments processed by the audio processing module and the video processing module, and sends the material set including all the material segments to a video synthesis application program.

[0175] The audio-video synthesis module is configured to display a video synthesis interface in response to a user operation, and to synthesize a corresponding video based on the material segments selected by the user in the video synthesis interface.

[0176] The video processing method provided by the embodiments of the present application will be described in detail below in combination with the electronic device described above. Among them, Figure 6 A flowchart of a video processing method according to an embodiment of the present application is shown. The video processing method provided by the embodiments of the present application can be executed by the electronic device described above, and the video processing method comprises:

[0177] 601: In response to a user operation, a recording group is created.

[0178] It can be understood that the electronic device can respond to some user operations for creating a recording group to achieve creating a group, or respond to some user operations for joining a recording group to achieve joining an existing group.

[0179] It can be understood that the electronic devices that can join the recording group can be divided into two categories. The first category of devices are mobile phones, tablet computers and other devices with wireless network functions, which can join or create a recording group by connecting to the same wireless network. The second category of devices are microphones, Bluetooth headsets and other devices that do not have wireless network functions but have Bluetooth functions. Among them, the second category of devices can join the recording group through the first category of devices as relay nodes, that is, the second category of devices can join the recording group by connecting to the first category of devices through Bluetooth.

[0180] In some embodiments, the electronic device can obtain the names of other electronic devices connected to the same wireless network as the device. And based on the user's selection of each electronic device in the same wireless network, a temporary recording group is created.

[0181] For example, as shown in Figure 7 , it is assumed that electronic device A1, electronic device B1 and electronic device C1 all have wireless network functions, and are connected to the wireless network with the account name "PH888" at the same time as shown in Figure 3 , electronic device D1 does not have wireless network function but has Bluetooth function and is connected to electronic device D1 through Bluetooth. Among them, electronic device A1, electronic device B1 and electronic device C1 can be mobile phones or microphones.

[0182] At this time, electronic device A1, electronic device B1, electronic device C1 and electronic device D1 first perform networking. For example, as shown in Figure 4a , the user can open the video synthesis application through electronic device A1. When the new video item control 4011 in the video item interface 401 of the video synthesis application is clicked, the electronic device A1 will jump to, for example, as shown in Figure 4bThe group selection interface 402 shown can display a new group control 4021 and a search for existing group controls 4022. Clicking the new group control will redirect to the new group interface 403. Clicking the search for existing group controls 4022 will display existing groups.

[0183] The newly created group interface 403 can display the WIFI control 4031 and the name of the current wireless network that electronic device A1 is connected to; when the WIFI control 4031 is clicked, electronic device A1 will be redirected to, for example, Figure 4d The WIFI settings interface 002 is described above. The new group interface 403 can also display the icons and names of electronic devices connected to the wireless network with the account name "PH888". For example, it can include icons 4032, 4033, and 4034 corresponding to electronic devices A1, B1, and C1, and their corresponding names PHA, PHB, and PHC, respectively. The new group interface 403 can also display the icons and names of other electronic devices connected to each electronic device via Bluetooth, such as icon 4035 and name PHD corresponding to electronic device D1 connected to electronic device B1. The new group interface 403 can also display a group name control 4036 and a new group control 4037. Clicking the group name control 4036 allows changing and setting the group name, for example, setting the group name to "Recording Group 1". Clicking the new group control 4037 allows creating a new recording group, i.e., setting up a network.

[0184] It is understood that in the embodiments of this application, Figure 4 illustrates the method of creating a recording group by taking the newly created recording group by electronic device A1 as an example. In some embodiments, electronic device A1 can also actively join an already created group.

[0185] For example, such as Figure 4b As shown, the group selection interface 402 displays a search function for existing group controls 4022. When the user clicks "Create New Group Control" or "Search Existing Group Controls 4022," the electronic device A1 can jump to... Figure 8 The existing group interface 801 shown can display the created groups under the wireless network currently connected to the electronic device A1, such as Recording Group 1. Users can join the created group "Recording Group 1" by clicking the "Join" control 8011.

[0186] It is understood that the above-described methods for creating groups and the interface of the electronic device in this application embodiment are merely illustrative examples, and this application may also adopt any other feasible methods for creating groups and setting up interfaces.

[0187] 602: record audio or video to obtain recording data, and obtain recording content tags corresponding to the recording data and time mark information.

[0188] It can be understood that the recording content tags described above can include category tags reflecting the recording category and scene tags reflecting the main shooting scene in the video data, etc. Among them, the category tags can include video tags or audio tags; the scene tags can include medium shot tags, close-up tags, long shot tags, and close-up tags, etc.

[0189] It can be understood that there can be various ways to obtain recording content tags in the embodiments of the present application, which will be introduced by way of example as follows:

[0190] In an implementable manner, the user can independently select the corresponding recording content tags on the group interface corresponding to the recording group. For example, when the user clicks the new group control 4037, the electronic device A1 can display the group interface 404 as shown in Figure 4e . Among them, the group interface can display the icons corresponding to the electronic device A1, the electronic device B1, the electronic device C1 and the electronic device D1, which are icon 4032, icon 4022, icon 4034 and icon 4035 respectively. The group interface 404 can also display the tag rows, such as close-up tag row 4044, medium shot tag row 4045, long shot tag row 4046, audio tag row 4047, etc., in addition to the receiving device row 4048 and the start recording control 4049. As shown in Figure 4f , the user can drag the icon corresponding to the electronic device to the corresponding tag row to select the recording content tag corresponding to the recording data, and drag the electronic device to the receiving device row 4048 to become a receiving device. For example, the user can drag the icon 4041 corresponding to the electronic device A1 to the close-up tag row 4044 and the receiving device row 4048 respectively, drag the icon 4042 corresponding to the electronic device B1 to the long shot tag row 4046, drag the icon 4042 corresponding to the electronic device C1 to the audio tag row 4047, and drag the icon 4042 corresponding to the electronic device D1 to the audio tag row 4047.

[0191] When the user clicks the start recording control 4049, the electronic device A1 can call the camera or microphone to start recording according to the selected recording content tags. For example, when the recording content tag of the electronic device A1 is the audio tag, the electronic device A1 can call the microphone to start recording, and when the recording content tag selected by the electronic device A1 is the medium shot tag, the long shot tag or other video category tags, the electronic device A1 can call the camera to start recording. And record the start time and end time of recording. And the recorded content can be stored in correspondence with the recording content tag corresponding to the tag row where the electronic device A1 is located.

[0192] In some embodiments, the electronic device can perform feature extraction or recognition on the recording data to generate recording content labels corresponding to the recording data after obtaining the recording data.

[0193] For example, the electronic device can use a neural network model or related algorithm to recognize the recording data to determine whether the recording data is audio data or video data, to generate corresponding audio labels or video labels. And recognize the video data to determine the scene labels corresponding to the video data.

[0194] It can be understood that the above-mentioned time mark information can include recording start time and recording end time. The time mark information is used to align the time of each video data or audio data to facilitate post-editing and synthesis. For example, if any electronic device captures a non-full time period video due to interruption, the recording start time and the recording end time of the non-full time period video can still be obtained during post-video processing. In this way, the non-full time period video can be aligned with the video of the corresponding time period in other complete videos.

[0195] In some embodiments, the electronic device can record the start time and end time of the recording when recording, wherein the recording start time and end time recorded by the electronic device can be aligned with the network time.

[0196] 603: In the case of the electronic device as a receiving device, the recording data sent by other devices in the recording group and the corresponding time mark information and recording content labels are obtained.

[0197] It can be understood that in the embodiments of the present application, there can be many ways for the electronic device to become a receiving device. In one implementable way, as shown in the foregoing Figure 4f The group interface 404 of the electronic device can display a receiving device row 4048, and the user can drag the icon corresponding to the electronic device to the receiving device row 4048 to make the electronic device become a receiving device. It can be understood that there can be any number of electronic devices in the recording group to become receiving devices.

[0198] In some embodiments, the electronic device can also only choose to become a receiving device to receive the recording data sent by other devices in the recording group, without selecting other label rows and without recording video and audio by itself.

[0199] It can be understood that in another embodiment, the electronic device can send a selection window of whether to become a receiving device to the recording interface after detecting the end of recording, and determine whether the electronic device becomes a receiving device based on the user's selection. For example, as Figure 9As shown, when the user selects the "Yes" control 9012 in the selection window 9011 of the recording interface 901 of the electronic device A1, the electronic device A1 will become a receiving device, which can be used to receive the recording data of other electronic devices in the recording group and the corresponding time mark information and recording content label.

[0200] When the user selects the "No" control 9013 in the selection window 9011 of the recording interface 901 of the electronic device A1, the electronic device A1 will become a sending device, which will actively send the recording data and the corresponding time mark information and recording content label to the receiving device after the recording is completed.

[0201] Figure 10 A schematic diagram of data transmission among the electronic device A1, the electronic device B1 and the electronic device C1 in the recording group is shown in FIG. 6. It can be understood that each video data can include video stream data and audio stream data. As shown in FIG. 6, Figure 10 As shown, the electronic device A1 is a receiving device, the electronic device B1, the electronic device C1 and the electronic device D1 are sending devices, the electronic device B1 can receive the video data recorded by the electronic device D1 through Bluetooth, and can send the received video data recorded by the electronic device D1 and the video data recorded by itself to the electronic device A1 through a wireless network, the electronic device C1 can send the recorded audio data to the electronic device A1 through a wireless network, the electronic device D1 can send the recorded audio data to the electronic device A1 through Bluetooth, and the electronic device A1 can perform video processing on the received data to obtain a synthesized video, as in embodiments 604-607 of the present application.

[0202] 604: Preprocess all recording data based on the time mark information and recording content label of each recording data to generate a material set, which includes material segments, material labels and time mark information corresponding to the material segments.

[0203] It can be understood that in the embodiments of the present application, all video data can be preprocessed based on the time mark information and recording content label of each recording data to generate material segments, material labels and time mark information corresponding to the material segments. The material segments can include segments corresponding to different shots and different target persons, and the material labels are labels marking the content of the material segments. For example, the material segments can include a medium shot corresponding to a target X1 and a long shot corresponding to a target X2, and the corresponding material labels can be medium shot target X1 and long shot target X2.

[0204] The time mark information corresponding to the material segments can include the start time and end time of the material segments.

[0205] The electronic device pre-processes all the video data based on the time mark information and the video recording label of each video data, generates the material segments and the material labels corresponding to the material segments; and can include:

[0206] The electronic device can preliminarily classify the recording data based on the recording content label, for example, can classify the recording data into an audio set and a video set according to the category label;

[0207] The audio processing is performed on each data in the audio set to obtain the corresponding sound material segments of different target persons, and the labels corresponding to the sound material segments are generated. For example, the material segments in the audio set can include the sound material segments corresponding to target X1 and the sound material segments corresponding to target X2, and the corresponding material labels can be target X1 audio and target X2 audio, respectively.

[0208] In the embodiments of the present application, the electronic device can also convert the audio content of the target person in each sound material segment into subtitles through a voice recognition algorithm. It can be understood that in the embodiments of the present application, the audio of non-target persons in the audio data can also be converted into bullet screens in the synthesized video. For example, in a speech recording scene, the video recorded includes the video and audio of target X1 and target X2, but also contains some audio of the audience on the scene. Therefore, the audio data corresponding to the audience on the scene can be converted into corresponding bullet screens displayed in the synthesized video.

[0209] In addition, the electronic device can divide each data in the video set into a long shot sub-set, a close-up sub-set, a close-up sub-set, and the like according to the scene label of each recording data in the video data set. Then, each recording data in each sub-set is identified and extracted to obtain the material segments corresponding to different target objects in each sub-set and the material labels. For example, the long shot sub-set in the video set can include the material segments corresponding to target X1 and the material segments corresponding to target X2, and the material labels can be long shot target X1 video and long shot target X2 video.

[0210] It can be understood that in some embodiments, the user may exist in some cases without shooting based on the selected scene, for example, the user may select the long shot as the scene label, but in the shooting process, the user shoots the close-up for part of the time. At this time, the electronic device will also extract the material segments of different scenes corresponding to the sub-set when identifying and extracting each recording data in each sub-set. For example, the electronic device will also extract the close-up material segments corresponding to different target persons when identifying and extracting each recording data in the long shot sub-set. The label of the material segment can be close-up target X1 or close-up target X2.

[0211] It can be understood that in the embodiments of the present application, the manner of audio processing each data can include pre-processing the data using a multi-channel joint processing algorithm, wherein the multi-channel joint processing algorithm can include a multi-channel noise reduction algorithm, a speech separation algorithm, and a speech recognition algorithm, etc. In some embodiments, the manner of audio processing each data can be to first perform noise reduction processing on the data by a multi-channel noise reduction algorithm, then perform separation processing on the data obtained after noise reduction processing by a speech separation algorithm to obtain the speech corresponding to each target person or non-target person, and then perform speech recognition processing on the speech corresponding to each target person or non-target person by a speech recognition algorithm to obtain the subtitles or barrage corresponding to each speech.

[0212] In the embodiments of the present application, the manner of processing each recording data in the video frequency set can include processing each recording data using a pedestrian re-identification algorithm to obtain the material segments corresponding to different targets of each sub-set, and generating the material labels and time mark information corresponding to the material segments.

[0213] The following describes the manner of the electronic device A1 in the scene of recording the speech site, pre-processing the recording content of the electronic device A1 itself and the recording content of other electronic devices in the recording group to obtain the material segments and the corresponding material labels.

[0214] In some embodiments, as shown in FIG. 4A, in the video project interface 401, the recording group material folder corresponding to each recording group can be displayed, for example, the recording group 1 material folder 4012 corresponding to the current recording group is displayed, when the recording group 1 material folder 4012 is clicked, the electronic device can display the material display interface 405 as shown in FIG. 4B, in the material display interface 405, the recording data recorded by the electronic device itself and the recording data received from other electronic devices can be displayed, for example, the close-up 1 recording data recorded by the electronic device A1, the long shot 1 recording data recorded by the electronic device B1, the audio 1 recording data recorded by the electronic device C1 and the close-up 2 recording data recorded by the electronic device D1 can be displayed. Figure 4g Figure 4h

[0215] ​​The preprocessing control 4051 can also be displayed in the material display interface 405. When the preprocessing control 4051 is clicked, a selection submenu 4052 can be displayed in the material display interface 405. The selection submenu 4052 can display a clip generation control 4053, a subtitle generation control 4054, and a bullet screen generation control 4055. The clip generation control 4053 is used to generate different scene material clips corresponding to each target person based on the video recording data. The subtitle generation control 4054 is used to generate a subtitle clip corresponding to each material clip based on the audio and the audio of the target person in the video recording data. The bullet screen generation control 4055 is used to generate a corresponding subtitle clip based on the audio of a non-target person in the audio and video recording data. The subtitle clip corresponding to the audio of the non-target person can be displayed as a bullet screen in the video.

[0216] When the clip generation control 4053, the subtitle generation control 4054, and the bullet screen generation control 4055 are selected, and the preprocessing control 4051 is clicked, the electronic device can preprocess all recording data in the recording group 1 material folder 4012 to generate different scene material clips corresponding to each target person, subtitle clips corresponding to each material clip, and bullet screen clips corresponding to non-target persons.

[0217] In an implementable manner, the preprocessing manner can be that the electronic device A1 can first store the long-range video recorded by the electronic device B1 and the close-range video recorded by the electronic device A1 and the electronic device D1 into a video set, and store the audio sent by the electronic device C1 into an audio set.

[0218] Then, according to the scene labels of each data in the video data set, each data in the video set is divided into a long-range sub-set and a close-range sub-set. Each recording data in each sub-set is identified and extracted to obtain material clips corresponding to different targets in each sub-set. For example, the long-range sub-set can include a long-range material clip corresponding to target X2. The long-range material clip corresponding to target X2 can have a material label of “long-range target X2”. The close-range sub-set can include close-range material clips corresponding to target X1 and target X2. The close-range material clips corresponding to target X1 and target X2 can have material labels of close-range target X1 and close-range target X2. In addition, the electronic device A1 can also perform audio processing on the audio stream in each video material clip to obtain sound material clips corresponding to target A and target X2 in each video material clip.

[0219] The electronic device A1 can also perform audio processing on each data in the audio set to obtain sound material clips corresponding to target A and target X2, and corresponding material labels “target X1 audio” and “target X2 audio”.

[0220] When the preprocessing control 4051 is clicked, the electronic device A1 can display a video processing interface 406 as shown in Figure 4i The video processing interface can include a material tab 4061, a timeline tab 4062, and a preview tab 4063. When the material tab 4061 is clicked, the electronic device A1 displays the preprocessed material segments of the electronic device A1 as shown in Figure 4i

[0221] 605: Synthesize a corresponding video based on the user-selected material segments.

[0222] It can be understood that in the embodiments of the present application, the electronic device can synthesize a corresponding video based on the user-selected material segments in the video synthesis interface, and preview and play in the preview area.

[0223] It can be understood that the video previewed and played in the preview area can include subtitles corresponding to the target person.

[0224] It can be understood that in the embodiments of the present application, the video previewed and played in the preview area can include a barrage corresponding to a non-target person. For example, in a live recording scene, the recorded video includes the video and audio of the target X1 and the target X2, but also contains some audio of the on-site audience. Therefore, the audio data corresponding to the on-site audience can be converted into corresponding barrage and displayed in the synthesized video.

[0225] As shown in Figure 4j When the user clicks the timeline tab 4062 in the video synthesis interface 406, the electronic device A1 displays a material segment arrangement area 4064, a timeline 4065, and a track area, which can include a video track 4066, an audio track 4067, a subtitle track 4068, a barrage track 4069, etc. In some embodiments, in order to better synthesize the video, the video track in the track area can be multiple, for example, including a long shot track for placing long shot videos and other tracks for placing videos other than long shots. In this way, the superposition of videos of multiple types can be realized, for example, any long shot video and close-up video can be superimposed in the same time period to present the superimposed video effect.

[0226] When the preview tab 4063 is clicked, the synthesized video can be previewed in the preview area 4060 as shown in Figure 4k The video can include corresponding subtitles and barrages.

[0227] It can be understood that in the embodiments of the present application, in order to make it more convenient for the user to drag the material tag to the corresponding track area, when the user clicks any target material segment, the electronic device can display the corresponding position of the target material segment in the track area according to the time mark information of the target material segment. For example, as shown in Figure 11 ​As shown, the time mark information of the material segment corresponding to the audio target X1 is 00:05-00:10, and when the audio target X1 is clicked, the region of the audio track region corresponding to the time axis 00:05-00:10 will be highlighted or prompted in a highlighted color to remind the user of the corresponding position of the material segment of the audio target X1.

[0228] In some embodiments, the position of the material segment can be arranged according to the time mark information and the time axis below. The user can directly drag the material displayed in the material segment arrangement region to the corresponding track region to complete video editing and synthesis.

[0229] Based on the above scheme, the recording data generated by each electronic device in the recording group can be aggregated on any electronic device in the recording group for video synthesis processing, so that the recording data does not need to be transferred to an additional dedicated synthesis device, saving the operation process. In addition, the electronic device can preprocess all video data according to the time mark information and recording content label of each video data, generate material segments and labels corresponding to the material segments, so that the post-processing step can be greatly simplified, and non-professional users can also select material segments to synthesize videos.

[0230] The embodiment of the present application provides an electronic device, comprising: a memory for storing instructions executed by one or more processors of the electronic device, and the processor is one of the one or more processors of the electronic device, and is used for executing the above-mentioned video processing method.

[0231] The embodiment of the present application provides a readable storage medium, and the readable medium stores instructions, which, when executed on an electronic device, cause the electronic device to execute the above-mentioned video processing method.

[0232] The embodiment of the present application provides a computer program product, comprising instructions, which, when executed on an electronic device, cause the electronic device to execute the above-mentioned video processing method.

[0233] The hardware structure of the electronic device provided by the embodiment of the present application will be described below taking the mobile phone 10 as an example.

[0234] As shown in the figure, Figure 12 The mobile phone 10 can include a processor 110, a power module 140, a memory 180, a mobile communication module 130, a wireless communication module 120, a sensor module 190, an audio module 150, a camera 170, an interface module 160, a key 101, and a display screen 102.

[0235] It is understood that the structure illustrated by the embodiments does not constitute a specific limitation on the mobile phone 10. In other embodiments of the present application, the mobile phone 10 can include more or fewer components than those shown, or combine some components, or split some components, or different arrangement of components. The components shown can be implemented in hardware, software or a combination of software and hardware.

[0236] The processor 110 can include one or more processing units, for example, can include a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro-programmed control unit (MCU), an artificial intelligence (AI) processor, or a field programmable gate array (FPGA) processing module or processing circuit, etc. Among them, different processing units can be independent devices, or can be integrated in one or more processors. The storage unit can be provided in the processor 110 for storing instructions and data. In some embodiments, the storage unit in the processor 110 is a cache memory 180.

[0237] It is understood that the video processing method in the embodiments of the present application can be executed by the processor 110.

[0238] The power module 140 can include a power supply, a power management component, etc. The power supply can be a battery. The power management component is used to manage the charging of the power supply and the power supply to other modules. In some embodiments, the power management component includes a charging management module and a power management module. The charging management module is used to receive charging input from the charger; the power management module is used to connect the power supply, the charging management module and the processor 110. The power management module receives the input of the power supply and / or the charging management module, and supplies power to the processor 110, the display screen 102, the camera 170, and the wireless communication module 120, etc.

[0239] The mobile communication module 130 can include, but is not limited to, an antenna, a power amplifier, a filter, an LNA (Low noise amplify), and the like. The mobile communication module 130 can provide a solution including 2G / 3G / 4G / 5G wireless communication applied to the mobile phone 10. The mobile communication module 130 can receive electromagnetic waves by the antenna, and perform filtering, amplification, and the like on the received electromagnetic waves, and transmit to the modem processor for demodulation. The mobile communication module 130 can also amplify the signal modulated by the modem processor, and radiate as electromagnetic waves through the antenna. In some embodiments, at least part of the function modules of the mobile communication module 130 can be disposed in the processor 110. In some embodiments, at least part of the function modules of the mobile communication module 130 can be disposed in the same device as at least part of the modules of the processor 110.

[0240] The wireless communication module 120 can include an antenna, and realize the transceiving of electromagnetic waves via the antenna. The wireless communication module 120 can provide a solution including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) network), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared (IR) technology, and the like, which are wireless communication applied to the mobile phone 10. The mobile phone 10 can communicate with the network and other devices through the wireless communication technology.

[0241] It can be understood that, in the embodiments of the present application, when the electronic device is a receiving device, the video data and the audio data from other electronic devices in the recording group received by the wireless communication module, and the recording content label and the time mark information corresponding to each data can be received. And when the electronic device is a sending device, the video data and the audio data, and the recording content label and the time mark information corresponding to each data can be sent to other electronic devices in the recording group through the wireless communication module.

[0242] In some embodiments, the mobile communication module 130 and the wireless communication module 120 of the mobile phone 10 can also be located in the same module.

[0243] The display screen 102 is configured to display a human-computer interaction interface, an image, a video, etc. The display screen 102 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light emitting diode (AMOLED), a flexible light-emitting diode (FLED), a Miniled, a MicroLed, a Micro-oLed, a quantum dot light emitting diode (QLED), etc.

[0244] The sensor module 190 can include a proximity light sensor, a pressure sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a distance sensor, a fingerprint sensor, a temperature sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc.

[0245] It can be understood that the ambient light sensor in the embodiments of the present application can be used to acquire illumination state information and send the illumination state information to the processor.

[0246] The audio module 150 is configured to convert digital audio information into an analog audio signal output, or convert an analog audio input into a digital audio signal. The audio module 150 can also be configured to encode and decode an audio signal. In some embodiments, the audio module 150 can be disposed in the processor 110, or some functional modules of the audio module 150 can be disposed in the processor 110. In some embodiments, the audio module 150 can include a speaker, a receiver, a microphone, and a headset jack.

[0247] The camera 170 is configured to capture a still image or a video. An object generates an optical image through a lens and projects the optical image onto a photosensitive element. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to an ISP (Image Signal Processing) to convert the electrical signal into a digital image signal. The mobile phone 10 can realize a shooting function through the ISP, the camera 170, a video codec, a GPU (Graphic Processing Unit), the display screen 102, and an application processor, etc.

[0248] The interface module 160 includes an external memory interface, a universal serial bus (USB) interface, and a subscriber identification module (SIM) card interface, etc. The external memory interface can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the mobile phone 10. The external memory card communicates with the processor 110 through the external memory interface to realize data storage functions. The universal serial bus interface is used for communication between the mobile phone 10 and other electronic devices. The subscriber identification module card interface is used to communicate with the SIM card installed in the mobile phone 10, such as reading the phone number stored in the SIM card, or writing the phone number into the SIM card.

[0249] In some embodiments, the mobile phone 10 also includes a button 101, a motor, and an indicator, etc. The button 101 can include a volume button, a power on / off button, etc. The motor is used to generate a vibration effect of the mobile phone 10, such as generating a vibration when the mobile phone 10 of the user is called to prompt the user to answer the call. The indicator can include a laser indicator, a radio frequency indicator, an LED indicator, etc.

[0250] Embodiments of the mechanisms disclosed herein can be implemented in hardware, software, firmware, or any combination thereof. Embodiments of the application can be implemented as computer programs or program code executing on programmable systems comprising at least one processor, a storage system (including volatile and non-volatile memory and / or storage elements), at least one input device, and at least one output device.

[0251] Program code can be applied to input instructions to perform the functions described herein and generate output information. The output information can be applied to one or more output devices, in known fashion. For purposes of this application, a processing system includes any system that has a processor, such as a digital signal processor (DSP), a microcontroller, an application specific integrated circuit (ASIC), or a microprocessor.

[0252] The program code can be implemented in a high level procedural or object oriented programming language to communicate with a processing system. The program code can be implemented in assembly or machine language, if desired. In fact, the mechanisms described herein are not limited in scope to any particular programming language. In any case, the language can be a compiled or interpreted language.

[0253] In some cases, the disclosed embodiments can be implemented in hardware, firmware, software, or any combination thereof. The disclosed embodiments can also be implemented as instructions carried by or stored on a transitory or non-transitory machine-readable (e.g., computer-readable) medium, which can be read and executed by one or more processors. For example, the instructions can be downloaded from a network or by way of another computer readable medium. Thus, a machine-readable medium can include any mechanism for storing or transmitting information in a form readable by a machine (e.g., a computer), including without limitation floppy disks, optical disks, optical disks, compact discs, read-only memory (CD-ROMs), magnetic disks, read-only memory (ROM), random access memory (RAM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), magnetic or optical cards, flash memory, or a tangible, machine-readable storage used in the transmission of information over the Internet with a propagated signal in electronic, electromagnetic, or optical form, such as carrier waves, infrared signals digital signals, etc. Accordingly, a machine-readable medium includes any type of medium suitable for storing or transmitting electronic instructions or information in a form readable by a machine (e.g., a computer).

[0254] In the drawings, some of the structural or methodological features can be shown in particular arrangements and / or orders. However, it should be understood that such particular arrangements and / or orders can not be required. Instead, these features can be arranged in a different manner and / or order than shown in the illustrative figures, in some embodiments. Additionally, the inclusion of a structural or methodological feature in a particular figure is not meant to imply that such feature is required in all embodiments, and in some embodiments, these features can not be included or can be combined with other features.

[0255] It should be noted that each unit / module mentioned in the device embodiments of the present application is a logical unit / module, in physical, one logical unit / module can be one physical unit / module, also can be a part of one physical unit / module, also can be realized in combination of multiple physical unit / modules, the physical realization of these logical units / modules is not the most important, the combination of the functions realized by these logical units / modules is the key to solve the technical problems proposed in the present application. In addition, in order to highlight the innovative part of the present application, the above-mentioned device embodiments of the present application do not introduce the units / modules which are not closely related to solving the technical problems proposed in the present application, which does not mean that the above-mentioned device embodiments do not have other units / modules.

[0256] In the embodiments of the present application, the method provided by the embodiments of the present application is introduced from the perspective of an electronic device (e.g., a mobile phone) as an execution subject. To implement each function in the method provided by the embodiments of the present application, the electronic device can include a hardware structure and / or a software module, and each function is implemented in the form of a hardware structure, a software module, or a hardware structure plus a software module. Whether a certain function in the above functions is implemented in the form of a hardware structure, a software module, or a hardware structure plus a software module depends on the specific application of the technical solution and design constraints.

[0257] In the above embodiments, the term "when" or "after" can be interpreted as "if" or "after" or "in response to determining" or "in response to detecting" according to the context. Similarly, the phrase "upon determining" or "if detecting (the stated condition or event)" can be interpreted as "if determining" or "in response to determining" or "upon detecting (the stated condition or event)" or "in response to detecting (the stated condition or event)" according to the context.

[0258] It should be noted that the relational terms herein such as first and second and the like can be used solely to distinguish one entity or action from another entity or action without necessarily requiring or implying any such actual relationship or order between such entities or actions. Moreover, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. Without more limitations, an element defined by the statement "comprising a" does not exclude the existence of additional identical elements in the process, method, article, or apparatus that includes the element.

[0259] In this specification, the reference to "one embodiment" or "some embodiments" etc. means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the application. Thus, the appearances of the phrases "in one embodiment", "in some embodiments", "in other embodiments", "in additional embodiments", etc. in various places in the specification are not necessarily all referring to the same embodiment, unless otherwise be clear from the context.

[0260] In the embodiments of the present application, "and / or" is only a kind of description of the association relationship of associated objects, which means that there can be three kinds of relationships, for example, A and / or B can represent the following three cases: A exists alone, A and B exist simultaneously, and B exists alone.In addition, the character " / " in this paper generally represents that the front and rear associated objects are a kind of "or" relationship.

[0261] Although the present application has been illustrated and described with reference to certain preferred embodiments thereof, it should be understood by those skilled in the art that various changes in form and details can be made therein without departing from the scope of the present application.

Claims

1. A method of video processing, the method comprising: The method comprises: A first processing device acquires a plurality of sets of recording data from a plurality of recording devices, wherein the plurality of recording devices belong to a same recording group; the first processing device belongs to a device in the recording group; A first interface of the first processing device displays the plurality of sets of recording data; In response to a first operation of a user, the first processing device pre-processes the plurality of sets of recording data to generate a set of material clips, a second interface of the first processing device displays each material clip in the set of material clips; and In response to a second operation of the user, a third interface of the first processing device displays a synthesized video; The first processing device pre-processes the plurality of sets of recording data to generate a set of material clips, comprising: The first processing device identifies and extracts video data in the plurality of sets of recording data to acquire video material clips of different scenes corresponding to different targets in the video data.

2. The video processing method of claim 1, wherein, The first operation is an operation of triggering a pre-processing control of the first interface.

3. The video processing method of claim 1, wherein, The second interface comprises a track area, and the second operation is an operation of the user selecting at least one material clip in the material clips to a corresponding position of the track area and triggering a preview control.

4. The video processing method of claim 3, wherein, The track area comprises a video track, an audio track, a subtitle track, and a bullet screen track. The video track comprises a plurality of video tracks of different scenes.

5. The video processing method of any of claims 1-4, wherein, The preview area of the third interface displays the synthesized video.

6. The video processing method of claim 5, wherein, The preview area comprises one or more display areas, and each display area displays a corresponding synthesized video according to a material clip of a corresponding track area.

7. The video processing method of claim 5, wherein, The preview area comprises a plurality of display areas, and a part of the plurality of display areas displays a synthesized video, and another part of the plurality of display areas displays a user-selected material clip.

8. The video processing method of any of claims 1, 6 or 7, wherein, The video of the same time point in the synthesized video comprises video of at least one scene.

9. The video processing method of any of claims 1, 6 or 7, wherein, In response to a third operation of the user, the synthesized video switches from a first synthesized video to a second synthesized video.

10. The video processing method of claim 9, wherein, The third operation is an operation of clicking a switching control on the third interface.

11. The video processing method of claim 3, wherein, In response to a fourth operation of the user on a target material clip, the first processing device prompts a region position corresponding to the target material clip in the track area.

12. The video processing method of claim 11, wherein, The fourth operation is a click operation.

13. The video processing method of claim 11, wherein, The prompt of the region position corresponding to the target material clip in the track area comprises: The first processing device highlights the region position corresponding to the target material clip in the track area.

14. The video processing method of claim 1, wherein, The first processing device receives time mark information and recording content labels corresponding to recording data sent by a plurality of recording devices.

15. The video processing method of claim 1, wherein, The first processing device identifies the plurality of sets of recording data to acquire recording content labels of the plurality of sets of recording data.

16. The video processing method of claim 1, wherein, The first processing device identifies the plurality of sets of recording data to acquire time mark information corresponding to the plurality of sets of recording data.

17. The video processing method of any of claims 14-16, wherein, The time mark information comprises a start time, an end time, and / or a special mark time corresponding to the recording data.

18. The video processing method of any of claims 14 or 15, wherein, The recording content labels comprise scene labels and recording category labels.

19. The video processing method of claim 18, wherein, The scene labels comprise a long shot label, a medium shot label, a close-up label, and a close-up label. The recording category label includes a video label and an audio label.

20. The video processing method of any of claims 1-4 or 6-7, wherein, The material segment set includes video material segments of different scenes corresponding to each target respectively, audio material segments corresponding to each target respectively, material labels corresponding to the video material segments and the audio material segments respectively, time mark information, and subtitle information.

21. The video processing method of claim 20, wherein, The material segment set further includes barrage information corresponding to audio of non-targets included in the multiple sets of recording data, and time mark information corresponding to the barrage information.

22. The video processing method of claim 20, wherein, The preprocessing of the multiple sets of recording data to generate a material segment set further includes: processing video data in the multiple sets of recording data to obtain material labels and time mark information corresponding to the video material segments; processing audio data in the multiple sets of recording data and audio stream data in video data of the multiple sets of recording data to obtain audio segments corresponding to each target respectively, and material labels, time mark information, and subtitle information corresponding to the audio segments.

23. The video processing method of claim 1, wherein, The multiple recording devices include the first processing device.

24. The video processing method of claim 1, wherein, The devices in the recording group include first-type devices, or first-type devices and second-type devices; The first-type devices are devices with wireless network functions, the first-type devices are connected to the same wireless network, and the first processing device is a first-type device. The second-type devices are devices without wireless network functions but with Bluetooth functions, and the second-type devices are connected to the first-type devices through Bluetooth functions.

25. The video processing method of claim 1, wherein, Further comprising: The first processing device sends the recording data obtained by the first processing device to a second processing device in the recording group.

26. An electronic device, comprising: Further comprising: a memory configured to store instructions for execution by one or more processors of the electronic device, and the processor is one of the one or more processors of the electronic device, configured to execute the video processing method of any of claims 1-25.

27. A readable storage medium characterized by, The readable storage medium has instructions stored thereon, which, when executed on an electronic device, cause the electronic device to perform the video processing method of any of claims 1-25.

28. A computer program product, characterised in that, The instructions, when executed on an electronic device, cause the electronic device to perform the video processing method of any of claims 1-25.

Citation Information

Patent Citations

  • Media editing with multi-camera media clips

    US20130121668A1

  • Multimedia editing method and intelligent terminal

    WO2019227283A1