A video generation method, apparatus, electronic device, and storage medium

By extracting and classifying video clips from the to-process videos and generating target videos, it solves the problem that users find it difficult to quickly understand video content, and achieves efficient video content display and user experience improvement.

CN114268848BActive Publication Date: 2025-06-10BEIJING DAJIA INTERNET INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111553032.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-17
Publication Date
2025-06-10
Estimated Expiration
2041-12-17

AI Technical Summary

Technical Problem

It is difficult for prior art to quickly understand video content by watching a complete video, especially under the limitations of video real-time and duration.

Method used

By determining multiple video clips from the to-process video, each video clip represents an image sequence that describes the preset category object, determines the clip type of each video clip, and selects the target video clip of the same clip type from it to generate the target video.

Benefits of technology

The type of target video clips in the generated target video is the same, displaying preset category objects in the pending video, improving the content correlation between the target video and the pending video, allowing users to quickly understand the original video content and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114268848B_ABST
    Figure CN114268848B_ABST
Patent Text Reader

Abstract

The present disclosure relates to a video generation method, apparatus, electronic device, and storage medium. The method includes: obtaining a video to be processed; determining a plurality of video segments from the video to be processed; each video segment in the plurality of video segments representing an image sequence depicting a preset category of objects; determining the segment type of each video segment; determining a plurality of target video segments from the plurality of video segments; the segment type of each target video segment in the plurality of target video segments being the same; and generating a target video based on the plurality of target video segments. The types of the target video segments in the target video are the same, and the preset category of objects in the video to be processed are displayed, so that the target video has a relatively high content correlation with the video to be processed, and the overall display effect of the target video is relatively unified, which can help users quickly understand the content of the original video and improve the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of Internet technologies, and in particular, to a video generation method, apparatus, electronic device, and storage medium. Background Art

[0002] With the development of network technologies, video applications have become very popular in people's daily lives. In some scenarios, due to the real-time nature and duration of videos, it is difficult for users to watch a complete video to understand its content. In related technologies, by selecting a continuous part from a video to generate a target video, the generated target video in this way cannot well help users quickly understand the content of the original video and cannot meet the user's needs. Summary of the Invention

[0003] The present disclosure provides a video generation method, apparatus, electronic device, and storage medium. The technical solution of the present disclosure is as follows:

[0004] According to a first aspect of an embodiment of the present disclosure, there is provided a video generation method, including:

[0005] Obtain a video to be processed;

[0006] Determine a plurality of video segments from the video to be processed; each video segment in the plurality of video segments represents an image sequence describing a preset category of objects;

[0007] Determine the segment type of each video segment; the segment type is determined according to the object information in each video segment, and the object information includes a preset category of objects;

[0008] Determine a plurality of target video segments from the plurality of video segments; the segment type of each target video segment in the plurality of target video segments is the same;

[0009] Generate a target video based on the plurality of target video segments.

[0010] In some possible embodiments, after determining the segment type of each video segment and before determining a plurality of target video segments from the plurality of video segments, the method further includes:

[0011] Obtain a target audio; the target audio includes a plurality of points;

[0012] Determining a plurality of target video segments from the plurality of video segments includes:

[0013] Determine a plurality of target video segments from the plurality of video segments that match the number of the plurality of points;

[0014] Generating a target video based on the plurality of target video segments includes:

[0015] Generate a target video based on multiple target video segments and a target audio.

[0016] In some possible embodiments, obtaining a video to be processed includes:

[0017] Process the obtained live video stream data to obtain a video to be processed.

[0018] In some possible embodiments, the duration of the video to be processed is a first preset duration; processing the obtained live video stream data to obtain a video to be processed includes:

[0019] When a first video generation request of a first account in a live state is received, use the moment when the first video generation request is received as the starting moment;

[0020] Start timing from the starting moment, and when the preset duration is reached, use the live video stream data obtained in real time as the video to be processed;

[0021] Use the current moment as the starting moment again to obtain a new video to be processed.

[0022] In some possible embodiments, processing the obtained live video stream data to obtain a video to be processed includes:

[0023] When a second video generation request of a first account in a non-live state is received, obtain the live video stream data corresponding to the live data identifier according to the live data identifier carried in the second video generation request;

[0024] Perform segmentation processing on the live video stream data according to a second preset duration to obtain multiple videos to be processed; wherein, the duration of each video to be processed is the second preset duration.

[0025] In some possible embodiments, determining multiple video segments from the video to be processed includes:

[0026] Perform video content understanding processing on the video to be processed to obtain video content description information; the video content description information includes multiple video segments and object description information corresponding to each video segment in the multiple video segments, and the category of the object in the object description information includes a preset category.

[0027] In some possible embodiments, determining the segment type of each video segment includes:

[0028] Perform object recognition on each video segment to obtain the object information of each video segment;

[0029] When the object information includes a target object and at least one preset category object, determine that the video segment is of the first segment type; or; when the object information only includes multiple preset category objects with the same preset attributes and does not include the target object, determine that the video segment is of the second segment type; or; when the object information includes at least two preset category objects with different preset attributes and does not include the target object, determine that the video segment is of the third segment type.

[0030] In some possible embodiments, determining multiple target video segments from multiple video segments includes:

[0031] Determine multiple candidate video segments from multiple video segments according to the target segment type;

[0032] Obtain multiple feature information of each candidate video segment among the multiple candidate video segments;

[0033] Sort the multiple candidate video segments according to the multiple feature information of each candidate video segment to obtain the sorted multiple candidate video segments;

[0034] Determine multiple target video segments that match the number of multiple points from the sorted multiple candidate video segments.

[0035] In some possible embodiments, determining multiple target video segments that match the number of multiple points from the sorted multiple candidate video segments includes:

[0036] Determine the first N candidate video segments among the sorted multiple candidate video segments as multiple target video segments; where the size of N is determined according to the number of multiple points.

[0037] In some possible embodiments, multiple target video segments are used to be fused with corresponding audio segments in sequence; the audio segments represent the audio segments between adjacent points among multiple points; generating a target video according to the multiple target video segments and the target audio includes:

[0038] For each target video segment among the multiple target video segments:

[0039] If the duration of the target video segment is greater than the duration of the corresponding audio segment, crop the target video segment according to the duration of the audio segment or play it at the first speed;

[0040] Or;

[0041] If the duration of the target video segment is less than the duration of the corresponding audio segment, replay the target video segment according to the duration of the audio segment or play it at the second speed;

[0042] Wherein, the first speed is greater than the second speed.

[0043] In some possible embodiments, determining multiple target video segments that match the number of multiple points from multiple sorted candidate video segments includes:

[0044] Regarding the first point among the multiple points as the current point;

[0045] According to the duration of the audio segment between the current point and the next adjacent point, determining, from the multiple sorted candidate video segments, a target video segment that matches the duration of the audio segment;

[0046] Regarding the next adjacent point as the current point again to obtain a new set of multiple target video segments.

[0047] In some possible embodiments, the multiple feature information includes multiple live broadcast attribute data;

[0048] Sorting the multiple candidate video segments according to the multiple feature information of each candidate video segment includes:

[0049] Sorting the multiple candidate video segments according to a preset live broadcast attribute priority order and the multiple live broadcast attribute data of each candidate video segment.

[0050] In some possible embodiments, the multiple live broadcast attribute data includes at least one of the peak number of viewers in the live broadcast room and the number of clicks on the object link.

[0051] In some possible embodiments, obtaining the target audio includes:

[0052] Determining the target duration of the target video to be generated;

[0053] Determining, from the audio resource pool, a target audio that matches the target duration.

[0054] According to the second aspect of the embodiments of the present disclosure, there is provided a video generation device, including:

[0055] A first acquisition module configured to acquire a video to be processed;

[0056] A first determination module configured to determine multiple video segments from the video to be processed; each video segment among the multiple video segments represents an image sequence depicting a preset category of object;

[0057] A second determination module configured to determine the segment type of each video segment; the segment type is determined according to the object information in each video segment, and the object information includes a preset category of object;

[0058] A third determination module, configured to determine a plurality of target video segments from a plurality of video segments; each of the plurality of target video segments has the same segment type;

[0059] A generation module, configured to generate a target video based on the plurality of target video segments.

[0060] In some possible embodiments, the apparatus further includes:

[0061] A second acquisition module, configured to acquire target audio; the target audio includes a plurality of points;

[0062] A third determination module, configured to determine a plurality of target video segments that match the number of the plurality of points from the plurality of video segments;

[0063] A generation module, configured to generate a target video according to the plurality of target video segments and the target audio.

[0064] In some possible embodiments, the first acquisition module is further configured to process the acquired live video stream data to obtain a video to be processed.

[0065] In some possible embodiments, the duration of the video to be processed is a first preset duration; the first acquisition module includes:

[0066] A first processing unit, configured to, when receiving a first video generation request of a first account in a live state, use the moment when the first video generation request is received as the start moment;

[0067] A second processing unit, configured to start timing from the start moment, and when the preset duration is reached, use the live video stream data obtained in real time as the video to be processed;

[0068] A third processing unit, configured to use the current moment as the start moment again to obtain a new video to be processed.

[0069] In some possible embodiments, the first acquisition module includes:

[0070] A fourth processing unit, configured to, when receiving a second video generation request of a first account in a non-live state, obtain the live video stream data corresponding to the live data identifier according to the live data identifier carried in the second video generation request;

[0071] A fifth processing unit, configured to perform segmentation processing on the live video stream data according to a second preset duration to obtain a plurality of videos to be processed; wherein, the duration of each video to be processed is the second preset duration.

[0072] In some possible embodiments, the first determination module is further configured to perform video content understanding processing on the video to be processed to obtain video content description information; the video content description information includes a plurality of video segments and object description information corresponding to each video segment in the plurality of video segments, and the categories of the objects in the object description information include preset categories.

[0073] In some possible embodiments, the second determination module includes:

[0074] An identification unit, configured to perform object identification on each video segment to obtain object information of each video segment;

[0075] A type determination unit, configured to determine that the video segment is of the first segment type when the object information includes a target object and at least one preset category object; or; when the object information only includes preset category objects with the same plurality of preset attributes and does not include the target object, determine that the video segment is of the second segment type; or; when the object information includes at least two preset category objects with different preset attributes and does not include the target object, determine that the video segment is of the third segment type.

[0076] In some possible embodiments, the third determination module includes:

[0077] A first determination unit, configured to determine a plurality of candidate video segments from the plurality of video segments according to the target segment type;

[0078] An acquisition unit, configured to acquire a plurality of feature information of each candidate video segment in the plurality of candidate video segments;

[0079] A sorting unit, configured to sort the plurality of candidate video segments according to the plurality of feature information of each candidate video segment to obtain the sorted plurality of candidate video segments;

[0080] A second determination unit, configured to determine a plurality of target video segments that match the number of a plurality of points from the sorted plurality of candidate video segments.

[0081] In some possible embodiments, the second determination unit is further configured to determine the top N candidate video segments in the sorted plurality of candidate video segments as the plurality of target video segments; where the size of N is determined according to the number of the plurality of points.

[0082] In some possible embodiments, the plurality of target video segments are used to be fused with the corresponding audio segments in sequence; the audio segments represent the audio segments between adjacent points among the plurality of points;

[0083] A generation module, configured to perform, for each of a plurality of target video segments: if the duration of a target video segment is greater than the duration of the corresponding audio segment, crop the target video segment according to the duration of the audio segment or play it at a first speed; or; if the duration of the target video segment is less than the duration of the corresponding audio segment, replay the target video segment according to the duration of the audio segment or play it at a second speed; wherein the first speed is greater than the second speed.

[0084] In some possible embodiments, the second determination unit is further configured to perform taking the first point among a plurality of points as the current point; determining, among the sorted plurality of candidate video segments, a target video segment that matches the duration of the audio segment according to the duration of the audio segment between the current point and the next adjacent point; and taking the next adjacent point as the current point again to obtain a new plurality of target video segments.

[0085] In some possible embodiments, the plurality of feature information includes a plurality of live broadcast attribute data;

[0086] The sorting unit is further configured to perform sorting the plurality of candidate video segments according to a preset live broadcast attribute priority order and the plurality of live broadcast attribute data of each candidate video segment.

[0087] In some possible embodiments, the plurality of live broadcast attribute data includes at least one of the peak number of viewers in the live broadcast room and the number of clicks on the object link.

[0088] In some possible embodiments, the first acquisition module is further configured to perform determining the target duration of the target video to be generated; and determining a target audio that matches the target duration from the audio resource pool.

[0089] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including:

[0090] A processor;

[0091] A memory for storing instructions executable by the processor;

[0092] Wherein, the processor is configured to execute instructions to implement the video generation method provided in the first aspect of the embodiments of the present disclosure.

[0093] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, when the instructions in the computer-readable storage medium are executed by the processor of the electronic device, enabling the electronic device to execute the video generation method provided in the first aspect of the embodiments of the present disclosure.

[0094] According to a fifth aspect of the embodiments of the present disclosure, a computer program product is provided. The computer program product includes a computer program stored in a readable storage medium. At least one processor of the computer device reads and executes the computer program, so that the computer device executes the video generation method provided in the first aspect of the embodiments of the present disclosure.

[0095] The technical solutions provided by the embodiments of the present disclosure at least bring the following beneficial effects:

[0096] By processing the video to be processed, multiple video segments related to the preset category object are obtained, the segment types of the video segments are determined, and multiple target video segments of the required segment types are selected from the multiple video segments to generate a target video; the types of the target video segments in the finally generated target video are the same, and the preset category object in the video to be processed is displayed, so that the content correlation degree between the target video and the video to be processed is relatively high, and the overall display effect of the target video is relatively unified, which can help users quickly understand the content of the original video and improve the user experience.

[0097] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0098] The accompanying drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation to the present disclosure.

[0099] Figure 1 is a schematic diagram of an application environment shown according to an exemplary embodiment;

[0100] Figure 2 is a flowchart of a video generation method shown according to an exemplary embodiment;

[0101] Figure 3 is a flowchart of obtaining a video to be processed shown according to an exemplary embodiment;

[0102] Figure 4 is a flowchart of obtaining another video to be processed shown according to an exemplary embodiment;

[0103] Figure 5 is a flowchart of determining the segment type of each video segment shown according to an exemplary embodiment;

[0104] Figure 6 is a flowchart of determining multiple target video segments shown according to an exemplary embodiment;

[0105] Figure 7Another flowchart for determining multiple target video segments shown according to an exemplary embodiment;

[0106] Figure 8 A schematic diagram for determining a target video segment shown according to an exemplary embodiment;

[0107] Figure 9 A block diagram of a video generation device shown according to an exemplary embodiment;

[0108] Figure 10 A block diagram of an electronic device for video generation shown according to an exemplary embodiment. Detailed implementation manners

[0109] To enable those of ordinary skill in the art to better understand the technical solutions of the present disclosure, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings.

[0110] It should be noted that the terms "first", "second", etc. in the specification and claims of the present disclosure and the above-mentioned drawings are used to distinguish similar first objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0111] It should be noted that the user information involved in the present disclosure (including but not limited to user device information, user personal information, etc.) is information that has been authorized by the user or fully authorized by all parties.

[0112] Please refer to Figure 1 , Figure 1 A schematic diagram of an application environment of a video generation method shown according to an exemplary embodiment. As Figure 1 shown, the application environment may include a server 01 and a client 02.

[0113] In some possible embodiments, the server 01 may obtain a video to be processed, and determine multiple video segments from the video to be processed; each video segment in the multiple video segments represents an image sequence depicting a preset category of objects; determine the segment type of each video segment; the segment type is determined according to the object information in each video segment, and the object information includes the preset category of objects; determine multiple target video segments from the multiple video segments; the segment type of each target video segment in the multiple target video segments is the same; generate a target video based on the multiple target video segments, and then send the target video to the client 02.

[0114] In some possible embodiments, the server 01 may include an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The operating system running on the server may include, but is not limited to, Android, IOS, Linux, Windows, Unix, etc.

[0115] In some possible embodiments, the above-mentioned client 02 may include, but is not limited to, types of clients such as smart phones, desktop computers, tablet computers, laptop computers, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices. It may also be software running on the above-mentioned clients, such as application programs, applets, etc. Optionally, the operating system running on the client may include, but is not limited to, Android, IOS, Linux, Windows, Unix, etc.

[0116] In addition, it should be noted that Figure 1 The shown is only an application environment of the video generation method provided by the present disclosure. In actual applications, there may also be other application environments.

[0117] Figure 2 is a flowchart of a video generation method shown according to an exemplary embodiment. As Figure 2 shown, the video generation method may be applied to a server or other node devices, and includes the following steps:

[0118] In step S201, obtain a video to be processed.

[0119] In some possible embodiments, the server may receive a video to be processed transmitted by other devices; or, the server may receive an original video transmitted by other devices, and after preprocessing the original video by the server, obtain the video to be processed. Among them, the other device is the device of the provider of the video to be processed or the original video.

[0120] In some possible embodiments, the server may also obtain the video to be processed or the original video from the local storage area or the remote storage area. When the obtained video is the original video, after preprocessing the original video by the server, obtain the video to be processed. Among them, the video to be processed or the original video is pre-stored by the server in the local storage area or the remote storage area.

[0121] In the embodiments of the present disclosure, the video to be processed may be a video shot by a client user using video interaction software. In a specific application scenario, such as a live broadcast scenario, the video to be processed may also be a video obtained by processing live video stream data.

[0122] In the related art, users can only know what type of goods the merchant sells after entering the video live broadcast room, resulting in that when users are looking for the live broadcast room of the goods they need, they may repeatedly enter multiple irrelevant video live broadcast rooms, with low efficiency. Therefore, the present disclosure can generate a target video based on the live video stream data, and only the segments related to the goods in the live broadcast room are displayed in the target video. In this way, the time for users to search for the live broadcast room can be saved, and the user experience can be improved.

[0123] Correspondingly, in some possible embodiments, in the above step S201, obtaining the video to be processed may specifically include: obtaining live video stream data; processing the obtained live video stream data to obtain the video to be processed.

[0124] Specifically, the server may process the live video stream data that is being broadcast or the live replay video that has been broadcast to obtain the video to be processed, and then further process the video to be processed through subsequent steps to generate the corresponding target video. The target video can be delivered to the promotion platform. When the user browses the promotion interface of the promotion platform through the client, the user can view the target video. By watching the target video, the user can quickly understand the objects displayed in the corresponding live broadcast room; according to different types of live broadcast rooms, correspondingly, the types of objects displayed in the live broadcast room are different. For example, when the live broadcast room is a shopping live broadcast room, the object displayed is a commodity; when the live broadcast room is a game live broadcast room, the objects displayed may include virtual characters in the game.

[0125] In the above embodiments, the live video stream data obtained is processed to obtain a video to be processed, and further combined with subsequent steps to generate a target video. On the one hand, for live users, promoting the target video through a promotion platform can attract more viewing users interested in the objects shown to them. On the other hand, for users watching the live broadcast, by watching the target video, they can quickly understand the products shown in the corresponding live broadcast room, so that users can quickly find the live broadcast room where the object they are looking for is located without frequently entering and exiting the live broadcast room, which can greatly improve the user experience.

[0126] In some possible embodiments, the duration of the video to be processed is a first preset duration; the processing of the obtained live video stream data to obtain the video to be processed may specifically include the following steps as Figure 3 shown:

[0127] In step S301, when a first video generation request of a first account in the live state is received, the moment when the first video generation request is received is taken as the starting moment.

[0128] In practical applications, the live user who is currently live, that is, the first account in the live state, can, according to needs, perform a preset operation on the live broadcast interface displayed on the client to trigger the generation of the first video generation request, and at the same time send the first video generation request to the server. The first video generation request is used to instruct the server to process the live video stream data corresponding to the first account to generate a corresponding target video.

[0129] In step S303, starting from the starting moment, when the preset duration is reached, the live video stream data obtained in real time is used as the video to be processed.

[0130] In step S305, the current moment is taken as the starting moment, and the steps are repeated: starting from the starting moment, when the first preset duration is reached, the live video stream data obtained in real time is used as the video to be processed.

[0131] For the first account in the live state, the server obtains the live video stream data corresponding to the first account. Correspondingly, the server can process the live video stream data of the first account every preset duration, use the live video stream data of the preset duration as the video to be processed, and obtain the target video through subsequent processing steps. In this way, the corresponding target video can be generated in a timely manner so that users can watch the target video in a timely manner, understand the objects being shown in the corresponding live broadcast room through the target video, reduce the repeated operations of users, and improve the user experience.

[0132] Among them, the first preset duration can be 10 minutes, 20 minutes, or 30 minutes, etc.; the specific duration can be set according to actual needs.

[0133] In a specific embodiment, the host of a certain shopping live broadcast room in the live broadcast state can trigger the generation of a first video generation request by clicking the target video generation button on the live broadcast interface of his client. After the server receives the first video generation request sent by the client, it starts timing and obtains the live video stream data corresponding to the host. When the first 10 minutes arrives, the live video stream data obtained within the first 10 minutes is used as the video to be processed, and the subsequent steps of processing are performed on the video to be processed to obtain the target video. At the same time, the timing is cleared, and a new round of timing starts. When the second 10 minutes arrives, a video to be processed can be obtained based on the live video stream data obtained within the second 10 minutes, and then the subsequent steps of processing are similarly performed on the video to be processed, and a new target video can be obtained.

[0134] In the above embodiment, the server can obtain a video to be processed every preset duration. The video to be processed contains the latest live data, so that the target video reflecting the current live object can be obtained based on the video to be processed, so that the object shown in the target video can be kept as consistent as possible with the object shown in the actual live broadcast room, so that the viewing user can understand the latest situation of the live broadcast room.

[0135] In some possible embodiments, the above processing of the obtained live video stream data to obtain the video to be processed may specifically include the following steps as Figure 4 shown:

[0136] In step S401, when receiving a second video generation request of a first account in the non-live state, the live video stream data corresponding to the live data identifier is obtained according to the live data identifier carried in the second video generation request.

[0137] In practical applications, for the past live replay videos of the first account, the server can also process them to obtain the video to be processed, and then obtain the target video through subsequent steps. Correspondingly, the live data identifier represents the identifier of the live replay video. Specifically, the first account can select the target live replay video in the live replay video list displayed on the client to trigger the generation of the second video generation request, and at the same time send the second video generation request to the server. The second video generation request carries the identifier of the target live replay video, and the second video generation request is used to instruct the server to obtain the target live replay video according to the identifier of the target live replay video, and process the target live replay video to obtain the target video.

[0138] In step S403, the live video stream data is segmented according to the second preset duration to obtain a plurality of videos to be processed; wherein, the duration of each video to be processed is the second preset duration.

[0139] Correspondingly, after obtaining the target live replay video, the server can perform segmentation processing on the target live replay video. Specifically, the server segments the target live replay video according to the time sequence of each image frame in the target live replay video to obtain multiple videos to be processed, and the duration of each video to be processed is the second preset duration; where the second preset duration can be 10 minutes, 20 minutes, 30 minutes, etc.; the specific duration can be set according to actual requirements.

[0140] The server can process each video to be processed among the multiple videos to be processed; or; the server can select a preset number of videos to be processed from the multiple videos to be processed for subsequent processing steps. Among them, the selection of the preset number of videos to be processed can be executed by the server, or the server can send a corresponding selection instruction to the client of the first account. The selection instruction instructs the first account to select a preset number of videos to be processed from the multiple videos to be processed. The server determines the preset number of videos to be processed according to the feedback information of the selection instruction returned by the client, and then processes each of the videos to be processed in subsequent steps. In this way, in the target video generated based on the live replay video, there are objects that have been shown in the past live broadcast room. By displaying the target video on the corresponding object description page in the client, the dynamic explanation effect of the object is increased.

[0141] For example, a cosmetics merchant selects a live replay video about cosmetics from its past live broadcasts. The server can segment the live replay video about cosmetics to obtain multiple videos to be processed. Each video to be processed may contain a segment of content unrelated to cosmetics, or may contain introductions of multiple different types of cosmetics. After the server processes the videos to be processed through subsequent steps, the content unrelated to cosmetics can be deleted, and the obtained target video only contains introductions of the same type of cosmetics. The server can display the target video on the shopping page of the corresponding cosmetics in the shopping software in the client to increase the dynamic explanation effect of the cosmetics, help the purchasing user quickly and intuitively understand the cosmetics, and improve the user experience.

[0142] In step S203, multiple video segments are determined from the video to be processed; each video segment among the multiple video segments represents an image sequence describing a preset category of objects.

[0143] In the embodiment of the present disclosure, the server processes the video to be processed and determines multiple video segments from the video to be processed. Each video segment includes a certain number of image frames, that is, each video segment is an image sequence, and the number of image frames included in each of the different video segments can be different. The preset category of objects refers to a certain category of objects preset in advance.

[0144] In some possible embodiments, the server processes the video to be processed through a preset algorithm, and identifies the image frames in the video to be processed that are related to the objects of a certain preset category. Among the adjacent two image frames in each video segment, in the original video to be processed, they can be two consecutive image frames or two image frames within a preset number of frames apart.

[0145] In some possible embodiments, the above determination of multiple video segments from the video to be processed may specifically include the following steps:

[0146] Perform video content understanding processing on the video to be processed to obtain video content description information; the video content description information includes multiple video segments and object description information corresponding to each video segment among the multiple video segments, and the category of the object in the object description information includes preset categories.

[0147] Among them, the preset category is the above-mentioned certain preset category. The main purpose of performing video content understanding processing on the video to be processed is to identify the segments in the video to be processed that are related to the objects of a certain preset category; correspondingly, the server can use a video content understanding algorithm or a related machine learning model to process the video to be processed, or, through one or a combination of methods such as image semantic extraction, video action recognition, and video scene understanding, extract the object description information related to the preset category from the video content description information to obtain multiple video segments.

[0148] In a specific embodiment, the preset category is a commodity; the video to be processed is obtained based on the live video stream data of a shopping live broadcast room. Assuming that the duration of the video to be processed is 10 minutes, after the video to be processed undergoes video content understanding, the obtained video content description information includes multiple video segments, namely video segment 1, video segment 2, video segment 3, video segment 4... and the object description information corresponding to each video segment. For example, the object description information of video segment 1 includes: the category of the object includes commodities and people, the object description information of video segment 2 includes: the category of the object includes commodities and people, the object description information of video segment 3 includes: the category of the object includes commodities, and the object description information of video segment 4 includes: the category of the object includes commodities. It should be noted that the video content understanding processing can only perform a coarse-grained recognition on the video to be processed, that is, only the general category of the object needs to be recognized. For example, in this specific embodiment, the category of the object in the object description information includes commodities. As for the specific attributes of the object, they can be recognized in fine-grained in subsequent steps.

[0149] In the above embodiments, through video content understanding, segments related to the preset category (commodity) objects in the video to be processed are extracted, and the multiple video segments obtained are all content related to the preset category (commodity) objects. In this way, invalid content can be eliminated, so that the content of the target video finally generated based on the multiple video segments is all content centered around the preset category (commodity) objects, facilitating the viewing user to obtain effective information therefrom.

[0150] In some possible embodiments, the duration of each video segment does not exceed a preset duration. Correspondingly, after the server performs video content understanding processing on the video to be processed and obtains multiple video segments, the following steps may further be included:

[0151] The server determines the number of image frames in each video segment;

[0152] If the number of image frames of the video segment is greater than the preset number of frames, frame reduction or frame rate increase processing is performed on the video segment.

[0153] Among them, frame reduction refers to reducing the number of image frames in the video segment; frame rate increase refers to increasing the number of image frames displayed per second. The preset number of frames corresponds to the preset duration of the video segment, and the two can be converted according to the frame rate of the original video segment. In this way, through frame reduction or frame rate increase processing, the duration of the video segment can be controlled not to exceed the preset duration; since multiple video segments need to be spliced and synthesized to generate a target video in practical applications, the preset duration can be set relatively short in actual applications, such as 2 seconds, 4 seconds, 5 seconds, or 10 seconds, etc.

[0154] In step S205, the segment type of each video segment is determined; the segment type is determined according to the object information in each video segment, and the object information includes preset category objects.

[0155] In the embodiments of the present disclosure, for each video segment, the server determines the segment type of the video segment according to the object information in the video segment.

[0156] In some possible embodiments, the object information may be obtained based on the object description information in the video content description information in the above embodiments; at this time, the recognition granularity requirements for the corresponding video content understanding algorithm or machine learning model used by the server in the above embodiments are relatively high, that is, the attribute information of each object in the video to be processed needs to be recognized. The advantage of doing this is that the step of recognizing each video segment can be omitted.

[0157] In some possible embodiments, the object information may be obtained after the server performs fine-grained recognition on each video segment. Correspondingly, the above determination of the segment type of each video segment may specifically include the following steps as Figure 5 shown:

[0158] In step S501, object recognition is performed on each video clip to obtain the object information of each video clip.

[0159] Specifically, the server can use an object recognition algorithm to obtain the object information of each video clip; among them, the object information includes the attribute information of all objects in the video clip; among them, the attribute information can include at least one of name, color, model, and brand. It can be understood that in a video clip, multiple preset category objects can be included, and the multiple preset category objects appear in different image frames of the video clip.

[0160] The server can determine the clip type of the video clip according to the object information of each video clip. That is, the corresponding clip type is determined according to the category and / or attribute information of the objects included in the video clip. For example, the following step S503 provides a specific determination method.

[0161] In step S503, when the object information includes a target object and at least one preset category object, it is determined that the video clip is of the first clip type; or; when the object information only includes multiple preset category objects with the same preset attributes and does not include the target object, it is determined that the video clip is of the second clip type; or; when the object information includes at least two preset category objects with different preset attributes and does not include the target object, it is determined that the video clip is of the third clip type.

[0162] Among them, the target object is a certain category of object preset in advance, and the category of the target object is different from the category of the preset category object. The preset attribute can be any one of the above name, color, model, and brand.

[0163] In a live broadcast scenario, the target object can be the live streamer, and the preset category object can be the commodity object that the live streamer is introducing. The preset attribute can be the object name. Based on the above embodiments, the server performs object recognition on video clip 1 and obtains the object information of video clip 1, including 4 objects. That is, the category of object 1 is human, and the name is the live streamer; the category of object 2 is a commodity, the name is lipstick, and the color is 201; the category of object 3 is a commodity, the name is lipstick, and the color is 202; the category of object 4 is a commodity, the name is powder puff, and the color is 01. Since the object information of video clip 1 contains the target object (object 1) and three preset category objects (the categories of object 2, object 3, and object 4 are all commodities), it is determined that video clip 1 is the first clip type. Similarly, the object information of video clip 2 includes the live streamer, lipstick 203, and lipstick 204. Accordingly, it is determined that video clip 2 is the first clip type. Similarly, the object information of video clip 3 only includes lipstick 205, lipstick 206, and lipstick 207. Accordingly, it is determined that video clip 3 is the second clip type. Similarly, the object information of video clip 4 includes lipstick 208 and powder puff 02. Accordingly, it is determined that video clip 4 is the third clip type.

[0164] In the above embodiments, by determining the clip type of each video clip, in subsequent steps, a target video is generated based on video clips of the same clip type, which can ensure the overall painting style of the target video is consistent and improve the user viewing experience. In practical applications, the corresponding clip type can also be selected according to different requirements to generate a target video that meets specific user needs. For example, considering the preferences of some users for a specific live broadcast room, the server can generate a target video based on the video clips under the first clip type. Or, some users have a high degree of interest in a specific type of commodity, and the server can generate a target video based on the video clips under the second clip type.

[0165] In step S207, multiple target video clips are determined from multiple video clips; the clip type of each target video clip in the multiple target video clips is the same.

[0166] In step S209, a target video is generated based on the multiple target video clips.

[0167] In the embodiments of the present disclosure, the types of the target video clips in the target video are the same, and the preset category objects in the video to be processed are displayed, so that the content correlation between the target video and the video to be processed is relatively high, and the overall display effect of the target video is relatively unified.

[0168] In some possible embodiments, after step S205 and before step S207, the video generation method of the embodiments of the present disclosure may further include:

[0169] In step S206, the target audio is obtained; the target audio includes multiple points.

[0170] In the embodiments of the present disclosure, the server obtains the target audio and obtains multiple points of the target audio. The points represent the moments where the beat points are located in the target audio. The server determines multiple target video segments that match the number of multiple points from multiple video segments. Each target video segment is used to fill between two adjacent points to generate a target video, and the background music of the target video is the target audio.

[0171] Correspondingly, determining multiple target video segments from multiple video segments may include: determining multiple target video segments that match the number of multiple points from multiple video segments.

[0172] Correspondingly, generating a target video based on multiple target video segments may include: generating a target video according to multiple target video segments and the target audio.

[0173] In the embodiments of the present disclosure, the server generates a target video according to multiple target video segments and the target audio. Specifically, the server may first remove the original background music of multiple target video segments, and then fuse each target video segment with the corresponding audio segment to obtain a target video; in this way, the moment of each point in the target video, that is, the moment where the beat point is located, can switch the target video segment, which can achieve accurate beat matching. Compared with the traditional video showing preset category objects, the display effect is richer, and the user viewing experience can be improved.

[0174] In some possible embodiments, the beat points are obtained in advance by a beat recognition model and stored in the server.

[0175] In some possible embodiments, the above-mentioned obtaining of the target audio may specifically include the following steps:

[0176] Determine the target duration of the target video to be generated;

[0177] Determine the target audio that matches the target duration from the audio resource pool.

[0178] Specifically, an audio resource pool is stored in the server, and the audio resource pool contains a large number of audio resources; the server can receive a user request sent by the client, and the user request carries the target duration of the target video required by the user; the server determines the target audio that matches the target duration from the audio resource pool according to the target duration required by the user; among them, the target audio that matches the target duration may include an audio with a duration of the target duration, or may include an audio with a duration close to the target duration.

[0179] In the above embodiments, the user requirements are given priority. Based on the user's duration requirement for the target video, the target audio that meets the duration requirement is first determined, and then, according to the number of points of the target audio, a plurality of target video segments with the same number are determined from a plurality of video segments.

[0180] In other embodiments of the present disclosure, the server may also, based on the user's requirements for the target video segments, determine the number of points according to the number of the target video segments selected by the user on the client, and then search for the audio that meets the number of points in the audio resource pool and use it as the target audio.

[0181] In some possible embodiments, the determining of a plurality of target video segments from a plurality of video segments may specifically include the following steps as Figure 6 shown:

[0182] In step S601, according to the target segment type, a plurality of candidate video segments are determined from a plurality of video segments.

[0183] Among them, each candidate video segment in the plurality of candidate video segments is of the target segment type. The target segment type may be any one of the first segment type, the second segment type, and the third segment type in the above embodiments; the server may generate a corresponding target video based on each segment type. Taking the target segment type as the first segment type as an example, at this time, video segment 1 and video segment 2 are candidate video segments.

[0184] In step S603, a plurality of feature information of each candidate video segment in the plurality of candidate video segments is obtained.

[0185] In step S605, according to the plurality of feature information of each candidate video segment, the plurality of candidate video segments are sorted to obtain the sorted plurality of candidate video segments.

[0186] In practical applications, the number of candidate video segments may be relatively large, while the number of video segments that can be accommodated in a target video is limited. Therefore, in the above embodiments, based on the plurality of feature information of each candidate video segment, the plurality of candidate video segments are sorted, and a plurality of target video segments matching the number of a plurality of points are determined from the sorted plurality of candidate video segments. In this way, it can be ensured that the candidate video segments with important features can appear in the target video, and the content quality of the target video can be improved.

[0187] When applied to a live broadcast scenario, the above-mentioned multiple feature information may include multiple live broadcast attribute data; correspondingly, the above-mentioned sorting of multiple candidate video segments according to the multiple feature information of each candidate video segment specifically includes: sorting multiple candidate video segments according to a preset live broadcast attribute priority order and the multiple live broadcast attribute data of each candidate video segment.

[0188] Among them, the live broadcast attribute data refers to the metric data related to the live broadcast room;

[0189] In a specific embodiment, the above-mentioned multiple live broadcast attribute data includes at least one of the peak number of viewers in the live broadcast room, the number of clicks on the object link, the number of likes in the live broadcast room, the number of comments, and the number of forwards; among them, the number of clicks on the object link represents the number of clicks on the shopping cart icon displayed in the live broadcast room screen.

[0190] In practical applications, the priority order of each live broadcast attribute data can be determined according to the importance of the above-mentioned live broadcast attribute data, and then the multiple candidate video segments can be sorted according to the priority order. For example, first select the candidate video segment with the largest peak number of viewers in the live broadcast room, and then select the candidate video segment with the largest number of clicks on the shopping cart icon. In this way, according to the preset live broadcast attribute priority order, the sorting of all candidate video segments is determined until all are determined, or the sorting stops when the preset ranking is reached. In order to ensure that there are enough target video segments to match the positions of the target audio in the subsequent process, the preset ranking should not be set too small.

[0191] In the above embodiment, sorting multiple candidate video segments according to the preset live broadcast attribute priority order can comprehensively consider the data under different dimensional indicators to select target video segments according to actual needs.

[0192] In step S607, determine multiple target video segments that match the number of multiple positions from the sorted multiple candidate video segments.

[0193] In some possible embodiments, after determining the sorting of all candidate video segments, the above-mentioned determining multiple target video segments that match the number of multiple positions from the sorted multiple candidate video segments may specifically include: determining the first N candidate video segments in the sorted multiple candidate video segments as multiple target video segments; where the size of N is determined according to the number of multiple positions.

[0194] Specifically, the server can determine the size of N as the number of positions in the target audio; that is, select an equal number of candidate video segments as target video segments according to the number of positions in the target audio.

[0195] Correspondingly, in subsequent steps, the multiple target video segments are sequentially filled into the corresponding positions in order to generate the target video.

[0196] In some possible embodiments, the server first executes the step of obtaining the target audio, and then executes the step of determining multiple target video segments; thus, the above-mentioned determining multiple target video segments that match the number of multiple positions from the sorted multiple candidate video segments may specifically further include the following steps as Figure 7 shown:

[0197] In step S701, the first position among the multiple positions is regarded as the current position.

[0198] In practical applications, for the convenience of calculation, the server uses the start time of the target audio as the time corresponding to the first position among the multiple positions, and the end time of the target audio as the last position among the multiple positions. Then, correspondingly, if the target audio includes n positions, then correspondingly, n - 1 target video segments need to be determined from the multiple candidate video segments.

[0199] For the convenience of understanding, the following is illustrated by a simple example. As Figure 8 shown, the target audio includes 10 positions, where the time corresponding to the first position (position 0) is 00'00", and the times corresponding to the remaining positions are as shown; assume there are 100 sorted candidate video segments, and the segment types of these 100 candidate video segments are the same, for example, the segment types are all the above-mentioned second segment type; the following describes how to select 9 target video segments from these 100 candidate video segments to generate the target video.

[0200] In step S703, according to the duration of the audio segment between the current position and the next adjacent position, among the sorted multiple candidate video segments, a target video segment that matches the duration of the audio segment is determined.

[0201] As Figure 8 shown, the time corresponding to the next adjacent position (position 1) is 00'02", then the duration of the audio segment is determined to be 2 seconds, and then among the sorted multiple candidate video segments, it is determined whether the duration of the currently highest-ranked candidate video segment is 2 seconds; if the duration of the currently highest-ranked candidate video segment is 2 seconds, then the currently highest-ranked candidate video segment is used as the target video segment; or; if the duration of the currently highest-ranked candidate video segment is not 2 seconds, then it is determined whether the duration of the candidate video segment in the next rank of the currently highest-ranked candidate video segment is 2 seconds, if the duration of the candidate video segment in the next rank is 2 seconds, then it is used as the target video segment; if the duration of the candidate video segment in the next rank is also not 2 seconds, then continue to search downwards until a candidate video segment with a duration of 2 seconds is found;

[0202] If there is no candidate video segment with a duration of 2 seconds among 100 candidate video segments, then adjust 2 seconds based on a preset adjustment threshold to obtain an adjusted duration, and then start from the highest-ranked candidate video segment among the 100 candidate video segments to determine whether its duration meets the adjusted duration; for example, the adjusted duration can be any duration within the range of 1 second to 3 seconds, that is, the preset adjustment threshold is 1 second; then as long as the duration of the candidate video segment is within the range of 1 second to 3 seconds, it can be determined as the target video segment.

[0203] In step S705, take the next adjacent point as the current point, and repeat step S703: According to the duration of the audio segment between the current point and the next adjacent point, determine the target video segment that matches the duration of the audio segment among the sorted multiple candidate video segments; until the current point is the last point among the multiple points, to obtain multiple target video segments.

[0204] As Figure 8 shown, take point 1 as the current point and repeat the above step S703; the time corresponding to the next adjacent point (point 2) is 00'06", then determine that the duration of the audio segment between the current point (point 1) and point 2 is 4 seconds, and among the sorted multiple candidate video segments, determine the target candidate video corresponding to the audio segment between point 1 and point 2 by referring to the method introduced in the above embodiment; until the current point is point 9, to obtain 9 target video segments.

[0205] In the above embodiment, after the server obtains the target audio, according to the duration of the audio segment between adjacent points among the multiple points of the target audio, preferentially use the candidate video segment whose duration matches the duration of the audio segment among the multiple candidate video segments as the target video segment. In this way, the target video segment can maintain the original playback rate of the picture and ensure the picture integrity.

[0206] In some possible embodiments, multiple target video segments are used to be fused with the corresponding audio segments in sequence; the audio segments represent the audio segments between adjacent points among the multiple points; the specific steps of generating the target video according to the multiple target video segments and the target audio may include the following:

[0207] For each target video segment among the multiple target video segments: If the duration of the target video segment is greater than the duration of the corresponding audio segment, crop the target video segment according to the duration of the audio segment or play it at the first speed; or; if the duration of the target video segment is less than the duration of the corresponding audio segment, replay the target video segment according to the duration of the audio segment or play it at the second speed; where the first speed is greater than the second speed.

[0208] Fuse the processed target video segment with the corresponding audio segment to obtain the target video.

[0209] Among them, cropping refers to performing frame reduction processing on the target video segment so that the duration of the target video segment after frame reduction is equal to the duration of the corresponding audio segment; the first speed ratio is greater than 1, and the second speed ratio is less than 1. In this way, when the duration of the target video segment is greater than the duration of the corresponding audio segment, playing the target video segment at a speed greater than 1 can shorten the target video segment; when the duration of the target video segment is less than the duration of the corresponding audio segment, playing the target video segment at a speed less than 1 can extend the duration of the target video segment.

[0210] In a specific embodiment, the multiple target video segments are the top N candidate video segments among the sorted multiple candidate video segments. For example, Figure 8 the top 9 candidate video segments selected from 100 sorted candidate video segments are used to fuse with the corresponding audio segments in sequence. Correspondingly, in the process of generating the target video according to the top 9 candidate video segments and the target audio, for each candidate video segment among the top 9 candidate video segments, determine the size relationship between its duration and the duration of the corresponding audio segment. Taking the first candidate video segment as an example, the duration of its corresponding audio segment is 2 seconds. Assuming the duration of the first candidate video segment is 4 seconds, the number of image frames in the first candidate video segment can be reduced, or it can be played at 2 times the speed to adjust the duration of the first candidate video segment to 2 seconds, making it equal to the duration of the corresponding audio segment; for another example, the second candidate video segment, the duration of its corresponding audio segment is 4 seconds. Assuming the duration of the second candidate video segment is 2 seconds, then, in the corresponding audio segment, the second candidate video segment can be played repeatedly. When playing for the second time, only the content of the first 1 second needs to be played, or the second candidate video segment can be played at 0.5 times the speed to adjust the duration of the second candidate video segment to 4 seconds.

[0211] The duration adjustment method of this embodiment is also applicable to the candidate video segments in the above embodiment with a duration in the range of 1 second to 3 seconds but not 2 seconds. By cropping, speed ratio adjustment, repeated playback, etc., the duration of the candidate video segment is trimmed to be the same as the duration of the corresponding audio segment, so as to facilitate fusing to obtain a target video that can accurately hit the beat.

[0212] In some alternative embodiments, auxiliary explanatory information, including the promotional language of the preset category object, can be added to the target video.

[0213] In the above embodiments, by processing the video to be processed, a plurality of video segments related to the preset category object are obtained, the segment type of the video segment is determined, and according to actual requirements, a plurality of target video segments of the required segment type are selected from the plurality of video segments, and combined with a plurality of points of the target audio to generate a target video; a target video segment can be displayed between adjacent points, and since the types of the target video segments are the same, the display effect of the overall target video is relatively unified; and the moment of each point in the target video, that is, the moment where the beat is captured, can switch the target video segment, achieving accurate beat capture. Compared with the traditional video for displaying the preset category object, the display effect is richer, which can improve the user viewing experience.

[0214] When applied to the live broadcast scenario, based on the live video stream data, the video to be processed is obtained. After being processed by the above steps, target videos of various styles can be obtained. Only the segments related to the product are displayed in the generated target video. For example, the target video only contains the product. The product can be products with different attributes or products with the same attribute.

[0215] The server can deliver the target video to the promotion platform. When the user browses the promotion interface of the promotion platform through the client, the target video can be viewed. By clicking to watch the target video, the user can quickly understand the object displayed in the corresponding live broadcast room. In this way, it can help the user quickly understand the content played in the live broadcast room, save the user's time to find the live broadcast room, and improve the user experience.

[0216] Figure 9 It is a block diagram of a video generation device shown according to an exemplary embodiment. Referring to Figure 9 this, the device includes a first acquisition module 901, a first determination module 902, a second determination module 903, a third determination module 904, and a generation module 905;

[0217] The first acquisition module 901 is configured to acquire the video to be processed;

[0218] The first determination module 902 is configured to determine a plurality of video segments from the video to be processed; each video segment in the plurality of video segments represents an image sequence describing the preset category object;

[0219] The second determination module 903 is configured to determine the segment type of each video segment; the segment type is determined according to the object information in each video segment, and the object information includes the preset category object;

[0220] The third determination module 904 is configured to determine a plurality of target video segments from the plurality of video segments; each target video segment in the plurality of target video segments has the same segment type;

[0221] A generation module 905, configured to generate a target video based on a plurality of target video segments.

[0222] In some possible embodiments, the apparatus further includes:

[0223] A second acquisition module, configured to acquire target audio; the target audio includes a plurality of points.

[0224] A third determination module 904, configured to determine a plurality of target video segments that match the number of the plurality of points from the plurality of video segments.

[0225] A generation module 905, configured to generate a target video according to the plurality of target video segments and the target audio.

[0226] In some possible embodiments, the first acquisition module 901 is further configured to process the acquired live video stream data to obtain a video to be processed.

[0227] In some possible embodiments, the duration of the video to be processed is a first preset duration; the first acquisition module 901 includes:

[0228] A first processing unit, configured to, when receiving a first video generation request of a first account in a live state, use the moment when the first video generation request is received as the start moment.

[0229] A second processing unit, configured to start timing from the start moment, and when the preset duration is reached, use the real-time acquired live video stream data as the video to be processed.

[0230] A third processing unit, configured to use the current moment as the start moment again to obtain a new video to be processed.

[0231] In some possible embodiments, the first acquisition module 901 includes:

[0232] A fourth processing unit, configured to, when receiving a second video generation request of a first account in a non-live state, acquire the live video stream data corresponding to the live data identifier according to the live data identifier carried in the second video generation request.

[0233] A fifth processing unit, configured to perform segmentation processing on the live video stream data according to a second preset duration to obtain a plurality of videos to be processed; wherein, the duration of each video to be processed is the second preset duration.

[0234] In some possible embodiments, the first determination module 902 is further configured to perform video content understanding processing on the video to be processed to obtain video content description information; the video content description information includes a plurality of video segments and object description information corresponding to each video segment in the plurality of video segments, and the categories of objects in the object description information include preset categories.

[0235] In some possible embodiments, the second determination module 903 includes:

[0236] An identification unit configured to perform object identification on each video segment to obtain object information of each video segment;

[0237] A type determination unit configured to perform determining that the video segment is a first segment type when the object information includes a target object and at least one preset category object; or; determining that the video segment is a second segment type when the object information only includes preset category objects with the same plurality of preset attributes and does not include the target object; or; determining that the video segment is a third segment type when the object information includes at least two preset category objects with different preset attributes and does not include the target object.

[0238] In some possible embodiments, the third determination module 905 includes:

[0239] A first determination unit configured to perform determining a plurality of candidate video segments from the plurality of video segments according to the target segment type;

[0240] An acquisition unit configured to perform acquiring a plurality of feature information of each candidate video segment in the plurality of candidate video segments;

[0241] A sorting unit configured to perform sorting the plurality of candidate video segments according to the plurality of feature information of each candidate video segment to obtain the sorted plurality of candidate video segments;

[0242] A second determination unit configured to perform determining a plurality of target video segments that match the number of a plurality of points from the sorted plurality of candidate video segments.

[0243] In some possible embodiments, the second determination unit is further configured to perform determining the first N candidate video segments in the sorted plurality of candidate video segments as the plurality of target video segments; where the size of N is determined according to the number of the plurality of points.

[0244] In some possible embodiments, the plurality of target video segments are used to be fused with corresponding audio segments in sequence; the audio segments represent the audio segments between adjacent points in the plurality of points;

[0245] A generation module 906, configured to perform, for each of a plurality of target video segments: if the duration of a target video segment is greater than the duration of the corresponding audio segment, crop the target video segment according to the duration of the audio segment or play it at a first speed; or; if the duration of the target video segment is less than the duration of the corresponding audio segment, replay the target video segment according to the duration of the audio segment or play it at a second speed; wherein, the first speed is greater than the second speed.

[0246] In some possible embodiments, the second determination unit is further configured to perform taking the first point among the plurality of points as the current point; determining, from the sorted plurality of candidate video segments, a target video segment that matches the duration of the audio segment according to the duration of the audio segment between the current point and the next adjacent point; taking the next adjacent point as the current point again to obtain a new plurality of target video segments.

[0247] In some possible embodiments, the plurality of feature information includes a plurality of live broadcast attribute data;

[0248] The sorting unit is further configured to perform sorting the plurality of candidate video segments according to a preset live broadcast attribute priority order and the plurality of live broadcast attribute data of each candidate video segment.

[0249] In some possible embodiments, the plurality of live broadcast attribute data includes at least one of the peak number of viewers in the live broadcast room and the number of clicks on the object link.

[0250] In some possible embodiments, the first acquisition module 901 is further configured to perform determining the target duration of the target video to be generated; and determining the target audio that matches the target duration from the audio resource pool.

[0251] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

[0252] Figure 10 It is a block diagram of an electronic device 1000 for video generation shown according to an exemplary embodiment.

[0253] The electronic device may be a server or a terminal device, and its internal structure diagram may be as Figure 10As shown. The electronic device includes a processor, a memory, and a network interface connected through a system bus. Among them, the processor of the electronic device is used to provide computing and control capabilities. The memory of the electronic device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the electronic device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a video generation method is implemented.

[0254] Those skilled in the art can understand that Figure 10 the structure shown in [figure reference] is only a block diagram of some structures related to the solution of the present disclosure, and does not constitute a limitation on the electronic device to which the solution of the present disclosure is applied. The specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0255] In an exemplary embodiment, an electronic device is further provided, including: a processor; a memory for storing executable instructions of the processor; wherein, the processor is configured to execute the instructions to implement the video generation method as in the embodiment of the present disclosure.

[0256] In an exemplary embodiment, a computer-readable storage medium is further provided. When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device can execute the video generation method in the embodiment of the present disclosure.

[0257] In an exemplary embodiment, a computer program product is further provided. The computer program product includes a computer program. The computer program is stored in a readable storage medium. At least one processor of the computer device reads and executes the computer program from the readable storage medium, so that the computer device executes the video generation method in the embodiment of the present disclosure.

[0258] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. This computer program can be stored in a non-volatile computer-readable storage medium. When this computer program is executed, it can include the processes of the embodiments of the above various methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0259] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present disclosure. This application is intended to cover any variations, uses, or adaptations of the present disclosure, which follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present disclosure are pointed out by the following claims.

[0260] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.

Claims

1. A video generation method, characterized in that, it includes: Obtain a video to be processed; Determine multiple video segments from the video to be processed; Each video segment in the multiple video segments represents an image sequence describing a preset category object; Determine the segment type of each video segment; the segment type is determined according to the object information in each video segment, and the object information includes the number of the preset category objects and / or target objects; Determine multiple target video segments from the multiple video segments; each target video segment in the multiple target video segments has the same segment type; Generate a target video based on the multiple target video segments; The determining multiple video segments from the video to be processed includes: Perform video content understanding processing on the video to be processed to obtain video content description information; the video content description information includes the multiple video segments and the object description information corresponding to each video segment in the multiple video segments, and the category of the object in the object description information includes the preset category; The determining the segment type of each video segment includes: Perform object recognition on each video segment to obtain the object information of each video segment; the granularity of the object recognition is finer than that of the video content understanding processing; Determine the segment type of each video segment based on the number of the preset category objects and / or target objects.

2. The video generation method according to claim 1, characterized in that, after determining the segment type of each video segment and before determining multiple target video segments from the multiple video segments, the method further includes: Obtain target audio; the target audio includes multiple points; The determining multiple target video segments from the multiple video segments includes: Determine multiple target video segments from the multiple video segments that match the number of the multiple points; The generating a target video based on the multiple target video segments includes: Generate the target video according to the multiple target video segments and the target audio.

3. The video generation method according to claim 1 or 2, characterized in that, the obtaining a video to be processed includes: Process the obtained live video stream data to obtain the video to be processed.

4. The video generation method according to claim 3, characterized in that, the duration of the video to be processed is a first preset duration; the processing the obtained live video stream data to obtain the video to be processed includes: When receiving a first video generation request of a first account in a live state, regard the moment when the first video generation request is received as the starting moment; Start timing from the starting moment, and when the preset duration is reached, use the live video stream data obtained in real time as the video to be processed; Regard the current moment as the starting moment again to obtain a new video to be processed.

5. The video generation method according to claim 3, characterized in that, the processing the obtained live video stream data to obtain the video to be processed includes: When receiving a second video generation request of a first account in a non-live state, obtain the live video stream data corresponding to the live data identifier according to the live data identifier carried in the second video generation request; Perform segmentation processing on the live video stream data according to a second preset duration to obtain a plurality of to-be-processed videos; wherein, the duration of each to-be-processed video is the second preset duration.

6. The video generation method according to claim 1, characterized in that, the determining the segment type of each video segment includes: Performing object recognition on each video segment to obtain the object information of each video segment; When the object information includes a target object and at least one of the preset category objects, determining that the video segment is a first segment type; or; when the object information only includes a plurality of preset category objects with the same preset attributes and does not include the target object, determining that the video segment is a second segment type; or; when the object information includes at least two preset category objects with different preset attributes and does not include the target object, determining that the video segment is a third segment type.

7. The video generation method according to claim 2, characterized in that, the determining a plurality of target video segments from the plurality of video segments includes: Determining a plurality of candidate video segments from the plurality of video segments according to the target segment type; Obtaining a plurality of feature information of each candidate video segment in the plurality of candidate video segments; Sorting the plurality of candidate video segments according to the plurality of feature information of each candidate video segment to obtain a plurality of sorted candidate video segments; Determining a plurality of target video segments that match the number of the plurality of points from the plurality of sorted candidate video segments.

8. The video generation method according to claim 7, characterized in that, the determining a plurality of target video segments that match the number of the plurality of points from the plurality of sorted candidate video segments includes: Determining the first N candidate video segments in the sorted plurality of candidate video segments as the plurality of target video segments; wherein, the size of N is determined according to the number of the plurality of points.

9. The video generation method according to claim 8, characterized in that, the plurality of target video segments are used to be fused with corresponding audio segments in sequence; the audio segments represent the audio segments between adjacent points in the plurality of points; the generating a target video according to the plurality of target video segments and the target audio includes: For each target video segment in the plurality of target video segments: If the duration of the target video segment is greater than the duration of the corresponding audio segment, crop the target video segment according to the duration of the audio segment or play it at a first speed; Or; If the duration of the target video segment is less than the duration of the corresponding audio segment, replay the target video segment according to the duration of the audio segment or play it at a second speed; wherein, the first speed is greater than the second speed.

10. The video generation method according to claim 7, wherein, determining a plurality of target video segments that match the number of the plurality of points from the sorted plurality of candidate video segments includes: regarding the first point among the plurality of points as the current point; determining, from the sorted plurality of candidate video segments, a target video segment that matches the duration of the audio segment between the current point and the next adjacent point; regarding the next adjacent point as the current point again to obtain a new plurality of target video segments.

11. The video generation method according to claim 7, wherein, the plurality of feature information includes a plurality of live broadcast attribute data; sorting the plurality of candidate video segments according to the plurality of feature information of each candidate video segment includes: sorting the plurality of candidate video segments according to a preset live broadcast attribute priority order and the plurality of live broadcast attribute data of each candidate video segment.

12. The video generation method according to claim 11, wherein, the plurality of live broadcast attribute data includes at least one of the peak number of viewers in the live broadcast room and the number of clicks on the object link.

13. The video generation method according to claim 2, wherein, obtaining the target audio includes: determining the target duration of the target video to be generated; determining, from the audio resource pool, a target audio that matches the target duration.

14. A video generation device, wherein, comprising: a first acquisition module configured to acquire a video to be processed; a first determination module configured to determine a plurality of video segments from the video to be processed; each video segment in the plurality of video segments represents an image sequence describing a preset category of object; a second determination module configured to determine the segment type of each video segment; the segment type is determined according to the object information in each video segment, and the object information includes the preset category of object and / or the number of target objects; a third determination module configured to determine a plurality of target video segments from the plurality of video segments; each target video segment in the plurality of target video segments has the same segment type; a generation module configured to generate a target video based on the plurality of target video segments; determining a plurality of video segments from the video to be processed includes: performing video content understanding processing on the video to be processed to obtain video content description information; the video content description information includes the plurality of video segments and the object description information corresponding to each video segment in the plurality of video segments, and the category of the object in the object description information includes a preset category; determining the segment type of each video segment includes: performing object recognition on each video segment to obtain the object information of each video segment; the granularity of the object recognition is finer than that of the video content understanding processing; determining the segment type of each video segment based on the preset category of object and / or the number of target objects.

15. The video generation device according to claim 14, wherein, the device further comprises: a second acquisition module configured to acquire target audio; the target audio includes multiple points; a third determination module configured to determine a plurality of target video segments from the plurality of video segments that match the number of the plurality of points; a generation module configured to generate the target video according to the plurality of target video segments and the target audio.

16. The video generation device according to claim 14 or 15, wherein, the first acquisition module is further configured to process the acquired live video stream data to obtain the video to be processed.

17. The video generation device according to claim 16, wherein, the duration of the video to be processed is a first preset duration; the first acquisition module includes: a first processing unit configured to, when receiving a first video generation request of a first account in a live state, use the moment when the first video generation request is received as the start moment; a second processing unit configured to start timing from the start moment, and when reaching the preset duration, use the live video stream data obtained in real time as the video to be processed; a third processing unit configured to use the current moment as the start moment again to obtain a new video to be processed.

18. The video generation device according to claim 16, wherein, the first acquisition module includes: a fourth processing unit configured to, when receiving a second video generation request of a first account in a non-live state, acquire the live video stream data corresponding to the live data identifier according to the live data identifier carried in the second video generation request; a fifth processing unit configured to perform segmentation processing on the live video stream data according to a second preset duration to obtain a plurality of the videos to be processed; wherein, the duration of each of the videos to be processed is the second preset duration.

19. The video generation device according to claim 14, wherein, the second determination module includes: an identification unit configured to perform object identification on each video segment to obtain object information of each video segment; a type determination unit configured to, when the object information includes a target object and at least one of the preset category objects, determine the video segment as a first segment type; or; when the object information only includes a plurality of preset category objects with the same preset attributes and does not include the target object, determine the video segment as a second segment type; or; when the object information includes at least two preset category objects with different preset attributes and does not include the target object, determine the video segment as a third segment type.

20. The video generation device according to claim 16, wherein, the third determination module includes: a first determination unit configured to determine a plurality of candidate video segments from the plurality of video segments according to the target segment type, An acquisition unit, configured to acquire multiple feature information of each candidate video segment among the multiple candidate video segments; A sorting unit, configured to sort the multiple candidate video segments according to the multiple feature information of each candidate video segment, and obtain the sorted multiple candidate video segments; A second determination unit, configured to determine multiple target video segments that match the number of the multiple points from the sorted multiple candidate video segments.

21. The video generation device according to claim 20, wherein, the second determination unit is further configured to determine the candidate video segments ranked in the top N positions among the sorted multiple candidate video segments as the multiple target video segments; wherein, the size of N is determined according to the number of the multiple points.

22. The video generation device according to claim 21, wherein, the multiple target video segments are used to be fused with corresponding audio segments in sequence; the audio segments represent the audio segments between adjacent points among the multiple points; the generation module is configured to perform, for each target video segment among the multiple target video segments: if the duration of the target video segment is greater than the duration of the corresponding audio segment, crop the target video segment according to the duration of the audio segment or play it at a first speed; or; if the duration of the target video segment is less than the duration of the corresponding audio segment, replay the target video segment according to the duration of the audio segment or play it at a second speed; wherein, the first speed is greater than the second speed.

23. The video generation device according to claim 20, wherein, the second determination unit is further configured to use the first point among the multiple points as the current point; determine a target video segment that matches the duration of the audio segment between the current point and the next adjacent point from the sorted multiple candidate video segments; use the next adjacent point as the current point again to obtain a new multiple target video segments.

24. The video generation device according to claim 20, wherein, the multiple feature information includes multiple live broadcast attribute data; the sorting unit is further configured to sort the multiple candidate video segments according to a preset live broadcast attribute priority order and the multiple live broadcast attribute data of each candidate video segment.

25. The video generation device according to claim 24, wherein, the multiple live broadcast attribute data includes at least one of the peak number of viewers in the live broadcast room and the number of clicks on the object link.

26. The video generation device according to claim 15, wherein, the first acquisition module is further configured to determine the target duration of the target video to be generated; determine a target audio that matches the target duration from the audio resource pool.

27. An electronic device, wherein, comprising: a processor; a memory for storing executable instructions of the processor; Wherein, the processor is configured to execute the instructions to implement the video generation method according to any one of claims 1 to 13.

28. A computer-readable storage medium, characterized in that, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the video generation method according to any one of claims 1 to 13.

29. A computer program product, characterized in that, the computer program product includes a computer program, the computer program is stored in a readable storage medium, and at least one processor of a computer device reads and executes the computer program, so that the computer device executes the video generation method according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Mixed shear video generation method and device, electronic equipment and computer readable storage medium

    CN111683209A

  • Video processing method and device

    CN111866585A