A video data processing method, apparatus, device and medium
By using automated video data processing methods, videos are stored in matching video groups based on the needs of the video model and the characteristics of the dataset. This solves the problems of low efficiency and insufficient accuracy in existing technologies, and improves the efficiency and accuracy of video model training.
Patent Information
- Application Number
- CN202411713393.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-27
- Publication Date
- 2026-01-30
- Estimated Expiration
- 2044-11-27
AI Technical Summary
Existing video data processing solutions have low processing efficiency, high labor and time costs, and difficulty in guaranteeing accuracy.
By acquiring the video dataset and video group requirements corresponding to the target video model, and based on the video requirements and the length, resolution, and aspect ratio of the video dataset, the system automatically stores the videos in the matching video group storage location, and retrieves the expected number of videos from the storage location to provide to the processor system.
It enables the rapid and accurate storage of video datasets into corresponding video groups, reducing manpower and time costs and improving the efficiency and accuracy of the video model training process.
Smart Images

Figure CN119520846B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer technology, and in particular to a video data processing method, apparatus, device, and medium. Background Technology
[0002] With the development of video generation technology, more and more companies are starting to use video models for video generation. A video model can generate videos with a specified length, resolution, and aspect ratio based on an input video generation request, and then output the generated video. Video models are typically trained on multiple groups of videos with lengths, resolutions, and aspect ratios within different numerical ranges. Each video group needs to contain multiple videos with lengths, resolutions, and aspect ratios within the specified numerical ranges. During the training process, a large number of pre-collected videos need to be processed to obtain multiple video groups with different lengths, resolutions, and aspect ratios for training the video model. Then, the videos from each video group are provided to the processor system used to train the video model.
[0003] In related technologies, a common video data processing scheme involves technicians processing a large number of pre-collected videos to obtain multiple video groups with varying lengths, resolutions, and aspect ratios for training the video model. The videos from each group are then provided to a processor system used for training the model. This method requires manual intervention to obtain these video groups, resulting in low processing efficiency, high labor and time costs, and difficulty in guaranteeing accuracy. Summary of the Invention
[0004] This invention provides a video data processing method, apparatus, device, and medium to solve the problems of low processing efficiency, high labor and time costs, and difficulty in guaranteeing accuracy in related video data processing schemes.
[0005] According to one aspect of the present invention, a video data processing method is provided, comprising:
[0006] Obtain the video dataset corresponding to the target video model;
[0007] Obtain video requirement information for each video group corresponding to the target video model; wherein, the video requirement information includes identification information, expected length, expected resolution, expected aspect ratio, probability of video length reduction, probability of video resolution reduction, storage location information, and expected number of videos;
[0008] Based on the video requirement information of each video group, the length, resolution and aspect ratio of each video in the video dataset, determine the video group that matches each video in the video dataset, and store each video in the video dataset in the storage location of the matched video group.
[0009] For each video group, the expected number of videos for the video group are retrieved from the storage location of the video group, and the retrieved videos of the expected number of videos for the video group are provided to the processor system corresponding to the target video model.
[0010] According to another aspect of the present invention, a video data processing apparatus is provided, comprising:
[0011] The dataset acquisition module is used to acquire the video dataset corresponding to the target video model;
[0012] The information acquisition module is used to acquire video requirement information for each video group corresponding to the target video model; wherein, the video requirement information includes identification information, expected length, expected resolution, expected aspect ratio, probability of video length reduction, probability of video resolution reduction, storage location information, and expected number of videos;
[0013] The video matching module is used to determine the video group that matches each video in the video dataset based on the video requirement information of each video group, the length, resolution and aspect ratio of each video in the video dataset, and to store each video in the video dataset in the storage location of the matched video group.
[0014] The video providing module is used to obtain the expected number of videos of each video group from the storage location of the video group, and provide the obtained expected number of videos of the video group to the processor system corresponding to the target video model.
[0015] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0016] At least one processor;
[0017] and a memory communicatively connected to the at least one processor;
[0018] The memory stores a computer program that is executed by the at least one processor, which enables the at least one processor to perform the video data processing method according to any embodiment of the present invention.
[0019] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the video data processing method according to any embodiment of the present invention.
[0020] The technical solution of this invention obtains a video dataset corresponding to a target video model and video requirement information for each video group corresponding to the target video model. The video requirement information includes identification information, expected length, expected resolution, expected aspect ratio, probability of video length reduction, probability of video resolution reduction, storage location information, and expected number of videos. Then, based on the video requirement information of each video group and the length, resolution, and aspect ratio of each video in the video dataset, a video group matching each video in the video dataset is determined, and each video in the video dataset is stored in the storage location of the matching video group. Finally, for each video group, the expected number of videos for the video group are retrieved from the storage location of the video group, and the retrieved videos of the expected number of videos for the video group are provided to the processor system corresponding to the target video model. This solves the problem of low processing efficiency in related video data processing schemes. The previous methods, which involved high labor and time costs and difficulty in guaranteeing accuracy, can automatically and accurately store each video in the video dataset into a video group to which it belongs, based on the expected length, expected resolution, expected aspect ratio, probability of video length reduction, probability of video resolution reduction, and the length, resolution, and aspect ratio of the videos in the video dataset. This results in multiple video groups corresponding to the video model, which are multiple video groups with different lengths, resolutions, and aspect ratios required during the training of the video model. The videos in each video group are then provided to the processor system used to train the video model. This method has lower labor and time costs, improves the accuracy and efficiency of video data processing, and can quickly and accurately meet the data requirements of the video model training process, thereby improving the training efficiency of the video model.
[0021] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart of a video data processing method provided in Embodiment 1 of the present invention.
[0024] Figure 2 This is a flowchart of a video data processing method provided in Embodiment 2 of the present invention.
[0025] Figure 3 This is a schematic diagram of the structure of a video data processing device provided in Embodiment 3 of the present invention.
[0026] Figure 4 A schematic diagram of the structure of an electronic device for implementing the video data processing method of this invention. Detailed Implementation
[0027] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0028] It should be noted that the terms "target," "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising," "including," and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0029] Example 1
[0030] Figure 1 This is a flowchart illustrating a video data processing method according to Embodiment 1 of the present invention. This embodiment is applicable to situations where a large number of pre-collected videos are processed during the training of a video model. The method can be executed by a video data processing device, which can be implemented in hardware and / or software and can be configured in an electronic device. The electronic device can be an electronic device installed in an enterprise for processing pre-collected videos. For example... Figure 1 As shown, the method includes:
[0031] Step 101: Obtain the video dataset corresponding to the target video model.
[0032] Optionally, the video model can be a model trained on a deep learning model using multiple groups of videos with varying lengths, resolutions, and aspect ratios as training samples. The video model can generate a video with a specified length, resolution, and aspect ratio based on an input video generation request, and output the generated video. The input to the video model is the video generation request, and the output is the video corresponding to the request. The video generation request can be text instructing the video model to generate a video with a specified length, resolution, and aspect ratio based on specified text. The video corresponding to the video generation request can be the video generated by the video model based on the request.
[0033] Optionally, the length of a video can refer to the number of video images it contains. For example, if a video contains 1 frame, its length is 1. If a video contains 8 frames, its length is 8. If a video contains 16 frames, its length is 16.
[0034] Optionally, the video resolution can refer to the resolution of the video images contained in the video. For example, if the resolution of the video images contained in the video is 480P, then the video resolution is 480P. If the resolution of the video images contained in the video is 360P, then the video resolution is 360P. If the resolution of the video images contained in the video is 720P, then the video resolution is 720P.
[0035] Optionally, the aspect ratio of a video can refer to the ratio between the width and height of the video images contained in the video. For example, if the ratio of the width to the height of the video images contained in the video is 16:9, then the aspect ratio of the video is 16:9. If the ratio of the width to the height of the video images contained in the video is 4:3, then the aspect ratio of the video is 4:3. If the ratio of the width to the height of the video images contained in the video is 1:1, then the aspect ratio of the video is 1:1. If the ratio of the width to the height of the video images contained in the video is 3:4, then the aspect ratio of the video is 3:4. If the ratio of the width to the height of the video images contained in the video is 9:16, then the aspect ratio of the video is 9:16.
[0036] Optionally, the target video model can be a video model that needs to be trained. The video dataset corresponding to the target video model consists of multiple pre-collected videos used to train the target video model. For example, the video dataset corresponding to the target video model consists of 2000 pre-collected videos used to train the target video model.
[0037] Optionally, acquiring the video dataset corresponding to the target video model includes: acquiring the video dataset corresponding to the target video model sent by the target user. The target user can be a technician responsible for managing the video model. The target user can send the video dataset corresponding to the target video model to an electronic device via a terminal device. The electronic device acquires the video dataset corresponding to the target video model sent by the target user.
[0038] Step 102: Obtain video requirement information for each video group corresponding to the target video model.
[0039] The video demand information includes identification information, expected length, expected resolution, expected aspect ratio, probability of video length reduction, probability of video resolution reduction, storage location information, and expected number of videos.
[0040] Optionally, each video group corresponding to the target video model is a group of videos with different lengths, resolutions, and aspect ratios that need to be used during the training of the target video model. Each video group needs to consist of multiple videos, and the length, resolution, and aspect ratio of each video in the video group need to be within a specified range. The video requirement information for a video group can be information related to the video group and the videos that can be attributed to the video group. The identification information for a video group can be a string used to uniquely identify the video group. The expected length of a video group can be the length that each video in the video group needs to be greater than. The expected resolution of a video group can be the resolution that each video in the video group needs to be greater than. The expected aspect ratio of a video group can be the aspect ratio that each video in the video group needs to have. The probability of reducing the length of a video group can be the probability of reducing the length of a video in the video group during model training to reduce computation. The probability of reducing the resolution of a video group can be the probability of reducing the resolution of a video in the video group during model training to reduce computation. The storage location information for a video group can be a string used to identify the storage location of the video group. The storage location of a video group can be a pre-set memory or storage file used to store the videos in the video group. The expected number of videos in a video group can be the number of videos in the video group that need to be used during model training.
[0041] Optionally, obtaining video demand information for each video group corresponding to the target video model includes: obtaining video demand information for each video group corresponding to the target video model sent by the target user. The target user can send the video demand information for each video group corresponding to the target video model to an electronic device via a terminal device. The electronic device obtains the video demand information for each video group corresponding to the target video model sent by the target user.
[0042] Step 103: Based on the video requirement information of each video group, the length, resolution, and aspect ratio of each video in the video dataset, determine the video group that matches each video in the video dataset, and store each video in the video dataset in the storage location of the matched video group.
[0043] Optionally, for each video in the video dataset, the video group that matches the video is the video group to which the video can be assigned corresponding to the target video model. The video is then stored in the storage location of the matched video group, thus storing the video in the video group to which it can be assigned corresponding to the target video model.
[0044] Optionally, based on the video requirement information of each video group, the length, resolution, and aspect ratio of each video in the video dataset, determine the video group that matches each video in the video dataset, and store each video in the video dataset in the storage location of the matched video group. This includes performing the following operations for each video in the video dataset: determining each video group as a candidate video group corresponding to the video; determining each basic candidate video group whose expected length is less than the length of the video and whose expected resolution is less than the resolution of the video; determining the maximum length among the expected lengths of each basic candidate video group and the maximum resolution among the expected resolutions of each basic candidate video group; and obtaining one basic candidate video group whose length is the maximum length and whose resolution is the maximum resolution as... A target candidate video group corresponding to the video is selected; a random number between 0 and 1 is generated, and it is determined whether the random number is greater than or equal to the probability of video length reduction of the target candidate video group; if the random number is greater than or equal to the probability of video length reduction of the target candidate video group, a new random number between 0 and 1 is generated, and it is determined whether the new random number is greater than or equal to the probability of video resolution reduction of the target candidate video group; if the new random number is greater than or equal to the probability of video resolution reduction of the target candidate video group, it is determined whether the aspect ratio of the video is equal to the expected aspect ratio of the target candidate video group; if the aspect ratio of the video is equal to the expected aspect ratio of the target candidate video group, the target candidate video group is determined as the video group matching the video, and the video is stored in the storage location of the matching video group according to the storage location information of the target candidate video group.
[0045] Optionally, for each video in the video dataset, the electronic device can detect the video to determine its length, resolution, and aspect ratio.
[0046] Optionally, each candidate video group corresponding to the video is a video group to which the video may belong. The basic candidate video groups are those candidate video groups whose expected length is less than the length of the video and whose expected resolution is less than the resolution of the video. Each video group is determined as a candidate video group corresponding to the video. The video is detected to determine its length, resolution, and aspect ratio. For each candidate video group, it is detected whether the expected length of the candidate video group is less than the length of the video and whether its expected resolution is less than the resolution of the video. The candidate video groups whose expected length is less than the length of the video and whose expected resolution is less than the resolution of the video are determined as basic candidate video groups, thus obtaining each basic candidate video group within the existing candidate video groups.
[0047] Optionally, the maximum length is the maximum expected length among all basic candidate video groups. The maximum resolution is the maximum expected resolution among all basic candidate video groups. The electronic device can detect the expected length and expected resolution of each basic candidate video group to determine the maximum expected length and the maximum expected resolution among all basic candidate video groups.
[0048] Optionally, after determining the maximum expected length and the maximum expected resolution of each basic candidate video group, a basic candidate video group whose length and resolution are both the maximum expected length and the maximum expected resolution are selected as the target candidate video group corresponding to the video. Then, a random number between 0 and 1 is generated using a random number generation component, and it is determined whether the random number is greater than or equal to the probability of video length reduction in the target candidate video group. The random number generation component can be a component in an electronic device used to generate random numbers between 0 and 1. If the random number is greater than or equal to the probability of video length reduction in the target candidate video group, a new random number between 0 and 1 is generated, and it is determined whether the new random number is greater than or equal to the probability of video resolution reduction in the target candidate video group. If the new random number is greater than or equal to the probability of video resolution reduction in the target candidate video group, it is determined whether the aspect ratio of the video is equal to the expected aspect ratio of the target candidate video group. If the aspect ratio of the video is equal to the expected aspect ratio of the target candidate video group, then the target candidate video group is determined as the video group that matches the video. Based on the storage location information of the target candidate video group, the video is stored in the storage location of the matched video group. The electronic device can determine the storage location of the target candidate video group based on the storage location information of the target candidate video group, and store the video in the storage location of the target candidate video group, thereby storing the video in the storage location of the matched video group.
[0049] Optionally, after determining whether the random number is greater than or equal to the video length reduction probability of the target candidate video group, the method further includes: if the random number is less than the video length reduction probability of the target candidate video group, and there is a basic candidate video group in each basic candidate video group whose expected length is less than the expected length of the target candidate video group and whose expected resolution is equal to the expected resolution of the target candidate video group, then after removing the target candidate video group from each basic candidate video group, the method returns to perform the operation of determining the maximum length among the expected lengths of each basic candidate video group and the maximum resolution among the expected resolutions of each basic candidate video group.
[0050] Optionally, if the random number is less than the probability of video length reduction in the target candidate video group, then it is checked whether there exists a basic candidate video group in each basic candidate video group whose expected length is less than the expected length of the target candidate video group and whose expected resolution is equal to the expected resolution of the target candidate video group. If there exists a basic candidate video group in each basic candidate video group whose expected length is less than the expected length of the target candidate video group and whose expected resolution is equal to the expected resolution of the target candidate video group, then the target candidate video group is removed from each basic candidate video group. After removing the target candidate video group from each basic candidate video group, the process returns to determining the maximum length among the expected lengths of each basic candidate video group and the maximum resolution among the expected resolutions of each basic candidate video group.
[0051] Optionally, it also includes: if the random number is less than the probability of video length reduction of the target candidate video group, then after removing the target candidate video group from each basic candidate video group, return to perform the operation of determining the maximum length among the expected lengths of each basic candidate video group and the maximum resolution among the expected resolutions of each basic candidate video group.
[0052] Optionally, after determining whether the new random number is greater than or equal to the video down-resolution probability of the target candidate video group, the method further includes: if the new random number is less than the video down-resolution probability of the target candidate video group, and there is a basic candidate video group in each basic candidate video group whose expected resolution is less than the expected resolution of the target candidate video group and whose expected length is equal to the expected length of the target candidate video group, then after removing the target candidate video group from each basic candidate video group, the method returns to perform the operation of determining the maximum length among the expected lengths of each basic candidate video group and the maximum resolution among the expected resolutions of each basic candidate video group.
[0053] Optionally, if the new random number is less than the video downresolution probability of the target candidate video group, then it is detected whether there exists a basic candidate video group in each basic candidate video group whose expected resolution is less than the expected resolution of the target candidate video group and whose expected length is equal to the expected length of the target candidate video group. If there exists a basic candidate video group in each basic candidate video group whose expected resolution is less than the expected resolution of the target candidate video group and whose expected length is equal to the expected length of the target candidate video group, then the target candidate video group is removed from each basic candidate video group. After removing the target candidate video group from each basic candidate video group, the process returns to determining the maximum length among the expected lengths of each basic candidate video group and the maximum resolution among the expected resolutions of each basic candidate video group.
[0054] Optionally, it also includes: if the new random number is less than the video down-resolution probability of the target candidate video group, then after removing the target candidate video group from each basic candidate video group, return to perform the operation of determining the maximum length among the expected lengths of each basic candidate video group and the maximum resolution among the expected resolutions of each basic candidate video group.
[0055] Optionally, if the aspect ratio of the video is not equal to the expected aspect ratio of the target candidate video group, the target candidate video group is removed from each basic candidate video group. After removing the target candidate video group from each basic candidate video group, the operation of determining the maximum length among the expected lengths of each basic candidate video group and the maximum resolution among the expected resolutions of each basic candidate video group is returned.
[0056] Therefore, based on the generated random numbers between 0 and 1, the expected length, expected resolution, expected aspect ratio, probability of video length reduction, probability of video resolution reduction, and the length, resolution, and aspect ratio of each video group corresponding to the target video model, each video in the video dataset can be randomly stored into a video group to which it can belong. This results in the various video groups corresponding to the target video model, i.e., multiple video groups with different lengths, resolutions, and aspect ratios that need to be used during the training of the target video model.
[0057] Step 104: For each video group, retrieve the expected number of videos for the video group from the storage location of the video group, and provide the retrieved expected number of videos for the video group to the processor system corresponding to the target video model.
[0058] Optionally, the processor system corresponding to the target video model can be a processor system used to train the deep learning model to obtain the target video model by using videos from each video group corresponding to the target video model as training samples. The processor system corresponding to the target video model can be composed of multiple graphics processing units (GPUs).
[0059] Optionally, for each video group, the expected number of videos for the video group are obtained from the storage location of the video group, and the obtained videos of the expected number of videos for the video group are provided to the processor system corresponding to the target video model. This includes: for each video group, obtaining the expected number of videos for the video group from the storage location of the video group, randomly arranging the obtained videos to form a video sequence, shuffling the video sequence to obtain a shuffled video sequence of the video group, and providing the shuffled video sequence of the video group to the processor system corresponding to the target video model.
[0060] Optionally, the shuffled video sequence of the video group can be a shuffled video sequence obtained by shuffling a video sequence formed by the expected number of videos in the video group. The expected number of videos for the video group are retrieved from the storage location of the video group, and the retrieved videos are randomly arranged to form a video sequence. Then, the video sequence is shuffled according to a preset shuffling rule to obtain the shuffled video sequence, thus obtaining the shuffled video sequence of the video group. Finally, the shuffled video sequence of the video group is provided to the processor system corresponding to the target video model. For example, the expected number of videos in the video group is 20. 20 videos are retrieved from the storage location of the video group, and the retrieved 20 videos are randomly arranged to form a video sequence. Then, the video sequence is shuffled according to a preset shuffling rule to obtain the shuffled video sequence, thus obtaining the shuffled video sequence of the video group. Finally, the shuffled video sequence of the video group is provided to the processor system corresponding to the target video model.
[0061] Optionally, the preset shuffling rule can be a pre-set rule for shuffling a video sequence to obtain a shuffled video sequence. The electronic device can shuffle the video sequence according to the preset shuffling rule to obtain a shuffled video sequence.
[0062] Optionally, providing the shuffled video sequence of the video group to the processor system corresponding to the target video model includes: sending the shuffled video sequence of the video group to the processor system corresponding to the target video model.
[0063] Optionally, in a specific instance, the processor system corresponding to the target video model includes a first GPU, a second GPU, and a third GPU. The first GPU, second GPU, and third GPU are three different GPUs. The first GPU is used to train the deep learning model using videos from a group containing videos with a length greater than 1, a resolution greater than 480P, and an aspect ratio of 16:9 as training samples. The second GPU is used to train the deep learning model using videos from a group containing videos with a length greater than 8, a resolution greater than 360P, and an aspect ratio of 1:1 as training samples. The third GPU is used to train the deep learning model using videos from a group containing videos with a length greater than 16, a resolution greater than 720P, and an aspect ratio of 3:4 as training samples.
[0064] The technical solution of this invention obtains a video dataset corresponding to a target video model and video requirement information for each video group corresponding to the target video model. The video requirement information includes identification information, expected length, expected resolution, expected aspect ratio, probability of video length reduction, probability of video resolution reduction, storage location information, and expected number of videos. Then, based on the video requirement information of each video group and the length, resolution, and aspect ratio of each video in the video dataset, a video group matching each video in the video dataset is determined, and each video in the video dataset is stored in the storage location of the matching video group. Finally, for each video group, the expected number of videos for the video group are retrieved from the storage location of the video group, and the retrieved videos of the expected number of videos for the video group are provided to the processor system corresponding to the target video model. This solves the problem of low processing efficiency in related video data processing schemes. The previous methods, which involved high labor and time costs and difficulty in guaranteeing accuracy, can automatically and accurately store each video in the video dataset into a video group to which it belongs, based on the expected length, expected resolution, expected aspect ratio, probability of video length reduction, probability of video resolution reduction, and the length, resolution, and aspect ratio of the videos in the video dataset. This results in multiple video groups corresponding to the video model, which are multiple video groups with different lengths, resolutions, and aspect ratios required during the training of the video model. The videos in each video group are then provided to the processor system used to train the video model. This method has lower labor and time costs, improves the accuracy and efficiency of video data processing, and can quickly and accurately meet the data requirements of the video model training process, thereby improving the training efficiency of the video model.
[0065] Example 2
[0066] Figure 2This is a flowchart illustrating a video data processing method according to Embodiment 2 of the present invention. Embodiments of the present invention can be combined with various optional solutions from one or more of the above embodiments. For example... Figure 2 As shown, the method includes:
[0067] Step 201: Obtain the video dataset corresponding to the target video model.
[0068] Step 202: Obtain video requirement information for each video group corresponding to the target video model.
[0069] The video demand information includes identification information, expected length, expected resolution, expected aspect ratio, probability of video length reduction, probability of video resolution reduction, storage location information, and expected number of videos.
[0070] Step 203: Based on the video requirement information of each video group, the length, resolution, and aspect ratio of each video in the video dataset, determine the video group that matches each video in the video dataset, and store each video in the video dataset in the storage location of the matched video group.
[0071] Step 204: For each video group, retrieve the expected number of videos from the storage location of the video group, randomly arrange the retrieved videos to form a video sequence, and shuffle the video sequence to obtain a shuffled video sequence of the video group. Provide the shuffled video sequence of the video group to the processor system corresponding to the target video model.
[0072] The technical solution of this invention can automatically and accurately store each video in a video dataset into a video group to which it can belong, based on the expected length, expected resolution, expected aspect ratio, probability of video length reduction, probability of video resolution reduction, and the length, resolution, and aspect ratio of the videos in the video dataset, according to the video requirement information of each video group corresponding to the video model. This results in multiple video groups corresponding to the video model, i.e., multiple video groups with different lengths, resolutions, and aspect ratios that need to be used in the training process of the video model. Based on the expected number of videos in each video group, a shuffled video sequence for each video group is obtained. The shuffled video sequence of the video group is then provided to the processor system corresponding to the video model. Thus, based on the video requirements of each video group, the videos in each video group are provided to the processor system used to train the video model. This reduces labor and time costs, improves the accuracy and efficiency of the video data processing process, and can quickly and accurately meet the data requirements of the video model training process, thereby improving the training efficiency of the video model.
[0073] Example 3
[0074] Figure 3 This is a schematic diagram of a video data processing device according to Embodiment 3 of the present invention. The device can be configured in an electronic device. Figure 3 As shown, the device includes: a dataset acquisition module 301, an information acquisition module 302, a video matching module 303, and a video providing module 304.
[0075] The system includes a dataset acquisition module 301, used to acquire a video dataset corresponding to the target video model; an information acquisition module 302, used to acquire video requirement information for each video group corresponding to the target video model, wherein the video requirement information includes identification information, expected length, expected resolution, expected aspect ratio, probability of video length reduction, probability of video resolution reduction, storage location information, and expected number of videos; a video matching module 303, used to determine the video group matching each video in the video dataset based on the video requirement information of each video group, the length, resolution, and aspect ratio of each video in the video dataset, and store each video in the video dataset in the storage location of the matched video group; and a video providing module 304, used to acquire the expected number of videos of each video group from the storage location of the video group, and provide the acquired expected number of videos of the video group to the processor system corresponding to the target video model.
[0076] The technical solution of this invention obtains a video dataset corresponding to a target video model and video requirement information for each video group corresponding to the target video model. The video requirement information includes identification information, expected length, expected resolution, expected aspect ratio, probability of video length reduction, probability of video resolution reduction, storage location information, and expected number of videos. Then, based on the video requirement information of each video group and the length, resolution, and aspect ratio of each video in the video dataset, a video group matching each video in the video dataset is determined, and each video in the video dataset is stored in the storage location of the matching video group. Finally, for each video group, the expected number of videos for the video group are retrieved from the storage location of the video group, and the retrieved videos of the expected number of videos for the video group are provided to the processor system corresponding to the target video model. This solves the problem of low processing efficiency in related video data processing schemes. The previous methods, which involved high labor and time costs and difficulty in guaranteeing accuracy, can automatically and accurately store each video in the video dataset into a video group to which it belongs, based on the expected length, expected resolution, expected aspect ratio, probability of video length reduction, probability of video resolution reduction, and the length, resolution, and aspect ratio of the videos in the video dataset. This results in multiple video groups corresponding to the video model, which are multiple video groups with different lengths, resolutions, and aspect ratios required during the training of the video model. The videos in each video group are then provided to the processor system used to train the video model. This method has lower labor and time costs, improves the accuracy and efficiency of video data processing, and can quickly and accurately meet the data requirements of the video model training process, thereby improving the training efficiency of the video model.
[0077] In an optional embodiment of the present invention, the video matching module 303 is specifically configured to: perform the following operations for each video in the video dataset: determine each video group as a candidate video group corresponding to the video; determine each basic candidate video group whose expected length is less than the length of the video and whose expected resolution is less than the resolution of the video; determine the maximum length among the expected lengths of each basic candidate video group and the maximum resolution among the expected resolutions of each basic candidate video group; obtain a basic candidate video group whose length is the maximum length and whose resolution is the maximum resolution among the basic candidate video groups as the target candidate video group corresponding to the video; generate a random number between 0 and 1, and determine the target candidate video group. The system checks whether the random number is greater than or equal to the probability of video length reduction in the target candidate video group; if the random number is greater than or equal to the probability of video length reduction in the target candidate video group, a new random number between 0 and 1 is generated, and it is determined whether the new random number is greater than or equal to the probability of video resolution reduction in the target candidate video group; if the new random number is greater than or equal to the probability of video resolution reduction in the target candidate video group, it is determined whether the aspect ratio of the video is equal to the expected aspect ratio of the target candidate video group; if the aspect ratio of the video is equal to the expected aspect ratio of the target candidate video group, the target candidate video group is determined as the video group matching the video, and the video is stored in the storage location of the matching video group according to the storage location information of the target candidate video group.
[0078] In an optional embodiment of the present invention, the video matching module 303 is further configured to: if the random number is less than the video length reduction probability of the target candidate video group, and there is a basic candidate video group in each basic candidate video group whose expected length is less than the expected length of the target candidate video group and whose expected resolution is equal to the expected resolution of the target candidate video group, then after removing the target candidate video group from each basic candidate video group, return to perform the operation of determining the maximum length among the expected lengths of each basic candidate video group and the maximum resolution among the expected resolutions of each basic candidate video group.
[0079] In an optional embodiment of the present invention, the video matching module 303 is further configured to: if the new random number is less than the video down-resolution probability of the target candidate video group, and there is a basic candidate video group in each basic candidate video group whose expected resolution is less than the expected resolution of the target candidate video group and whose expected length is equal to the expected length of the target candidate video group, then after removing the target candidate video group from each basic candidate video group, return to perform the operation of determining the maximum length among the expected lengths of each basic candidate video group and the maximum resolution among the expected resolutions of each basic candidate video group.
[0080] In an optional embodiment of the present invention, the video providing module 304 is specifically configured to: for each video group, obtain the expected number of videos of the video group from the storage location of the video group, randomly arrange the obtained videos to form a video sequence, shuffle the video sequence to obtain a shuffled video sequence of the video group, and provide the shuffled video sequence of the video group to the processor system corresponding to the target video model.
[0081] In an optional embodiment of the present invention, the dataset acquisition module 301 is specifically used to: acquire the video dataset sent by the target user that corresponds to the target video model.
[0082] In an optional embodiment of the present invention, the information acquisition module 302 is specifically used to: acquire video demand information of each video group corresponding to the target video model sent by the target user.
[0083] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0084] The video data processing device described above can execute the video data processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the video data processing method.
[0085] Example 4
[0086] Figure 4 A schematic diagram of the structure of an electronic device 10 that can be used to implement the video data processing method of embodiments of the present invention is shown. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the invention described and / or claimed herein.
[0087] like Figure 4As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executed by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or a computer program constructed from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0088] Multiple components in electronic device 10 are connected to input / output (I / O) interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of monitors, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0089] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as video data processing methods.
[0090] In some embodiments, the video data processing method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via read-only memory (ROM) 12 and / or communication unit 19. When the computer program is built into random access memory (RAM) 13 and executed by processor 11, one or more steps of the video data processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the video data processing method by any other suitable means (e.g., by means of firmware).
[0091] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0092] Computer programs for implementing the video data processing method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The computer programs can be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0093] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0094] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0095] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0096] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0097] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0098] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.
Claims
1. A method of processing video data, the method comprising: The method comprises: acquiring a video dataset corresponding to a target video model; acquiring video requirement information of each video group corresponding to the target video model; wherein the video requirement information comprises identification information, expected length, expected resolution, expected aspect ratio, video length reduction probability, video resolution reduction probability, storage location information, and expected number of videos, the expected length of a video group is a length that each video in the video group needs to be greater than, the expected resolution of a video group is a resolution that each video in the video group needs to be greater than, the expected aspect ratio of a video group is an aspect ratio that each video in the video group needs to have, the video length reduction probability of a video group is a probability that the length of a video in the video group is reduced in a model training process in order to reduce the amount of calculation, the video resolution reduction probability of a video group is a probability that the resolution of a video in the video group is reduced in a model training process in order to reduce the amount of calculation, and the expected number of videos of a video group is a number of videos in the video group that need to be used in a model training process; determining a video group that matches each video in the video dataset according to the video requirement information of each video group, the length, resolution, and aspect ratio of each video in the video dataset, and storing each video in the video dataset into the storage location of the matched video group; wherein the following operations are performed for each video in the video dataset: determining each video group as each candidate video group corresponding to the video, and determining each basic candidate video group in each candidate video group whose expected length is less than the length of the video and whose expected resolution is less than the resolution of the video; determining the maximum value of the length in the expected length of each basic candidate video group and the maximum value of the resolution in the expected resolution of each basic candidate video group; acquiring a basic candidate video group whose length is the maximum value of the length and whose resolution is the maximum value of the resolution in each basic candidate video group as a target candidate video group corresponding to the video; generating a random number between 0 and 1, and determining whether the random number is greater than or equal to the video length reduction probability of the target candidate video group; if the random number is greater than or equal to the video length reduction probability of the target candidate video group, generating a new random number between 0 and 1, and determining whether the new random number is greater than or equal to the video resolution reduction probability of the target candidate video group; if the new random number is greater than or equal to the video resolution reduction probability of the target candidate video group, determining whether the aspect ratio of the video is equal to the expected aspect ratio of the target candidate video group; if the aspect ratio of the video is equal to the expected aspect ratio of the target candidate video group, determining the target candidate video group as the video group that matches the video, and storing the video into the storage location of the matched video group according to the storage location information of the target candidate video group; for each video group, acquiring the expected number of videos of the video group from the storage location of the video group, and providing the acquired expected number of videos of the video group to a processor system corresponding to the target video model.
2. The video data processing method of claim 1, wherein, After judging whether the random number is greater than or equal to the video length probability of the target candidate video group, further comprising: If the random number is less than the video length probability of the target candidate video group, and there is a basic candidate video group in each basic candidate video group whose expected length is less than the expected length of the target candidate video group and whose expected resolution is equal to the expected resolution of the target candidate video group, after removing the target candidate video group from each basic candidate video group, the operation of determining the maximum length in the expected length of each basic candidate video group and the maximum resolution in the expected resolution of each basic candidate video group is returned to be executed.
3. The method of Claim 1, wherein, After judging whether the new random number is greater than or equal to the video resolution probability of the target candidate video group, further comprising: If the new random number is less than the video resolution probability of the target candidate video group, and there is a basic candidate video group in each basic candidate video group whose expected resolution is less than the expected resolution of the target candidate video group and whose expected length is equal to the expected length of the target candidate video group, after removing the target candidate video group from each basic candidate video group, the operation of determining the maximum length in the expected length of each basic candidate video group and the maximum resolution in the expected resolution of each basic candidate video group is returned to be executed.
4. The method of Claim 1, wherein, For each video group, the expected number of videos of the video group is obtained from the storage location of the video group, and the obtained expected number of videos of the video group is provided to the processor system corresponding to the target video model, comprising: For each video group, the expected number of videos of the video group is obtained from the storage location of the video group, and the obtained expected number of videos of the video group is provided to the processor system corresponding to the target video model, comprising:
5. The method of Claim 1, wherein, Obtaining a video data set corresponding to a target video model, comprising: Obtaining a video data set corresponding to a target video model sent by a target user.
6. The method of Claim 1, wherein, Obtaining video demand information of each video group corresponding to the target video model, comprising: Obtaining video demand information of each video group corresponding to the target video model sent by a target user.
7. A video data processing apparatus, comprising: Comprising: A data set acquisition module for obtaining a video data set corresponding to a target video model; A data set acquisition module for obtaining a video data set corresponding to a target video model; An information obtaining module is configured to obtain video requirement information of each video group corresponding to the target video model, wherein the video requirement information comprises identification information, expected length, expected resolution, expected aspect ratio, video length reduction probability, video resolution reduction probability, storage location information, and expected video quantity, the expected length of a video group is a length that each video in the video group needs to be greater than, the expected resolution of the video group is a resolution that each video in the video group needs to be greater than, the expected aspect ratio of the video group is an aspect ratio that each video in the video group needs to have, the video length reduction probability of the video group is a probability that the length of the video in the video group is reduced in a model training process in order to reduce calculation amount, the video resolution reduction probability of the video group is a probability that the resolution of the video in the video group is reduced in the model training process in order to reduce calculation amount, and the expected video quantity of the video group is a quantity of videos in the video group that need to be used in the model training process; A video matching module is configured to determine a video group matched with each video in the video data set according to the video requirement information of each video group, the length, resolution, and aspect ratio of each video in the video data set, and store each video in the video data set into the storage location of the matched video group, wherein the following operations are performed for each video in the video data set: each video group is determined as each candidate video group corresponding to the video, and each basic candidate video group in which the expected length is less than the length of the video and the expected resolution is less than the resolution of the video is determined; a maximum value of the length in the expected length of each basic candidate video group and a maximum value of the resolution in the expected resolution of each basic candidate video group are determined; a basic candidate video group in which the length is the maximum value of the length and the resolution is the maximum value of the resolution is obtained from each basic candidate video group as a target candidate video group corresponding to the video; a random number between 0 and 1 is generated, and it is determined whether the random number is greater than or equal to the video length reduction probability of the target candidate video group; if the random number is greater than or equal to the video length reduction probability of the target candidate video group, a new random number between 0 and 1 is generated, and it is determined whether the new random number is greater than or equal to the video resolution reduction probability of the target candidate video group; if the new random number is greater than or equal to the video resolution reduction probability of the target candidate video group, it is determined whether the aspect ratio of the video is equal to the expected aspect ratio of the target candidate video group; if the aspect ratio of the video is equal to the expected aspect ratio of the target candidate video group, the target candidate video group is determined as the video group matched with the video, and the video is stored into the storage location of the matched video group according to the storage location information of the target candidate video group; A video providing module is configured to obtain, for each video group, the expected video quantity of videos of the video group from the storage location of the video group, and provide the obtained expected video quantity of videos of the video group to a processor system corresponding to the target video model.
8. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected in communication with the at least one processor; wherein the memory stores a computer program executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the video data processing method of any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for causing the processor to implement the video data processing method of any one of claims 1-6 when executed.
Citation Information
Patent Citations
System and method of digital copyright control applied to multi-media file transmission
CN103458273A
Method and apparatus for assigning data sets to virtual volumes in a mass store
US4310883A