Video sub-screen node prediction method, device, terminal and storage medium

By performing feature extraction and sliding form model on video data combined with non-maximum suppression operations, the problem of low accuracy in prediction of video segmentation points is solved, and more efficient identification and prediction of split-screen nodes is achieved.

CN113987265BActive Publication Date: 2025-07-18特赞(上海)信息科技有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111244239.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-25
Publication Date
2025-07-18
Estimated Expiration
2041-10-25

AI Technical Summary

Technical Problem

In the prior art, the prediction accuracy of video segmentation points is low, making it difficult to effectively divide video content.

Method used

By extracting the video data features, obtaining picture sequences and audio sequences, combining sliding form models and non-maximum suppression operations, identifying and eliminating redundant data, improving the prediction accuracy of split-screen nodes.

Benefits of technology

The sliding form model's recognition effectiveness and prediction accuracy of split-screen nodes are improved, and the robustness of split-screen nodes is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113987265B_ABST
    Figure CN113987265B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device, terminal and storage medium for predicting split-screen nodes of a video. The method includes: extracting features from video data to obtain a picture sequence and an audio sequence; determining sample data based on the picture sequence and the audio sequence; extracting prediction sample data from the sample data and inputting the prediction sample data into a trained sliding window model to obtain an initial prediction sequence of split-screen nodes corresponding to the prediction sample data; performing a non-maximum suppression operation on the initial prediction sequence of split-screen nodes to eliminate the split-screen nodes corresponding to redundant data in the prediction sample data, so as to obtain a target prediction sequence of split-screen nodes corresponding to the target data. The present invention can improve the effectiveness of the sliding window model in identifying split-screen nodes and the accuracy of prediction, and further improve the robustness of predicting split-screen nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of video sub-screening, and more specifically, to a method, device, terminal, and storage medium for predicting sub-screen nodes of a video. Background Art

[0002] Video is a type of data with a temporal structure. More effective information can be extracted only by analyzing video content. However, dividing video content is a prerequisite task for analyzing video content, and how to effectively divide video content has become an urgent problem to be solved.

[0003] Currently, in the video analysis scenario, video is generally divided based on inter-frame pictures to determine the splitting points (i.e., sub-screen nodes) of the pictures.

[0004] However, using the above method to predict the splitting points has the problem of low accuracy. Summary of the Invention

[0005] The main purpose of this application is to provide a method, device, terminal, and storage medium for predicting sub-screen nodes of a video, so as to solve the problem of low accuracy in predicting splitting points in related technologies.

[0006] To achieve the above purpose, in a first aspect, this application provides a method for predicting sub-screen nodes of a video, including:

[0007] Performing feature extraction on video data to obtain a picture sequence and an audio sequence;

[0008] Determining sample data based on the picture sequence and the audio sequence;

[0009] Extracting prediction sample data from the sample data and inputting the prediction sample data into a trained sliding window model to obtain an initial prediction sequence of sub-screen nodes corresponding to the prediction sample data;

[0010] Performing non-maximum suppression operation on the initial prediction sequence of sub-screen nodes to eliminate sub-screen nodes corresponding to redundant data in the prediction sample data, and obtaining a target prediction sequence of sub-screen nodes corresponding to the target data.

[0011] In a possible implementation manner, performing feature extraction on video data to obtain a picture sequence and an audio sequence includes:

[0012] Respectively using a C3D recognition model and a VGGISH pre-trained model to perform 3D convolutional feature and audio feature extraction on video data at a preset interval to obtain a picture sequence and an audio sequence.

[0013] In a possible implementation manner, determining sample data based on the picture sequence and the audio sequence includes:

[0014] Splice the video frame sequence and the audio sequence to obtain the spliced video data;

[0015] Extract window data from the spliced video data by using a preset window at a preset step length to obtain window data;

[0016] Based on the window data and the corresponding tag value sequence of the window data, construct sample data.

[0017] In a possible implementation, the initial prediction sequence of the scene segmentation node includes a first initial prediction sequence for predicting whether a scene segmentation node is included and a second initial prediction sequence for predicting the position of the scene segmentation node;

[0018] Extract prediction sample data from the sample data, and input the prediction sample data into the trained sliding window model to obtain the initial prediction sequence of the scene segmentation node corresponding to the prediction sample data, including:

[0019] Randomly select a preset number of samples without marked scene segmentation node data in the sample data as prediction sample data;

[0020] Input the prediction sample data into the trained sliding window model. When the confidence level of the scene segmentation node in the preset loss function reaches the first preset confidence threshold, output the first initial prediction sequence and the second initial prediction sequence corresponding to the prediction sample data.

[0021] In a possible implementation, the target prediction sequence of the scene segmentation node includes: a first target prediction sequence for predicting whether a scene segmentation node is included and a second target prediction sequence for predicting the position of the scene segmentation node;

[0022] Perform non-maximum suppression operation on the initial prediction sequence of the scene segmentation node to eliminate the scene segmentation nodes corresponding to the redundant data in the prediction sample data, and obtain the target prediction sequence of the scene segmentation node corresponding to the target data, including:

[0023] Extract the first initial prediction sequence in the initial prediction sequence of the scene segmentation node, and sort the first initial prediction sequence in descending order of confidence level to obtain a sorting result;

[0024] Traverse the data in the sorting result in turn with a preset radius, and eliminate the data within the preset radius to obtain the first target prediction sequence for whether a scene segmentation node is included;

[0025] Based on the first target prediction sequence, determine the second target prediction sequence.

[0026] In a possible implementation, before extracting the prediction sample data from the sample data and inputting the prediction sample data into the trained sliding window model to obtain the initial prediction sequence of the scene segmentation nodes corresponding to the prediction sample data, it further includes:

[0027] Extract training sample data from the sample data;

[0028] Obtain an initial sliding window model, and use the training sample data to train the initial sliding window model to obtain the trained sliding window model.

[0029] In a possible implementation, extracting training sample data from the sample data includes:

[0030] Select a preset number of samples marked with scene segmentation node data in the sample data as the training sample data.

[0031] In a second aspect, an embodiment of the present invention provides a scene segmentation node prediction device for a video, including:

[0032] A feature extraction module, configured to extract features from video data to obtain a picture sequence and an audio sequence;

[0033] A sample determination module, configured to determine sample data based on the picture sequence and the audio sequence;

[0034] An initial sequence determination module, configured to extract prediction sample data from the sample data and input the prediction sample data into the trained sliding window model to obtain the initial prediction sequence of the scene segmentation nodes corresponding to the prediction sample data;

[0035] A target sequence determination module, configured to perform a non-maximum suppression operation on the initial prediction sequence of the scene segmentation nodes to remove the scene segmentation nodes corresponding to the redundant data in the prediction sample data, and obtain the target prediction sequence of the scene segmentation nodes corresponding to the target data.

[0036] In a third aspect, an embodiment of the present invention provides a terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of any one of the above-mentioned scene segmentation node prediction methods for a video are implemented.

[0037] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned scene segmentation node prediction methods for a video are implemented.

[0038] An embodiment of the present invention provides a method, apparatus, terminal, and storage medium for predicting split-screen nodes of a video, including: extracting features from video data to obtain a picture sequence and an audio sequence, then determining sample data based on the picture sequence and the audio sequence, extracting prediction sample data from the sample data, and inputting the prediction sample data into a trained sliding window model to obtain an initial prediction sequence of split-screen nodes corresponding to the prediction sample data. Finally, a non-maximum suppression operation is performed on the initial prediction sequence of split-screen nodes to remove the split-screen nodes corresponding to redundant data in the prediction sample data, and a target prediction sequence of split-screen nodes corresponding to the target data is obtained. By extracting the picture sequence and the audio sequence, the present invention can effectively identify split-screen nodes, and then set a loss function in the sliding window model to make the sliding window model more sensitive to split-screen nodes, improve the effectiveness of the sliding window model in identifying split-screen nodes and the accuracy of prediction, and further improve the robustness of predicting split-screen nodes. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] The drawings constituting a part of this application are used to provide a further understanding of this application, making other features, objectives, and advantages of this application more obvious. The schematic embodiments and descriptions thereof of this application are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0040] Figure 1 is a flowchart of the implementation of a method for predicting split-screen nodes of a video provided by an embodiment of the present invention;

[0041] Figure 2 is a schematic diagram of the principle of window extraction provided by an embodiment of the present invention;

[0042] Figure 3 is a schematic diagram of the structure of sample features provided by an embodiment of the present invention;

[0043] Figure 4 is a flowchart of the implementation of establishing an initial sliding window model provided by an embodiment of the present invention;

[0044] Figure 5 is a schematic diagram of the principle of removing redundant split-screen nodes provided by an embodiment of the present invention;

[0045] Figure 6 is a schematic diagram of the structure of an apparatus for a method for predicting split-screen nodes of a video provided by an embodiment of the present invention;

[0046] Figure 7 is a schematic diagram of a terminal provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part rather than all of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0048] In the description of the present invention and the claims as well as the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order different from those illustrated or described herein.

[0049] It should be understood that in various embodiments of the present invention, the magnitude of the serial numbers of the processes does not mean the order of execution, and the execution order of each process should be determined by its function and internal logic and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0050] It should be understood that in the present invention, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0051] It should be understood that in the present invention, "a plurality of" means two or more. "And / or" is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. "Including A, B, and C" and "including A, B, C" mean that all of A, B, and C are included. "Including A, B, or C" means including any one of A, B, and C. "Including A, B, and / or C" means including any one or any two or all three of A, B, and C.

[0052] It should be understood that in the present invention, "B corresponding to A", "B corresponding to A relatively", "A corresponding to B relatively", or "B corresponding to A relatively" means that B is associated with A, and B can be determined according to A. Determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information. The matching of A and B means that the similarity between A and B is greater than or equal to a preset threshold.

[0053] Depending on the context, as used herein, "if" may be interpreted as "when", "while", "in response to determining", or "in response to detecting".

[0054] The technical solution of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0055] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described below with reference to the accompanying drawings through specific embodiments.

[0056] In one embodiment, as Figure 1 shown, a method for predicting sub-screen nodes of a video is provided, including the following steps:

[0057] Step S101: Extract features from video data to obtain a sequence of frames and an audio sequence;

[0058] Step S102: Determine sample data based on the sequence of frames and the audio sequence;

[0059] Step S103: Extract predicted sample data from the sample data and input the predicted sample data into a trained sliding window model to obtain an initial prediction sequence of sub-screen nodes corresponding to the predicted sample data;

[0060] Step S104: Perform non-maximum suppression operation on the initial prediction sequence of sub-screen nodes to remove sub-screen nodes corresponding to redundant data in the predicted sample data, and obtain a target prediction sequence of sub-screen nodes corresponding to the target data.

[0061] Among them, the video includes long videos and short videos, and the types of long videos and short videos may include videos such as advertisements and movies, which are not specifically limited here. Taking the video data sourced from advertisement video data (with a duration between 30s and 2min) on a short video platform as an example, this solution will be described. Specifically, a sequence of frames and an audio sequence are extracted from the advertisement video data, and the sequence of frames and the audio sequence are spliced. Among them, the sequence of frames includes multiple frame data, and the audio sequence includes multiple audio data, and the multiple frame data and the multiple audio data correspond one by one. One frame data and one audio data are spliced to form a sample, and the sample data of the present invention is constituted by the splicing of the above-mentioned sequence of frames and audio sequence, and the sample data includes multiple samples. Then, predicted sample data is extracted from the sample data, and the predicted sample data is input into a trained sliding window model to obtain an initial prediction sequence of sub-screen nodes corresponding to the predicted sample data. Finally, a non-maximum suppression operation is performed on the initial prediction sequence of sub-screen nodes to remove sub-screen nodes corresponding to redundant data in the predicted sample data, and a target prediction sequence of sub-screen nodes corresponding to the target data is obtained.

[0062] An embodiment of the present invention provides a method for predicting split-screen nodes of a video, including: extracting features from video data to obtain a frame sequence and an audio sequence, then determining sample data based on the frame sequence and the audio sequence, extracting predicted sample data from the sample data, and inputting the predicted sample data into a trained sliding window model to obtain an initial prediction sequence of split-screen nodes corresponding to the predicted sample data. Finally, a non-maximum suppression operation is performed on the initial prediction sequence of split-screen nodes to remove the split-screen nodes corresponding to redundant data in the predicted sample data, and a target prediction sequence of split-screen nodes corresponding to the target data is obtained. By extracting the frame sequence and the audio sequence, the present invention can effectively identify split-screen nodes, and then set a loss function in the sliding window model to make the sliding window model more sensitive to split-screen nodes, improve the effectiveness of the sliding window model in identifying split-screen nodes and the accuracy of prediction, and further improve the robustness of predicting split-screen nodes.

[0063] In one embodiment, step S101 includes:

[0064] Step S201: Use a C3D recognition model and a VGGISH pre-trained model to extract 3D convolutional features and audio features from video data at a preset interval to obtain a frame sequence and an audio sequence.

[0065] Specifically, for the extraction of the frame sequence, the frame rate of the video needs to be converted first and uniformly set to 16fps, and then a C3D recognition model pre-trained based on the UCF101 dataset is used to extract 3D convolutional features from the video data. Among them, a certain extraction interval (i.e., the preset interval) needs to be set for video feature extraction. The extraction interval in the present invention is 0.2s, and one 3D convolutional feature is extracted per second (16 frame images), and its frame feature size is [1*1024]. The audio sequence mainly uses a VGGISH pre-trained model to extract audio features from the video data, and its extraction interval is the same as that of frame feature extraction, both being 0.2s. One audio feature is extracted per second, and the audio feature size is [1*128].

[0066] In one embodiment, step S102 includes:

[0067] Step S301: Concatenate the frame sequence and the audio sequence to obtain the concatenated video data;

[0068] Step S302: Use a preset window to extract window data from the concatenated video data at a preset step size;

[0069] Step S303: Based on the window data and the label value sequence corresponding to the window data, form sample data.

[0070] Specifically, by splicing the video sequence obtained in the previous embodiment with the audio sequence, multiple pieces of data with a feature size of [1*1152] can be obtained. The above data constitutes the spliced video data, and the spliced video data is presented in the form of feature segments, where each feature segment represents a piece of data with a feature size.

[0071] Combined with Figure 2 , the present invention uses a preset window form with a sliding step of 0.4 s (the interval between two feature segments, i.e., the preset step) to slide over the obtained feature segments, obtaining multiple window forms, and forming window form data from the multiple window forms. The length of the preset window form is set to the length of 5 feature segments, and its corresponding duration is 1.8 s (i.e., 4 intervals of 0.2 s plus the length of 1 feature segment).

[0072] In addition, the video data in the present invention has been manually labeled for each frame of the video before feature extraction, and each frame of the video carries a label value. The above method processes the video data, and each window form in the obtained window form data also carries a label value. Therefore, the entire window form data corresponds to a sequence of label values. Then, based on the window form data and the sequence of label values corresponding to the window form data, sample data is formed. It should be noted that the window form data is obtained by performing feature extraction on the video data, and it includes a video sequence and an audio sequence. In addition, the video data has also been marked with split-screen nodes before feature extraction, that is, the split-screen nodes in the video frames containing split-screen nodes are marked, and this standard is reflected in the window form data in the form of a split-screen node sequence and a split-screen node position sequence.

[0073] After determining the sample data through the above embodiment, it is also necessary to further extract the sample data to determine the training sample data and the prediction sample data, and train the initial sliding window form model based on the training sample data to determine the trained sliding window form model. Then, based on the prediction sample data and the trained sliding window form model, an initial prediction sequence of the split-screen nodes corresponding to the prediction sample data is obtained.

[0074] In one embodiment, the specific steps for further extracting the sample data to determine the training sample data and the prediction sample data are as follows: Select a preset number of samples marked with split-screen node data in the sample data as the training sample data (i.e., positive sample data), and then randomly select a preset number of samples not marked with split-screen node data in the sample data as the prediction sample data (i.e., negative sample data).

[0075] Specifically, combined with Figure 3 , each sample marked with split-screen node data includes five features: video features, audio features, label values ( Figure 3Not shown in the figure), whether it contains a scene division node and the position of the scene division node. When there are multiple samples, the above features consist of a video sequence, an audio sequence, a tag value sequence, a sequence indicating whether a scene division node is included, and a sequence of the positions of the scene division nodes to form a file, which is the training sample data. Each sample without marked scene division node data includes three features: a video feature, an audio feature, and a tag value. When there are multiple samples, the above features consist of a video sequence, an audio sequence, and a tag value sequence to form a file, which is the prediction sample data.

[0076] Based on the training sample data determined through the above embodiments, it is necessary to train the initial sliding window model based on the training sample data to determine the trained sliding window model, which specifically includes:

[0077] 1) Obtain the initial sliding window model, where Figure 4 In the initial sliding window model of the present invention, the Loss function in the initial sliding window model uses a custom function, which is specifically expressed as follows:

[0078]

[0079] where λ is a parameter for adjusting whether a scene division node can be detected. The larger the value of λ, the more sensitive it is to the appearance of a scene division node, and the higher the detection accuracy; m is the number of samples, y represents the true value, represents the predicted value corresponding to y, where y0 is the confidence level that the window contains a scene division node, the predicted confidence level that the window contains a scene division node; y1 is the offset, is the predicted offset. is a conditional judgment formula. When y0 = 1 and are satisfied, the result of this formula is 1, otherwise it is 0. threshold is the confidence level threshold for determining whether the output result of the initial sliding window model is a scene division node.

[0080] 2) Use the training sample data to train the initial sliding window model to obtain the trained sliding window model.

[0081] After the prediction sample data and the trained sliding window model determined through the above embodiments, step S103 is executed, where step S103 includes:

[0082] Step S401: Randomly select a preset number of samples without marked scene division node data from the sample data as the prediction sample data;

[0083] Step S402: Input the predicted sample data into the trained sliding window model. Wait until the confidence level of the segmentation node in the preset loss function reaches the first preset confidence threshold, and output the first initial prediction sequence and the second initial prediction sequence corresponding to the predicted sample data.

[0084] Among them, the initial prediction sequence of the segmentation node includes the first initial prediction sequence for predicting whether a segmentation node is included and the second initial prediction sequence for predicting the position of the segmentation node. Input the picture sequence, audio sequence, and label value sequence in the predicted sample into the trained sliding window model. When the confidence level of the segmentation node in the preset loss function reaches the first preset confidence threshold, output the first initial prediction sequence for predicting whether a segmentation node is included and the second initial prediction sequence for predicting the position of the segmentation node. Since the predicted sample data is randomly selected, some samples include segmentation nodes, but the relevant data of the segmentation nodes is not included in the predicted sample, and some do not include segmentation nodes. Therefore, the first initial prediction sequence for predicting whether a segmentation node is included in the output includes data with segmentation nodes and data without segmentation nodes. The data with segmentation nodes and the data without segmentation nodes are represented by different numbers or letters.

[0085] In one embodiment, step S104 includes:

[0086] Step S501: Extract the first initial prediction sequence in the initial prediction sequence of the segmentation node, and sort the first initial prediction sequence in descending order of confidence level to obtain a sorting result;

[0087] Step S502: Traverse the data in the sorting result in turn with a preset radius, and remove the data within the preset radius to obtain the first target prediction sequence;

[0088] Step S503: Determine the second target prediction sequence based on the first target prediction sequence.

[0089] Among them, the target prediction sequence of the segmentation node includes: the first target prediction sequence for predicting whether a segmentation node is included and the second target prediction sequence for predicting the position of the segmentation node. In order to remove the redundant segmentation nodes in the initial prediction sequence of the segmentation node, the present invention first sorts the first initial prediction sequence in the initial prediction sequence of the segmentation node according to the confidence level; then specifies an inhibition radius (i.e., the preset radius, the preset radius can be 1s or 2s, etc., which is not specifically limited here) to traverse the data in the sorting result (i.e., the prediction result) in turn, and removes the redundant segmentation nodes that appear within the inhibition radius. By Figure 5It can be known that the preset radius is set to 1s, and the redundant sub-screen nodes 0.6 and 0.7 in the prediction result are removed to obtain the NMS (Non-Maximum Suppression) processing result; finally, 0.8 and 0.9 constitute the first target prediction sequence, and then the second target prediction sequence is determined.

[0090] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not imply the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0091] The following is the device embodiment of the present invention. For the details not described in detail, reference can be made to the corresponding method embodiments above.

[0092] Figure 6 The structural schematic diagram of a sub-screen node prediction device for a video provided by an embodiment of the present invention is shown. For the sake of convenience of description, only the parts related to the embodiments of the present invention are shown. A sub-screen node prediction device for a video includes a feature extraction module 61, a sample determination module 62, an initial sequence determination module 63, and a target sequence determination module 64, specifically as follows:

[0093] The feature extraction module 61 is used to extract features from video data to obtain a picture sequence and an audio sequence;

[0094] The sample determination module 62 is used to determine sample data based on the picture sequence and the audio sequence;

[0095] The initial sequence determination module 63 is used to extract prediction sample data from the sample data and input the prediction sample data into the trained sliding window model to obtain an initial prediction sequence of the sub-screen nodes corresponding to the prediction sample data;

[0096] The target sequence determination module 64 is used to perform non-maximum suppression operation on the initial prediction sequence of the sub-screen nodes to remove the sub-screen nodes corresponding to the redundant data in the prediction sample data, and obtain a target prediction sequence of the sub-screen nodes corresponding to the target data.

[0097] In a possible implementation manner, the feature extraction module 61 includes:

[0098] The feature extraction sub-module is used to extract 3D convolution features and audio features from video data at a preset interval by using a C3D recognition model and a VGGISH pre-trained model respectively to obtain a picture sequence and an audio sequence.

[0099] In a possible implementation manner, the sample determination module 62 includes:

[0100] The splicing sub-module is used to splice the picture sequence and the audio sequence to obtain the spliced video data;

[0101] A window extraction sub-module, which is used to extract window data from the spliced video data by using a preset window at a preset step length to obtain window data;

[0102] A sample determination sub-module, which is used to construct sample data based on the window data and the corresponding tag value sequence of the window data.

[0103] In a possible implementation manner, the initial prediction sequence of the scene division node includes a first initial prediction sequence for predicting whether a scene division node is included and a second initial prediction sequence for predicting the position of the scene division node;

[0104] The initial sequence determination module 63 includes:

[0105] A sample selection sub-module, which is used to randomly select a preset number of samples without marked scene division node data in the sample data as prediction sample data;

[0106] An initial sequence determination sub-module, which is used to input the prediction sample data into the trained sliding window model, and when the confidence level of the scene division node in the preset loss function reaches the first preset confidence threshold, output the first initial prediction sequence and the second initial prediction sequence corresponding to the prediction sample data.

[0107] In a possible implementation manner, the target prediction sequence of the scene division node includes: a first target prediction sequence for predicting whether a scene division node is included and a second target prediction sequence for predicting the position of the scene division node;

[0108] The target sequence determination module 64 includes:

[0109] A sorting sub-module, which is used to extract the first initial prediction sequence in the initial prediction sequence of the scene division node, and sort the first initial prediction sequence in descending order of confidence level to obtain a sorting result;

[0110] A data elimination sub-module, which is used to sequentially traverse the data in the sorting result with a preset radius, and eliminate the data within the preset radius to obtain the first target prediction sequence;

[0111] A target sequence determination sub-module, which is used to determine the second target prediction sequence based on the first target prediction sequence.

[0112] In a possible implementation manner, before the initial sequence determination module 63, it further includes:

[0113] A training sample extraction module, which is used to extract training sample data from the sample data;

[0114] A model training module, configured to obtain an initial sliding window model and train the initial sliding window model using training sample data to obtain a trained sliding window model.

[0115] In a possible implementation, the training sample extraction module includes:

[0116] A training sample extraction sub-module, configured to select a preset number of samples marked with scene segmentation node data from the sample data as training sample data.

[0117] Figure 7 is a schematic diagram of a terminal provided by an embodiment of the present invention. As Figure 7 shown, the terminal 7 of this embodiment includes: a processor 70, a memory 71, and a computer program 72 stored in the memory 71 and executable on the processor 70. When the processor 70 executes the computer program 72, the steps in the embodiments of the above-mentioned scene segmentation node prediction method for each video are implemented, such as Figure 1 the steps 101 to 104 shown. Alternatively, when the processor 70 executes the computer program 72, the functions of each module / unit in the above-mentioned device embodiments are implemented, such as Figure 6 the functions of the modules / units 61 to 64 shown.

[0118] The present invention also provides a readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it is used to implement the methods provided by the above various implementation manners.

[0119] Among them, the readable storage medium may be a computer storage medium or a communication medium. The communication medium includes any medium that facilitates the transmission of a computer program from one place to another. The computer storage medium may be any available medium that can be accessed by a general-purpose or special-purpose computer. For example, the readable storage medium is coupled to the processor, so that the processor can read information from the readable storage medium and write information to the readable storage medium. Of course, the readable storage medium may also be a component of the processor. The processor and the readable storage medium may be located in an application specific integrated circuit (ASIC). In addition, the ASIC may be located in a user device. Of course, the processor and the readable storage medium may also exist as discrete components in a communication device. The readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0120] The present invention also provides a program product, which includes execution instructions stored in a readable storage medium. At least one processor of the device can read the execution instructions from the readable storage medium, and the execution of the execution instructions by the at least one processor causes the device to implement the methods provided by the above various embodiments.

[0121] In the embodiment of the above device, it should be understood that the processor may be a central processing unit (English: Central Processing Unit, abbreviated: CPU), and may also be other general-purpose processors, digital signal processors (English: Digital Signal Processor, abbreviated: DSP), application specific integrated circuits (English: Application Specific Integrated Circuit, abbreviated: ASIC), etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the present invention may be directly embodied as being completed by the execution of a hardware processor, or completed by a combination of hardware and software modules in the processor.

[0122] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included in the protection scope of the present invention.

Claims

1. A method for predicting split-screen nodes of a video, characterized in that Including: Performing feature extraction on video data to obtain a sequence of video frames and an audio sequence; Concatenating the sequence of video frames and the audio sequence to obtain concatenated video data; extracting window data from the concatenated video data using a preset window form with a preset step size; constructing sample data based on the window data and the corresponding sequence of label values; Randomly selecting a preset number of samples without marked scene segmentation node data from the sample data as prediction sample data; inputting the prediction sample data into a trained sliding window model, and when the confidence level of the scene segmentation node in a preset loss function reaches a first preset confidence threshold, outputting an initial prediction sequence corresponding to the prediction sample data, where the initial prediction sequence includes a first initial prediction sequence for predicting whether a scene segmentation node is included and a second initial prediction sequence for predicting the position of the scene segmentation node; Performing non-maximum suppression on the initial prediction sequence to remove the scene segmentation nodes corresponding to redundant data in the prediction sample data, and obtaining a target prediction sequence of the scene segmentation nodes of the target data.

2. The method for predicting split-screen nodes of a video according to claim 1, wherein The performing feature extraction on video data to obtain a sequence of video frames and an audio sequence includes: Respectively using a C3D recognition model and a VGGISH pre-trained model to perform 3D convolutional feature and audio feature extraction on the video data at a preset interval to obtain the sequence of video frames and the audio sequence.

3. The method for predicting split-screen nodes of a video according to claim 1, wherein, The target prediction sequence of the scene segmentation node includes a first target prediction sequence for predicting whether a scene segmentation node is included and a second target prediction sequence for predicting the position of the scene segmentation node; The performing non-maximum suppression on the initial prediction sequence of the scene segmentation node to remove the scene segmentation nodes corresponding to redundant data in the prediction sample data and obtaining a target prediction sequence of the scene segmentation nodes of the target data includes: Extracting the first initial prediction sequence from the initial prediction sequence of the scene segmentation node, and sorting the first initial prediction sequence in descending order of confidence level to obtain a sorting result; Successively traversing the data in the sorting result with a preset radius, and removing the data within the preset radius to obtain the first target prediction sequence; Determining the second target prediction sequence based on the first target prediction sequence.

4. The method for predicting split-screen nodes of a video according to claim 1 or 3, wherein Before extracting prediction sample data from the sample data and inputting the prediction sample data into a trained sliding window model to obtain an initial prediction sequence of the scene segmentation nodes corresponding to the prediction sample data, it further includes: Extracting training sample data from the sample data; Obtaining an initial sliding window model, and training the initial sliding window model using the training sample data to obtain the trained sliding window model.

5. The method for predicting split-screen nodes of a video according to claim 4, characterized in that, The extracting training sample data from the sample data includes: Selecting the preset number of samples marked with scene segmentation node data from the sample data as the training sample data.

6. A split-screen node prediction device for a video, characterized in that, Including: A feature extraction module for performing feature extraction on video data to obtain a sequence of video frames and an audio sequence; A sample determination module, configured to splice the picture sequence and the audio sequence to obtain spliced video data; extract window data from the spliced video data at a preset step size by using a preset window form; and constitute sample data based on the window data and the tag value sequence corresponding to the window data. An initial sequence determination module, configured to randomly select a preset number of samples without marked scene segmentation node data in the sample data as predicted sample data; input the predicted sample data into a trained sliding window model, and when the confidence level of the scene segmentation node in a preset loss function reaches a first preset confidence threshold, output an initial predicted sequence corresponding to the predicted sample data, where the initial predicted sequence includes a first initial predicted sequence for predicting whether a scene segmentation node is included and a second initial predicted sequence for predicting the position of the scene segmentation node. A target sequence determination module, configured to perform a non-maximum suppression operation on the initial predicted sequence to remove the scene segmentation nodes corresponding to redundant data in the predicted sample data, and obtain a target predicted sequence of the scene segmentation nodes corresponding to the target data.

7. A terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the scene segmentation node prediction method for a video according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, the steps of the scene segmentation node prediction method for a video according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Poster CTR prediction method and device based on aesthetic features

    CN112767038A

  • Training data acquisition method and device

    CN113450774A