Video processing method and device, electronic equipment and computer readable storage medium
By encoding and decoding the video frames and encoding and decoding the visual language model, and combining the self-attention model to predict the highlights, the problem of unknown impact of images on highlights in the prior art is solved, and the accuracy of the highlights prediction of the video is improved.
Patent Information
- Application Number
- CN202510250630.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-05-30
AI Technical Summary
Existing video highlight prediction technology cannot effectively reflect the impact of images in video on highlight prediction, resulting in low prediction accuracy.
A video processing method is proposed, by extracting frames on the target video, encoding and decoding each image frame using a preset visual language model's vision encoder and language decoder to obtain image feature sequences and language highlight feature sequences. Then, the highlights are predicted by the self-attention model to determine the highlights image position of the video.
It improves the accuracy of video highlight prediction, can more effectively capture image highlight information in video, and provides more accurate video content and services.
Smart Images

Figure CN120071222A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology and is applied to the field of fintech. In particular, it relates to a video processing method and apparatus, an electronic device, and a computer-readable storage medium. Background Art
[0002] Video highlight prediction in video processing technology refers to using algorithms to predict the content points or critical moments that viewers will be interested in or pay attention to when watching a video. In the field of fintech, video highlight prediction can help video platforms, content creators, marketers, etc. better understand the interests of viewers, optimize video content and promotion strategies, and improve the viewing rate and user engagement of videos. For example, in financial scenarios such as digital marketing, customer service, and employee training, video highlight prediction technology can help enterprises better understand the needs of users and employees, provide more personalized and accurate video content and services, thereby enhancing the competitiveness of enterprises and user satisfaction. In medical scenarios such as psychotherapy and preoperative training, video highlight prediction technology can help doctors or nurses better understand the treatment needs and concerns of patients, provide more personalized and accurate video content and services, thereby improving patient cooperation and enhancing the success rate of treatment.
[0003] Currently, there are already some algorithms or models for video processing. For example, in one algorithm, the speech in the video is first converted into text to obtain the text content, and then the highlight video segments are predicted based on the text content. Another example is that in another algorithm, the video is segmented, and then the visual features of the video segments are directly extracted, and then the excitement level of the video segments is predicted based on the visual features.
[0004] The disadvantage of the related technology is that neither the text content in the video nor the visual features of the video segments can effectively reflect the impact of the images in the video on highlight prediction, resulting in low accuracy of video highlight prediction. Summary of the Invention
[0005] The main purpose of the embodiments of this application is to propose a video processing method and apparatus, an electronic device, and a computer-readable storage medium, which can perform video highlight prediction based on the images in the video and improve the accuracy of highlight prediction.
[0006] To achieve the above object, a first aspect of the embodiments of this application proposes a video processing method, and the method includes:
[0007] Obtain a target video;
[0008] Extract frames from the target video to obtain at least two image frames;
[0009] Perform visual encoding on each of the image frames through the visual encoder of a preset vision-language model to obtain an image feature sequence;
[0010] Perform language decoding on the image feature sequence of each said image frame through the language decoder of the said visual language model to obtain the language highlight feature sequence of each said image frame;
[0011] Perform highlight prediction on the language highlight feature sequences of the at least two image frames and a preset initial highlight prediction task description feature vector through a preset first self-attention model to obtain a first highlight prediction task description feature vector;
[0012] Determine the highlight image position of the target video according to the first highlight prediction task description feature vector.
[0013] Optionally, the determining the highlight image position of the target video according to the first highlight prediction task description feature vector includes:
[0014] Perform image highlight evaluation on the image frame through a preset image highlight evaluation model to obtain a highlight evaluation score;
[0015] Screen the image frames according to the highlight evaluation score to obtain at least one reference image frame;
[0016] Perform highlight prediction on the language highlight feature sequences of the at least one reference image frame and a preset video highlight prediction task description through a preset second self-attention model to obtain a second highlight prediction task description feature vector;
[0017] Screen out a first highlight start timestamp from the timestamps of the image frames according to the first highlight prediction task description feature vector;
[0018] Screen out a second highlight start timestamp from the timestamps of the reference image frames according to the second highlight prediction task description feature vector;
[0019] Determine a highlight start timestamp according to the first highlight start timestamp and the second highlight start timestamp, and determine the highlight start timestamp as the highlight image position.
[0020] Optionally, the determining the highlight start timestamp according to the first highlight start timestamp and the second highlight start timestamp includes:
[0021] Perform difference calculation on the second highlight start timestamp and each of the first highlight start timestamps to obtain a timestamp difference;
[0022] Sort the at least two first highlight start timestamps according to the timestamp difference to obtain a sorting result;
[0023] Determine the start timestamp of the first highlight with the smallest timestamp difference according to the sorting result, and determine the start timestamp of the first highlight as the start timestamp of the highlight.
[0024] Optionally, the language decoder of the visual language model performs language decoding on the image feature sequence of each image frame to obtain the language highlight feature sequence of each image frame, including:
[0025] Obtain the text information associated with the target video to obtain the video highlight text information;
[0026] Obtain the text information associated with the image frame to obtain the image text information of the image frame;
[0027] Perform highlight evaluation on the image text information of each image frame according to the video highlight text information to obtain the image text highlight score of each image frame;
[0028] Screen the image text information of each image frame according to the image text highlight score to obtain the selected image text information;
[0029] Perform language decoding on the selected image text information and the image feature sequence of each image frame through the language decoder to obtain the language highlight feature sequence of each image frame.
[0030] Optionally, the performing highlight evaluation on the image text information of each image frame according to the video highlight text information to obtain the image text highlight score includes:
[0031] Calculate the similarity between the video highlight text information and the image text information of the image frame to obtain the text similarity;
[0032] If the text similarity is less than the first similarity threshold and greater than the second similarity threshold, use the preset score value as the image text highlight score;
[0033] If the text similarity is greater than or equal to the first similarity threshold, perform score mapping on the text similarity according to the preset first mapping function to obtain the first mapping score, and determine the first mapping score as the image text highlight score; wherein, the first mapping function is an increasing function;
[0034] If the text similarity is less than or equal to the second similarity threshold, perform score mapping on the text similarity according to the preset second mapping function to obtain the second mapping score, and determine the second mapping score as the image text highlight score; wherein, the second mapping function is a decreasing function;
[0035] Wherein, the first mapping score and the second mapping score are greater than the preset score value.
[0036] Optionally, after visually encoding each of the image frames through the visual encoder of a preset visual language model to obtain an image feature sequence, the method further includes:
[0037] Obtaining text information associated with the target video to obtain video highlight text information;
[0038] Generating a relevance based on the image feature sequence and the video highlight text information to obtain an image highlight relevance; the image highlight relevance indicates the degree of relevance between the image feature sequence and the video highlight text information;
[0039] Adding the image highlight relevance to the image feature sequence to update the image feature sequence.
[0040] Optionally, the extracting at least two image frames from the target video includes:
[0041] Obtaining a fixed interval duration;
[0042] Randomly extracting one frame from the candidate image frames of the target video to obtain a first image frame;
[0043] Calculating at least one extraction timestamp based on the timestamp of the first image frame and the fixed interval duration;
[0044] Extracting one frame from the candidate image frames of the target video according to the extraction timestamp to obtain at least one second image frame;
[0045] Combining the first image frame and the at least one second image frame to obtain the at least two image frames.
[0046] To achieve the above object, a second aspect of the embodiments of the present application provides a video processing device, the device includes:
[0047] A video acquisition module, configured to acquire a target video;
[0048] An image extraction module, configured to extract at least two image frames from the target video;
[0049] A visual encoding module, configured to visually encode each of the image frames through the visual encoder of a preset visual language model to obtain an image feature sequence;
[0050] A language decoding module, configured to perform language decoding on the image feature sequence of each of the image frames through the language decoder of the visual language model to obtain a language highlight feature sequence of each of the image frames;
[0051] A highlight prediction module, configured to perform highlight prediction on the language highlight feature sequence of the at least two image frames and a preset initial highlight prediction task description feature vector through a preset first self-attention model, so as to obtain a first highlight prediction task description feature vector;
[0052] A position determination module, configured to determine the highlight image position of the target video according to the first highlight prediction task description feature vector.
[0053] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, where the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the video processing method described in the first aspect above is implemented.
[0054] To achieve the above object, a fourth aspect of the embodiments of the present application provides a storage medium, where the storage medium is a computer-readable storage medium, the storage medium stores a computer program, and when the computer program is executed by a processor, the video processing method described in the first aspect above is implemented.
[0055] The video processing method, video processing device, electronic device, and computer-readable storage medium provided by the present application do not directly perform highlight prediction on the entire video or video segment using a single model. Instead, the target video is first frame-extracted to obtain at least two image frames, so that each image frame can be processed separately in a targeted manner. Further, each image frame is encoded and decoded through the visual encoder and language decoder of the vision-language model to obtain the language highlight feature sequence of each image frame, which can fully extract the feature information related to highlights within the image frame. Further, a learnable initial highlight prediction task description vector is introduced, and then the language highlight feature sequence of each image frame and the initial highlight prediction task description feature vector are used for highlight prediction through the first self-attention model to obtain a first highlight prediction task description feature vector; in this way, the first highlight prediction task description feature vector contains the learned position information related to highlights, and the accuracy of this feature vector is relatively high. Finally, according to the first highlight prediction task description feature vector, the highlight image position of the target video is determined to achieve video highlight prediction. In summary, the present application can achieve video highlight prediction and improve the accuracy of highlight prediction. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 is a flowchart of the video processing method provided by the embodiments of the present application;
[0057] Figure 2 is Figure 1 a flowchart of step 102 in
[0058] Figure 3 It is a flowchart of a video processing method provided by another embodiment of the present application;
[0059] Figure 4 It is Figure 1 a flowchart of step 104 in
[0060] Figure 5 It is Figure 1 another flowchart of step 106 in
[0061] Figure 6 It is Figure 5 a flowchart of step 506 in
[0062] Figure 7 It is a block diagram of the module structure of a video processing device provided by an embodiment of the present application;
[0063] Figure 8 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed Embodiments
[0064] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0065] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the specification, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0066] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0067] First, several nouns involved in the present application are analyzed:
[0068] Artificial Intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence; Artificial intelligence is a branch of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. It also uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results in terms of theories, methods, technologies, and application systems.
[0069] Natural Language Processing (NLP): NLP uses computers to process, understand, and apply human languages (such as Chinese, English, etc.). NLP is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, and is often referred to as computational linguistics. Natural language processing includes syntactic analysis, semantic analysis, discourse understanding, etc. Natural language processing is commonly used in technical fields such as machine translation, recognition of handwritten and printed characters, speech recognition and text-to-speech conversion, information image processing, information extraction and filtering, text classification and clustering, public opinion analysis, and opinion mining. It involves data mining related to language processing, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research, and linguistic research related to language computing, etc.
[0070] Currently, video highlight prediction algorithms have the potential to improve the user experience and video recommendation effects, but in practical applications, some challenges and drawbacks still need to be overcome to improve prediction accuracy and user satisfaction.
[0071] Based on this, the embodiments of this application propose a video processing method, a video processing device, an electronic device, and a computer-readable storage medium, which can perform video highlight prediction based on the images in the video and improve the accuracy of highlight prediction.
[0072] The video processing method provided by the embodiments of this application can be applied to terminals and server sides, and can also be software running on the server side. The server side can be configured as an independent physical server, or as a server cluster or distributed system composed of multiple physical servers, or as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the video processing method, etc., but is not limited to the above forms.
[0073] This application can be used in numerous general-purpose or special-purpose computer system environments or configurations. For example: server computers, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0074] Embodiments of this application provide a video processing method, a video processing device, an electronic device, and a computer-readable storage medium, which will be specifically described through the following embodiments. First, the video processing method in the embodiments of this application will be described.
[0075] It should be noted that in each specific implementation manner of this application, when it comes to performing relevant processing based on data related to the user's identity or characteristics, such as the user's voice data, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards.
[0076] Referring to Figure 1 , Figure 1 is an optional flowchart of the video processing method provided by the embodiments of this application, which may include but is not limited to steps 101 to 106.
[0077] Step 101, obtain a target video;
[0078] Step 102, extract frames from the target video to obtain at least two image frames;
[0079] Step 103, perform visual encoding on each image frame through the visual encoder of a preset vision-language model to obtain an image feature sequence;
[0080] Step 104, perform language decoding on the image feature sequence of each image frame through the language decoder of the vision-language model to obtain the language key-point feature sequence of each image frame;
[0081] Step 105, perform key-point prediction on the language key-point feature sequences of at least two image frames and a preset initial key-point prediction task description feature vector through a preset first self-attention model to obtain a first key-point prediction task description feature vector;
[0082] Step 106: Determine the position of the highlight image of the target video according to the first highlight prediction task description feature vector.
[0083] In the embodiments of the present application, steps 101 to 106 do not directly perform highlight prediction on the entire video or video segment using a single model. Instead, the target video is first frame-extracted to obtain at least two image frames, so that each image frame can be processed separately. Further, each image frame is encoded and decoded through the visual encoder and language decoder of the vision-language model to obtain the language highlight feature sequence of each image frame, so as to fully extract the feature information related to highlights within the image frame. Further, a learnable initial highlight prediction task description vector is introduced, and then the first self-attention model is used to perform highlight prediction on the language highlight feature sequence of each image frame and the initial highlight prediction task description feature vector to obtain the first highlight prediction task description feature vector. In this way, the first highlight prediction task description feature vector contains the learned position information related to highlights, and the accuracy of this feature vector is relatively high. Finally, according to the first highlight prediction task description feature vector, the position of the highlight image of the target video is determined to achieve video highlight prediction. In summary, the present application can achieve video highlight prediction and improve the accuracy of highlight prediction.
[0084] When the embodiments of the present application are applied to the digital marketing business in the financial scenario, the digital marketing video can be used as the target video, and then the target video is frame-extracted to obtain marketing image frames. Subsequently, highlight prediction is performed on the marketing image frames through the visual encoder, language decoder, and first attention model (corresponding to steps 103 to 105 above) to obtain the marketing highlight prediction feature vector, which can indicate the image where the marketing highlight is located. Finally, according to the marketing highlight prediction feature vector, the position of the marketing highlight image of the digital marketing video is determined. For example, the marketing image frame corresponding to the position of the marketing highlight image contains content such as the original price of the marketing product, the discount strength, and the after-sales service introduction. In this way, based on the position of the marketing highlight image, the user's needs regarding digital marketing can be better understood, and more personalized and accurate digital marketing video content and services can be provided, thereby enhancing the competitiveness of the enterprise and user satisfaction.
[0085] When the embodiments of the present application are applied to the customer service business in the financial scenario, the customer service video can be used as the target video. When the embodiments of the present application are applied to the employee training business in the financial scenario, the employee training video can be used as the target video. After obtaining the target video, the subsequent processing process is similar to that of the above digital marketing business, and the beneficial effects achieved are also similar, which will not be elaborated here.
[0086] When the embodiments of the present application are applied to the preoperative training service in the medical scenario, the preoperative training video can be used as the target video, and then the target video is frame-extracted to obtain preoperative training image frames. Subsequently, the visual encoder, language decoder, and the first attention model are used to perform key-point prediction on the preoperative training image frames (corresponding to steps 103 to 105 above), obtaining a preoperative training key-point prediction feature vector, which can indicate the image where the preoperative training key points are located. Finally, based on the preoperative training key-point prediction feature vector, the position of the training key-point image of the preoperative training video is determined. For example, the preoperative training image frames corresponding to the position of the training key-point image include content such as the operation duration, operation site, and matters that patients should pay attention to during the operation. In this way, based on the position of the preoperative training key-point image, the needs of patients for preoperative training can be better understood, and more personalized and accurate preoperative training video content and services can be provided, thereby enhancing the patient's cooperation during the operation and improving the user experience.
[0087] In step 101 of some embodiments, a target video is obtained. The target video refers to a video that provides specific content for a specified object. For example, in the digital marketing service in the financial scenario, the target video refers to a video that provides marketing content for users, that is, a digital marketing video. The marketing content may include product name, product price, discount rate, after-sales service, product sales volume, product target population, etc. Another example is that in the preoperative training service in the medical scenario, the target video refers to a video that provides preoperative training content for patients, that is, a preoperative training video. The preoperative training content may include operation duration, operation name, the part of the patient to be operated on (such as the liver), preoperative precautions (such as fasting), intraoperative precautions (such as staying awake), etc.
[0088] In one example, for a predetermined service (such as digital marketing service, preoperative training service), a target video is generated using a video production tool (such as a plugin, a small program, etc.). In another example, a video generation model is used to generate a target video from business text content. The business text content refers to text content associated with the predetermined service, such as the aforementioned marketing content and preoperative training content, etc.
[0089] In step 102 of some embodiments, the target video is frame-extracted to obtain at least two image frames. The target video can be frame-extracted according to a fixed number of frames at intervals or a fixed time interval.
[0090] In one embodiment, referring to Figure 2 , step 102 may include:
[0091] Step 201, obtaining a fixed time interval;
[0092] Step 202, randomly extracting one frame from the candidate image frames of the target video to obtain a first image frame;
[0093] Step 203: Calculate at least one extraction timestamp according to the timestamp of the first image frame and the fixed interval duration.
[0094] Step 204: Extract one frame from the candidate image frames of the target video according to the extraction timestamp to obtain at least one second image frame.
[0095] Step 205: Merge the first image frame and at least one second image frame to obtain at least two image frames.
[0096] In step 201, the fixed interval duration means that within a specific time period of the target video, the time interval between image frames remains constant or consistent.
[0097] In one embodiment, step 201 may include: obtaining the duration of the target video to get the video duration; evenly dividing the video duration to get the segment duration; and determining the segment duration as the fixed interval duration. In this way, the fixed interval duration can be dynamically determined according to the target video, with relatively high flexibility.
[0098] In one embodiment, the fixed interval duration can also be preset according to experience, but the fixed interval duration needs to be less than the duration of the target video. For example, if the duration of the target video is 120 seconds, the fixed interval duration can be 2 seconds, 5 seconds, 10 seconds, etc.
[0099] In step 202, randomly extract one frame from the candidate image frames of the target video. It may extract the first candidate image frame, the second candidate image frame, the third candidate image frame, etc. of the target video, and determine the extracted candidate image frame as the first image frame.
[0100] In step 203, the duration between the extraction timestamp and the timestamp of the first image frame is equal to N * the fixed interval duration, where N is a positive integer. For example, if the timestamp of the first image frame is 00:00:01 and the fixed interval duration is 2 seconds, then the obtained extraction timestamps may include: 00:00:03, 00:00:05, 00:00:07, 00:00:09, etc.
[0101] In step 204, according to the extraction timestamp, the candidate image frame with this extraction timestamp can be extracted from the target video to obtain the second image frame. For example, the second image frame may include: the candidate image frame with the timestamp of 00:00:03, the candidate image frame with the timestamp of 00:00:05, the candidate image frame with the timestamp of 00:00:07, the candidate image frame with the timestamp of 00:00:09, etc.
[0102] In step 205, the union of the first image frame and at least one second image frame is taken to obtain at least two image frames. For example, the union of the first image frame with a timestamp of 00:00:01 and the second image frames with timestamps of 00:00:03, 00:00:05, 00:00:07, and 00:00:09 respectively is taken to obtain 5 image frames.
[0103] The benefits of the embodiments of the above steps 201 to 205 are that, when extracting frames, by virtue of the variability of random extraction, the extraction flexibility is improved, and by virtue of the fixed interval duration, the extraction stability is enhanced. In this way, it can well adapt to the complex and changeable video content in financial scenarios and medical scenarios, and has a high universality.
[0104] In step 103 of some embodiments, each image frame is visually encoded by a visual encoder of a preset vision-language model to obtain an image feature sequence.
[0105] The vision-language model is widely defined as a multi-modal model that can learn from images and texts, and can accept image and text inputs and generate text outputs. The vision-language model can be a model based on the Qwen-VL architecture. In this embodiment, mainly the vision encoder of the vision-language model is used to visually encode the image frames to obtain an image feature sequence. The image feature sequence includes at least two image features, and the at least two image features jointly indicate various information carried on an image frame. The vision encoder can also dynamically process image frames of different resolutions.
[0106] In one embodiment, referring to Figure 3 , after step 103, the video processing method may further include:
[0107] Step 301, obtaining text information associated with the target video to obtain video highlight text information;
[0108] Step 302, generating a relevance degree based on the image feature sequence and the video highlight text information to obtain an image highlight relevance degree; the image highlight relevance degree indicates the degree of relevance between the image feature sequence and the video highlight text information;
[0109] Step 303, adding the image highlight relevance degree to the image feature sequence to update the image feature sequence.
[0110] In step 301, the text information associated with the target video may include a video content summary, a video applicable service (such as a digital marketing service or a preoperative training service), etc. At least one of the video content summary and the video applicable service may be determined as the video highlight text information, and the video highlight text information is used to evaluate whether the text information of the image frame has highlights or the degree of having highlights.
[0111] In step 302, a relevance can be determined for the image feature sequence and the video highlight text information through a predetermined relevance function or a predetermined relevance prediction model. The higher the relevance, the higher the degree of correlation between the image feature sequence and the video highlight text information. The relevance function can be a cosine similarity function. The relevance prediction model can be obtained by fine-tuning a single-task prediction model.
[0112] In step 303, since the image feature sequence is a vector, the image highlight relevance can be added as a vector element to the image feature sequence to update the image feature sequence.
[0113] The benefits of the embodiments of the above steps 301 to 303 are that the information richness of the image feature sequence is improved by means of the relevance, and the relevance can help the subsequent language decoder to be personalized to adapt to image frames with different relevance, further improving the accuracy of highlight prediction.
[0114] In step 104 of some embodiments, the image feature sequence of each image frame is decoded linguistically through the language decoder of the vision-language model to obtain the language highlight feature sequence of each image frame. The language decoder can be a large language model (LLM). In addition to inputting the image feature sequence into the language decoder for linguistic decoding, text information such as the labels of the target video and the title of the target video can also be input into the language decoder for linguistic decoding to enhance the attention to the video highlights during the decoding process, thereby improving the decoding accuracy.
[0115] In one embodiment, referring to Figure 4 , step 104 may include:
[0116] Step 401, obtaining the text information associated with the target video to obtain the video highlight text information;
[0117] Step 402, obtaining the text information associated with the image frame to obtain the image text information of the image frame;
[0118] Step 403, evaluating the highlights of the image text information of each image frame according to the video highlight text information to obtain the image text highlight score of each image frame;
[0119] Step 404, screening the image text information of each image frame according to the image text highlight score to obtain the selected image text information;
[0120] Step 405, decoding the selected image text information and the image feature sequence of each image frame linguistically through the language decoder to obtain the language highlight feature sequence of each image frame.
[0121] In step 401, the text information associated with the target video may include a video content summary, the applicable services of the video (such as digital marketing services or preoperative training services), etc. At least one of the video content summary and the applicable services of the video may be determined as the video highlight text information, which is used to evaluate whether the text information of the image frame has highlights or the degree of having highlights.
[0122] In step 402, the text information associated with the image frame may include an image content summary and image text content. At least one of the image content summary and the image text content may be determined as the image text information. The image content summary may be, for example, "This frame of image is used to show the appearance of XX product", or "This frame of image is used to show the product introduction of XX product". The image text content may be, for example, "XX product is manufactured by XX manufacturer, and is now sold at XX price, and a 60% discount can be obtained if you place an order now..."
[0123] In step 403, the image text highlight score indicates the degree to which the image text information has highlights. The higher the image text highlight score, the higher the degree to which the image text information has highlights. The lower the image text highlight score, the lower the degree to which the image text information has highlights.
[0124] In one embodiment, step 303 may include:
[0125] Calculate the similarity between the video highlight text information and the image text information of the image frame to obtain a text similarity;
[0126] If the text similarity is less than the first similarity threshold and greater than the second similarity threshold, then use a preset score value as the image text highlight score;
[0127] If the text similarity is greater than or equal to the first similarity threshold, then perform score mapping on the text similarity according to a preset first mapping function to obtain a first mapping score, and determine the first mapping score as the image text highlight score;
[0128] If the text similarity is less than or equal to the second similarity threshold, then perform score mapping on the text similarity according to a preset second mapping function to obtain a second mapping score, and determine the second mapping score as the image text highlight score.
[0129] The text similarity can indicate the similarity between the video highlight text information and the image text information. However, it is found in practice that the relationship between the text similarity and the image text highlight score does not always conform to a linear relationship. In this embodiment, a first similarity threshold and a second similarity threshold are preset, and the first similarity threshold is less than the second similarity threshold. If the text similarity is less than the first similarity threshold and greater than the second similarity threshold, a preset score value is used as the image text highlight score, and this preset score value is usually small, such as zero. This embodiment also presets a first mapping function and a second mapping function. Moreover, the first mapping function is an increasing function, and the second mapping function is a decreasing function. If the text similarity is greater than or equal to the first similarity threshold, the text similarity is score-mapped according to the first mapping function. If the text similarity is less than or equal to the second similarity threshold, the text similarity is score-mapped according to the preset second mapping function to obtain the image text highlight score. In this way, the image text information with higher or lower text similarity will be given a higher image text highlight score, so as to fully excavate the image frames that may have highlights and improve the accuracy of highlight evaluation.
[0130] It should be noted that both the first mapping score and the second mapping score are greater than the preset score value.
[0131] The benefits of the embodiments of steps 401 to 405 above are that by introducing the image text information, the language decoding accuracy is further improved, thereby improving the highlight prediction accuracy.
[0132] In step 105 of some embodiments, a first highlight prediction task description feature vector is obtained by performing highlight prediction on the language highlight feature sequences of at least two image frames and a preset initial highlight prediction task description feature vector through a preset first self-attention model. The first self-attention model is a model based on the Transformer architecture, and the self-attention mechanism is used inside the Transformer to implement the highlight prediction task. The initial highlight prediction task description feature vector is a learnable vector and is a query object related to the video highlight prediction task. The initial highlight prediction task description feature vector is modified through the first self-attention model to obtain the first highlight prediction task description feature vector, and this feature vector can indicate the position where the highlight image is located.
[0133] In step 106 of some embodiments, according to the feature vector of the first highlight prediction task description, the highlight image position of the target video is determined. For example, according to the feature vector of the first highlight prediction task description, the first highlight start timestamp is filtered out from the timestamps of the image frames, and then the first highlight start timestamp is determined as the highlight image position. The specific filtering process may be to perform vector decoding according to the feature vector of the first highlight prediction task description to obtain the first highlight start timestamp. It can also be vector look-up table, etc. The specific filtering process can be set according to actual needs, and this embodiment does not make specific limitations on this.
[0134] In one embodiment, referring to Figure 5 , step 106 may include:
[0135] Step 501, performing image highlight evaluation on the image frames through a preset image highlight evaluation model to obtain highlight evaluation scores;
[0136] Step 502, filtering the image frames according to the highlight evaluation scores to obtain at least one reference image frame;
[0137] Step 503, performing highlight prediction on the language highlight feature sequence of at least one reference image frame and the preset video highlight prediction task description through a preset second self-attention model to obtain a second highlight prediction task description feature vector;
[0138] Step 504, filtering out the first highlight start timestamp from the timestamps of the image frames according to the first highlight prediction task description feature vector;
[0139] Step 505, filtering out the second highlight start timestamp from the timestamps of the reference image frames according to the second highlight prediction task description feature vector;
[0140] Step 506, determining the highlight start timestamp according to the first highlight start timestamp and the second highlight start timestamp, and determining the highlight start timestamp as the highlight image position.
[0141] In step 501, the image highlight evaluation model is a deep learning network for evaluating the highlight evaluation scores of images. The image highlight evaluation model can adopt a Convolutional Neural Network. When training the image highlight evaluation model, the samples are images, the labels are highlight scores, and the loss function can choose the cross-entropy loss function. The higher the highlight evaluation score, the higher the degree of the image frame having highlights. The lower the highlight evaluation score, the lower the degree of the image frame having highlights.
[0142] In step 502, image frames with a highlight evaluation score greater than a predetermined score threshold can be retained to obtain at least one reference image frame. For example, for image frames with timestamps 00:00:01, 00:00:03, 00:00:05, 00:00:07, and 00:00:09 respectively, perhaps only the image frames at 00:00:03, 00:00:05, and 00:00:09 have highlight evaluation scores greater than the predetermined score threshold. Then, the image frames at 00:00:03, 00:00:05, and 00:00:09 are used as reference image frames, that is, 3 reference image frames are obtained.
[0143] In step 503, the second self-attention model is a model based on the Transformer architecture. The self-attention mechanism is used inside the Transformer to implement the highlight prediction task. The initial highlight prediction task description feature vector is a learnable vector and is a query object related to the video highlight prediction task. The second self-attention model modifies the initial highlight prediction task description feature vector to obtain a second highlight prediction task description feature vector, which can indicate the position of the highlight image.
[0144] For the specific processing procedures of steps 504 and 505, reference can be made to the detailed introduction of step 106 in the above text, and details will not be elaborated here.
[0145] In step 506, according to the first highlight start timestamp and the second highlight start timestamp, the highlight start timestamp is determined, and the highlight start timestamp is determined as the highlight image position.
[0146] The advantage of the embodiments of the above steps 501 to 506 is that the second highlight prediction task description feature vector output by the second attention model can be used to further describe the video highlight task, and then the two highlight prediction task description feature vectors are jointly used to improve the accuracy of determining the highlight position of the image.
[0147] In one embodiment, referring to Figure 6 , step 506 may include:
[0148] Step 601, calculate the difference between the second highlight start timestamp and each first highlight start timestamp to obtain a timestamp difference;
[0149] Step 602, sort at least two first highlight start timestamps according to the timestamp difference to obtain a sorting result;
[0150] Step 603, determine the first highlight start timestamp with the smallest timestamp difference according to the sorting result, and determine the first highlight start timestamp as the highlight start timestamp.
[0151] In step 601, for example, if the start timestamp of the second key point is 00:00:03 and the start timestamp of the first key point is 00:00:01, then the timestamp difference is 2 seconds. If there is more than one start timestamp of the second key point, for each start timestamp of the first key point, select the smallest timestamp difference. For example, if there is another start timestamp of the second key point as 00:00:05, then the timestamp difference between it and the start timestamp of the first key point 00:00:01 is 4 seconds. Since 4 > 2, the 4 - second value will be deleted, and the timestamp difference corresponding to the start timestamp of the first key point will continue to be maintained as 2 seconds, unless a timestamp difference smaller than 2 appears.
[0152] In step 602, it can be sorted in descending order or ascending order according to the timestamp difference. For example, sorting in ascending order according to the timestamp difference, the sorting result is {00:00:03 (timestamp difference is 0), 00:00:05 (timestamp difference is 0), 00:00:01 (timestamp difference is 2), 00:00:07 (timestamp difference is 2), 00:00:09 (timestamp difference is 4)}.
[0153] In step 603, taking the above example, the start timestamps of the first key points 00:00:03 and 00:00:05 can both be determined as the start timestamps of the key points.
[0154] The above steps 601 to 603 can reduce the number of start timestamps of the key points and improve the concentration degree of the video key points.
[0155] Based on the above embodiments, the present application can at least achieve the following beneficial effects: It can better integrate features and make better use of information. By introducing a learnable key point prediction task to describe the feature vector, it can improve the accuracy of key point prediction and more flexibly handle the video key point prediction task.
[0156] Please refer to Figure 7 , the embodiment of the present application also provides a video processing device, which can implement the above video processing method. Figure 7 It is a block diagram of the module structure of the video processing device provided by the embodiment of the present application. The device includes:
[0157] A video acquisition module 701, configured to acquire a target video;
[0158] An image extraction module 702, configured to extract frames from the target video to obtain at least two image frames;
[0159] A visual encoding module 703, configured to perform visual encoding on each image frame through the visual encoder of a preset visual - language model to obtain an image feature sequence;
[0160] A language decoding module 704, configured to perform language decoding on the image feature sequence of each image frame through a language decoder of a vision-language model, so as to obtain a language key-point feature sequence of each image frame;
[0161] A key-point prediction module 705, configured to perform key-point prediction on the language key-point feature sequences of at least two image frames and a preset initial key-point prediction task description feature vector through a preset first self-attention model, so as to obtain a first key-point prediction task description feature vector;
[0162] A position determination module 706, configured to determine the key-point image position of the target video according to the first key-point prediction task description feature vector.
[0163] In one embodiment, the video processing device further includes a feature update module. After performing visual encoding on each of the image frames through a visual encoder of a preset vision-language model to obtain an image feature sequence, the feature update module is configured to: acquire text information associated with the target video to obtain video key-point text information; generate a relevance degree according to the image feature sequence and the video key-point text information to obtain an image key-point relevance degree, where the image key-point relevance degree indicates the degree of relevance between the image feature sequence and the video key-point text information; and add the image key-point relevance degree to the image feature sequence to update the image feature sequence.
[0164] It should be noted that the specific implementation manner of this video processing device is basically the same as the specific embodiments of the above video processing method, and will not be elaborated here.
[0165] An embodiment of the present application further provides an electronic device, which includes: a memory, a processor, a program stored on the memory and executable on the processor, and a data bus for implementing connection communication between the processor and the memory. When the program is executed by the processor, the above video processing method is implemented. The electronic device may be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0166] Please refer to Figure 8 , Figure 8 which schematically shows the hardware structure of an electronic device in another embodiment. The electronic device includes:
[0167] A processor 801, which may be implemented in a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is configured to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0168] The memory 802 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 802 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 802, and the processor 801 is used to call and execute the video processing method of the embodiments of this application;
[0169] The input / output interface 803 is used to implement information input and output;
[0170] The communication interface 804 is used to implement communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.);
[0171] The bus 805 transmits information between various components of the device (such as the processor 801, the memory 802, the input / output interface 803, and the communication interface 804);
[0172] Among them, the processor 801, the memory 802, the input / output interface 803, and the communication interface 804 achieve communication connections with each other inside the device through the bus 805.
[0173] The embodiments of this application also provide a storage medium. The storage medium is a computer-readable storage medium for computer-readable storage. The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the above-mentioned video processing method.
[0174] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory can optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0175] The embodiments described in the embodiments of this application are for more clearly explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are equally applicable to similar technical problems.
[0176] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.
[0177] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0178] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0179] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0180] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or a similar expression means any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0181] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical, or other forms.
[0182] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0183] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0184] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store programs.
[0185] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the rights of the embodiments of the present application. Any modification, equivalent replacement, and improvement made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.
Claims
1. A video processing method, characterized in that: The method comprises: Get the target video; Extracting frames from the target video to obtain at least two image frames; Visually encoding each of the image frames using a visual encoder of a preset visual language model to obtain an image feature sequence; Performing language decoding on the image feature sequence of each image frame by using the language decoder of the visual language model to obtain a language highlight feature sequence of each image frame; Performing highlight prediction on the language highlight feature sequences of the at least two image frames and a preset initial highlight prediction task description feature vector through a preset first self-attention model to obtain a first highlight prediction task description feature vector; Determine a highlight image position of the target video according to the first highlight prediction task description feature vector.
2. The method according to claim 1, characterized in that: The step of determining the position of the highlight image of the target video according to the first highlight prediction task description feature vector includes: Performing image viewpoint evaluation on the image frame using a preset image viewpoint evaluation model to obtain a viewpoint evaluation score; Screening the image frames according to the viewing point evaluation scores to obtain at least one reference image frame; Performing a highlight prediction on the language highlight feature sequence of the at least one reference image frame and a preset video highlight prediction task description through a preset second self-attention model to obtain a second highlight prediction task description feature vector; According to the first highlight prediction task description feature vector, filtering out a first highlight start timestamp from timestamps of the image frames; According to the second highlight prediction task description feature vector, filtering out a second highlight start timestamp from the timestamp of the reference image frame; A viewing point starting timestamp is determined according to the first viewing point starting timestamp and the second viewing point starting timestamp, and the viewing point starting timestamp is determined as the viewing point image position.
3. The method according to claim 2, characterized in that The determining the viewing point starting timestamp according to the first viewing point starting timestamp and the second viewing point starting timestamp includes: Calculate the difference between the second viewing point start timestamp and each of the first viewing point start timestamps to obtain a timestamp difference; Sorting the at least two first viewing point start timestamps according to the timestamp difference to obtain a sorting result; The first viewing point starting timestamp with the smallest timestamp difference is determined according to the sorting result, and the first viewing point starting timestamp is determined as the viewing point starting timestamp.
4. The method according to claim 1, characterized in that The language decoding of the image feature sequence of each image frame by the language decoder of the visual language model to obtain the language highlight feature sequence of each image frame includes: Acquire text information associated with the target video to obtain video highlights text information; Acquire text information associated with the image frame to obtain image text information of the image frame; Performing an appreciation evaluation on the image text information of each image frame according to the video appreciation text information to obtain an image text appreciation score of each image frame; Filtering the image text information of each image frame according to the image text highlight score to obtain selected image text information; The selected image text information and the image feature sequence of each image frame are language decoded by the language decoder to obtain a language highlight feature sequence of each image frame.
5. The method according to claim 4, characterized in that The step of evaluating the image text information of each image frame according to the video highlight text information to obtain an image text highlight score includes: Calculating the similarity between the video highlight text information and the image text information of the image frame to obtain text similarity; If the text similarity is less than the first similarity threshold and greater than the second similarity threshold, a preset score value is used as the image text highlight score; If the text similarity is greater than or equal to the first similarity threshold, score mapping is performed on the text similarity according to a preset first mapping function to obtain a first mapping score, and the first mapping score is determined as the image text highlight score; wherein the first mapping function is an increasing function; If the text similarity is less than or equal to the second similarity threshold, score mapping is performed on the text similarity according to a preset second mapping function to obtain a second mapping score, and the second mapping score is determined as the image text highlight score; wherein the second mapping function is a subtraction function; The first mapping score and the second mapping score are greater than the preset score value.
6. The method according to any one of claims 1 to 5, characterized in that: After visually encoding each of the image frames by using a visual encoder of a preset visual language model to obtain an image feature sequence, the method further includes: Acquire text information associated with the target video to obtain video highlights text information; Generate a correlation degree according to the image feature sequence and the video highlight text information to obtain an image highlight correlation degree; the image highlight correlation degree indicates the correlation degree between the image feature sequence and the video highlight text information; The image viewing point relevance is added to the image feature sequence to update the image feature sequence.
7. The method according to any one of claims 1 to 5, characterized in that: The step of extracting frames from the target video to obtain at least two image frames includes: Get the fixed interval duration; Randomly extracting a frame from the candidate image frames of the target video to obtain a first image frame; Calculating at least one extraction timestamp according to the timestamp of the first image frame and the fixed interval duration; Extracting a frame from the candidate image frames of the target video according to the extraction timestamp to obtain at least one second image frame; The first image frame and the at least one second image frame are merged to obtain the at least two image frames.
8. A video processing device, characterized in that: The device comprises: A video acquisition module is used to acquire a target video; An image extraction module, used to extract frames from the target video to obtain at least two image frames; A visual encoding module, used for visually encoding each of the image frames through a visual encoder of a preset visual language model to obtain an image feature sequence; A language decoding module, configured to perform language decoding on the image feature sequence of each image frame through a language decoder of the visual language model to obtain a language highlight feature sequence of each image frame; A highlight prediction module is used to perform highlight prediction on the language highlight feature sequences of the at least two image frames and a preset initial highlight prediction task description feature vector through a preset first self-attention model to obtain a first highlight prediction task description feature vector; A position determination module is used to determine the position of the highlight image of the target video according to the first highlight prediction task description feature vector.
9. An electronic device, characterized in that: The electronic device comprises a memory and a processor, the memory stores a computer program, and the processor implements the video processing method according to any one of claims 1 to 7 when executing the computer program.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the video processing method according to any one of claims 1 to 7 is implemented.