An image extraction model construction, image query and video generation method and device

By encoding and data augmenting video frame images, an image extraction model is constructed and trained, which solves the problem of inaccurate video frame feature extraction in existing technologies and improves the accuracy of video frame image matching.

CN116595220BActive Publication Date: 2026-02-24SHENZHEN UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202310468982.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-19
Publication Date
2026-02-24
Estimated Expiration
2043-04-19

AI Technical Summary

Technical Problem

Existing technologies cannot correctly extract features from video frames, resulting in low accuracy in video frame image matching.

Method used

After encoding and data augmentation of video frame images, an image extraction model is constructed. The pre-trained image extraction model is then used for training, ensuring that the similarity between the encoded sample of each target video frame image and its corresponding sample is greater than the similarity between other samples.

Benefits of technology

It improves the accuracy of video frame image matching, ensuring that the image extraction model can accurately extract the features of video frames.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116595220B_ABST
    Figure CN116595220B_ABST
Patent Text Reader

Abstract

The application discloses a kind of image extraction model construction, image query and video generation method, device, the method includes obtaining all video frame images in video library;Each of all video frame images is carried out encoding processing, and first video frame image encoding sample set is obtained;Each of all video frame images is carried out data enhancement processing, and the image after enhancement is encoded to obtain second video frame image encoding sample set;First video frame image encoding sample set, second video frame image encoding sample set are input into pre-trained image extraction model and are trained, so that image extraction model can accurately extract the feature of video frame, improve the accuracy of video frame image matching.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of video retrieval, and in particular to an image extraction model construction method, an image query method and a video generation method and device. BACKGROUND

[0002] With the development of computer technology, the related technology of big data has made great progress. In today's information overload, people have various search needs, and searching for video clips has become an even more urgent need. Users often search for videos by taking screenshots. However, in the search process, if the user input is a cropped video frame, the same pixel point may be divided into different blocks from the complete video frame, resulting in misalignment of the image blocks and the inability to extract the correct features. Therefore, there is an urgent need to propose an image extraction model construction method that can accurately extract the features of the video frame and improve the accuracy of video frame image matching. SUMMARY

[0003] Therefore, the technical problem to be solved by the present application is to overcome the defects of the prior art that cannot correctly extract the features of the video frame, resulting in low accuracy of video frame image matching, thereby providing an image extraction model construction method, an image query method and a video generation method and device.

[0004] According to a first aspect, an image extraction model construction method is disclosed, the method comprising: obtaining all video frame images in a video library; encoding each of the video frame images to obtain a first video frame image encoding sample set; performing data enhancement processing on each of the video frame images, and encoding the enhanced images to obtain a second video frame image encoding sample set; inputting the first video frame image encoding sample set and the second video frame image encoding sample set into a pre-trained image extraction model for training, so that the similarity between each target video frame image encoding sample in the first video frame image encoding sample set and the corresponding sample in the second video frame image encoding sample set is greater than the similarity between the target video frame image encoding sample and other samples in the second video frame image encoding sample set.

[0005] According to a second aspect, an image query method is also disclosed, the method comprising: when a target query image input by a user is received, encoding the target query image; using an image extraction model to match the encoded target query image data with all video frame image data in a preset video library, wherein the image extraction model is constructed using the image extraction model construction method of the first aspect; and determining the video frame images in the preset database that meet the requirements according to the matching result.

[0006] According to a third aspect, the embodiments of the present application further disclose a video generation method, which comprises: when receiving a target query image input by a user and containing text information, matching the text information contained in the target query image with text information contained in a video library to obtain a plurality of target text information meeting requirements; performing time consistency comparison on time stamps corresponding to the plurality of target text information and time stamps corresponding to a plurality of target video frame images, wherein the plurality of target video frame images are obtained by querying the video library using the image query method of the second aspect; combining the target text information meeting the time consistency requirement with the corresponding target video frame image, and generating a video feedback to the user end.

[0007] Optionally, the text information contained in the video library is obtained by the following steps: separating audio track information of all videos in the video library; and performing speech recognition on the audio track information to extract the text information.

[0008] According to a fourth aspect, the embodiments of the present application further disclose an image extraction model construction device, a video frame image acquisition module, configured to acquire all video frame images in a video library; a first image encoding module, configured to encode each of the video frame images to obtain a first video frame image encoding sample set; a second image encoding module, configured to perform data enhancement processing on each of the video frame images, and encode the enhanced image to obtain a second video frame image encoding sample set; and a model training module, configured to input the first video frame image encoding sample set and the second video frame image encoding sample set into a pre-trained image extraction model for training, so that the similarity between each target video frame image encoding sample in the first video frame image encoding sample set and a corresponding sample in the second video frame image encoding sample set is greater than the similarity between the target video frame image encoding sample and other samples in the second video frame image encoding sample set.

[0009] According to a fifth aspect, the embodiments of the present application further disclose an image query device, which comprises: a query image encoding module, configured to encode a target query image input by a user when receiving the target query image; a data matching module, configured to match the encoded target query image data with all video frame image data in a preset video library using an image extraction model, wherein the image extraction model is constructed using the image extraction model construction method of the first aspect; and a video frame image determination module, configured to determine video frame images meeting requirements in the preset database according to a matching result.

[0010] According to a sixth aspect, the embodiments of the present application further disclose a video generation device, which comprises: a text information matching module, configured to match text information contained in a target query image input by a user with text information contained in a video library to obtain a plurality of target text information meeting requirements when the target query image contains the text information; a time comparison module, configured to compare time stamps corresponding to the plurality of target text information with time stamps corresponding to a plurality of target video frame images in terms of time consistency, wherein the plurality of target video frame images are obtained by querying the video library using the image query method according to the second aspect; and a video generation module, configured to combine the target text information meeting the time consistency requirement with the corresponding target video frame images, and generate a video and feed back the video to a user terminal.

[0011] According to a seventh aspect, the embodiments of the present application further disclose an electronic device, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the image extraction model construction method according to the first aspect, or the steps of the image query method according to the second aspect, or the steps of the video generation method according to the third aspect or any optional implementation manner of the third aspect.

[0012] According to an eighth aspect, the embodiments of the present application further disclose a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the image extraction model construction method according to the first aspect, or the steps of the image query method according to the second aspect, or the steps of the video generation method according to the third aspect or any optional implementation manner of the third aspect.

[0013] The technical scheme of the present application has the following advantages:

[0014] The image extraction model construction method provided by the present application encodes each video frame image to obtain a first video frame image encoding sample set, encodes each video frame image after data enhancement to obtain a second video frame image encoding sample set, inputs the first video frame image encoding sample set and the second video frame image encoding sample set into the image extraction model for training, so that the similarity between each target video frame image encoding sample in the first video frame image encoding sample set and the corresponding sample in the second video frame image encoding sample set is greater than the similarity between the target video frame image encoding sample and other samples in the second video frame image encoding sample set, so that the image extraction model constructed can accurately extract the features of the video frame, and then the accuracy of the video frame image matching can be improved. BRIEF DESCRIPTION OF DRAWINGS

[0015] In order to more clearly illustrate the technical solutions in the specific embodiments of the present application or the prior art, the following will briefly introduce the drawings needed to be used in the specific embodiments or prior art description. Obviously, the drawings described below are some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor based on these drawings.

[0016] Figure 1 A flow chart of a specific example of the image extraction model construction method in the embodiments of the present application;

[0017] Figure 2 A flow chart of a specific example of the image query method in the embodiments of the present application;

[0018] Figure 3 A flow chart of a specific example of the video generation method in the embodiments of the present application;

[0019] Figure 4 A principle block diagram of a specific example of the image extraction model construction device in the embodiments of the present application;

[0020] Figure 5 A principle block diagram of a specific example of the image query device in the embodiments of the present application;

[0021] Figure 6 A principle block diagram of a specific example of the video generation device in the embodiments of the present application;

[0022] Figure 7 A specific example of the electronic device in the embodiments of the present application. DETAILED DESCRIPTION

[0023] The technical solutions of the present application will be described clearly and completely below in conjunction with the drawings. Obviously, the described embodiments are some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the protection scope of the present application.

[0024] In the description of the present application, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the device or element referred to must have a particular orientation, be constructed and operated in a particular orientation, therefore it cannot be understood as a limitation of the present application. In addition, the terms "first", "second", "third" are only for the purpose of description, and cannot be understood as indicating or implying relative importance.

[0025] In the description of this invention, it should be noted that, unless otherwise explicitly specified and limited, the terms "installation," "connection," and "linking" should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can also refer to the internal connection of two components; and they can refer to a wireless connection or a wired connection. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0026] Furthermore, the technical features involved in the different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0027] This invention discloses a method for constructing an image extraction model, such as... Figure 1 As shown, the method includes the following steps:

[0028] Step S101: Obtain all video frame images in the video library; for example, the video library may contain all types of videos in the scene to be searched. In this embodiment of the application, the video screenshot function of the FFmpeg computer program is used to generate a series of video frame images from the videos in the video library, which is only an example.

[0029] Step S102: Encode each video frame image in all video frame images to obtain the first video frame image encoding sample set;

[0030] For example, in this embodiment of the application, all video frame images obtained in step S101 are encoded to obtain a first video frame image encoded sample set. This embodiment of the application does not limit the encoding processing method, and those skilled in the art can determine it according to actual needs. In a specific embodiment, a query encoder is constructed. For any video frame image A in all video frame images, it is used as a sample, and the sample is encoded to obtain query(A). This is only an example.

[0031] Step S103: Perform data augmentation on each video frame image in all video frame images, and encode the augmented images to obtain the second video frame image encoding sample set;

[0032] For example, in this embodiment, all video frame images obtained in step S101 are subjected to data augmentation processing, such as cropping, blurring, and rotating the images. This is just an example to ensure accurate image matching even if the received user-input video frame image is not clear or is incomplete. All the augmented video frame images are encoded to obtain a second video frame image encoding sample set. In a specific embodiment, any video image A is subjected to data augmentation processing to obtain sample A′, and the remaining sample set A is used as negative samples. A keyword encoder is constructed to encode A′ to obtain key(A′), and the negative sample A is encoded to obtain key(A). This is just an example.

[0033] Step S104: Input the first video frame image encoding sample set and the second video frame image encoding sample set into the pre-trained image extraction model for training, so that the similarity between each target video frame image encoding sample in the first video frame image encoding sample set and the corresponding sample in the second video frame image encoding sample set is greater than the similarity between the target video frame image encoding sample and other samples in the second video frame image encoding sample set.

[0034] For example, in this embodiment, the existing ResNet residual neural network model is used as the pre-trained image extraction model. Using the first video frame image encoding sample set and the second video frame image encoding sample set obtained in the above steps, self-supervised training is performed based on the MoCo contrastive learning method. This makes the similarity between each target video frame image encoding sample in the first video frame image encoding sample set and the corresponding data-enhanced sample in the second video frame image encoding sample set greater than the similarity between the target video frame image encoding sample and other samples in the second video frame image encoding sample set. In this embodiment, the residual neural network model is only an example. AlexNet, VGG, and Vision Transformer models can also be used, as long as they can achieve the image extraction function. SimCLR can also be used instead of MoCo to train the image extraction model.

[0035] In a specific embodiment, the query(A), key(A), and key(A′) obtained in the above steps are input into a pre-trained image extraction model for training. By automatically updating the parameters of the neural network model, the Euclidean distance between query(A) and key(A′) in the vector space is made as close as possible, and the Euclidean distance between query(A) and key(A) is made as far as possible, thereby improving the image extraction capability of the image extraction model. In this specific embodiment, Euclidean distance is used to represent the similarity between the target video frame image encoding sample and other samples. This is only an example; cosine similarity, Manhattan distance, etc., can also be used.

[0036] The image extraction model construction method provided by this invention involves encoding each video frame image to obtain a first video frame image encoded sample set, performing data augmentation on each video frame image and then encoding it again to obtain a second video frame image encoded sample set. The first and second video frame image encoded sample sets are then input into an image extraction model for training. This ensures that the similarity between each target video frame image encoded sample in the first video frame image encoded sample set and its corresponding sample in the second video frame image encoded sample set is greater than the similarity between the target video frame image encoded sample and other samples in the second video frame image encoded sample set. This allows the constructed image extraction model to accurately extract features from video frames, thereby improving the accuracy of video frame image matching.

[0037] This invention also discloses an image query method, such as... Figure 2 As shown, the method includes the following steps:

[0038] Step S201: When the target query image input by the user is received, the target query image is encoded. For example, after receiving the target query image input by the user, a trained image extraction model or other integrated encoding function modules can be used to extract and encode features of the target query image to obtain the target query vector.

[0039] Step S202 involves matching the encoded target query image data with all video frame image data in a preset video library using an image extraction model. The image extraction model is constructed using the image extraction model construction method described in the above embodiments. For example, in this application embodiment, a pre-trained image extraction model or other encoding function modules can be used to extract and encode features from all video frame images in the video library to obtain a feature vector corresponding to each video frame image. All generated feature vectors are then stored in a vector database, which can be a Qdrant database, as an example only. The image extraction model constructed using the above image extraction model matches the target query vector obtained in step S201 with all feature vectors in the vector database. When storing all generated feature vectors in the vector database, the mapping from feature vectors to video information needs to be preserved. This ensures that the video frame images that meet the requirements in the vector database can obtain their corresponding video information, such as the timestamp corresponding to the video frame.

[0040] Step 203: Determine the video frame images that meet the requirements in the preset database based on the matching results. For example, in this embodiment, an image extraction model can be used to quickly match the N1 feature vectors closest to the target query vector in the vector database. This means matching the video frame images closest to the user-input target query image, and each video frame image has a corresponding score for its proximity to the target query image, such as 10, 9, 8, etc. This is just an example. Alternatively, in this embodiment, the matched video frame images closest to the user-input target query image can be directly input into the video generation module, and video clips can be generated based on the timestamps corresponding to the video frame images and fed back to the user for viewing.

[0041] The image query method provided by this invention encodes the target query image input by the user and then uses an image extraction model to match it with all video frame images in the video library. Based on the matching results, the method determines the video frame images in the video library that meet the requirements, thereby retrieving the video frame images most similar to the target query image more accurately and completely.

[0042] This invention also discloses a video generation method that can be applied to a video query system. This system integrates a search service module, which receives user video query operations and displays the query results through a user interface. Figure 3 As shown, the method includes the following steps:

[0043] Step S301: When the user inputs a target query image containing text information, the text information contained in the target query image is matched with the text information contained in the video library to obtain multiple target text information that meet the requirements.

[0044] For example, in this embodiment, the user input may be text information included in the target query image, which needs to be identified to obtain the text information. Alternatively, the user may directly input text information corresponding to the target query image, which is then matched with text information contained in the video library. The text information in the video library includes subtitle text of dialogue and corresponding timestamps. The text information extracted from the video library is pre-stored in the dialogue storage and matching module. When the user's query text information is received, the ElasticSearch inverted index search engine performs keyword matching with the text information in the video library in the dialogue storage and matching module to obtain the N2 sentences of text information most similar to the query text information. The N2 sentences of text information most similar to the query text information can be obtained by setting a similarity threshold. The matched N2 sentences of text information can be sorted and scored according to their similarity with the query text information, such as 10 points, 9 points, 8 points, etc., as an example only.

[0045] Step S302 involves comparing the timestamps corresponding to the multiple target text information with the timestamps corresponding to the multiple target video frame images, wherein the multiple target video frame images are obtained using the image query method described in the above embodiment. For example, in this embodiment, the video timestamps corresponding to the N1 video frame images most adjacent to the target query image obtained through the above image query method are compared with the video timestamps corresponding to the N2 sentences of text information obtained in step 301. The N1 video frame images most adjacent to the target query image can also be obtained using a preset similarity threshold.

[0046] Step S303: Combine the target text information that meets the time consistency requirement with the corresponding target video frame image, and generate a video to be fed back to the user terminal.

[0047] For example, in this embodiment of the application, if the video timestamp corresponding to a certain video frame image is substantially consistent with the video timestamp corresponding to a certain text information, they are combined. A weighted scoring and sorting method is adopted, scoring is performed according to the weights corresponding to the video frame image and the text information respectively. The top N combination results are taken. For example, if the weight corresponding to the video frame image is 0.5 and the weight corresponding to the text information is 0.5, the combination of the text information and the video frame image according to the timestamp results in a combination of a certain video frame image (score of 8 points) and a certain text information (score of 8 points), and the score of this combination is 0.5×8+0.5×8=8. The scores of all combinations are obtained sequentially, and they are sorted according to the scores. The top N combination results are input into the video generation module, and the FFmpeg computer program generates the corresponding video segment according to the corresponding video timestamp. The generated video segment is then fed back to the user terminal for preview.

[0048] The video generation method provided by this invention matches the received user-inputted text information with the text information in a pre-saved video library to obtain multiple target text information that meet the requirements. It then performs a time consistency comparison between the timestamps corresponding to the target text information and the timestamps corresponding to the multiple target video frame images. The target text information that meets the time consistency requirements is combined with the corresponding target video frame images to generate a video that is fed back to the user. By combining image and text methods for video segment retrieval, video retrieval can be performed more effectively.

[0049] As an optional embodiment of the present invention, the text information contained in the video library is obtained through the following steps: separating the audio track information of all videos in the video library; and extracting text information from the audio track information using speech recognition. Exemplarily, in this embodiment, the audio track information of all videos in the video library is separated using moviepy in Python, and then the text information in the audio track information is extracted using speech recognition via ApiSpeech. The extracted text information is then stored in the dialogue matching module. This eliminates the need to rely on image recognition or the internet to obtain text information, ensuring more efficient and comprehensive acquisition of text information corresponding to all videos in the video library, thus broadening its applicability. This is merely an example.

[0050] This invention also discloses an image extraction model construction device, such as... Figure 4As shown, the device includes: a video frame image acquisition module 401, used to acquire all video frame images in a video library; a first image encoding module 402, used to encode each video frame image in all video frame images to obtain a first video frame image encoding sample set; a second image encoding module 403, used to perform data augmentation processing on each video frame image in all video frame images and encode the augmented image to obtain a second video frame image encoding sample set; and a model training module 404, used to input the first video frame image encoding sample set and the second video frame image encoding sample set into a pre-trained image extraction model for training, such that the similarity between each target video frame image encoding sample in the first video frame image encoding sample set and the corresponding sample in the second video frame image encoding sample set is greater than the similarity between the target video frame image encoding sample and other samples in the second video frame image encoding sample set.

[0051] The image extraction model construction apparatus provided by this invention obtains a first video frame image encoded sample set by encoding each video frame image, and obtains a second video frame image encoded sample set by performing data augmentation processing on each video frame image and then encoding it. The first and second video frame image encoded sample sets are input into an image extraction model for training, such that the similarity between each target video frame image encoded sample in the first video frame image encoded sample set and the corresponding sample in the second video frame image encoded sample set is greater than the similarity between the target video frame image encoded sample and other samples in the second video frame image encoded sample set. This enables the constructed image extraction model to accurately extract features from video frames, thereby improving the accuracy of video frame image matching.

[0052] This invention also discloses an image query device, such as... Figure 5 As shown, the device includes: a query image encoding module 501, used to encode the target query image when a user inputs a target query image; a data matching module 502, used to match the encoded target query image data with all video frame image data in a preset video library using an image extraction model, wherein the image extraction model is constructed using the image extraction model construction method described in the above embodiment; and a video frame image determination module 503, used to determine the video frame images in the preset database that meet the requirements based on the matching results.

[0053] The image query device provided by the present invention encodes the target query image input by the user and then uses an image extraction model to match it with all video frame images in the video library. Based on the matching results, it determines the video frame images in the video library that meet the requirements, and can retrieve the video frame images most similar to the target query image more accurately and completely.

[0054] This invention also discloses a video generation apparatus, such as... Figure 6 As shown, the device includes: a text information matching module 601, used to match the text information contained in the target query image with the text information contained in the video library when a user inputs a target query image containing text information to obtain multiple target text information that meet the requirements; a time comparison module 602, used to perform time consistency comparison between the timestamps corresponding to the multiple target text information and the timestamps corresponding to the multiple target video frame images, wherein the multiple target video frame images are obtained by querying using the image query method described in the above embodiment; and a video generation module 603, used to combine the target text information that meets the time consistency requirements with the corresponding target video frame images and generate a video to be fed back to the user terminal.

[0055] The video generation device provided by this invention obtains multiple target text information that meet the requirements by matching the received user-inputted text information with the text information in a pre-saved video library, and compares the timestamps corresponding to the target text information with the timestamps corresponding to the multiple target video frame images for time consistency. The target text information that meets the time consistency requirements is combined with the corresponding target video frame images to generate a video and feed it back to the user. By combining image and text methods for video segment retrieval, video retrieval can be performed more effectively.

[0056] As an optional embodiment of the present invention, the text information matching module includes: an audio track information separation submodule, used to separate the audio track information of all videos in the video library; and a text information recognition submodule, used to extract text information from the audio track information through speech recognition.

[0057] This invention also provides an electronic device, such as... Figure 7 As shown, the electronic device may include a processor 701 and a memory 702, wherein the processor 701 and the memory 702 may be connected via a bus or other means. Figure 7 Taking the example of a connection between China and Israel via a bus.

[0058] Processor 701 can be a central processing unit (CPU). Processor 701 can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations of the above types of chips.

[0059] The memory 702, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the illegal behavior detection method in the embodiments of the present invention. The processor 701 executes various functional applications and data processing by running the non-transitory software programs, instructions, and modules stored in the memory 702, thereby implementing the image extraction model construction method, image query method, or video generation method in the above method embodiments.

[0060] The memory 702 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created by the processor 701, etc. Furthermore, the memory 702 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, the memory 702 may optionally include memory remotely located relative to the processor 701, and these remote memories may be connected to the processor 701 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0061] The one or more modules are stored in the memory 702, and when executed by the processor 701, they perform actions such as... Figure 1 The image extraction model construction method or in the embodiments shown Figure 2 The image query method in the illustrated embodiment or as shown Figure 3 The video generation method in the illustrated embodiment.

[0062] For specific details regarding the aforementioned electronic devices, please refer to the relevant documentation. Figure 1 or Figure 2 or Figure 3 The relevant descriptions and effects in the illustrated embodiments are for understanding purposes only and will not be repeated here.

[0063] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk drive (HDD), or solid-state drive (SSD), etc.; the storage medium can also include combinations of the above types of memory.

[0064] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the defined scope.

Claims

1. A method for constructing an image extraction model, characterized in that, The method includes: Retrieve all video frame images from the video library; Encode each video frame image in all video frame images to obtain the first video frame image encoded sample set; Data augmentation is performed on each video frame image in all video frame images, and the augmented images are encoded to obtain the second video frame image encoding sample set; The first video frame image encoding sample set and the second video frame image encoding sample set are input into a pre-trained image extraction model for training, such that the similarity between each target video frame image encoding sample in the first video frame image encoding sample set and the corresponding data-enhanced video frame image sample in the second video frame image encoding sample set is greater than the similarity between the target video frame image encoding sample and other samples in the second video frame image encoding sample set.

2. An image query method, characterized in that, The method includes: When a target query image is received from the user, the target query image is encoded; The encoded target query image data is matched with all video frame image data in a preset video library using an image extraction model, wherein the image extraction model is constructed using the image extraction model construction method described in claim 1. Based on the matching results, determine the video frame images in the preset video library that meet the requirements.

3. A video generation method, characterized in that, The method includes: When a user inputs a target query image containing text information, the text information in the target query image is matched with the text information in the video library to obtain multiple target text information that meet the requirements. The timestamps corresponding to the plurality of target text information are compared with the timestamps corresponding to the plurality of target video frame images, wherein the plurality of target video frame images are obtained by querying using the image query method described in claim 2. The target text information that meets the time consistency requirement is combined with the corresponding target video frame image, and a video is generated and fed back to the user.

4. The video generation method according to claim 3, characterized in that, The text information contained in the video library is obtained through the following steps: Extract the audio track information from all videos in the video library; The audio track information is used to extract text information through speech recognition.

5. An image extraction model construction device, characterized in that, The device includes: The video frame image acquisition module is used to acquire all video frame images in the video library; The first image encoding module is used to encode each video frame image in all video frame images to obtain the first video frame image encoding sample set; The second image encoding module is used to perform data augmentation processing on each video frame image in all video frame images, and encode the augmented image to obtain the second video frame image encoding sample set; The model training module is used to input the first video frame image encoding sample set and the second video frame image encoding sample set into a pre-trained image extraction model for training, and to input the first video frame image encoding sample set and the second video frame image encoding sample set into the pre-trained image extraction model for training, such that the similarity between each target video frame image encoding sample in the first video frame image encoding sample set and the corresponding data-enhanced video frame image sample in the second video frame image encoding sample set is greater than the similarity between the target video frame image encoding sample and other samples in the second video frame image encoding sample set.

6. An image query device, characterized in that, The device includes: The query image encoding module is used to encode the target query image when it is received by the user. The data matching module is used to match the encoded target query image data with all video frame image data in the preset video library using an image extraction model, wherein the image extraction model is constructed using the image extraction model construction method described in claim 1. The video frame image determination module is used to determine the video frame images that meet the requirements in the preset video library based on the matching results.

7. A video generation apparatus, characterized in that, The device includes: The text information matching module is used to match the text information contained in the target query image with the text information contained in the video library when it receives text information from the user input target query image to obtain multiple target text information that meet the requirements. The time comparison module is used to perform a time consistency comparison between the timestamps corresponding to the plurality of target text information and the timestamps corresponding to the plurality of target video frame images, wherein the plurality of target video frame images are obtained by querying using the image query method described in claim 2. The video generation module is used to combine target text information that meets the time consistency requirements with the corresponding target video frame image, and generate a video to be fed back to the user.

8. The apparatus according to claim 7, characterized in that, The text information matching module includes: The audio track information separation submodule is used to separate the audio track information of all videos in the video library; The text information recognition submodule is used to extract text information from the audio track information through speech recognition.

9. An electronic device, characterized in that, include: At least one processor; And a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to cause the at least one processor to perform the steps of the image extraction model construction method as claimed in claim 1, the steps of the image query method as claimed in claim 2, or the steps of the video generation method as claimed in any one of claims 3-4.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program is executed by a processor using the steps of the image extraction model construction method as described in claim 1, the image query method as described in claim 2, or the video generation method as described in any of claims 3-4.

Citation Information

Patent Citations

  • Video clip query method and apparatus

    CN107122439A

  • Video retrieval method and device, electronic device and storage medium

    CN110913241A

  • Training method and device, electronic equipment and computer readable storage medium

    CN112307883A

  • Video subtitle processing method and device, equipment and storage medium

    CN112995749A