Target video searching method and device, equipment, medium and product
By extracting keyframe images from historical videos and performing facial or feature recognition, the problems of inefficient and low accuracy of traditional target video searches are solved, and efficient and accurate target video search and management are achieved.
Patent Information
- Application Number
- CN202510154098.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-05-23
AI Technical Summary
In the field of security monitoring, traditional target video search methods rely on manual query, resulting in low search efficiency and low accuracy.
By obtaining the target object information, extracting multi-frame keyframe images from the historical video based on preset time intervals, and matching the target object information through face recognition or feature recognition, thereby obtaining the target video clip.
Improve the efficiency and accuracy of target video search, allowing users to quickly locate and obtain video clips that match the target object, simplifying the management and retrieval of video clips.
Smart Images

Figure CN120030179A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the field of image recognition technology. More specifically, the embodiments of the present invention relate to a method, device, equipment, medium and product for searching a target video. Background Art
[0002] In the field of security monitoring, querying and retrieving historical videos has always been an important challenge faced by end users. Traditional methods of searching for target videos are mainly manual, searching by date and time and searching from event lists to search through a large number of videos. Although these two methods can meet basic video retrieval needs to a certain extent, the disadvantages of low search efficiency and low accuracy are becoming increasingly prominent. Summary of the invention
[0003] In this context, embodiments of the present invention are intended to provide a method, apparatus, device, medium and product for searching a target video.
[0004] In a first aspect of an embodiment of the present invention, a method for searching a target video is provided, comprising: Get target object information; Based on a preset time interval, a plurality of key frame images are obtained from the pre-acquired historical video recordings; Acquire a target key frame image matching the target object information from the multiple key frame images; A target video segment corresponding to the target key frame is obtained from the historical video.
[0005] In an example of this implementation manner, the target object information includes at least a time interval, the historical video includes a plurality of video clips, the duration of each video clip is a preset duration, the plurality of video clips are sorted in the order of recording time, and the acquiring of a plurality of key frame images from the pre-acquired historical video based on the preset time interval includes: Obtaining the start time and end time of the time interval; Determine a starting video segment in the historical video corresponding to the starting time; Determine a termination video segment in the historical video corresponding to the termination time; Acquire an intermediate video segment between the start video segment and the end video segment from the historical video; Based on a preset time interval, a plurality of key frame images are acquired from the start video segment, the end video segment and the middle video segment.
[0006] In an embodiment of this implementation manner, the target object information further includes a target object image. Obtaining a target key frame image that matches the target object information from the multiple key frame images includes: Perform type analysis on the target object image to determine the image type of the target object image; wherein, the image type includes a person type and an object type; If the image type is the person type, perform face recognition on the target object image to obtain the target face feature of the target object image; Filter from the multiple key frame images to obtain face key frame images containing face features; Perform face recognition on the face key frame images to obtain candidate face features of the face key frame images; Compare the candidate face features with the target face features to obtain key face features; Determine the face key frame image corresponding to the key face features as the target key frame image.
[0007] In an embodiment of this implementation manner, if the image type is the object type, the method further includes: Perform feature recognition on the target object image to obtain the target object feature of the target object image; Filter from the multiple key frame images to obtain object key frame images containing object features; Perform feature recognition on the object key frame images to obtain candidate object features of the object key frame images; Compare the candidate object features with the target object features to obtain key object features; Determine the object key frame image corresponding to the key object features as the target key frame image.
[0008] In an embodiment of this implementation manner, obtaining a target key frame image that matches the target object information from the multiple key frame images includes: Detect whether there is a target key frame image that matches the target object information in the multiple key frame images to obtain a detection result; If the detection result indicates that there is a target key frame image that matches the target object information in the multiple key frame images, then perform the step of obtaining the target video segment corresponding to the target key frame from the historical video; If the detection result indicates that there is no target key frame image matching the target object information in the multiple key frame images, then obtaining a target video segment other than the start video segment, the end video segment and the middle video segment in the historical video; Based on a preset time interval, acquiring multiple frames of current key frame images from the target video clip; A target key frame image matching the target object information is acquired from the multiple frames of current key frame images.
[0009] In one embodiment of this implementation, the method further includes: Obtain the video identification of the target video clip; Sorting the video identifications of the target video clips in the order of video recording time to obtain a video identification sequence; The video identification sequence is outputted so that a user can obtain a target video segment corresponding to the video identification based on the video identification sequence.
[0010] In a second aspect of the embodiments of the present invention, a device for searching a target video is provided, comprising: A first acquisition unit, used to acquire target object information; A second acquisition unit, configured to acquire a plurality of key frame images from pre-acquired historical video recordings based on a preset time interval; A third acquisition unit, configured to acquire a target key frame image matching the target object information from the multiple key frame images; The fourth acquisition unit is used to acquire a target video segment corresponding to the target key frame from the historical video.
[0011] In a third aspect of an embodiment of the present invention, a computing device is provided, comprising: at least one processor, a memory, and an input-output unit; wherein the memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute any one of the methods described in the first aspect.
[0012] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, which includes instructions, and when the instructions are executed on a computer, the computer executes the method described in any one of the first aspects.
[0013] In a fifth aspect of the embodiments of the present invention, a computer program product is provided, comprising a computer program, which implements any one of the methods in the first aspect when executed by a processor.
[0014] According to the target video search method, device, equipment, medium and product of the embodiment of the present invention, it is possible to obtain target object information; based on a preset time interval, obtain multiple key frame images from pre-acquired historical videos; obtain target key frame images matching the target object information from the multiple key frame images; and obtain target video clips corresponding to the target key frames from the historical videos. It can be seen that this method can meet basic video retrieval requirements and improve search efficiency and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The above and other objects, features and advantages of the exemplary embodiments of the present invention will become readily understood by reading the detailed description below with reference to the accompanying drawings. In the accompanying drawings, several embodiments of the present invention are shown in an exemplary and non-limiting manner, in which: Figure 1 A schematic diagram of a flow chart of a method for searching a target video provided by an embodiment of the present invention; Figure 2 A schematic diagram of the structure of a target video search device provided by an embodiment of the present invention; Figure 3 A schematic diagram of the structure of a medium according to an embodiment of the present invention is schematically shown; Figure 4 A schematic diagram of the structure of a computing device according to an embodiment of the present invention is schematically shown.
[0016] In the drawings, the same or corresponding reference numerals represent the same or corresponding parts. DETAILED DESCRIPTION
[0017] The principles and spirit of the present invention will be described below with reference to several exemplary embodiments. It should be understood that these embodiments are provided only to enable those skilled in the art to better understand and implement the present invention, and are not intended to limit the scope of the present invention in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0018] Those skilled in the art know that the embodiments of the present invention can be implemented as a system, device, apparatus, method or computer program product. Therefore, the present disclosure can be specifically implemented in the following forms, namely: complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0019] According to an embodiment of the present invention, a method, device, equipment, medium and product for searching a target video are proposed.
[0020] It should be noted that any number of elements in the drawings is for illustration rather than limitation, and any naming is only for distinction and does not have any limiting meaning.
[0021] The principle and spirit of the present invention are explained in detail below with reference to several representative embodiments of the present invention.
[0022] Exemplary Methods Reference below Figure 1 , Figure 1 This is a flow chart of a method for searching a target video provided by an embodiment of the present invention. It should be noted that the embodiments of the present invention can be applied to any applicable scenario.
[0023] Figure 1 The process of the target video search method provided by an embodiment of the present invention includes: Step S101, obtaining target object information.
[0024] In the embodiment of the present invention, the target object information includes at least a time interval and a target object image. The time interval may be a time interval that has already passed. For example: 15:00-17:00, January 17, 2025. The target object information may also include target object text information, and the text information may include the name and description information of the target object, which is not limited in the embodiment of the present invention.
[0025] In the embodiment of the present invention, the target object information can be input by the user using a terminal device, and the target object information input by the user can be saved in the cloud.
[0026] Step S102: acquiring multiple key frame images from pre-acquired historical video recordings based on a preset time interval.
[0027] In the embodiment of the present invention, the historical video includes a plurality of video clips, the duration of each video clip is a preset duration, and the plurality of video clips are arranged in the order of recording time.
[0028] In the embodiment of the present invention, historical video can be obtained by shooting with a camera. When the camera captures a motion detection event, the video recorded by the camera at that time is saved in segments every 10 seconds and uploaded to the cloud through an http request.
[0029] For example, the historical video to be searched may be extracted frames, for example, one key frame image may be extracted every 2 seconds, thereby obtaining multiple key frame images.
[0030] As an optional implementation, based on a preset time interval, a method of acquiring multiple key frame images from pre-acquired historical video recordings may specifically be: Obtain the start time and end time of the time interval; Determine the start video segment in the historical video corresponding to the start time; Determine the end video segment in the historical video corresponding to the end time; Obtain the intermediate video segment between the start video segment and the end video segment from the historical video; Based on a preset time interval, obtain multiple key frame images from the start video segment, the end video segment, and the intermediate video segment.
[0031] Among them, implementing this implementation method, by determining the start and end times of the time interval and accurately locating the corresponding start video segment and end video segment, this method ensures the timeliness of key frame extraction. Further, by including the processing of the intermediate video segment, this method can comprehensively cover the entire time interval of interest and avoid information omission. Extracting key frame images from all relevant video segments based on a preset time interval not only improves the processing efficiency but also ensures the representativeness of the key frames, enabling the finally obtained key frame images to accurately reflect the important content and changes in the historical video and providing strong support for subsequent analysis, monitoring, or review work.
[0032] Step S103, obtain the target key frame image that matches the target object information from the multiple key frame images.
[0033] In the embodiment of the present invention, the target recognition model obtained through pre-training can be used to match the multiple key frame images with the target object information, thereby making the obtained target key frames more accurate.
[0034] As an optional implementation method, the method of obtaining the target key frame image that matches the target object information from the multiple key frame images can specifically be: Conduct type analysis on the target object image to determine the image type of the target object image; wherein, the image type includes the human type and the object type; If the image type is the human type, perform face recognition on the target object image to obtain the target face feature of the target object image; Screen from the multiple key frame images to obtain the face key frame images containing face features; Perform face recognition on the face key frame images to obtain the candidate face features of the face key frame images; Compare the candidate face features with the target face features to obtain the key face features; Determine the face key frame image corresponding to the key face features as the target key frame image.
[0035] Among them, by implementing this implementation method, by performing type analysis on the target object image, the method can intelligently identify whether the target object belongs to a person or an object, so as to take subsequent processing measures in a targeted manner. This type analysis not only improves the accuracy of processing, but also avoids unnecessary resource consumption. Secondly, when the target object is a person, the target facial features are extracted by face recognition technology, and compared with the facial features in multiple key frame images. The method can accurately screen out the face key frame images that match the target object. This process is not only efficient, but also greatly improves the accuracy of obtaining the target key frame images. In addition, by comparing the candidate facial features with the target facial features, the method can further confirm the key facial features, thereby ensuring that the final target key frame image is highly matched with the target object. This meticulous comparison process enhances the reliability of the results, so that the target key frame image can truly reflect the actual situation of the target object.
[0036] To sum up, this implementation method realizes the efficient and accurate acquisition of target key frame images matching the target object information from multiple key frame images through a series of intelligent processing steps such as type analysis, face recognition and feature comparison, providing strong support for subsequent image analysis, monitoring and tracking and other tasks.
[0037] Optionally, if the image type is the object type, performing feature recognition on the target object image to obtain target object features of the target object image; Screening the multiple key frame images to obtain an object key frame image containing object features; Performing feature recognition on the object key frame image to obtain candidate object features of the object key frame image; Comparing the candidate object features with the target object features to obtain key object features; The object key frame image corresponding to the key object feature is determined as the target key frame image.
[0038] Among them, by implementing this implementation method, the target object features in the target object image are accurately extracted through feature recognition technology. This step ensures the baseline accuracy of subsequent analysis. Subsequently, images containing similar object features are screened in multiple key frame images, which not only narrows the processing scope but also improves the processing efficiency. Furthermore, feature recognition is performed on the screened object key frame images and a detailed comparison is made with the target object features. This process ensures that the final target key frame image is highly consistent with the target object. Therefore, this implementation method not only improves the accuracy of acquiring the target key frame image of the object type, but also expands the application scenarios of the method, so that it can be widely used in various fields such as object tracking, scene reconstruction, object recognition, etc., showing strong practical value and generalization ability.
[0039] As an optional implementation manner, a method of acquiring a target key frame image matching the target object information from the multiple key frame images may specifically be: Detecting whether there is a target key frame image matching the target object information among the multiple key frame images, and obtaining a detection result; If the detection result indicates that there is a target key frame image matching the target object information in the multiple key frame images, executing step S104; If the detection result indicates that there is no target key frame image matching the target object information in the multiple key frame images, then obtaining a target video segment other than the start video segment, the end video segment and the middle video segment in the historical video; Based on a preset time interval, acquiring multiple frames of current key frame images from the target video clip; A target key frame image matching the target object information is acquired from the multiple frames of current key frame images.
[0040] Among them, by implementing this implementation method, by preliminarily detecting whether there is a target key frame image matching the target object information in the multiple key frame images, the method can quickly determine whether further processing is required, thereby avoiding unnecessary resource consumption. When a match is detected, the subsequent steps are directly executed to obtain the target key frame image, thereby improving processing efficiency. When the preliminary detection fails to find a match in the multiple key frame images, the method does not give up directly, but further mines potential information from historical videos. By obtaining the target video clips other than the start, end and middle video clips, and extracting multiple frames of current key frame images therefrom based on a preset time interval, the method expands the search range and increases the possibility of finding the target key frame image.
[0041] Step S104: acquiring a target video segment corresponding to the target key frame from the historical video.
[0042] As an alternative implementation, after step S104, the following steps may also be executed: Obtain the video identifier of the target video clip; Sort the video identifiers of the target video clips in chronological order of video time to obtain a video identifier sequence; Output the video identifier sequence so that the user can obtain the target video clip corresponding to the video identifier based on the video identifier sequence.
[0043] Among them, by implementing this implementation method, by obtaining the video identifiers of the target video clips and sorting these identifiers in chronological order of video time to obtain a video identifier sequence, this method not only helps the user clearly understand the time sequence of each target video clip, but also greatly simplifies the process for the user to search for and obtain specific video clips. The user only needs to quickly locate the video identifier corresponding to the required video clip according to the output video identifier sequence, and then obtain the target video clip. The addition of this step not only improves the user experience, but also makes the management and retrieval of video clips more intuitive and efficient, saving a large amount of time and effort for the user.
[0044] The present invention can meet the basic video retrieval requirements, improve the search efficiency and accuracy. In addition, the present invention can also make the finally obtained key frame images accurately reflect the important content and changes in the historical video. In addition, the present invention can also efficiently and accurately obtain the target key frame images that match the target object information from multiple key frame images. In addition, the present invention can also improve the accuracy of obtaining the object type target key frame images. In addition, the present invention can also increase the possibility of finding the target key frame images. In addition, the present invention can also improve the user experience, making the management and retrieval of video clips more intuitive and efficient, saving a large amount of time and effort for the user.
[0045] Exemplary Devices After introducing the method of the exemplary embodiment of the present invention, next, reference is made to Figure 2 A search device for a target video of an exemplary embodiment of the present invention will be described. The device includes: A first acquisition unit 201 for acquiring target object information; A second acquisition unit 202 for acquiring multiple key frame images from the pre-acquired historical video based on a preset time interval; In an embodiment of the present invention, at least the time interval and the target object image are included in the target object information, multiple video clips are included in the historical video, the duration of each video clip is a preset duration, and the multiple video clips are sorted in chronological order of recording time.
[0046] A third acquisition unit 203 is used to acquire a target key frame image matching the target object information from the multiple key frame images; The fourth acquisition unit 204 is configured to acquire a target video segment corresponding to the target key frame from the historical video.
[0047] As an optional implementation manner, the second acquisition unit 202 acquires multiple key frame images from the pre-acquired historical video based on the preset time interval in the following manner: Obtaining the start time and end time of the time interval; Determine a starting video segment in the historical video corresponding to the starting time; Determine a termination video segment in the historical video corresponding to the termination time; Acquire an intermediate video segment between the start video segment and the end video segment from the historical video; Based on a preset time interval, a plurality of key frame images are acquired from the start video segment, the end video segment and the middle video segment.
[0048] Among them, by implementing this implementation method, by determining the start and end time of the time interval and accurately locating the corresponding start video segment and end video segment, the method ensures the timeliness of key frame extraction. Furthermore, by including the processing of intermediate video segments, the method can fully cover the entire time interval of interest and avoid information omission. Extracting key frame images from all relevant video segments based on preset time intervals not only improves processing efficiency, but also ensures the representativeness of key frames, so that the key frame images finally obtained can accurately reflect the important content and changes in the historical video, providing strong support for subsequent analysis, monitoring or review work.
[0049] As an optional implementation manner, the third acquisition unit 203 acquires the target key frame image matching the target object information from the multiple key frame images in the following manner: Performing type analysis on the target object image to determine the image type of the target object image; wherein the image type includes a person type and an object type; If the image type is the person type, performing face recognition on the target object image to obtain target face features of the target object image; Screening from the multiple key frame images to obtain a face key frame image containing face features; Performing face recognition on the face key frame image to obtain candidate face features of the face key frame image; Comparing the candidate facial features with the target facial features to obtain key facial features; The face key frame image corresponding to the key face feature is determined as the target key frame image.
[0050] Among them, implementing this embodiment, By performing type analysis on the target object image, the method can intelligently identify whether the target object is a person or an object, so as to take targeted follow-up processing measures. This type analysis not only improves the accuracy of processing, but also avoids unnecessary resource consumption. Secondly, when the target object is a person, the target facial features are extracted by face recognition technology, and compared with the facial features in multiple key frame images. The method can accurately screen out the facial key frame images that match the target object. This process is not only efficient, but also greatly improves the accuracy of the target key frame image acquisition. In addition, by comparing the candidate facial features with the target facial features, the method can further confirm the key facial features, thereby ensuring that the final target key frame image is highly matched with the target object. This meticulous comparison process enhances the reliability of the results, so that the target key frame image can truly reflect the actual situation of the target object.
[0051] To sum up, this implementation method realizes the efficient and accurate acquisition of target key frame images matching the target object information from multiple key frame images through a series of intelligent processing steps such as type analysis, face recognition and feature comparison, providing strong support for subsequent image analysis, monitoring and tracking and other tasks.
[0052] As an optional implementation manner, the third obtaining unit 203 is further configured to: If the image type is the object type, performing feature recognition on the target object image to obtain target object features of the target object image; Screening the multiple key frame images to obtain an object key frame image containing object features; Performing feature recognition on the object key frame image to obtain candidate object features of the object key frame image; Comparing the candidate object features with the target object features to obtain key object features; The object key frame image corresponding to the key object feature is determined as the target key frame image.
[0053] Among them, by implementing this implementation method, the target object features in the target object image are accurately extracted through feature recognition technology. This step ensures the baseline accuracy of subsequent analysis. Subsequently, images containing similar object features are screened in multiple key frame images, which not only narrows the processing scope but also improves the processing efficiency. Furthermore, feature recognition is performed on the screened object key frame images and a detailed comparison is made with the target object features. This process ensures that the final target key frame image is highly consistent with the target object. Therefore, this implementation method not only improves the accuracy of acquiring the target key frame image of the object type, but also expands the application scenarios of the method, so that it can be widely used in various fields such as object tracking, scene reconstruction, object recognition, etc., showing strong practical value and generalization ability.
[0054] As an optional implementation manner, the third acquisition unit 203 acquires the target key frame image matching the target object information from the multiple key frame images in the following manner: Detecting whether there is a target key frame image matching the target object information among the multiple key frame images, and obtaining a detection result; If the detection result indicates that there is a target key frame image matching the target object information in the multiple key frame images, then executing the step of acquiring a target video clip corresponding to the target key frame from the historical video; If the detection result indicates that there is no target key frame image matching the target object information in the multiple key frame images, then obtaining a target video segment other than the start video segment, the end video segment and the middle video segment in the historical video; Based on a preset time interval, acquiring multiple frames of current key frame images from the target video clip; A target key frame image matching the target object information is acquired from the multiple frames of current key frame images.
[0055] Among them, by implementing this implementation method, by preliminarily detecting whether there is a target key frame image matching the target object information in the multiple key frame images, the method can quickly determine whether further processing is required, thereby avoiding unnecessary resource consumption. When a match is detected, the subsequent steps are directly executed to obtain the target key frame image, thereby improving processing efficiency. When the preliminary detection fails to find a match in the multiple key frame images, the method does not give up directly, but further mines potential information from historical videos. By obtaining the target video clips other than the start, end and middle video clips, and extracting multiple frames of current key frame images therefrom based on a preset time interval, the method expands the search range and increases the possibility of finding the target key frame image.
[0056] As an optional implementation manner, the fourth obtaining unit 204 is further configured to: Obtain the video identification of the target video clip; Sorting the video identifications of the target video clips in the order of video recording time to obtain a video identification sequence; The video identification sequence is outputted so that a user can obtain a target video segment corresponding to the video identification based on the video identification sequence.
[0057] Among them, by implementing this implementation method, by obtaining the video identification of the target video clip, and sorting these identifications according to the order of the video recording time, a video identification sequence is obtained. This method not only helps the user to clearly understand the time sequence of each target video clip, but also greatly simplifies the process of the user to find and obtain a specific video clip. The user only needs to quickly locate the video identification corresponding to the required video clip according to the output video identification sequence, and then obtain the target video clip. The addition of this step not only improves the user experience, but also makes the management and retrieval of video clips more intuitive and efficient, saving a lot of time and energy for the user.
[0058] The present invention can meet basic video retrieval requirements and improve search efficiency and accuracy. In addition, the present invention can also enable the key frame images finally obtained to accurately reflect the important content and changes in historical videos. In addition, the present invention can also efficiently and accurately obtain target key frame images that match target object information from multiple frames of key frame images. In addition, the present invention can also improve the accuracy of obtaining target key frame images of object types. In addition, the present invention can also increase the possibility of finding target key frame images. In addition, the present invention can also improve the user experience, making the management and retrieval of video clips more intuitive and efficient, saving users a lot of time and energy.
[0059] Exemplary Media After introducing the method and apparatus of the exemplary embodiment of the present invention, next, reference is made to Figure 3 For a description of a computer-readable storage medium according to an exemplary embodiment of the present invention, please refer to Figure 3 , the computer-readable storage medium shown is a CD 30, on which a computer program (i.e., a program product) is stored. When the computer program is executed by the processor, it will implement the steps recorded in the above method implementation, for example, it can obtain target object information; based on a preset time interval, obtain multiple key frame images from pre-acquired historical video recordings; obtain target key frame images matching the target object information from the multiple key frame images; obtain a target video clip corresponding to the target key frame from the historical video recording; the specific implementation method of each step will not be repeated here.
[0060] It should be noted that examples of the computer-readable storage medium may also include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical or magnetic storage media, which are not listed here one by one.
[0061] Exemplary Computing Devices After introducing the method, apparatus and medium of the exemplary embodiments of the present invention, next, reference is made to Figure 4 A computing device for searching a target video according to an exemplary embodiment of the present invention.
[0062] Figure 4 A block diagram of an exemplary computing device 40 suitable for implementing embodiments of the present invention is shown, and the computing device 40 may be a computer system or a server. Figure 4 The computing device 40 shown is only an example and should not bring any limitation to the functionality and scope of use of the embodiments of the present invention.
[0063] like Figure 4 As shown, the components of the computing device 40 may include, but are not limited to: one or more processors or processing units 401 , a system memory 402 , and a bus 403 connecting different system components (including the system memory 402 and the processing unit 401 ).
[0064] The computing device 40 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computing device 40, including volatile and non-volatile media, removable and non-removable media.
[0065] The system memory 402 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 4021 and / or cache memory 4022. The computing device 40 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, the ROM 4023 may be used to read and write non-removable, non-volatile magnetic media ( Figure 4 is not shown in the Figure 4As shown in FIG. 4 , a disk drive for reading and writing a removable non-volatile disk (e.g., a “floppy disk”) and an optical disk drive for reading and writing a removable non-volatile optical disk (e.g., a CD-ROM, a DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 403 via one or more data medium interfaces. System memory 402 may include at least one program product having a set (e.g., at least one) of program modules that are configured to perform the functions of various embodiments of the present invention.
[0066] A program / utility 4025 having a set (at least one) of program modules 4024 may be stored, for example, in system memory 402, and such program modules 4024 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment. Program modules 4024 generally perform the functions and / or methods of the embodiments described herein.
[0067] The computing device 40 may also communicate with one or more external devices 404 (e.g., a keyboard, a pointing device, a display, etc.). Such communication may be performed via an input / output (I / O) interface 405. Furthermore, the computing device 40 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 406. Figure 4 As shown, the network adapter 406 communicates with other modules (such as the processing unit 401, etc.) of the computing device 40 via the bus 403. It should be understood that although Figure 4 Not shown, other hardware and / or software modules may be used in conjunction with computing device 40 .
[0068] The processing unit 401 executes various functional applications and data processing by running the program stored in the system memory 402, for example, it can obtain target object information; obtain multiple key frame images from the pre-acquired historical video based on a preset time interval; obtain the target key frame image matching the target object information from the multiple key frame images; obtain the target video clip corresponding to the target key frame from the historical video. The specific implementation method of each step is not repeated here. It should be noted that although several units / modules or sub-units / sub-modules of the target video search device are mentioned in the above detailed description, this division is only exemplary and not mandatory. In fact, according to an embodiment of the present invention, the features and functions of two or more units / modules described above can be concretized in one unit / module. Conversely, the features and functions of one unit / module described above can be further divided into multiple units / modules to be concretized.
[0069] In the description of the present invention, it should be noted that the terms “first”, “second” and “third” are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0070] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0071] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some communication interfaces, and the indirect coupling or communication connection of devices or units can be electrical, mechanical or other forms.
[0072] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0073] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0074] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a non-volatile computer-readable storage medium that can be executed by a processor. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., various media that can store program codes.
[0075] Finally, it should be noted that the above-described embodiments are only specific implementations of the present invention, which are used to illustrate the technical solutions of the present invention, rather than to limit them. The protection scope of the present invention is not limited thereto. Although the present invention is described in detail with reference to the above-described embodiments, ordinary technicians in the field should understand that any technician familiar with the technical field can still modify the technical solutions recorded in the above-described embodiments within the technical scope disclosed by the present invention, or can easily think of changes, or make equivalent replacements for some of the technical features therein; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
[0076] In addition, although the operations of the method of the present invention are described in a specific order in the drawings, this does not require or imply that the operations must be performed in this specific order, or that all the operations shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.
[0077] In an exemplary embodiment, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, the steps in the above method embodiments are implemented.
Claims
1. A method for searching a target video, characterized in that: include: Get target object information; Based on a preset time interval, a plurality of key frame images are obtained from the pre-acquired historical video recordings; Acquire a target key frame image matching the target object information from the multiple key frame images; A target video segment corresponding to the target key frame is obtained from the historical video.
2. The method for searching a target video according to claim 1, characterized in that: The target object information includes at least a time interval, the historical video includes a plurality of video clips, the duration of each video clip is a preset duration, the plurality of video clips are sorted in the order of recording time, and the acquisition of a plurality of key frame images from the pre-acquired historical video based on the preset time interval includes: Obtaining the start time and end time of the time interval; Determine a starting video segment in the historical video corresponding to the starting time; Determine a termination video segment in the historical video corresponding to the termination time; Acquire an intermediate video segment between the start video segment and the end video segment from the historical video; Based on a preset time interval, a plurality of key frame images are acquired from the start video segment, the end video segment and the middle video segment.
3. The method for searching a target video according to claim 2, characterized in that: The target object information also includes a target object image, and acquiring a target key frame image matching the target object information from the multiple key frame images includes: Performing type analysis on the target object image to determine the image type of the target object image; wherein the image type includes a person type and an object type; If the image type is the person type, performing face recognition on the target object image to obtain target face features of the target object image; Screening from the multiple key frame images to obtain a face key frame image containing face features; Performing face recognition on the face key frame image to obtain candidate face features of the face key frame image; Comparing the candidate facial features with the target facial features to obtain key facial features; The face key frame image corresponding to the key face feature is determined as the target key frame image.
4. The method for searching a target video according to claim 3, characterized in that: If the image type is the object type, the method further includes: Performing feature recognition on the target object image to obtain target object features of the target object image; Screening the multiple key frame images to obtain an object key frame image containing object features; Performing feature recognition on the object key frame image to obtain candidate object features of the object key frame image; Comparing the candidate object features with the target object features to obtain key object features; The object key frame image corresponding to the key object feature is determined as the target key frame image.
5. The method for searching a target video according to claim 2, characterized in that: The step of acquiring a target key frame image matching the target object information from the multiple key frame images includes: Detecting whether there is a target key frame image matching the target object information among the multiple key frame images, and obtaining a detection result; If the detection result indicates that there is a target key frame image matching the target object information in the multiple key frame images, then executing the step of acquiring a target video clip corresponding to the target key frame from the historical video; If the detection result indicates that there is no target key frame image matching the target object information in the multiple key frame images, then obtaining a target video segment other than the start video segment, the end video segment and the middle video segment in the historical video; Based on a preset time interval, acquiring multiple frames of current key frame images from the target video clip; A target key frame image matching the target object information is acquired from the multiple frames of current key frame images.
6. The method for searching a target video according to claim 1, characterized in that: The method further comprises: Obtain the video identification of the target video clip; Sorting the video identifications of the target video clips in the order of video recording time to obtain a video identification sequence; The video identification sequence is outputted so that a user can obtain a target video segment corresponding to the video identification based on the video identification sequence.
7. A device for searching a target video, characterized in that: include: A first acquisition unit, used to acquire target object information; A second acquisition unit, configured to acquire a plurality of key frame images from pre-acquired historical video recordings based on a preset time interval; A third acquisition unit, configured to acquire a target key frame image matching the target object information from the multiple key frame images; The fourth acquisition unit is used to acquire a target video segment corresponding to the target key frame from the historical video.
8. A computing device, comprising: at least one processor, memory, and input-output unit; The memory is used to store a computer program, and the processor is used to call the computer program stored in the memory to execute the method according to any one of claims 1 to 6.
9. A computer-readable storage medium comprising instructions, which, when executed on a computer, enables the computer to execute the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.