Video retrieval method and device, electronic equipment and storage medium

By binding facial features from different orientations and matching them with pre-generated video content, the problem of accuracy and completeness in facial feature retrieval under complex environments is solved, enabling real-time and accurate acquisition of video content.

CN121479014APending Publication Date: 2026-02-06ZHEJIANG UNIVIEW TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511150923.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-18
Publication Date
2026-02-06

AI Technical Summary

Technical Problem

In complex spatial environments, due to limitations in camera installation angle, height, and coverage, it is difficult to guarantee the acquisition of standard facial images in all scenarios, resulting in reduced accuracy and completeness of video clip retrieval based on facial features.

Method used

By acquiring image information of the target to be retrieved, detecting and binding facial features from different orientations, generating multi-dimensional facial feature combinations, and matching them using pre-generated video content, the feature matching failure caused by angle differences is reduced, ensuring the comprehensiveness and accuracy of image retrieval.

Benefits of technology

It improves the robustness of facial feature matching, ensures the integrity and coherence of video retrieval, avoids video content mismatch or omission, and achieves real-time and accurate video content acquisition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121479014A_ABST
    Figure CN121479014A_ABST
Patent Text Reader

Abstract

The invention discloses a video retrieval method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring first information, wherein the first information comprises an image of a snapshot target to be retrieved; detecting the feature similarity between the snapshot target facial feature in the first information and the snapshot target facial feature in at least one group of second information associated with the first shooting equipment; in response to the fact that the feature similarity between the snapshot target facial feature in the third information and the snapshot target facial feature in the first information in the at least one group of second information meets a preset similarity condition, determining second video content from the multiple pieces of first video content according to the snapshot target facial feature in the third information, each first video content of the plurality of first video contents is pre-generated according to a video clip obtained by snapshot of the same snapshot target by at least one second shooting device. According to the scheme, the problem of low accuracy during image retrieval can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of image processing technology, and in particular to a video retrieval method, apparatus, electronic device, and storage medium. Background Technology

[0002] With the development of short video technology, facial image-based search and matching technology has been widely applied in the field of video clip retrieval. Specifically, it uses self-taken frontal facial images as the retrieval basis to quickly locate video clips containing specific facial features. However, in practical applications, due to limitations in scene deployment conditions, in complex spatial environments, due to limitations in camera installation angle, height, and coverage, it is difficult to guarantee that standard facial images can be captured in all scenarios. Non-standard facial image data has feature differences from self-taken facial images, which greatly reduces the accuracy of facial recognition when performing image retrieval based on facial features. This, in turn, seriously affects the completeness and accuracy of video clip retrieval, making it impossible to effectively retrieve a large amount of image material containing the target face. Summary of the Invention

[0003] This invention provides a video retrieval method, apparatus, electronic device, and storage medium to solve the problem of low accuracy in image retrieval.

[0004] According to one aspect of the present invention, a video retrieval method is provided, the method comprising:

[0005] Obtain first information, which is an image including the target to be captured;

[0006] The similarity of features between the facial features of the captured target in the first information and the facial features of the captured target in at least one set of second information associated with the first shooting device is detected. The same set of second information includes a first type of captured target facial features and a second type of captured target facial features for the same captured target. The angle between the facial orientation of the captured target and the optical axis of the shooting device corresponding to the first type of captured target facial features is greater than a preset angle. The angle between the facial orientation of the captured target and the optical axis of the shooting device corresponding to the second type of captured target facial features is less than a preset angle. The second type of captured target facial features are facial features generated by simulating and adjusting the facial orientation of the captured target based on the first type of captured target facial features.

[0007] In response to the fact that the facial features of the captured target in the third information and the facial features of the captured target in the first information satisfy a preset similarity condition in the at least one set of second information, a second video content is determined from multiple first video contents based on the facial features of the captured target in the third information. Each of the multiple first video contents is pre-generated based on video segments obtained by capturing the same captured target by at least one second shooting device. Each first video content is associated with a set of captured target facial features in the second information. The second shooting device is used to capture the captured target before the first shooting device.

[0008] According to another aspect of the present invention, a video retrieval device is provided, the device comprising:

[0009] The acquisition module is used to acquire first information, which includes the target to be captured.

[0010] The detection module is used to detect the feature similarity between the facial features of the captured target in the first information and the facial features of the captured target in at least one set of second information associated with the first shooting device. The same set of second information includes a first type of captured target facial features and a second type of captured target facial features for the same captured target. The angle between the facial orientation of the captured target and the optical axis of the shooting device corresponding to the first type of captured target facial features is greater than a preset angle. The angle between the facial orientation of the captured target and the optical axis of the shooting device corresponding to the second type of captured target facial features is less than a preset angle. The second type of captured target facial features are facial features generated by simulating and adjusting the facial orientation of the captured target based on the first type of captured target facial features.

[0011] The retrieval module is configured to, in response to the fact that the facial features of the captured target in the third information and the facial features of the captured target in the first information satisfy a preset similarity condition in the at least one set of second information, determine second video content from multiple first video contents based on the facial features of the captured target in the third information. Each of the multiple first video contents is pre-generated based on video segments obtained by capturing the same captured target by at least one second shooting device. Each first video content is associated with a set of captured target facial features in the second information. The second shooting device is used to capture the captured target before the first shooting device.

[0012] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:

[0013] At least one processor; and

[0014] A memory communicatively connected to the at least one processor; wherein,

[0015] The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the video retrieval method according to any embodiment of the present invention.

[0016] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the video retrieval method according to any embodiment of the present invention.

[0017] According to another aspect of the present invention, a computer program product is provided, the computer program product comprising a computer program that, when executed by a processor, implements the video retrieval method described in any embodiment of the present invention.

[0018] The technical solution of this embodiment acquires first information including the target to be captured. When performing image retrieval by feature comparison on the captured target, each set of second information associated with the first capturing device simultaneously includes facial features of both the first type and the second type of captured target. By distinguishing captured images with different facial orientations, the coverage of facial feature comparison is expanded, and the robustness of facial feature matching is improved. Even if captured targets with varying facial orientations appear in the first information, feature complementarity can be achieved through captured targets with different facial orientations in each set of second information, thereby improving the accuracy of feature similarity detection. To reduce feature matching failures caused by angle differences and ensure the comprehensiveness of image retrieval results; since the pre-synthesized video has integrated the first video content including the same captured target based on the capture of the second shooting device before the first shooting device, there is no need for manual screening or editing. The corresponding pre-generated second video content can be directly matched from multiple first video contents through feature similarity. This enables the rapid and accurate location of the second video content matching the first information from the pre-synthesized video content after obtaining the first information, directly obtaining complete and coherent video content, avoiding video content mismatch or omission, and ensuring the consistency between the obtained video content and the movement process of the captured target.

[0019] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0020] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0021] Figure 1 This is a flowchart of a video retrieval method provided according to an embodiment of the present invention;

[0022] Figure 2 This is a flowchart of another video retrieval method provided according to an embodiment of the present invention;

[0023] Figure 3 This is a flowchart illustrating the pre-synthesis of video clips captured by a shooting device, applicable to embodiments of the present invention.

[0024] Figure 4 This is a schematic diagram of a prompt message template applicable to an embodiment of the present invention;

[0025] Figure 5 This is a schematic diagram illustrating the generation and feature binding of different types of facial features of captured targets according to an embodiment of the present invention;

[0026] Figure 6 This is a schematic diagram of the structure of a video retrieval device according to an embodiment of the present invention;

[0027] Figure 7 This is a schematic diagram of the structure of an electronic device that implements the video retrieval method of this invention. Detailed Implementation

[0028] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0029] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0030] Figure 1 The present invention provides a flowchart of a video retrieval method. This embodiment is applicable to situations where facial features of a first type of captured target and facial features of a second type of captured target are bound together for joint use in image retrieval containing captured targets. This method can be executed by a video retrieval device, which can be implemented in hardware and / or software and can be configured in any electronic device with network communication capabilities.

[0031] like Figure 1 As shown, the video retrieval method provided in this embodiment of the invention may include the following process:

[0032] S110. Obtain first information, which is an image including the target to be captured.

[0033] Multiple shooting devices can be configured in the task operation area. For each target to be captured, the task can be completed by passing through different shooting devices sequentially within the task operation area. When a target enters the shooting field of view of each shooting device, the device will capture the target within its own shooting field of view. This allows for the capture of the same target through different shooting devices. The target to be captured can be a specific object that is to be captured and recorded in the image or video acquisition scene. The target to be captured needs to have unique features to ensure that it can be detected and distinguished. For example, the target to be captured can be a pedestrian or animal that has been pre-authorized for image acquisition.

[0034] When a task is completed in the task operation area, if you want to obtain a video of a captured target captured by different shooting devices, you can obtain the first information uploaded. The first information contains the captured target to be retrieved for image retrieval. In the subsequent image retrieval, the facial features of the captured target to be retrieved can be extracted from the first information. Based on the extracted facial features, the video content related to the captured target can be searched from the videos acquired by various shooting devices.

[0035] S120. Detect the feature similarity between the facial features of the captured target in the first information and the facial features of the captured target in at least one set of second information associated with the first shooting device. The same set of second information includes a first type of captured target facial features and a second type of captured target facial features for the same captured target. The angle between the facial orientation of the captured target and the optical axis of the shooting device corresponding to the first type of captured target facial features is greater than a preset angle, and the angle between the facial orientation of the captured target and the optical axis of the shooting device corresponding to the second type of captured target facial features is less than a preset angle. The second type of captured target facial features are facial features generated by simulating and adjusting the facial orientation of the captured target based on the first type of captured target facial features.

[0036] Automatically generating videos of the same captured target during task execution within the task operation area has become an important feature for improving the user experience. However, relying on the real-time capture target retrieval of video clips often leads to queuing delays due to excessive server processing pressure. This is especially true when a large number of capture targets need to be simultaneously triggered to generate video clips. The video clips of the corresponding capture targets need to be retrieved segment by segment from a massive amount of video clips, and then edited and rendered, which can take several minutes or even longer. This seriously affects the real-time sharing experience of videos during task execution.

[0037] Considering that among the multiple shooting devices in the task operation area, the same target appears sequentially within the shooting angle of each shooting device, meaning that the same target appears in a certain order within the shooting angle of each shooting device, the facial features of the target associated with the first shooting device can be selected from multiple shooting devices to retrieve the video content when the second shooting device captures the target. The second shooting device is the shooting device that captures the target before the first shooting device according to the shooting order.

[0038] Different targets will appear within the shooting angle range of the first shooting device. Therefore, at least one set of second information associated with the first shooting device corresponds to different targets. That is, the facial features of the targets in different sets of second information are the facial features of the targets of different targets, while the facial features of the targets included in the same set of second information are the facial features of the targets of the same target. The facial features of the targets can refer to the key attributes or information that can be used to accurately locate, identify and distinguish the targets in an image or video capture scene. For example, the facial features of the targets can be facial features, and different targets have different facial features. Taking the facial features of the targets as an example, the facial features of the targets refer to the unique and recognizable physiological structures, shapes and details on the face. They are the key features for distinguishing individuals and realizing identity recognition. The facial features of the targets can include macroscopic facial contour features and / or microscopic facial features and their combination relationships.

[0039] In scenarios with a high density of targets, the capture accuracy of cameras is often affected by complex environments. Due to factors such as shooting angle deviations and changes in the target's posture, some cameras may capture images of targets that are not facing the camera directly. Although these images contain some facial features of the targets, the lack of these features makes it difficult to accurately match them with the targets in the initial information when used directly for target retrieval, resulting in a significant decrease in the accuracy of target retrieval.

[0040] Therefore, for capture images where the angle between the target's face and the optical axis of the capturing device is greater than a preset angle, this case addresses the issue of adjusting the facial features of the target to generate facial features highly consistent with the original target. This preserves the uniqueness of the target while supplementing key facial features. Consequently, for the same target, both Type I and Type II facial features will exist simultaneously. Type I facial features correspond to an angle greater than a preset angle between the target's face and the optical axis of the capturing device. In this case, due to the target's lateral rotation, some facial features are obscured and missing. Type II facial features correspond to an angle less than a preset angle between the target's face and the optical axis of the capturing device. In this case, the target faces the capturing device, and key facial features are not missing.

[0041] The facial features of the first type of captured target and the facial features of the second type of captured target generated for the same captured target are bound together to form a complete set of multi-feature combination data, forming a three-dimensional characterization of the captured target's facial features. This can significantly improve the retrieval accuracy of captured targets. Specifically, the binding process is not a simple image stitching, but rather the establishment of a captured target facial feature association index. Through the captured target facial feature matching algorithm, it is confirmed that the generated first type of captured target facial features and second type of captured target facial features belong to the same captured target. Subsequently, a unique identifier is assigned to each set of second information that can contain the first type of captured target facial features and second type of captured target facial features, and an association storage is established in the database to ensure that multi-dimensional features can be extracted synchronously during subsequent calls.

[0042] When the first information is received to retrieve the target to be captured, the facial features of the target to be captured are extracted from the first information. Then, a multi-dimensional feature similarity comparison is performed on the facial features of the target in at least one set of second information associated with the first shooting device in the database. That is, if the facial features of the target in the first information belong to the second type of facial features, the feature similarity comparison is performed with the second type of facial features of the target in at least one set of second information associated with the first shooting device. If the facial features of the target in the first information belong to the first type of facial features, the feature similarity comparison is performed with the first type of facial features of the target in at least one set of second information associated with the first shooting device. This avoids misjudgment caused by the angle deviation of the facial features of the target in the first information. This synergistic effect of the first type of facial features and the second type of facial features can effectively cover the changes in facial features of different types of targets and solve the limitations of facial feature matching caused by a single shooting angle.

[0043] From the perspective of practical application, the advantages of the scheme that combines the facial features of the first type and the second type of captured targets for target retrieval are as follows: First, the facial features of the captured targets are complementary. By generating facial features of the second type of captured targets, key information missing from the facial features of the first type of captured targets is supplemented, while retaining the true features of the captured targets. Second, robustness is improved. Even if the facial orientation angle of the captured target differs significantly from the capture angle of the shooting device, the combination of multi-dimensional facial features of the captured targets can still find matching anchor points, reducing the failure rate of target retrieval due to changes in shooting angle. Third, there is dynamic adaptability. As the number of captured images of the captured targets under different shooting devices increases, the facial features of the first type of captured targets will be dynamically updated, and the generated facial features of the second type of captured targets can also be iteratively optimized based on the new facial features of the captured targets.

[0044] S130. In response to the fact that the facial features of the target captured in the third information and the facial features of the target captured in the first information satisfy a preset similarity condition in at least one set of second information, the second video content is determined from multiple first video contents based on the facial features of the target captured in the third information. Each of the multiple first video contents is pre-generated based on video segments obtained by capturing the same target by at least one second shooting device. Each first video content is associated with a set of facial features of the target captured in the second information. The second shooting device is used to capture the target before the first shooting device.

[0045] Considering that when using the facial features of the captured target in the first information for target retrieval, if video generation is only triggered upon receiving the first information, the server would need to temporarily perform multiple intensive video processing tasks on multiple video segments acquired by various secondary shooting devices, including video segment extraction, target facial feature matching, and editing and rendering, which would significantly increase the target retrieval time. Therefore, this solution chooses pre-configured video generation and binding matching of multi-dimensional target facial features. This distributes the computational pressure generated by multiple intensive video processing tasks throughout the process of the captured target performing tasks in the task operation area, retaining only the final lightweight target facial feature comparison step, which is performed when the first information is received for target retrieval.

[0046] As each target enters the task operation area to perform its task, it sequentially enters the shooting field of view of each of the secondary shooting devices. Each secondary shooting device captures video clips of the target entering its field of view. The system continuously tracks the secondary shooting devices along the route the target travels, recording in real time the video clips captured when each target appears within the shooting field of view of each secondary shooting device (e.g., video footage from popular photo spots or amusement park rides). This allows the system to utilize the facial features of the target included in each set of secondary information to obtain video clips of the same target captured by at least one secondary shooting device. These video clips are then pre-synthesized, edited, and rendered to create a complete first video content, which is stored in the database. Crucially, when the same target passes the last shooting device in the task operation area (i.e., when it passes the first shooting device), the system can associate the first-type and second-type facial features of the same target captured by the first shooting device with the pre-synthesized, edited, and rendered complete first video content for the same target.

[0047] If the feature similarity between the captured facial features in the third information and the captured facial features in the first information meets the preset similarity condition, it means that the feature similarity between the captured facial features in the third information and the captured facial features in the first information is greater than the preset similarity. If the feature similarity between the captured facial features in the third information and the captured facial features in the first information does not meet the preset similarity condition, it means that the feature similarity between the captured facial features in the third information and the captured facial features in the first information is not greater than the preset similarity.

[0048] When at least one set of second information contains facial features of the target captured in the third information that satisfy a preset similarity condition with facial features of the target captured in the first information, it indicates that the facial features of the target captured in the third information and the facial features of the target captured in the first information are likely to belong to the same target captured. Therefore, the second video content containing the target captured by the target captured in the third information can be obtained directly from the first video content associated with the facial features of the target captured in the third information from the multiple pre-generated first video content.

[0049] When the feature similarity between the facial features of the captured target in the third information and the facial features of the captured target in the first information does not meet the preset similarity condition in at least one set of second information, it indicates that the facial features of the captured target in the second information do not belong to the same captured target as the facial features of the captured target in the first information. At this time, there is no captured target corresponding to the facial features of the captured target in the first information in the multiple first video contents generated in advance. Therefore, it is necessary to temporarily perform multiple intensive video processing tasks such as video segment extraction, captured target facial feature matching, and editing and rendering on the multiple video segments acquired by each second shooting device to obtain video content containing the captured target corresponding to the facial features of the captured target in the first information.

[0050] Using the above scheme, the pre-synthesis of multiple video clips belonging to the same target can be distributed throughout the target's task execution process. The final facial feature matching of the target requires only a short time, achieving instant matching even when a large number of targets are simultaneously being matched. Pre-generated video content has more time for refined editing, whereas traditional real-time synthesis often simplifies processing due to computational limitations, resulting in abrupt transitions between video clips. Video clip retrieval and synthesis are transformed from repetitive calculations to one-time calculations and multiple reuses. The pre-synthesized video only retains the video bound to the target's facial features, eliminating the need to back up the original video clips and reducing storage redundancy. Most importantly, the time spent from the first shooting device to receiving the first information is greater than the time for video pre-synthesis, preventing server queuing and allowing real-time viewing of the target's video content upon receiving the first information.

[0051] The technical solution of this embodiment acquires first information including the target to be captured. When performing image retrieval by feature comparison on the captured target, each set of second information associated with the first capturing device simultaneously includes facial features of both the first type and the second type of captured target. By distinguishing captured images with different facial orientations, the coverage of facial feature comparison is expanded, and the robustness of facial feature matching is improved. Even if captured targets with varying facial orientations appear in the first information, feature complementarity can be achieved through captured targets with different facial orientations in each set of second information, thereby improving the accuracy of feature similarity detection. To reduce feature matching failures caused by angle differences and ensure the comprehensiveness of image retrieval results; since the pre-synthesized video has integrated the first video content including the same captured target based on the capture of the second shooting device before the first shooting device, there is no need for manual screening or editing. The corresponding pre-generated second video content can be directly matched from multiple first video contents through feature similarity. This enables the rapid and accurate location of the second video content matching the first information from the pre-synthesized video content after obtaining the first information, directly obtaining complete and coherent video content, avoiding video content mismatch or omission, and ensuring the consistency between the obtained video content and the movement process of the captured target.

[0052] Figure 2 This is a flowchart illustrating another video retrieval method provided by an embodiment of the present invention. The technical solution of this embodiment is an optimization of the aforementioned embodiments based on the technical solutions of the above embodiments. This embodiment can be combined with various optional solutions in one or more of the above embodiments.

[0053] like Figure 2 As shown, the video retrieval method provided in this embodiment of the invention may include the following process:

[0054] S210. For each set of second information associated with the first shooting device, based on the first type of capture target facial features in each set of second information, determine at least one first capture target facial feature associated with each set of second information from the capture target facial features associated with each second shooting device. Different first capture target facial features in the at least one first capture target facial feature are capture target facial features associated with different second shooting devices. The feature similarity between each first capture target facial feature and the first type of capture target facial feature in each set of second information is greater than the feature similarity between the second capture target facial feature and the first type of capture target facial feature in each set of second information. The second capture target facial feature is the remaining capture target facial features other than the first capture target facial feature among the capture target facial features associated with the second shooting device associated with the first capture target facial feature.

[0055] For at least one set of second information associated with the first shooting device, each set of second information includes facial features of the first type of capture target and facial features of the second type of capture target. These are retrieved from the facial features of the capture targets captured by each of the previous second shooting devices. First, the facial features of the first type of capture target in the second information are used to search for facial features of the capture targets associated with each second shooting device. Then, based on the facial feature similarity, the facial feature with the highest feature similarity among the facial features associated with each second shooting device is selected and retained. At this point, multiple first-type capture target facial features are obtained by retrieving them from each second shooting device using different first-type capture target facial features in each set of second information. The result is as follows: Figure 3 Channl1_RecordID1_search, Channl1_RecordID2_search, Channl1_RecordID3_search, etc., belong to the search for multiple first captured target facial features by using different first-type captured target facial features in each group of second information to retrieve the captured target facial features associated with each second shooting device marked Channl1.

[0056] As an optional but not limited implementation, each set of second information includes multiple first-type target facial features; based on the first-type target facial features in each set of second information, at least one first target facial feature associated with each set of second information is determined from the target facial features associated with each second shooting device, including but not limited to the following steps A1-A2:

[0057] Step A1: For each first type of captured target facial feature in each group of second information, determine at least one fifth captured target facial feature associated with each first type of captured target facial feature from the captured target facial features associated with each second shooting device. Different fifth captured target facial features in the at least one fifth captured target facial feature are captured target facial features associated with different second shooting devices. The feature similarity between each fifth captured target facial feature and the first type of captured target facial feature in each group of second information is greater than the feature similarity between the sixth captured target facial feature and the first type of captured target facial feature in each group of second information. The sixth captured target facial feature is the remaining captured target facial features other than the fifth captured target facial feature among the captured target facial features associated with the second shooting device associated with the fifth captured target facial feature.

[0058] Step A2: Determine the seventh capture target facial feature of each first group from the different fifth capture target facial features of each first group. The different fifth capture target facial features of the same first group are capture target facial features associated with the same second shooting device selected from at least one fifth capture target facial feature associated with each of the multiple first type capture target facial features. The feature similarity between the seventh capture target facial feature of each first group and the first type capture target facial feature of each group of second information is greater than the feature similarity between the eighth capture target facial feature and the first type capture target facial feature of each group of second information. The eighth capture target facial feature is the remaining capture target facial feature of each first group other than the seventh capture target facial feature.

[0059] See Figure 3 Each set of second information includes multiple first-type target facial features. Using each first-type target facial feature, a search for target facial features is performed among the target facial features associated with each second shooting device. Each fifth target facial feature must originate from a different second shooting device; that is, each second shooting device contributes at most one fifth target facial feature. The feature similarity between the fifth target facial feature and the first-type target facial features in each set of second information must be higher than the feature similarity between other unselected sixth target facial features in the same second shooting device and the first-type target facial features in each set of second information. Only the target facial feature associated with each second shooting device that is most similar to the first-type target facial features in each set of second information can become the fifth target facial feature.

[0060] See Figure 3 The first group refers to grouping the selected fifth-type target facial features. Each first group must contain features from the same second imaging device, selected from the matching results of multiple first-type target facial features. In other words, it represents the optimal capture result from the same second imaging device for multiple first-type target facial features in the second information. From the multiple fifth-type target facial features in each first group, a seventh-type target facial feature is selected. The feature similarity between the seventh-type target facial feature and the first-type target facial features in each group of second information must be higher than the feature similarity between the other unselected eighth-type target facial features in the same first group and the first-type target facial features in each group of second information.

[0061] The above scheme employs a two-layer screening process: first, the optimal fifth captured facial feature from each second capturing device is selected; then, the optimal sixth captured facial feature from each first group is selected. This ensures that the final associated captured facial features are those with the highest similarity, reducing interference from irrelevant features. Furthermore, the fifth captured facial feature is required to originate from facial features associated with different second capturing devices, ensuring data coverage across multiple scenes and angles and minimizing misjudgments caused by data bias from a single second capturing device. Simultaneously, a seventh captured facial feature is selected through grouping, further focusing on the most critical, highly correlated captured facial features and reducing data processing volume.

[0062] S220. Based on the second type of target facial features in each group of second information, determine at least one third target facial feature associated with each group of second information from the target facial features associated with each second shooting device. Different third target facial features in the at least one third target facial feature are target facial features associated with different second shooting devices. The feature similarity between each third target facial feature and the second type of target facial features in each group of second information is greater than the feature similarity between the fourth target facial feature and the second type of target facial features in each group of second information. The fourth target facial feature is the remaining target facial features other than the third target facial features among the target facial features associated with the second shooting device associated with the third target facial features.

[0063] See Figure 3 In addition to the first type of facial features captured in the second information, the second type of facial features captured in each group of the second information are also used to search for the first type of facial features captured in each group of the second shooting devices and the second type of facial features captured in each group of the second information. This process is repeated for all the second shooting devices. The feature similarity between all facial features captured in each group of the second information and the second type of facial features captured in each group of the second information is calculated. Based on the facial feature similarity, the search result with the highest feature similarity between the first type of facial features captured in each group of the second shooting devices and the second type of facial features captured in each group of the second information is selected and retained. This results in two sets of search results, as shown below. Figure 3 Channl1_CreatID_search_Z, Channl1_CreatID_search_S, etc., belong to the category of using each second type of captured target facial feature in each group of second information to retrieve multiple third captured target facial features from the corresponding first type captured target facial features and second type captured target facial features associated with each second shooting device marked Channl1.

[0064] By adopting the above scheme, only the result that is most similar to the facial features of the second type of captured target in each second shooting device is retained, while interference items with lower similarity within the same second shooting device are excluded, thereby reducing the impact of redundant information on the matching results and ensuring the reliability of the associated features.

[0065] As an optional but not limited implementation, based on the second type of target facial features in each set of second information, at least one third target facial feature associated with each set of second information is determined from the target facial features associated with each second shooting device, including but not limited to the following steps B1-B3:

[0066] Step B1: Based on the second type of target facial features in each group of second information, determine at least one ninth target facial feature from the first type of target facial features associated with each second shooting device. The angle between the facial orientation of the target corresponding to the first type of target facial feature and the optical axis of the shooting device is greater than a preset angle. Different ninth target facial features in at least one ninth target facial feature are first type target facial features associated with different second shooting devices. The feature similarity between each ninth target facial feature and the second type of target facial features in each group of second information is greater than the feature similarity between the tenth target facial feature and the second type of target facial features in each group of second information. The tenth target facial feature is the remaining target facial features other than the ninth target facial feature in the first type of target facial features associated with the second shooting device associated with the ninth target facial feature.

[0067] Step B2: Based on the second type of target facial features in each group of second information, determine at least one eleventh target facial feature from the second type of target facial features associated with each second shooting device. The angle between the facial orientation of the target corresponding to the second type of target facial feature and the optical axis of the shooting device is less than a preset angle. The different eleventh target facial features in the at least one eleventh target facial feature are second type target facial features associated with different second shooting devices. The feature similarity between each eleventh target facial feature and the second type of target facial features in each group of second information is greater than the feature similarity between the twelfth target facial feature and the second type of target facial features in each group of second information. The twelfth target facial feature is the remaining target facial features other than the eleventh target facial feature among the second type of target facial features associated with the second shooting device associated with the eleventh target facial feature.

[0068] Step B3: Based on at least one ninth capture target facial feature and at least one eleventh capture target facial feature of the second type of capture target facial features in each group of second information, determine at least one third capture target facial feature associated with each group of second information. The at least one third capture target facial feature is the intersection result of the capture target facial features of at least one ninth capture target facial feature and at least one eleventh capture target facial feature.

[0069] See Figure 3 The first type of captured target facial features has a clear orientation; the angle between the target's facial orientation and the optical axis of the shooting device is greater than a preset angle, meaning that the first type of captured target facial features are not typical facial features from frontal or other conventional angles. The selection of the ninth captured target facial features follows strict standards. At least one different feature in the ninth captured target facial feature comes from the first type of captured target facial features associated with different second shooting devices, ensuring the diversity and independence of feature sources. Furthermore, the feature similarity between each ninth captured target facial feature and the second type of captured target facial features in each set of second information must be greater than the feature similarity between the tenth captured target facial feature and that second type of captured target facial feature. The tenth captured target facial feature is the remaining captured target facial features other than the ninth target facial feature among the first type of captured target facial features associated with the second shooting device associated with the ninth target facial feature. This screening process is equivalent to selecting the facial features of the target captured by the first type of each relevant second shooting device that are most similar to the facial features of the target captured by the second type of each set of second information, thus eliminating the interference of other facial features of the target captured by the same second shooting device that have low similarity, and providing a basis for the accuracy of subsequent search results.

[0070] See Figure 3The second type of captured target facial features is characterized by the fact that the angle between the facial orientation of the captured target and the optical axis of the shooting device is smaller than a preset angle, making it closer to capturing a frontal face or other conventional angles. Similar to the ninth captured target facial features, at least one eleventh captured target facial feature comes from different second-type captured target facial features associated with different second shooting devices. Furthermore, the feature similarity between each eleventh captured target facial feature and the second-type captured target facial features in each group of second information is greater than the feature similarity between the twelfth captured target facial feature and that second-type captured target facial feature. Here, the twelfth captured target facial feature is the remaining captured target facial features other than the eleventh captured target facial feature among the second-type captured target facial features associated with the second shooting device associated with the eleventh captured target facial feature. This process selects the feature with the highest similarity from the second-type captured features of each related second shooting device, further enriching the source of highly similar captured target facial features.

[0071] See Figure 3 Based on at least one ninth and at least one eleventh facial feature corresponding to the second type of captured facial features in each set of second information, at least one third facial feature associated with each set of second information is determined. This at least one third facial feature is the intersection of the facial features of at least one ninth and at least one eleventh captured facial feature. This means that only facial features present in both the ninth and eleventh captured facial features can be considered third facial features. This intersection method can be considered a double verification, excluding facial features that perform well in one type of captured facial feature comparison but poorly in another. This maximizes the matching between the determined third facial features and the second type of captured facial features in each set of second information, further improving the accuracy and reliability of the entire retrieval process.

[0072] S230. Based on at least one first capture target facial feature associated with each group of second information and at least one third capture target facial feature associated with each group of second information, generate first video content associated with each group of second information; there is an association and binding relationship between the capture target facial features associated with each second shooting device and the video segments acquired by each second shooting device.

[0073] When processing the retrieval results for generated capture targets, the retrieval result with the highest feature similarity is first selected, and the intersection of the two sets of results is performed. The core value of this operation is that, since the generated capture target is not an image of the real capture target, its retrieval results will inevitably have some differences from the real capture target. Taking the intersection can accurately retain the capture targets that both appeared in the previous cross-search, which is equivalent to double confirmation. Those results that differed in the cross-search are excluded, thus effectively ensuring the accuracy of the retrieval results. After such filtering, the retrieval results for different capture targets are finally obtained in two sets. For these last two sets of results, the retrieval result with the highest feature similarity is selected and retained, and then the union is performed. Retaining the retrieval result of the capture target with the highest feature similarity is to further improve the accuracy of the retrieval, while taking the union allows the retrieval results of the two sets of capture targets to complement each other, maximizing the completeness of the retrieval.

[0074] As an optional but not limited implementation, the first video content associated with each set of second information is generated based on at least one first captured target facial feature associated with each set of second information and at least one third captured target facial feature associated with each set of second information, including but not limited to the following steps C1-C3:

[0075] Step C1: Determine the first and third capture target facial features of each second group from at least one first capture target facial feature associated with each group of second information and at least one third capture target facial feature associated with each group of second information. The first and third capture target facial features of the same second group are capture target facial features associated with the same second shooting device.

[0076] Step C2: Determine the thirteenth target facial feature of each second group from the first and third target facial features of each second group. The feature similarity between the thirteenth target facial feature of each second group and the target facial features in each group of second information is greater than the feature similarity between the fourteenth target facial feature and the target facial features in each group of second information. The fourteenth target facial feature is the remaining target facial features of each second group excluding the thirteenth target facial feature.

[0077] Step C3: Generate the first video content associated with each group of second information based on the video clips obtained by the second shooting device that are associated with the facial features of the thirteenth captured target in each second group.

[0078] See Figure 3From at least one first-capture target facial feature associated with each group of second information and at least one third-capture target facial feature associated with each group of second information, the first-capture target facial features and the third-capture target facial features of each second group are accurately determined. Here, the first-capture target facial features and the third-capture target facial features of the same second group are respectively from the capture target facial features associated with the same second shooting device. This setting ensures the consistency and relevance of the source of the capture target facial features.

[0079] See Figure 3 Next, from the first and third captured facial features of each second group, the thirteenth captured facial feature of each second group is further determined. The criterion is that the feature similarity between the thirteenth captured facial feature of each second group and the captured facial features in each group of second information must be greater than the feature similarity between the remaining fourteenth captured facial features (excluding the thirteenth captured facial feature) of that second group and the captured facial features in each group of second information. This means that the thirteenth captured facial feature is the feature that best matches the target feature in the second group and can more accurately reflect the feature information of the target.

[0080] See Figure 3 Based on video clips acquired by a second shooting device that are associated with the facial features of the thirteenth captured target in each second group, a first video content associated with each group of second information is generated. By selecting the video clip corresponding to the most matching feature for generation, a high degree of fit between the first video content and the second information is ensured, further improving the accuracy and relevance of the video content.

[0081] S240. Obtain first information, which is an image including the target to be captured.

[0082] S250. Detect the feature similarity between the facial features of the captured target in the first information and the facial features of the captured target in at least one set of second information associated with the first shooting device. The same set of second information includes a first type of captured target facial features and a second type of captured target facial features for the same captured target. The angle between the facial orientation of the captured target and the optical axis of the shooting device corresponding to the first type of captured target facial features is greater than a preset angle, and the angle between the facial orientation of the captured target and the optical axis of the shooting device corresponding to the second type of captured target facial features is less than a preset angle. The second type of captured target facial features are facial features generated by simulating and adjusting the facial orientation of the captured target based on the first type of captured target facial features.

[0083] S260. In response to the fact that the facial features of the target captured in the third information and the facial features of the target captured in the first information satisfy a preset similarity condition in at least one set of second information, the second video content is determined from multiple first video contents based on the facial features of the target captured in the third information. Each of the multiple first video contents is pre-generated based on video segments obtained by capturing the same target by at least one second shooting device. Each first video content is associated with a set of facial features of the target captured in the second information. The second shooting device is used to capture the target before the first shooting device.

[0084] As an optional but not limited implementation, the video retrieval method provided in this embodiment of the invention may include the following steps D1-D3:

[0085] Step D1: Obtain the first captured image by a third shooting device, wherein the third shooting device includes the first shooting device and / or the second shooting device.

[0086] Step D2: In response to the fact that the facial features of the captured target in the first captured image captured by the third shooting device belong to the first type of captured target facial features, a prompt message for the first captured image is generated. Based on the prompt message of the first captured image and the first captured image, an image generation model is used to obtain a second captured image that belongs to the second type of captured target facial features. The prompt message of the first captured image is used to guide the image generation model to adjust the captured target in the first captured image so that the captured target meets the preset capture conditions. The captured target meeting the preset capture conditions includes at least one of the following: the angle between the face orientation of the captured target and the optical axis of the shooting device is less than a preset angle; the motion characteristics of the reference part of the captured target meet the preset motion characteristics; and there are no preset obstructions on the captured target. The motion characteristics of the reference part are used to indicate the dynamic or static characteristics presented by the surface morphology changes of the reference part.

[0087] Step D3: Determine the facial features of the target being captured by the third shooting device based on the first captured image and the second captured image.

[0088] The third shooting device includes the first shooting device and / or the second shooting device. Each third shooting device performs target detection on the target passing through the shooting field of view, and acquires multiple first capture images. For example, it can capture three optimal images (Top-3) that include the target, and then store the Top-3 capture results containing the target in the database. The image of the target to be detected is used to retrieve the corresponding target in the database and is used for subsequent video synthesis. However, in some scenarios, because the third shooting device is not installed directly facing the target, the target turns its head when passing the third shooting device, resulting in the angle between the target's face and the optical axis of the shooting device being greater than a preset angle. This reduces the similarity of the target's features in the later comparison, which can easily lead to retrieval errors or the target's feature similarity being less than the threshold and thus being filtered out.

[0089] Therefore, when the facial features of the captured target in the first captured image obtained by the third shooting device belong to the first type of captured target facial features, a prompt message for the first captured image is automatically generated. The core function of the prompt message for the first captured image is to clarify the adjustment direction that the image generation model needs to make for the captured target, so as to guide the image generation model to optimize the captured target included in the first captured image so that the captured target meets the preset capture conditions. Among them, the preset capture conditions include: the angle between the face orientation of the captured target and the optical axis of the shooting device is less than a preset angle (for example, the face is facing the lens of the shooting device); the motion characteristics of the reference parts of the captured target meet the preset standards (such as no blur in static state, and the dynamic amplitude is within a reasonable range); there are no preset obstructions on the captured target (such as no mask on the face, hands covering the face, etc.). The image generation model can be an image generation tool that can dynamically convert input text and images. At the same time, the image generation model also has element recognition capabilities, which can fill in the missing content in the image to ensure that the image is complete and natural.

[0090] See Figure 4 The first captured image and the prompt information can be input into the image generation model. The image generation model adjusts the captured target in the first captured image according to the prompt information and outputs a second captured image that meets the preset capture conditions. Then, by combining the original first captured image and the optimized second captured image, multiple facial features of the captured target are extracted, thereby obtaining the first type of captured target facial features and the second type of captured target facial features associated with the third shooting device for the same captured target.

[0091] The above solution addresses the poor image quality issues caused by the pose, occlusion, and motion of the target in traditional image capture. By optimizing the generative model, it ensures that the target meets the preset capture conditions, reducing invalid captures (such as side profiles, blurry images, and heavily occluded images) and increasing the proportion of valid data. Furthermore, it eliminates the need to upgrade the shooting equipment (e.g., increasing frame rate or sensor accuracy) to improve capture performance; instead, it optimizes the output of existing equipment through software algorithms (image generation model), reducing hardware costs. Simultaneously, by combining the original and optimized capture images to determine the final target's facial features, it preserves the authentic information of the original image while compensating for its deficiencies through optimization. This results in more comprehensive and accurate facial features in the final capture, achieving intelligent optimization and feature extraction of capture images, significantly improving the usability of capture data while maintaining efficiency.

[0092] Optionally, see Figure 5 The method involves determining the facial features of the target being captured, which are associated with the third shooting device and are for the same target, based on the first captured image and the second captured image. This includes extracting a first type of facial feature of the target being captured from the first captured image and extracting a second type of facial feature of the target being captured from the second captured image, and associating and binding the extracted first type of facial feature of the target being captured and the second type of facial feature of the target being captured with the third shooting device as different dimensions of facial features of the target being captured for the same target.

[0093] As an optional but not limited implementation, before generating the prompt information for the first captured image, the following steps may also be included:

[0094] The key point detection results of the captured target in the first captured image are identified, and the distance from the facial edge to the midface of the captured target is determined based on the key point detection results. Based on the distance from the facial edge to the midface of the captured target in the first captured image, it is determined whether the first captured image captured by the third shooting device belongs to the facial features of the first type of captured target.

[0095] As an optional but not limited implementation, the prompt information for generating the first captured image includes, but is not limited to, the following steps:

[0096] A first orientation of the target in the first captured image is determined, which indicates the directional change of the target's facial orientation relative to the central axis of the shooting device. A second orientation is determined based on the opposite direction of the first orientation, which is used to generate cue information for the first captured image to guide the image generation model to adjust the facial orientation of the target in the first captured image.

[0097] See Figure 4The method for selecting the prompt information for the first captured image is as follows: Select the top-3 images containing the captured target acquired by each third-channel camera, and choose the face image with the highest quality score from these top-3 images (quality score can be evaluated using image clarity and image content richness). First, perform key point detection on the captured target in the first captured image. Using the key point detection results, determine if the angle between the captured target's facial orientation and the optical axis of the camera is greater than a preset angle. Determine if the angle between the captured target's facial orientation and the optical axis of the camera is greater than a preset angle by measuring the distance from the edge of the captured target's face to the midface. If the captured target in the first captured image exhibits facial features of the second type of captured target, no further generation is required, and the generation stage is skipped. If the captured target in the first captured image exhibits facial features of the first type of captured target, determine the first turning direction of the captured target in the first captured image using the same point detection results, and generate the corresponding prompt in the opposite direction. That is, if the captured target's face is detected to be turning left, the corresponding prompt is "Turn right".

[0098] As an optional but not limited implementation, the prompt information for generating the first captured image includes, but is not limited to, the following steps:

[0099] Determine the outline size of the reference part of the captured target in the first captured image; determine the motion state of the reference part of the captured target based on the outline size of the reference part of the captured target, and use the motion state of the reference part of the captured target to form prompt information of the first captured image to guide the image generation model to adjust the motion characteristics of the reference part of the captured target in the first captured image.

[0100] See Figure 4 For generating prompts for the first captured image, the contour dimensions of the reference parts of the captured target are determined from the key point detection results. For example, taking the face as the target, the contour dimensions can be represented by the opening and closing of the eyes and the opening of the mouth. The motion state of the reference parts of the captured target is determined by using these contour dimensions. For instance, the opening and closing of the eyes and the opening of the mouth are used to determine whether the face is in an exaggerated expression state (e.g., when excited while riding a roller coaster, the face shows wide-open eyes and an open mouth). If the face is in an exaggerated expression state, the prompt word "with a calm expression" is selected, resulting in the prompt word "with a calm expression". If the face is not in an exaggerated expression state, the prompt word is "maintain the current expression".

[0101] As an optional but not limited implementation, the prompt information for generating the first captured image includes, but is not limited to, the following steps:

[0102] The first captured image is segmented, and the occlusion detection results of each segmented region are detected. The occlusion detection results are used to indicate whether there are preset occlusions in each segmented region. The occlusion detection results are used to generate prompt information for the first captured image to guide the image generation model to remove the preset occlusions in the first captured image.

[0103] See Figure 4 For prompts related to pre-defined obstructions such as helmets and raincoats, the system performs image segmentation on the first captured image including the target and classifies the segmentation results. The classification results determine whether the target in the first captured image is obstructed by a helmet or raincoat. If the segmentation results indicate that the target in the first captured image is obstructed by a helmet or raincoat, the corresponding prompts "remove helmet" or "remove raincoat" are added to the prompt.

[0104] As an optional but not limited implementation, a second capture image belonging to the facial features of the second type of capture target is obtained by using an image generation model based on the prompt information of the first capture image and the first capture image, including but not limited to the following steps:

[0105] The prompt information of the first captured image and the first captured image are input into the image generation model, and the third video content is output through the image generation model. The captured targets included in at least some image frames of the third video content belong to the captured targets that are adjusted from the captured targets in the first captured image and meet the preset captured conditions. Based on the image content and image style of the first captured image, a second captured image belonging to the facial features of the second type of captured target is determined from the third video content.

[0106] See Figure 4 The prompt information of the first captured image is used to guide the image generation model to adjust and orient the captured target in the captured image to generate a video containing the captured target that meets the preset capture conditions. In addition to the angle between the face of the captured target and the optical axis of the shooting device being greater than the preset angle, there are also issues such as the target being obscured by wearing a helmet or raincoat, and the motion characteristics of the reference part of the captured target not meeting the conditions due to excessive speed. These will all lead to a decrease in similarity when the captured targets are compared in the later stage. Therefore, this solution proposes to automatically generate prompt information corresponding to each first captured image to guide the image generation model to generate a second captured image that contains the captured target and the angle between the face of the captured target and the optical axis of the shooting device is less than the preset angle.

[0107] like Figure 4As shown, the template for the prompt message of the first captured image is "Based on the current face image," + "{_A_} expression," + "Turn head to {_B_}," + "Remove helmet," + "Remove raincoat," + "Generate a high-resolution video of about 5 seconds," where A can be either {keeping the current expression} or {calm}, and B can be either {left} or {right}. By automatically selecting corresponding prompt words to form the prompt message for different captured images, the image generation model is guided to generate an accurate video containing the captured target.

[0108] For example, if the target's face is captured by a third camera and the captured image shows the target turning their head to the right, with wide eyes and an open mouth due to the rapid movement, and the target is wearing a helmet and raincoat obscuring their face, the above method will automatically detect the corresponding result. Therefore, the prompt will automatically form the message "Based on the current facial image, with a calm expression, turn your head to the left, remove the helmet and raincoat, and generate a high-resolution video of about 5 seconds," to guide the subsequent image generation model to generate the video. If the tourist is not detected to be wearing a mask and raincoat, the prompt will be "Based on the current facial image, with a calm expression, turn your head to the left, and generate a high-resolution video of about 5 seconds."

[0109] As an optional but not limited implementation, based on the image content and style of the first captured image, a second captured image belonging to the facial features of the second type of captured target is determined from the third video content, including but not limited to the following steps:

[0110] At least one reference image frame is determined from the multiple image frames included in the third video content, wherein the angle between the face orientation of the captured target and the optical axis of the shooting device in the reference image frame is less than a preset angle; a second captured image belonging to the facial features of the second type of captured target is determined from the at least one reference image frame, wherein the difference between the second captured image and the first captured image in terms of image content and image style is less than the difference between the remaining image frames (excluding the second captured image) in terms of image content and image style and the first captured image.

[0111] For the first captured image, the optimal reference image frame is selected from at least one reference image frame. To ensure that the generated captured target is consistent with the first captured image in terms of image content and style, this scheme can use multi-layer feature map comparison of neural networks to compare the similarity of the generated reference image frame and the first captured image in terms of image content and style. During the process of multi-layer convolution of the image by the neural network, the feature map results of the shallow network are the texture, edge, and color features of the image, while the feature map of the deep network often focuses on the global semantic information and global features of the image. In summary, after the image containing the captured target is processed by the convolutional neural network, the output of the shallow network can represent the texture, edge, and color features of the captured target, i.e., the image style, while the output of the deep network represents the global features and semantic information of the captured target, i.e., the image content. After normal capture targets are generated, their image style and image content should be as small as possible compared with the first captured image. This method is used to optimize the images containing the captured targets.

[0112] Optionally, a pre-trained ResNet50 network is used. The reference image frame and the first captured image are simultaneously fed into the ResNet50 network. The shallow layer output is the feature map from ResNet50-Block 1, and the deep layer output is the feature map from ResNet50-Block 4. The two sets of feature maps are then compared using a Gram matrix to obtain the similarity between the generated reference image frame and the first captured image in terms of image content and style.

[0113] Meanwhile, to ensure that the generated second capture image is one where the angle between the target's face and the optical axis of the shooting device is less than a preset angle, the key point detection of the target in the reference image frame mentioned above can be performed. For each generated reference image frame, it can be determined that the angle between the target's face and the optical axis of the shooting device is less than a preset angle. Among the reference image frames where the angle between the target's face and the optical axis of the shooting device is less than a preset angle, the image content and image style most similar to the first capture image are selected to generate the capture image, thus obtaining the final second capture image.

[0114] The technical solution of this embodiment acquires first information including the target to be captured. When performing image retrieval by feature comparison on the captured target, each set of second information associated with the first capturing device simultaneously includes facial features of both the first type and the second type of captured target. By distinguishing captured images with different facial orientations, the coverage of facial feature comparison is expanded, and the robustness of facial feature matching is improved. Even if captured targets with varying facial orientations appear in the first information, feature complementarity can be achieved through captured targets with different facial orientations in each set of second information, thereby improving the accuracy of feature similarity detection. To reduce feature matching failures caused by angle differences and ensure the comprehensiveness of image retrieval results; since the pre-synthesized video has integrated the first video content including the same captured target based on the capture of the second shooting device before the first shooting device, there is no need for manual screening or editing. The corresponding pre-generated second video content can be directly matched from multiple first video contents through feature similarity. This enables the rapid and accurate location of the second video content matching the first information from the pre-synthesized video content after obtaining the first information, directly obtaining complete and coherent video content, avoiding video content mismatch or omission, and ensuring the consistency between the obtained video content and the movement process of the captured target.

[0115] Figure 6 This invention provides a structural diagram of a video retrieval device. This embodiment is applicable to situations where facial features of a first type of captured target and facial features of a second type of captured target are bound together for joint use in image retrieval containing captured targets. The video retrieval device can be implemented in hardware and / or software and can be configured in any electronic device with network communication capabilities.

[0116] like Figure 6 As shown, the video retrieval device provided in this embodiment of the invention may include the following:

[0117] The acquisition module 610 is used to acquire first information, wherein the first information is an image including the target to be captured;

[0118] The detection module 620 is used to detect the feature similarity between the facial features of the captured target in the first information and the facial features of the captured target in at least one set of second information associated with the first shooting device. The same set of second information includes a first type of captured target facial features and a second type of captured target facial features for the same captured target. The angle between the facial orientation of the captured target and the optical axis of the shooting device corresponding to the first type of captured target facial features is greater than a preset angle. The angle between the facial orientation of the captured target and the optical axis of the shooting device corresponding to the second type of captured target facial features is less than a preset angle. The second type of captured target facial features are facial features generated by simulating and adjusting the facial orientation of the captured target based on the first type of captured target facial features.

[0119] The retrieval module 630 is configured to, in response to the fact that the facial features of the captured target in the third information and the facial features of the captured target in the first information satisfy a preset similarity condition in the at least one set of second information, determine second video content from multiple first video contents based on the facial features of the captured target in the third information. Each of the multiple first video contents is pre-generated based on video segments obtained by capturing the same captured target by at least one second shooting device. Each first video content is associated with facial features of the captured target in a set of second information. The second shooting device is used to capture the captured target before the first shooting device.

[0120] Optionally, based on the above embodiments, before determining at least one second video content from multiple first video contents according to the third information, the method further includes:

[0121] For each set of the at least one set of second information, based on the first type of capture target facial features in each set of second information, at least one first capture target facial feature associated with each set of second information is determined from the capture target facial features associated with each second shooting device. The different first capture target facial features in the at least one first capture target facial features are capture target facial features associated with different second shooting devices. The feature similarity between each first capture target facial feature and the first type of capture target facial feature in each set of second information is greater than the feature similarity between the second capture target facial feature and the first type of capture target facial feature in each set of second information. The second capture target facial feature is the remaining capture target facial features other than the first capture target facial feature among the capture target facial features associated with the second shooting device associated with the first capture target facial feature.

[0122] Based on the second type of facial features of the captured target in each group of second information, at least one third facial feature of the captured target associated with each group of second information is determined from the facial features of the captured target associated with each second shooting device. The different third facial features of the at least one third facial feature are facial features of the captured target associated with different second shooting devices. The feature similarity between each third facial feature and the second type of facial feature of each group of second information is greater than the feature similarity between the fourth facial feature and the second type of facial feature of each group of second information. The fourth facial feature of the captured target is the remaining facial features of the captured target associated with the second shooting device associated with the third facial feature, excluding the third facial feature.

[0123] Based on at least one first captured target facial feature associated with each set of second information and at least one third captured target facial feature associated with each set of second information, a first video content associated with each set of second information is generated; there is an association and binding relationship between the captured target facial features associated with each second shooting device and the video segments acquired by each second shooting device.

[0124] Based on the above embodiments, optionally, each group of second information includes multiple first-type captured target facial features;

[0125] Based on the first type of facial features of the captured target in each group of second information, at least one first facial feature of the captured target associated with each group of second information is determined from the facial features of the captured target associated with each second shooting device, including:

[0126] For each first-type capture target facial feature in each group of second information, at least one fifth capture target facial feature associated with each first-type capture target facial feature is determined from the capture target facial features associated with each second shooting device. Different fifth capture target facial features in the at least one fifth capture target facial feature are capture target facial features associated with different second shooting devices. The feature similarity between each fifth capture target facial feature and the first-type capture target facial feature in each group of second information is greater than the feature similarity between the sixth capture target facial feature and the first-type capture target facial feature in each group of second information. The sixth capture target facial feature is the remaining capture target facial features other than the fifth capture target facial feature among the capture target facial features associated with the second shooting device associated with the fifth capture target facial feature.

[0127] The seventh capture target facial feature of each first group is determined from the different fifth capture target facial features of each first group. The different fifth capture target facial features of the same first group are capture target facial features associated with the same second shooting device selected from at least one fifth capture target facial feature associated with each of the multiple first type capture target facial features. The feature similarity between the seventh capture target facial feature of each first group and the first type capture target facial feature of each group of second information is greater than the feature similarity between the eighth capture target facial feature and the first type capture target facial feature of each group of second information. The eighth capture target facial feature is the remaining capture target facial features of each first group other than the seventh capture target facial feature.

[0128] Based on the above embodiments, optionally, according to the second type of target facial features in each group of second information, at least one third target facial feature associated with each group of second information is determined from the target facial features associated with each second shooting device, including:

[0129] Based on the second type of target facial features in each group of second information, at least one ninth target facial feature is determined from the first type of target facial features associated with each second shooting device. The angle between the facial orientation of the target corresponding to the first type of target facial feature and the optical axis of the shooting device is greater than a preset angle. Different ninth target facial features in the at least one ninth target facial feature are first type target facial features associated with different second shooting devices. The feature similarity between each ninth target facial feature and the second type of target facial features in each group of second information is greater than the feature similarity between the tenth target facial feature and the second type of target facial features in each group of second information. The tenth target facial feature is the remaining target facial features other than the ninth target facial feature in the first type of target facial features associated with the second shooting device associated with the ninth target facial feature.

[0130] Based on the second type of target facial features in each group of second information, at least one eleventh target facial feature is determined from the second type of target facial features associated with each second shooting device. The angle between the facial orientation of the target corresponding to the second type of target facial feature and the optical axis of the shooting device is less than a preset angle. The different eleventh target facial features in the at least one eleventh target facial features are second type target facial features associated with different second shooting devices. The feature similarity between each eleventh target facial feature and the second type of target facial features in each group of second information is greater than the feature similarity between the twelfth target facial feature and the second type of target facial features in each group of second information. The twelfth target facial feature is the remaining target facial features other than the eleventh target facial feature in the second type of target facial features associated with the second shooting device associated with the eleventh target facial feature.

[0131] Based on at least one ninth capture target facial feature and at least one eleventh capture target facial feature of the second type of capture target facial features in each group of second information, at least one third capture target facial feature associated with each group of second information is determined. The at least one third capture target facial feature is the intersection result of the capture target facial features of the at least one ninth capture target facial feature and the at least one eleventh capture target facial feature.

[0132] Based on the above embodiments, optionally, a first video content associated with each set of second information is generated according to at least one first captured target facial feature associated with each set of second information and at least one third captured target facial feature associated with each set of second information, including:

[0133] From at least one first capture target facial feature associated with each group of second information and at least one third capture target facial feature associated with each group of second information, determine the first capture target facial feature and the third capture target facial feature of each second group. The first capture target facial feature and the third capture target facial feature of the same second group are capture target facial features associated with the same second shooting device.

[0134] The thirteenth target facial feature of each second group is determined from the first and third target facial features of each second group. The feature similarity between the thirteenth target facial feature of each second group and the target facial features in each group of second information is greater than the feature similarity between the fourteenth target facial feature and the target facial features in each group of second information. The fourteenth target facial feature is the remaining target facial features of each second group excluding the thirteenth target facial feature.

[0135] The first video content associated with each group of second information is generated based on the video clips obtained by the second shooting device that are associated with the facial features of the thirteenth captured target in each second group.

[0136] Optionally, based on the above embodiments, the device further includes:

[0137] Acquire a first captured image obtained by a third capturing device, wherein the third capturing device includes the first capturing device and / or the second capturing device;

[0138] In response to the fact that the facial features of the captured target in the first captured image captured by the third shooting device belong to the first type of captured target facial features, a prompt message for the first captured image is generated. Based on the prompt message of the first captured image and the first captured image, an image generation model is used to obtain a second captured image that belongs to the second type of captured target facial features. The prompt message of the first captured image is used to guide the image generation model to adjust the captured target in the first captured image so that the captured target meets the preset capture conditions. The captured target meeting the preset capture conditions includes at least one of the following: the angle between the face orientation of the captured target and the optical axis of the shooting device is less than a preset angle; the motion characteristics of the reference part of the captured target meet the preset motion characteristics; and there are no preset obstructions on the captured target. The motion characteristics of the reference part are used to indicate the dynamic or static characteristics presented by the surface morphology changes of the reference part.

[0139] Based on the first captured image and the second captured image, the facial features of the captured target associated with the third shooting device for the same captured target are determined.

[0140] Optionally, based on the above embodiments, before generating the prompt information for the first captured image, the method further includes:

[0141] Identify the key point detection results of the captured target in the first captured image, and determine the distance from the facial edge to the midface of the captured target based on the key point detection results;

[0142] The distance from the facial edge to the midface of the captured target in the first captured image is used to determine whether the first captured image captured by the third shooting device belongs to the first type of captured target facial features.

[0143] Based on the above embodiments, optionally, the prompt information generated for the first captured image includes:

[0144] Determine the first orientation of the target in the first captured image, the first orientation being used to indicate the change in direction of the target's face relative to the central axis of the capturing device;

[0145] A second turn is determined based on the opposite direction of the first turn. The second turn is used to generate prompt information for the first captured image to guide the image generation model to adjust the facial orientation of the captured target in the first captured image.

[0146] Based on the above embodiments, optionally, the prompt information generated for the first captured image includes:

[0147] Determine the outline size of the reference part of the target in the first captured image;

[0148] The motion state of the reference part of the captured target is determined based on the outline size of the reference part of the captured target. The motion state of the reference part of the captured target is used to generate prompt information for the first captured image to guide the image generation model to adjust the motion characteristics of the reference part of the captured target in the first captured image.

[0149] Based on the above embodiments, optionally, the prompt information generated for the first captured image includes:

[0150] The first captured image is segmented, and the occlusion detection results of each segmented image region are detected. The occlusion detection results are used to indicate whether there are preset occlusions in each segmented image region. The occlusion detection results are used to generate prompt information for the first captured image to guide the image generation model to remove the preset occlusions in the first captured image.

[0151] Based on the above embodiments, optionally, a second capture image belonging to the facial features of the second type of capture target is obtained by using an image generation model based on the prompt information of the first capture image and the first capture image, including:

[0152] The prompt information of the first captured image and the first captured image are input into the image generation model, and the third video content is output through the image generation model. The captured target included in at least a portion of the image frames of the third video content belongs to the captured target that is adjusted from the captured target in the first captured image and meets the preset captured conditions.

[0153] Based on the image content and style of the first captured image, a second captured image belonging to the facial features of the second type of captured target is determined from the third video content.

[0154] Based on the above embodiments, optionally, according to the image content and image style of the first captured image, a second captured image belonging to the facial features of the second type of captured target is determined from the third video content, including:

[0155] At least one reference image frame is determined from the multiple image frames included in the third video content, wherein the angle between the face orientation of the captured target and the optical axis of the shooting device in the reference image frame is less than a preset angle;

[0156] A second captured image belonging to the second type of captured target facial features is determined from at least one reference image frame. The difference between the second captured image and the first captured image in terms of image content and image style is less than the difference between the remaining image frames (excluding the second captured image) and the first captured image in terms of image content and image style among the at least one reference image frames.

[0157] The video retrieval device provided in the embodiments of the present invention can execute the video retrieval method provided in any of the embodiments of the present invention, and has the corresponding functions and beneficial effects of executing the video retrieval method. For details, please refer to the relevant operations of the video retrieval method in the foregoing embodiments.

[0158] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of the present invention.

[0159] Figure 7A schematic diagram of an electronic device 10 that can be used to implement embodiments of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (e.g., helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.

[0160] like Figure 7 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 may also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0161] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0162] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as video retrieval methods.

[0163] In some embodiments, the video retrieval method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the video retrieval method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the video retrieval method by any other suitable means (e.g., by means of firmware).

[0164] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0165] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0166] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0167] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0168] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or computing systems that include middleware components (e.g., application servers), or computing systems that include frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.

[0169] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.

[0170] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0171] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. A video retrieval method, characterized in that, The method includes: Obtain first information, which is an image including the target to be captured; The similarity of features between the facial features of the captured target in the first information and the facial features of the captured target in at least one set of second information associated with the first shooting device is detected. The same set of second information includes a first type of captured target facial features and a second type of captured target facial features for the same captured target. The angle between the facial orientation of the captured target and the optical axis of the shooting device corresponding to the first type of captured target facial features is greater than a preset angle. The angle between the facial orientation of the captured target and the optical axis of the shooting device corresponding to the second type of captured target facial features is less than a preset angle. The second type of captured target facial features are facial features generated by simulating and adjusting the facial orientation of the captured target based on the first type of captured target facial features. In response to the fact that the facial features of the captured target in the third information and the facial features of the captured target in the first information satisfy a preset similarity condition in the at least one set of second information, a second video content is determined from multiple first video contents based on the facial features of the captured target in the third information. Each of the multiple first video contents is pre-generated based on video segments obtained by capturing the same captured target by at least one second shooting device. Each first video content is associated with a set of captured target facial features in the second information. The second shooting device is used to capture the captured target before the first shooting device.

2. The method according to claim 1, characterized in that, Before determining at least one second video content from a plurality of first video content based on the third information, the method further includes: For each set of the at least one set of second information, based on the first type of capture target facial features in each set of second information, at least one first capture target facial feature associated with each set of second information is determined from the capture target facial features associated with each second shooting device. The different first capture target facial features in the at least one first capture target facial features are capture target facial features associated with different second shooting devices. The feature similarity between each first capture target facial feature and the first type of capture target facial feature in each set of second information is greater than the feature similarity between the second capture target facial feature and the first type of capture target facial feature in each set of second information. The second capture target facial feature is the remaining capture target facial features other than the first capture target facial feature among the capture target facial features associated with the second shooting device associated with the first capture target facial feature. Based on the second type of facial features of the captured target in each group of second information, at least one third facial feature of the captured target associated with each group of second information is determined from the facial features of the captured target associated with each second shooting device. The different third facial features of the at least one third facial feature are facial features of the captured target associated with different second shooting devices. The feature similarity between each third facial feature and the second type of facial feature of each group of second information is greater than the feature similarity between the fourth facial feature and the second type of facial feature of each group of second information. The fourth facial feature of the captured target is the remaining facial features of the captured target associated with the second shooting device associated with the third facial feature, excluding the third facial feature. Based on at least one first captured target facial feature associated with each set of second information and at least one third captured target facial feature associated with each set of second information, a first video content associated with each set of second information is generated; there is an association and binding relationship between the captured target facial features associated with each second shooting device and the video segments acquired by each second shooting device.

3. The method according to claim 2, characterized in that, Each set of second information includes multiple facial features of the first type of captured target; Based on the first type of facial features of the captured target in each group of second information, at least one first facial feature of the captured target associated with each group of second information is determined from the facial features of the captured target associated with each second shooting device, including: For each first-type capture target facial feature in each group of second information, at least one fifth capture target facial feature associated with each first-type capture target facial feature is determined from the capture target facial features associated with each second shooting device. Different fifth capture target facial features in the at least one fifth capture target facial feature are capture target facial features associated with different second shooting devices. The feature similarity between each fifth capture target facial feature and the first-type capture target facial feature in each group of second information is greater than the feature similarity between the sixth capture target facial feature and the first-type capture target facial feature in each group of second information. The sixth capture target facial feature is the remaining capture target facial features other than the fifth capture target facial feature among the capture target facial features associated with the second shooting device associated with the fifth capture target facial feature. The seventh capture target facial feature of each first group is determined from the different fifth capture target facial features of each first group. The different fifth capture target facial features of the same first group are capture target facial features associated with the same second shooting device selected from at least one fifth capture target facial feature associated with each of the multiple first type capture target facial features. The feature similarity between the seventh capture target facial feature of each first group and the first type capture target facial feature of each group of second information is greater than the feature similarity between the eighth capture target facial feature and the first type capture target facial feature of each group of second information. The eighth capture target facial feature is the remaining capture target facial features of each first group other than the seventh capture target facial feature.

4. The method according to claim 2, characterized in that, Based on the second type of facial features of the captured target in each set of second information, at least one third facial feature of the captured target associated with each set of second information is determined from the facial features of the captured target associated with each second shooting device, including: Based on the second type of target facial features in each group of second information, at least one ninth target facial feature is determined from the first type of target facial features associated with each second shooting device. The angle between the facial orientation of the target corresponding to the first type of target facial feature and the optical axis of the shooting device is greater than a preset angle. Different ninth target facial features in the at least one ninth target facial feature are first type target facial features associated with different second shooting devices. The feature similarity between each ninth target facial feature and the second type of target facial features in each group of second information is greater than the feature similarity between the tenth target facial feature and the second type of target facial features in each group of second information. The tenth target facial feature is the remaining target facial features other than the ninth target facial feature in the first type of target facial features associated with the second shooting device associated with the ninth target facial feature. Based on the second type of target facial features in each group of second information, at least one eleventh target facial feature is determined from the second type of target facial features associated with each second shooting device. The angle between the facial orientation of the target corresponding to the second type of target facial feature and the optical axis of the shooting device is less than a preset angle. The different eleventh target facial features in the at least one eleventh target facial features are second type target facial features associated with different second shooting devices. The feature similarity between each eleventh target facial feature and the second type of target facial features in each group of second information is greater than the feature similarity between the twelfth target facial feature and the second type of target facial features in each group of second information. The twelfth target facial feature is the remaining target facial features other than the eleventh target facial feature in the second type of target facial features associated with the second shooting device associated with the eleventh target facial feature. Based on at least one ninth capture target facial feature and at least one eleventh capture target facial feature of the second type of capture target facial features in each group of second information, at least one third capture target facial feature associated with each group of second information is determined. The at least one third capture target facial feature is the intersection result of the capture target facial features of the at least one ninth capture target facial feature and the at least one eleventh capture target facial feature.

5. The method according to claim 2, characterized in that, Based on at least one first captured target facial feature associated with each set of second information and at least one third captured target facial feature associated with each set of second information, first video content associated with each set of second information is generated, including: From at least one first capture target facial feature associated with each group of second information and at least one third capture target facial feature associated with each group of second information, determine the first capture target facial feature and the third capture target facial feature of each second group. The first capture target facial feature and the third capture target facial feature of the same second group are capture target facial features associated with the same second shooting device. The thirteenth target facial feature of each second group is determined from the first and third target facial features of each second group. The feature similarity between the thirteenth target facial feature of each second group and the target facial features in each group of second information is greater than the feature similarity between the fourteenth target facial feature and the target facial features in each group of second information. The fourteenth target facial feature is the remaining target facial features of each second group excluding the thirteenth target facial feature. The first video content associated with each group of second information is generated based on the video clips obtained by the second shooting device that are associated with the facial features of the thirteenth captured target in each second group.

6. The method according to claim 2, characterized in that, The method further includes: Acquire a first captured image obtained by a third capturing device, wherein the third capturing device includes the first capturing device and / or the second capturing device; In response to the fact that the facial features of the captured target in the first captured image captured by the third shooting device belong to the first type of captured target facial features, a prompt message for the first captured image is generated. Based on the prompt message of the first captured image and the first captured image, an image generation model is used to obtain a second captured image that belongs to the second type of captured target facial features. The prompt message of the first captured image is used to guide the image generation model to adjust the captured target in the first captured image so that the captured target meets the preset capture conditions. The captured target meeting the preset capture conditions includes at least one of the following: the angle between the face orientation of the captured target and the optical axis of the shooting device is less than a preset angle; the motion characteristics of the reference part of the captured target meet the preset motion characteristics; and there are no preset obstructions on the captured target. The motion characteristics of the reference part are used to indicate the dynamic or static characteristics presented by the surface morphology changes of the reference part. Based on the first captured image and the second captured image, the facial features of the captured target associated with the third shooting device for the same captured target are determined.

7. The method according to claim 6, characterized in that, Based on the prompt information in the first captured image and the image generation model used in the first captured image, a second captured image belonging to the facial features of the second type of captured target is obtained, including: The prompt information of the first captured image and the first captured image are input into the image generation model, and the third video content is output through the image generation model. The captured target included in at least a portion of the image frames of the third video content belongs to the captured target that is adjusted from the captured target in the first captured image and meets the preset captured conditions. Based on the image content and style of the first captured image, a second captured image belonging to the facial features of the second type of captured target is determined from the third video content.

8. A video retrieval device, characterized in that, The device includes: The acquisition module is used to acquire first information, which includes the target to be captured. The detection module is used to detect the feature similarity between the facial features of the captured target in the first information and the facial features of the captured target in at least one set of second information associated with the first shooting device. The same set of second information includes a first type of captured target facial features and a second type of captured target facial features for the same captured target. The angle between the facial orientation of the captured target and the optical axis of the shooting device corresponding to the first type of captured target facial features is greater than a preset angle. The angle between the facial orientation of the captured target and the optical axis of the shooting device corresponding to the second type of captured target facial features is less than a preset angle. The second type of captured target facial features are facial features generated by simulating and adjusting the facial orientation of the captured target based on the first type of captured target facial features. The retrieval module is configured to, in response to the fact that the facial features of the captured target in the third information and the facial features of the captured target in the first information satisfy a preset similarity condition in the at least one set of second information, determine second video content from multiple first video contents based on the facial features of the captured target in the third information. Each of the multiple first video contents is pre-generated based on video segments obtained by capturing the same captured target by at least one second shooting device. Each first video content is associated with a set of captured target facial features in the second information. The second shooting device is used to capture the captured target before the first shooting device.

9. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the video retrieval method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the video retrieval method according to any one of claims 1-7.