A video material retrieval and generation method and system for target person identification
Patent Information
- Application Number
- CN202611079939.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-21
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2046-07-21
AI Technical Summary
[0002]当前线下商圈、文旅园区、会展场馆等场景均部署大规模分布式监控采集设备,不同点位设备搭载多路独立视频通道,每日产生海量时长的监控视频文件,行业内普遍存在针对特定人物检索监控片段并合成视频的业务需求,传统检索模式不会依托场景标识划分检索范围,执行人脸识别检索时需要遍历园区内全部监控设备与全部视频通道的存量人脸数据,无法通过空间区域前置过滤无关数据,多设备并发检索场景下产生巨大算力消耗,服务器内存占用与数据读取时长将大幅增加,检索响应速度缓慢,难以支撑游客集中时段的批量检索请求
本发明根据场景标识匹配采集设备关联信息的前置过滤机制,能够大幅缩小检索数据范围,无需加载全域全部监控通道数据,有效降低服务器内存占用与数据读取耗时,主检索分区搭配邻域关联分区的双分区检索架构,服务端同步调取双分区存量特征完成跨分区联动校验,匹配核心区域特征再核验毗邻区域时空一致性,依靠双层特征匹配过滤路人相似人脸带来的误匹配,有效提升人脸识别匹配准确率,避免丢失目标人物移动过程中的监控片段,整合有效特征的设备点位、空间位置、时序信息生成标准化匹配结果,让视频素材截取环节能够定位每一段目标人物出镜画面,根据匹配结果自动匹配视频模板完成视频合成,自动拼接时序片段并叠加片头片尾、转场、背景音乐,兼顾检索精准度与视频产出效率,同时适配不同场景点位输出风格适配的视频内容。
Smart Images

Figure CN122594540B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of video generation technology, and more specifically, to a method and system for retrieving and generating video footage for target person identification. Background Technology
[0002] Currently, large-scale distributed monitoring and acquisition equipment is deployed in offline business districts, cultural and tourism parks, and convention and exhibition venues. Different locations are equipped with multiple independent video channels, generating massive amounts of monitoring video files daily. There is a widespread business need within the industry to retrieve monitoring clips and synthesize videos for specific individuals. Traditional retrieval methods do not rely on scene identifiers to define the search scope. When performing facial recognition retrieval, it is necessary to traverse all existing facial data from all monitoring devices and video channels within the park, failing to filter irrelevant data through spatial regions. This results in huge computational consumption in multi-device concurrent retrieval scenarios, significantly increasing server memory usage and data reading time, leading to slow retrieval response speeds and difficulty in supporting batch retrieval requests during peak tourist seasons. Furthermore, conventional solutions rely solely on facial data from a single device for matching and comparison, lacking the linkage verification logic of a main retrieval partition combined with neighboring related partitions. They can only identify the appearance of a single target individual in the video feed, unable to link with adjacent, interconnected devices for cross-regional feature verification. If the target individual moves to an adjacent monitoring area, the system will lose subsequent video footage, easily leading to gaps in the detection of the individual's movement.
[0003] Current face matching processes lack cross-regional consistency verification mechanisms, relying solely on single-frame face vector similarity for matching. This fails to incorporate spatial location and continuous temporal constraints to filter out mismatched features caused by similar passersby. Consequently, the output matching information is mixed with numerous irrelevant points and incorrect timestamps, leading to the incorporation of a large amount of irrelevant footage when subsequently cropping video segments. Traditional video compositing relies entirely on manual operation. Staff must retrieve video from corresponding devices one by one based on scattered face records, manually filter, crop, and splice video segments, and then add additional packaging materials such as intros, transitions, and background music. This entire manual process is time-consuming and labor-intensive. Furthermore, manual screening is prone to time interval cropping errors, resulting in truncated footage of the target person, chaotic video narrative logic, and an inability to match the person's actual movement trajectory. Summary of the Invention
[0004] In view of the shortcomings of the existing technology, the purpose of this invention is to provide a method and system for retrieving and generating video materials for target person identification.
[0005] To achieve the above objectives, the present invention provides the following technical solution: A method for retrieving and generating video footage for target person identification includes the following steps: Receive target biometric data and scene identifiers of the target person; Obtain the acquisition device association information for the corresponding scene based on the scene identifier; the acquisition device association information includes the mapping relationship between the device and the video channel; Send the target biometric data and the associated information of the acquisition device to the server; Based on the information associated with the acquisition devices, the main retrieval partition and the neighboring associated partition are divided. The server synchronously retrieves the existing biometric data in the main retrieval partition and the neighboring associated partition. Based on the existing biometric data, the face feature vector of the target biometric data is cross-partitioned and verified to obtain valid feature data. The matching result is obtained by integrating the device location, spatial location, and continuous temporal information of the effective feature data; the matching result includes device identifier, location identifier, and associated timestamp; The target video is generated after filtering the corresponding video templates based on the matching results.
[0006] Preferably, the acquisition device in the acquisition device association information is bound to a video channel with an independent encoding identifier.
[0007] Preferably, the method further includes the following steps: Block detection and non-maximum suppression are performed on the target biometric data to obtain face detection bounding boxes; A cross-frame face trajectory is established based on the face detection bounding box, and the face feature vector of the target biometric data is extracted based on the face trajectory.
[0008] Preferably, the target biometric data is processed by block detection and non-maximum suppression to obtain a face detection box, specifically including the following steps: The target biometric data is divided into multiple sub-blocks with overlapping regions, and face detection is performed on each sub-block to obtain a set of candidate detection boxes; Non-maximum suppression is performed on the candidate detection box set to remove duplicate detection boxes and obtain the face detection box.
[0009] Preferably, the target video is generated after filtering the corresponding video template based on the matching results, specifically including the following steps: Extract the corresponding video segment based on the matching results; All video clips are spliced together in chronological order, and the intro, outro, transition effects, and background music associated with the video template are overlaid to generate the target video.
[0010] Preferably, the main retrieval partition and the neighboring association partition are divided based on the association information of the acquisition device, specifically including the following steps: Determine the spatial layout of the acquisition devices and the video surveillance coverage area based on the associated information of the acquisition devices; Based on the spatial deployment location of the acquisition devices, the target acquisition devices for matching points of target biometric data are selected, and the target acquisition devices and their corresponding video surveillance coverage areas are defined as the main search partitions. Device areas that are interconnected with the main search partition, spatially adjacent, and have continuously connected trajectories are designated as neighborhood association partitions.
[0011] Preferably, effective feature data is obtained by performing cross-regional linkage verification on the facial feature vector of the target biometric data based on existing biometric data, specifically including the following steps: Based on the existing biometric data, feature matching and comparison are performed within the main retrieval partition, and feature content that matches the facial feature vector is selected to obtain the main partition matching feature data. Neighborhood matching association feature data is obtained based on the primary partition matching feature data; The consistency of the primary partition matching feature data and the neighboring matching associated feature data is verified to obtain valid feature data.
[0012] Preferably, the matching result is obtained by integrating the device location, spatial location, and continuous temporal information of the effective feature data, specifically including the following steps: By combining the correlation between equipment locations, spatial positions, and temporal arrangements in the effective feature data, feature correlation information is obtained by performing global correlation fitting on the feature sequence; Based on the feature association information, the device's ownership, spatial location, and time correspondence are solidified to generate matching results.
[0013] Preferably, neighborhood matching association feature data is obtained based on the primary partition matching feature data, specifically as follows: Using the main partition matching feature data as the comparison benchmark, cross-domain association matching is performed on the existing biometric data corresponding to the neighboring related partitions to obtain the neighboring matching association feature data.
[0014] A video material retrieval and generation system for target person identification includes: Receiving module: Receives target biometric data and scene identifiers sent by the user terminal; Acquisition module: Acquires the acquisition device association information corresponding to the scene based on the scene identifier; the acquisition device association information includes the mapping relationship between the device and the video channel; Sending module: Sends target biometric data and acquisition device association information to the server; Verification module: Based on the association information of the acquisition device, the main retrieval partition and the neighboring association partition are divided. The server synchronously retrieves the existing biometric data in the main retrieval partition and the neighboring association partition. Based on the existing biometric data, the face feature vector of the target biometric data is verified across partitions to obtain valid feature data. Processing module: Integrates device location, spatial location, and continuous temporal information from valid feature data to obtain matching results; the matching results include device identifier, location identifier, and associated timestamp; Generation module: Generates the target video after filtering the corresponding video templates based on the matching results.
[0015] Compared with the prior art, the present invention has the following beneficial effects: This invention employs a pre-filtering mechanism based on scene identifier matching and acquisition device association information. This significantly narrows the scope of search data, eliminating the need to load data from all monitoring channels across the entire domain. This effectively reduces server memory usage and data reading time. The dual-partition search architecture, combining a main search partition with neighboring related partitions, allows the server to synchronously retrieve existing features from both partitions to complete cross-partition linkage verification. It matches core area features and then verifies the spatiotemporal consistency of adjacent areas. By relying on dual-layer feature matching to filter out false matches caused by similar faces of passersby, it effectively improves the accuracy of face recognition matching and avoids losing monitoring segments during the movement of the target person. It integrates effective features such as device location, spatial location, and temporal information to generate standardized matching results. This allows the video material extraction stage to locate each segment of the target person's appearance. Based on the matching results, it automatically matches video templates to complete video synthesis, automatically splices temporal segments, and overlays intros, outros, transitions, and background music. This balances search accuracy and video output efficiency while adapting to different scene locations and outputting video content with appropriate styles. Attached Figure Description
[0016] Figure 1 This is a flowchart illustrating a method for retrieving and generating video footage for target person identification, provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of a video material retrieval and generation system for target person identification provided in an embodiment of the present invention. Detailed Implementation
[0017] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0018] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0019] Secondly, the term "an embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places throughout this specification does not necessarily refer to the same embodiment, nor is it a single embodiment or an embodiment selectively excluded from other embodiments.
[0020] Reference Figures 1-2 As shown.
[0021] The embodiments further illustrate the video material retrieval and generation method and system for target person identification proposed in this invention.
[0022] A method for retrieving and generating video footage for target person identification includes the following steps: Receive target biometric data and scene identifiers of the target person; The target biometric data includes the original facial image pixel information of the target person, the grayscale value matrix of the face, and the basic data of the coordinates of the key points of the facial features. The scene identifier is a unique digital code of the offline monitoring area. For example, a large commercial complex is divided into 1 to 12 independent scene identifiers. Scene 1 corresponds to the atrium area on the first floor, scene 2 corresponds to the outdoor pedestrian street area, and different scene identifiers correspond to completely independent monitoring equipment clusters.
[0023] Obtain the acquisition device association information for the corresponding scene based on the scene identifier; the acquisition device association information includes the mapping relationship between the device and the video channel; The acquisition device association information is bound to a video channel with an independent encoding identifier.
[0024] Scene identifiers are fixed numerical codes used to distinguish different offline monitoring areas. Each independent offline activity area is assigned a unique and non-repeating scene identifier value. For example, a large amusement park can be divided into 8 sets of scene identifiers from 1 to 8. 1 corresponds to the park entrance plaza area, 3 corresponds to the indoor amusement park area, and 6 corresponds to the park's food and beverage street area. Different scene identifiers are isolated from each other, and there will be no duplicate code values.
[0025] The information associated with the acquisition devices carries a one-to-one mapping relationship between the acquisition devices and video channels. The mapping relationship records all the video channels that each acquisition device in the scene can call. At the same time, it stores the device's own hardware parameters, device spatial installation coordinates, video channel storage path, and channel real-time bitrate and other supporting data. The total number of video channels in a single scene is equal to the sum of the number of channels bound to all acquisition devices in the scene. For example, if a scene contains 2 acquisition devices, the first device is bound to 6 video channels, and the second device is bound to 8 video channels, then the total number of video channels in a single scene = 6 + 8 = 14, which means that there are 14 independent video acquisition data streams available for retrieval and calling in this scene.
[0026] The internal definition of the acquisition device association information defines the binding rules between acquisition devices and video channels. Each acquisition device included in the mapping relationship is bound to several groups of video channels with independent encoding identifiers. The independent encoding identifiers are generated by continuously arranging a numerical sequence. The encoding values of video channels belonging to the same acquisition device increase continuously. The channel encoding intervals corresponding to different acquisition devices are disconnected and do not overlap. The total number of channels bound to a single device = maximum channel encoding value - minimum channel encoding value + 1. Taking an acquisition device with a bound channel encoding interval of 2 to 9 as an example, the total number of channels bound to a single device = 9 - 2 + 1 = 8, which means that this acquisition device can output 8 independent video images that do not interfere with each other.
[0027] The independent encoding identifier gives each video channel the ability to distinguish its identity. In the retrieval process, a single video stream can be directly locked by the encoding identifier without having to traverse all channel data under the device, which greatly reduces the computing resources consumed by data retrieval. When it is necessary to retrieve a certain monitoring screen, it is only necessary to match the independent encoding identifier of the corresponding video channel to directly locate the video storage file corresponding to that channel, and there will be no problem of data confusion and wrong retrieval from multiple channels.
[0028] The binding storage mode of scene identifiers and acquisition device association information enables the pre-limitation of the retrieval space. The system will not read the device data of all scenes in the park, but only retrieve the device information of the corresponding area where the scene identifier is successfully matched. The scene retrieval data loading amount = the total storage data capacity of the target scene channel. Device association information without matching scene identifiers will not be loaded into memory, and the total data loading amount will be effectively compressed. For example, if the storage data capacity of a single channel video is 20GB per hour and the total number of scene channels is 14, if the retrieval time is 2 hours, then the total data loading amount = 14 × 20 × 2 = 560GB. Compared with loading the massive video data of all 8 scenes in the park, this mechanism can significantly reduce memory usage and reading time.
[0029] The scene information index table in the storage medium is traversed using the scene identifier value as the search key. Each record in the index table is bound to a set of scene identifiers and the corresponding acquisition device association information. After a successful search and match, the complete device channel mapping data is output. Each acquisition device entry in the mapping data is accompanied by its own bound multi-channel independent encoded video channel information.
[0030] Send the target biometric data and the associated information of the acquisition device to the server; It also includes the following steps: To obtain face detection bounding boxes, block detection and non-maximum suppression are performed on the target biometric data. This process includes the following steps: The target biometric data is divided into multiple sub-blocks with overlapping regions, and face detection is performed on each sub-block to obtain a set of candidate detection boxes; Non-maximum suppression is performed on the candidate detection box set to remove duplicate detection boxes and obtain the face detection box; A cross-frame face trajectory is established based on the face detection bounding box, and the face feature vector of the target biometric data is extracted based on the face trajectory.
[0031] This process divides the original face image corresponding to the target biometric data into multiple sub-blocks with overlapping areas. The carrier of the target biometric data is the pixel matrix of the complete image captured by surveillance. The image suffers from a technical flaw where small, distant faces are easily missed. The overall block-segmentation mechanism can magnify local image areas and improve the detection probability of small face targets. The block division follows fixed size and overlap standards. Each sub-block is uniformly set with a horizontal pixel length of 320 pixels and a vertical pixel height of 320 pixels. A horizontal overlap width of 64 pixels and a vertical overlap height of 64 pixels are reserved between two adjacent horizontally arranged sub-blocks.
[0032] There are fixed conversion formulas between the total horizontal pixel length, total vertical pixel height, and number of blocks in the complete original image. The total horizontal pixel length = number of horizontal sub-blocks × length of a single horizontal block - total overlapping pixels of horizontal sub-blocks, and the total overlapping pixels of horizontal sub-blocks = number of horizontal sub-blocks minus 1 × value of a single horizontal overlapping pixel segment. The total vertical pixel height = number of vertical sub-blocks × height of a single vertical block - total overlapping pixels of vertical sub-blocks, and the total overlapping pixels of vertical sub-blocks = number of vertical sub-blocks minus 1 × value of a single vertical overlapping pixel segment. Taking an image with a total horizontal pixel length of 1200 pixels as an example, substituting the values and calculating in reverse, 1200 = number of horizontal sub-blocks × 320 - number of horizontal sub-blocks minus 1 × 64. After solving, we get that the number of horizontal sub-blocks is 4, and the corresponding total overlapping pixels of horizontal sub-blocks = (4-1) × 64 = 192 pixels. Four sub-blocks with a width of 320 pixels each overlap a total overlapping area of 192 pixels, which just fills the complete image of 1200 pixels.
[0033] For each independent sub-block, the face detection algorithm is called individually to perform traversal recognition. All regions that match the facial contour features are identified within the sub-block's image. For each successfully identified region, a set of candidate detection boxes is generated, which includes the horizontal coordinates of the upper left corner, the vertical coordinates of the upper left corner, the horizontal coordinates of the lower right corner, the vertical coordinates of the lower right corner, and the face confidence value. All candidate detection boxes output by all sub-blocks are summarized and integrated to form a complete set of candidate detection boxes. The set contains a large number of overlapping candidate boxes generated by the same facial target, which need to be processed by maximum value suppression to remove redundancy.
[0034] Non-maximum suppression (NMS) is performed on the candidate bounding box set. The core of the operation relies on the Cross-Union Ratio (CIRR) to determine duplicate bounding boxes. The CIRR measures the degree of overlap in the image coverage of two sets of bounding boxes. CIRR = Total number of pixels in the overlapping area of the two sets of bounding boxes ÷ Total number of pixels in the merged area of the two sets of bounding boxes. The total number of pixels in the overlapping area refers to the total number of pixels simultaneously covered by both sets of boxes, and the total number of pixels in the merged area refers to the total number of non-overlapping pixels occupied by both sets of boxes. A fixed CIRR threshold of 0.4 is pre-set. All pairwise paired bounding boxes in the candidate bounding box set are traversed, and the CIRR is calculated for each pair. When the CIRR of any two sets of bounding boxes is greater than 0.4, it means that the two sets of boxes point to the same face target in the image, and are considered duplicate bounding boxes. The face confidence scores of the two sets of duplicate bounding boxes are compared. The confidence score represents the reliability of the algorithm in determining that the region is a face. The set of bounding boxes with the lower confidence score is directly discarded, and only the set with the higher confidence score is retained as a valid bounding box. After a complete traversal and comparison of all candidate boxes and the elimination of redundancy, the remaining non-repeating and non-overlapping detection boxes are the accurate face detection boxes. Each set of face detection boxes can accurately correspond to the unique face target in the image, eliminating the interference of repeated recognition caused by block detection and providing basic coordinate data for cross-frame trajectory construction.
[0035] After obtaining a stable and non-redundant face detection bounding box in a single frame, the process proceeds to cross-frame face trajectory construction. "Cross-frame" refers to multiple consecutive frames within the video stream that exist in chronological order. The face trajectory is used to connect the positions of the same target person appearing in different frames, achieving continuous cross-frame binding of the person's identity. The trajectory matching determination has a coordinate difference constraint standard. The center coordinates of two sets of face detection bounding boxes in two adjacent frames are taken. The horizontal coordinate of the detection box center = the horizontal coordinate of the top-left corner of the detection box plus the horizontal coordinate of the bottom-right corner of the detection box ÷ 2; the vertical coordinate of the detection box center = the vertical coordinate of the top-left corner of the detection box plus the vertical coordinate of the bottom-right corner of the detection box ÷ 2. The difference between the horizontal and vertical coordinates of the corresponding face bounding boxes in the preceding and following frames is calculated separately. If both differences are less than 80 pixels, the two sets of detection boxes are determined to belong to the same target person, and a connection is established between the two sets of detection boxes using temporal markers. By continuously matching face detection bounding boxes of multiple consecutive frames in chronological order, the face bounding boxes of multiple frames belonging to the same person are sequentially linked to form a complete and continuous cross-frame face trajectory. A complete face trajectory records the face position, image acquisition timestamp, and face confidence score of each frame in the entire process from the appearance to the disappearance of the target person, thus avoiding the problem of broken identity of the person caused by brief missed detection in single-frame face detection.
[0036] Facial feature vectors corresponding to the target biometric data are extracted based on cross-frame facial trajectories. The extraction operation is performed only on the facial image regions corresponding to all valid face detection boxes within the trajectory. Three preprocessing operations are performed on the facial regions: alignment of facial key points, illumination normalization, and uniform scaling to eliminate feature deviations caused by shooting angle and lighting conditions. After processing, a facial feature vector with a fixed length of 512 floating-point numbers is output, where each independent number corresponds to the weight value of a local texture feature on the face. All facial feature vectors extracted from the same facial trajectory are averaged and fused. The fused result is a unique set of standard facial feature vectors used as the identification identifier for the target biometric data. These facial feature vectors can quantify and distinguish the differences in facial features between different individuals.
[0037] Based on the associated information of the data collection devices, the main retrieval partition and the neighboring associated partition are divided, which specifically includes the following steps: Determine the spatial layout of the acquisition devices and the video surveillance coverage area based on the associated information of the acquisition devices; Based on the spatial deployment location of the acquisition devices, the target acquisition devices for matching points of target biometric data are selected, and the target acquisition devices and their corresponding video surveillance coverage areas are defined as the main search partitions. The device regions that are interconnected with the main search partition, spatially adjacent, and have continuously connected trajectories are designated as neighborhood association partitions. The system reads and analyzes the associated information of the acquisition devices, extracting the spatial deployment location and video surveillance coverage range for each device. The spatial deployment location is defined by horizontal and vertical coordinates, used to mark the device installation points within the overall scene plane. The video surveillance coverage range includes the core values of the device's maximum horizontal and vertical shooting distances. The maximum horizontal shooting distance is calculated as: Lens focal length (in millimeters) ÷ 1000 × Wide-angle magnification factor × Device installation height (in meters). During the calculation, the focal length in millimeters is converted to meters to ensure consistent units for all calculated values. For example, using an acquisition device with a lens focal length of 8 millimeters, a wide-angle magnification factor of 12, and an installation height of 3 meters, the maximum horizontal shooting distance is calculated as: 8 ÷ 1000 × 12 × 3 = 0.288 meters. This value represents the furthest horizontal spatial range that this device can capture. All two sets of coordinates and coverage distance data for all acquisition devices are stored in a spatial device index table.
[0038] The system filters and identifies acquisition devices that match target biometric locations and defines the main search partition. It reads the coordinates of the initial appearance of the target person obtained from the previous face trajectory matching. These coordinates are then compared to the monitoring coverage area of each acquisition device in the index table. The criteria are: the absolute value of the difference between the target person's horizontal coordinate and the device's horizontal coordinate is less than the device's maximum horizontal shooting distance; simultaneously, the absolute value of the difference between the target person's vertical coordinate and the device's vertical coordinate is less than the device's maximum vertical shooting distance. Acquisition devices that satisfy both of these constraints are considered target acquisition devices. For example, if the target person's initial appearance location has a horizontal coordinate of 40 meters and a vertical coordinate of 25 meters, and a certain acquisition device has a horizontal coordinate of 38 meters and a vertical coordinate of 24 meters, a maximum horizontal shooting distance of 3 meters, and a maximum vertical shooting distance of 2 meters, the absolute value of the horizontal coordinate difference is 40 - 38 = 2 meters, and the absolute value of the vertical coordinate difference is 25 - 24 = 1 meter. Since both differences are less than the corresponding coverage distances, this device is identified as the target acquisition device. All the selected target acquisition devices, along with the video surveillance coverage area corresponding to each target acquisition device, are integrated into a complete closed space. This closed space is designated as the main retrieval zone, which is the core area where the target person is first identified. All face retrieval operations are performed preferentially within this zone.
[0039] The system filters surrounding device areas that meet three constraints to generate a neighborhood association partition. These constraints are: the main search partition and the device to be judged have mutually accessible views, the two spatial locations are adjacent, and the movement trajectory of a person is continuous. Mutually accessible views mean that the monitoring ranges of the two devices overlap by at least 1 meter. Spatially adjacent means that the straight-line distance between the device to be judged and any target acquisition device within the main search partition is less than or equal to the average monitoring coverage distance of the two devices. Continuously connected trajectories mean that there are unobstructed passageways such as passageways, escalators, or corridors between the two areas, allowing the target person to move directly from the main search partition to the coverage area of the device to be judged without passing through areas outside the scene. Taking the main search partition of a shopping mall atrium as an example, acquisition devices connecting the atrium to escalator passageways and store outdoor areas simultaneously meet all three constraints: a 2-meter overlap between the escalator devices and the atrium devices, a straight-line distance of 6 meters between the two points, and an average monitoring coverage distance of 10 meters between the two devices. These values meet the adjacentity criteria, and there is an unobstructed passageway between the atrium and the escalator. Therefore, the monitoring areas corresponding to these devices will all be included in the neighborhood association partition. The neighborhood association partition is used to cover all the surrounding spaces that the target person may move to after leaving the main search partition. The server synchronously retrieves the existing biometric data of the main search partition and the neighborhood association partition for cross-partition linkage verification, fully capturing the facial features of the target person that appear continuously throughout the process, avoiding the loss of video footage of the person's movement when only the core partition is searched.
[0040] The server synchronously retrieves existing biometric data from the main search partition and neighboring related partitions; The effective feature data is obtained by performing cross-regional linkage verification on the facial feature vector of the target biometric data based on existing biometric data. The specific steps include: Based on the existing biometric data, feature matching and comparison are performed within the main retrieval partition, and feature content that matches the facial feature vector is selected to obtain the main partition matching feature data. The neighborhood matching association feature data is obtained based on the primary partition matching feature data, specifically: Using the main partition matching feature data as the comparison benchmark, cross-domain association matching is performed by linking the existing biometric data corresponding to the neighboring related partitions to obtain the neighboring matching association feature data. The consistency of the primary partition matching feature data and the neighboring matching associated feature data is verified to obtain valid feature data. The similarity calculation is performed one by one between the target face feature vector and the existing face feature vectors in the main search partition. The calculation uses cosine similarity as the matching criterion. The cosine similarity calculation formula is: Vector similarity value = (Sum of the product of corresponding values of the two feature vectors) ÷ (Sum of the squares of all values of the first feature vector) × (Sum of the squares of all values of the second feature vector). A matching threshold of 0.75 is set in advance. When the vector similarity value calculated between any existing face feature vector and the target face feature vector is greater than 0.75, the two face feature vectors are considered to match. The matching existing biometric data is completely retained. All the selected and retained matching feature content is integrated to form the main partition matching feature data. The main partition matching feature data completely records the corresponding device location, appearance time, and face similarity value for each appearance of the target person in the main search partition, completely eliminating irrelevant passerby feature data in the main partition that have significant differences from the target person's facial features. Taking a face feature vector with a length of 512 dimensions as an example, the sum of the two sets of vectors obtained by multiplying each digit is equal to 1260. The square root of the first set of vectors is equal to 42, and the square root of the second set of vectors is equal to 33. Substituting into the formula, we can get the vector similarity value = 1260 ÷ 42 × 33 ≈ 0.909. This value is greater than the judgment threshold of 0.75, and the corresponding existing features are included in the main partition matching feature data.
[0041] Based on the main partition matching feature data, neighboring matching associated feature data is extracted. Using the main partition matching feature data as the global comparison benchmark, two key constraint parameters are extracted from the main partition matching feature data: the time interval of the target person's first appearance and the core spatial coordinates. This is then used in conjunction with reading all existing biometric data from the neighboring associated partition cache for matching operations, overlaying spatiotemporal constraints. The image capture time corresponding to the existing feature to be compared within the neighboring partition must be within the range of the start time of the person's appearance recorded in the main partition matching feature minus 30 seconds to the end time of the person's appearance plus 30 seconds. Simultaneously, the straight-line distance between the spatial coordinates of the device bound to the existing feature and the core spatial coordinates of the main partition cannot exceed 1.5 times the average monitoring coverage distance of the devices in the main retrieval partition. Neighboring existing features that simultaneously meet the similarity threshold constraint and spatiotemporal range constraint are integrated and summarized to form neighboring matching associated feature data. This set of data completely records the process of the target person moving from the main retrieval partition to the surrounding adjacent areas, including all facial feature information appearing at each device location in the neighboring partition. Passerby feature data within the neighboring partition that does not match the spatiotemporal trajectory of the target person are filtered out.
[0042] A consistency verification was conducted between the main partition matching feature data and the neighboring partition matching associated feature data. The verification consisted of two parallel judgment dimensions: temporal continuity verification and spatial coherence verification. The temporal continuity verification read all timestamp sequences in both sets of data and calculated the difference between the end time of the disappearance of the person recorded in the main partition and the start time of the appearance of the person recorded in the neighboring partition. The temporal continuity difference = the start time of the person in the neighboring partition - the end time of the person in the main partition. When the temporal continuity difference is greater than or equal to 0 seconds and less than or equal to 120 seconds, it is determined that there is no discontinuity in the temporal sequence, which means that there is no long-term disappearance gap in the process of the target person moving from the main partition to the neighboring partition. The spatial coherence verification extracted the coordinates of all device points in both sets of data. The straight-line distance between points corresponding to adjacent temporal sequences cannot exceed twice the maximum monitoring coverage distance of a single device to avoid the two sets of matching data pointing to two unrelated points that are extremely far apart in the scene. Only the main partition matching features and the neighborhood matching associated features that pass both temporal continuity verification and spatial coherence verification will be marked and integrated into valid feature data. Matching data with temporal discontinuities or excessive spatial spans will be directly removed to eliminate cross-partition false matching interference caused by passersby with similar appearances.
[0043] The matching results are obtained by integrating effective feature data, including device location, spatial location, and continuous temporal information. This process specifically includes the following steps: By combining the correlation between equipment locations, spatial positions, and temporal arrangements in the effective feature data, feature correlation information is obtained by performing global correlation fitting on the feature sequence; Based on the feature association information, the correspondence between device ownership, spatial location and time is solidified to generate matching results; The matching results include device identifier, location identifier, and associated timestamp; The device location parameters, spatial location parameters, and temporal arrangement parameters built into each set of valid feature data are extracted. The device location parameters correspond to the digital code of the acquisition device, which is used to distinguish different monitoring hardware. The spatial location parameters consist of two sets of values: horizontal metric coordinates and vertical metric coordinates, which are used to mark the physical installation points of the device in the scene plane. The temporal arrangement parameters are the image acquisition timestamps, which are used to record the order in which the target person appears in the monitoring screen. There is a natural binding relationship between the three types of parameters. Within the same set of valid feature data, the three types of parameters belong to the same round of face recognition records. Multiple sets of valid feature data form a messy and disordered feature sequence. The global correlation fitting operation sorts out the internal logic of the three types of parameters within all feature sequences and reconstructs a coherent and unified trajectory data of the person.
[0044] Using temporal arrangement parameters as the primary sorting criterion, all valid feature data are sorted in ascending order according to their associated timestamp values. After sorting, spatial continuity verification is performed on adjacent feature records. The verification is calculated using the formula for the straight-line distance between two points in a plane. The preset spatial continuity judgment threshold is equal to 1.5 times the average monitoring coverage distance of a single device. When the calculated straight-line distance between two adjacent feature records is less than this threshold, the two records are determined to be identification data generated by continuous movement of a person, and a temporal connection mark is established between the two records. Taking a shopping mall scenario as an example, the timestamp of the first valid feature record is equal to 100 seconds, the horizontal coordinate of the device is equal to 50 meters, and the vertical coordinate is equal to 30 meters. The timestamp of the adjacent second record is equal to 112 seconds, the horizontal coordinate of the device is equal to 56 meters, and the vertical coordinate is equal to 34 meters. Substituting into the distance calculation formula, the straight-line distance between the two points in a plane is approximately 7.21 meters. If the average monitoring coverage distance of a single device in this scenario is equal to 6 meters, the continuity judgment threshold is equal to 6 × 1.5 = 9 meters. Since 7.21 meters is less than 9 meters, the two records are marked as continuously associated data. Traverse all sorted feature records and complete the connection marking, link the scattered and independent single effective features into a complete and unbroken spatiotemporal movement link of the person, and output standardized feature association information after integrating all link information.
[0045] The process solidifies the one-to-one correspondence between three types of information and outputs the matching results. The solidification operation binds three fixed sets of correspondences to each independent spatiotemporal recognition record within the feature association information. The first set of correspondences is device affiliation binding, where each record is matched with a unique digital code corresponding to the acquisition device, clearly identifying which hardware device acquired the face recognition data. The second set of correspondences is spatial orientation binding, locking the device's unique horizontal and vertical coordinates to this record, marking the physical area where the recognition occurred. The third set of correspondences is time binding, binding the start and end timestamps of the face's appearance to this record, locking the complete duration of the target person's continuous appearance within the device's view. These three sets of correspondences are interdependent and inseparable, avoiding chaotic mappings where one device affiliation corresponds to multiple sets of spatial orientations or multiple timestamps. After all binding logic is solidified, the matching results are encapsulated using a unified data structure.
[0046] The matching results have a fixed data structure. Each matching result entry carries three core identifier fields. The first field is the device identifier, which directly uses the digital code of the acquisition device and can directly index the video data stream stored on the corresponding device. The second field is the location identifier, which integrates the horizontal and vertical coordinates of the device to form a spatial positioning code, used to automatically match video templates of the corresponding area style. The third field is the associated timestamp, which contains the start and end seconds of the target person's appearance, used to extract surveillance video clips for the corresponding time period. The originally scattered and disordered effective feature data is transformed into a clearly structured matching result with complete spatiotemporal logic, eliminating the data fragmentation problem caused by different partitions and different devices. This ensures that the video material extraction and video synthesis processes can retrieve all surveillance video data of the target person at once, thus avoiding problems such as missing clips or misplaced locations.
[0047] After filtering the corresponding video templates based on the matching results, the target video is generated, which includes the following steps: Extract the corresponding video segment based on the matching results; All video clips are spliced together in chronological order, and the intro, outro, transition effects, and background music associated with the video template are overlaid to generate the target video.
[0048] The device identifier is used to locate the corresponding video stream file of the acquisition device within the storage medium. The associated timestamp contains the start and end seconds of the target person appearing on the monitoring screen. To avoid truncated footage where the person is just entering or about to leave the frame, the captured time interval is extended with a delay. The extension rule is that the start second of the segment capture equals the original start second of the matching record minus 5, the end second of the segment capture equals the original end second of the matching record plus 5, and the total duration of a single video segment equals the end second of the segment capture minus the start second of the segment capture. Taking a matching record as an example, if the original start second is 120 and the original end second is 150, the segment capture start second = 120 - 5 = 115, the segment capture end second = 150 + 5 = 155, and the total duration of a single video segment = 155 - 115 = 40, meaning that this matching record will capture a monitoring video segment with a total duration of 40 seconds. Based on the extended time interval corresponding to each matching record, the video decoding tool is called to extract independent and lossless video segments from the original video stream of the corresponding device. Each matching record will generate a video segment. After all segments are extracted, they are stored in a temporary cache space, waiting to enter the splicing and synthesis process.
[0049] Once all corresponding video segments have been extracted, the time-series splicing and multi-layer material overlay synthesis operations are initiated. The primary sorting criterion for the splicing operation is the starting value of the original associated timestamp corresponding to each video segment. All cached video segments are linearly spliced in ascending order of their starting seconds. The splicing process eliminates subtle differences in image resolution and bitrate between segments, uniformly converting them to the preset image resolution and bitrate parameters of the video template, ensuring a smooth and stutter-free basic video stream after splicing. After the basic video stream is spliced, the video template bound to the matching result position identifier is retrieved, and the four types of multimedia materials pre-stored in the template are read and layered onto the basic video stream. The four types of materials are intro material, outro material, transition effect material, and background music audio material. The layering of the four types of materials has a fixed layer arrangement order. The bottom layer is the spliced basic monitoring image video stream, the second layer intersperses transition effect material, the third layer overlays the background music audio track throughout, the intro material is overlaid at the very beginning of the video, and the outro material is overlaid at the very end of the video.
[0050] Transition effects are automatically filled into the gaps between two adjacent video clips. Each transition effect occupies a fixed 2-second duration. If the spliced base video stream contains N independent video clips, the total number of transition effect segments to be inserted is equal to N minus 1, and the total duration of the transition effects is equal to N minus 1 multiplied by 2. Taking four clipped video clips in the cache as an example, the total number of transition effect segments to be inserted is 4 - 1 = 3, and the total duration of the transition effects is 3 × 2 = 6 seconds. The three transition effects are filled sequentially at the transition positions between the first and second clips, the second and third clips, and the third and fourth clips, eliminating the abrupt disconnect caused by direct switching between two monitoring screens. The background music audio is uniformly adapted to the total playback time of the complete spliced video. The audio start time is synchronized with the start time of the opening clip, and the audio end time is synchronized with the end time of the closing clip. The audio volume is fixed at 30% of the original human voice volume value monitored, so as to preserve the original sound of the scene in the monitoring screen and prevent the background music from masking the human voice content in the screen.
[0051] After all the layered material overlay rendering operations are completed, the complete video data stream is packaged according to the preset encoding format of the video template. The packaged complete file is the final target video. The entire process relies on the matching results to complete the precise and complete extraction of video materials. Then, through standardized splicing and overlay of multiple layers of multimedia materials, the scattered monitoring clips are transformed into a finished video with a coherent narrative and complete packaging effect.
[0052] A video material retrieval and generation system for target person identification includes: Receiving module: Receives target biometric data and scene identifiers sent by the user terminal; Acquisition module: Retrieves the acquisition device association information for the corresponding scene based on the scene identifier; the acquisition device association information includes the mapping relationship between the device and the video channel; Sending module: Sends target biometric data and acquisition device association information to the server; Verification module: Based on the association information of the acquisition device, the main retrieval partition and the neighboring association partition are divided. The server synchronously retrieves the existing biometric data in the main retrieval partition and the neighboring association partition. Based on the existing biometric data, the face feature vector of the target biometric data is verified across partitions to obtain valid feature data. Processing module: Integrates valid feature data, including device location, spatial location, and continuous temporal information, to obtain matching results; the matching results include device identifier, location identifier, and associated timestamp; Generation module: Generates the target video after filtering the corresponding video templates based on the matching results.
[0053] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0054] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0055] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for retrieving and generating video footage for target person identification, characterized in that, Includes the following steps: Receive target biometric data and scene identifiers of the target person; Obtain the acquisition device association information for the corresponding scene based on the scene identifier; the acquisition device association information includes the mapping relationship between the device and the video channel; Send the target biometric data and the associated information of the acquisition device to the server; Based on the associated information of the data collection devices, the main retrieval partition and the neighboring associated partition are divided, which specifically includes the following steps: Determine the spatial layout of the acquisition devices and the video surveillance coverage area based on the associated information of the acquisition devices; Based on the spatial deployment location of the acquisition devices, the target acquisition devices for matching points of target biometric data are selected, and the target acquisition devices and their corresponding video surveillance coverage areas are defined as the main search partitions. The device regions that are interconnected with the main search partition, spatially adjacent, and have continuously connected trajectories are designated as neighborhood association partitions. The server synchronously retrieves existing biometric data from the main search partition and neighboring related partitions, and performs cross-partition linkage verification on the face feature vector of the target biometric data based on the existing biometric data to obtain valid feature data. The matching result is obtained by integrating the device location, spatial location, and continuous temporal information of the effective feature data; the matching result includes device identifier, location identifier, and associated timestamp; The target video is generated after filtering the corresponding video templates based on the matching results.
2. The method for retrieving and generating video footage for target person identification according to claim 1, characterized in that, The acquisition device association information is bound to a video channel with an independent encoding identifier.
3. The method for retrieving and generating video footage for target person identification according to claim 1, characterized in that, It also includes the following steps: Block detection and non-maximum suppression are performed on the target biometric data to obtain face detection bounding boxes; A cross-frame face trajectory is established based on the face detection bounding box, and the face feature vector of the target biometric data is extracted based on the face trajectory.
4. The method for retrieving and generating video footage for target person identification according to claim 3, characterized in that, To obtain face detection bounding boxes, block detection and non-maximum suppression are performed on the target biometric data. This process includes the following steps: The target biometric data is divided into multiple sub-blocks with overlapping regions, and face detection is performed on each sub-block to obtain a set of candidate detection boxes; Non-maximum suppression is performed on the candidate detection box set to remove duplicate detection boxes and obtain the face detection box.
5. The method for retrieving and generating video footage for target person identification according to claim 1, characterized in that, After filtering the corresponding video templates based on the matching results, the target video is generated, which includes the following steps: Extract the corresponding video segment based on the matching results; All video clips are spliced together in chronological order, and the intro, outro, transition effects, and background music associated with the video template are overlaid to generate the target video.
6. The method for retrieving and generating video footage for target person identification according to claim 5, characterized in that, The effective feature data is obtained by performing cross-regional linkage verification on the facial feature vector of the target biometric data based on existing biometric data. The specific steps include: Based on the existing biometric data, feature matching and comparison are performed within the main retrieval partition, and feature content that matches the facial feature vector is selected to obtain the main partition matching feature data. Neighborhood matching association feature data is obtained based on the primary partition matching feature data; The consistency of the primary partition matching feature data and the neighboring matching associated feature data is verified to obtain valid feature data.
7. The method for retrieving and generating video footage for target person recognition according to claim 6, characterized in that, The matching results are obtained by integrating effective feature data, including device location, spatial location, and continuous temporal information. This process specifically includes the following steps: By combining the correlation between equipment locations, spatial positions, and temporal arrangements in the effective feature data, feature correlation information is obtained by performing global correlation fitting on the feature sequence; Based on the feature association information, the device's ownership, spatial location, and time correspondence are solidified to generate matching results.
8. The method for retrieving and generating video footage for target person identification according to claim 7, characterized in that, The neighborhood matching association feature data is obtained based on the primary partition matching feature data, specifically: Using the main partition matching feature data as the comparison benchmark, cross-domain association matching is performed on the existing biometric data corresponding to the neighboring related partitions to obtain the neighboring matching association feature data.
9. A video material retrieval and generation system for target person recognition, applied to the video material retrieval and generation method for target person recognition as described in any one of claims 1-8, characterized in that, include: Receiving module: Receives target biometric data and scene identifiers sent by the user terminal; Acquisition module: Acquires the acquisition device association information corresponding to the scene based on the scene identifier; the acquisition device association information includes the mapping relationship between the device and the video channel; Sending module: Sends target biometric data and acquisition device association information to the server; Verification module: Based on the associated information of the acquisition devices, the main retrieval partition and the neighboring associated partition are divided, specifically including the following steps: Determine the spatial layout of the acquisition devices and the video surveillance coverage area based on the associated information of the acquisition devices; Based on the spatial deployment location of the acquisition devices, the target acquisition devices for matching points of target biometric data are selected, and the target acquisition devices and their corresponding video surveillance coverage areas are defined as the main search partitions. The device regions that are interconnected with the main search partition, spatially adjacent, and have continuously connected trajectories are designated as neighborhood association partitions. The server synchronously retrieves existing biometric data from the main search partition and neighboring related partitions, and performs cross-partition linkage verification on the face feature vector of the target biometric data based on the existing biometric data to obtain valid feature data. Processing module: Integrates device location, spatial location, and continuous temporal information from valid feature data to obtain matching results; the matching results include device identifier, location identifier, and associated timestamp; Generation module: Generates the target video after filtering the corresponding video templates based on the matching results.
Citation Information
Patent Citations
Camera monitoring video real-time identification data processing method and system
CN121482696A
Cross-video person location tracking method and system, and device
WO2021196294A1