Face information-based multi-video target fusion method and related equipment

By identifying overlapping areas in the field of view of adjacent cameras and merging targets with the same facial information, the problem of real-time and continuous positioning of personnel targets in multi-camera monitoring is solved, and stable fusion and display of personnel identity and trajectory are achieved in multi-camera scenarios.

CN121921822APending Publication Date: 2026-04-24BEIJING SINOITS TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
BEIJING SINOITS TECH
Filing Date
2025-12-25
Publication Date
2026-04-24

Smart Images

  • Figure CN121921822A_ABST
    Figure CN121921822A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-video target fusion method based on face information and related equipment, and relates to the technical field of target fusion, and the method comprises the steps: recognizing a camera with a view range having an overlapping region from each camera deployed in a preset space region as an adjacent camera; filtering out the trajectory information of the person target which does not carry the effective face information; for the filtered trajectory information of the personnel targets carrying the effective face information, identifying the personnel targets with the same face information in the trajectory information of the personnel targets from different adjacent cameras, combining the personnel targets into the same global personnel target, and then determining the spatial position coordinates of the global personnel target; and carrying out front-end map display on the spatial position coordinates of the global personnel target. According to the invention, a clear, unique and continuous real-time personnel position and motion track view can be provided for a user, and the problems of tracking loss and display interruption caused by shielding or face departure in a single camera scheme are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of target fusion technology, and in particular to a multi-video target fusion method and related equipment based on facial information. Background Technology

[0002] With the increasing maturity of facial recognition technology, the demand for accurate, real-time location of individuals' coordinates in three-dimensional space is becoming increasingly prominent. Especially in multi-camera monitoring scenarios, when multiple cameras simultaneously detect the same person, the front-end map system needs to effectively integrate location information (latitude and longitude coordinates or meter coordinates) from different data sources that may represent the same individual to ensure the uniqueness and continuity of the target displayed on the front-end map. Therefore, solving the problem of personnel location fusion in multi-camera scenarios, achieving unique identification of individuals on the map, and providing stable and reliable data support for historical trajectory tracking has become a key technical challenge.

[0003] Existing technical solutions include a simplified approach. To reduce system complexity and processing costs, this approach typically deploys a single face recognition camera in a single room or a limited, pre-defined spatial area. This camera captures video streams, performs face detection and recognition, and attempts to track and locate the target.

[0004] However, the aforementioned existing technical solutions have significant drawbacks. Due to the limited field of view of a single camera, when a person moves to the edge of the camera's field of view, is obscured by obstacles, or has their face turned away from the camera, the system is prone to face detection failure. This directly leads to the loss of tracking of the person and interruption of trajectory information. Therefore, it is difficult to achieve real-time, continuous, and accurate positioning of people in complex spaces, and it is even more unable to meet the requirement of unified fusion of target identity and trajectory in multi-camera collaborative scenarios.

[0005] To overcome the limitations of single-camera monitoring and improve the robustness and accuracy of personnel target localization and display in multi-camera scenarios, a technical solution that can effectively integrate multi-channel video information is needed. Summary of the Invention

[0006] The technical problem to be solved by the present invention is to address the shortcomings of the prior art, specifically by providing a multi-video target fusion method and related equipment based on facial information, as detailed below: 1) In a first aspect, the present invention provides a multi-video target fusion method based on facial information, the specific technical solution of which is as follows: The system identifies cameras with overlapping fields of view from each camera deployed within a predefined spatial area as adjacent cameras; it filters out trajectory information of personnel targets that do not carry valid facial information; for the trajectory information of personnel targets carrying valid facial information after filtering, it identifies personnel targets with the same facial information from different adjacent cameras; it merges the identified personnel targets with the same facial information into a single global personnel target, and uses the target identifier in the trajectory information of the personnel target with the longest trajectory length before merging as the identifier of the global personnel target, and uses the spatial geographic coordinates in the trajectory information of the personnel target with the longest trajectory length as the spatial location coordinates of the global personnel target; and it displays the spatial location coordinates of the global personnel target on a front-end map.

[0007] The beneficial effects of the multi-video target fusion method based on facial information provided by this invention are as follows: By identifying cameras with overlapping fields of view from each camera deployed within a predefined spatial area as adjacent cameras, a collaborative relationship network between cameras is constructed, laying the foundation for multi-source data fusion. By filtering out trajectory information of personnel targets lacking valid facial information, the quality of the input fusion process is ensured, avoiding interference from invalid or low-reliability information. By identifying personnel targets with identical facial information from different adjacent cameras, it is possible to accurately determine whether targets in different video streams belong to the same individual, thus solving the problem of personnel identity confusion in multi-camera scenarios. By merging identified personnel targets with identical facial information into a single global personnel target, and using the target identifier and spatial geographic coordinates from the trajectory information of the personnel target with the longest trajectory length before merging, a unique and stable identifier and location determination for the same person is achieved at the system level, effectively ensuring the continuity of target identity and trajectory consistency. Finally, by displaying the spatial location coordinates of the global personnel target on a front-end map, a clear, unique, and continuous real-time view of personnel location and motion trajectory is provided to users, overcoming the tracking loss and display interruption problems caused by occlusion or facial misalignment in single-camera solutions.

[0008] Based on the above scheme, the multi-video target fusion method based on face information of the present invention can be further improved as follows.

[0009] Furthermore, it also includes: establishing a coordinate mapping relationship between the image pixel coordinates and spatial geographic coordinates of each camera; processing the video stream acquired in real time by each camera frame by frame, detecting each person target contained in each image frame, and obtaining the image pixel coordinates and corresponding facial information of each person target in that image frame; for each camera, for each person target in the image frame acquired by the camera at each moment, using the coordinate mapping relationship established for the camera, converting the image pixel coordinates of the person target in the current image frame into the spatial geographic coordinates of the person target at the current moment; for each detected person target, generating and sending the trajectory information of the person target in real time based on the spatial geographic coordinates and corresponding facial information of the person target at consecutive moments; wherein, the trajectory information of the person target includes: target identifier, spatial geographic coordinates, facial information, and the camera device identifier for acquiring the facial information.

[0010] The beneficial effects of adopting the above-mentioned further scheme are as follows: By establishing a coordinate mapping relationship from image pixel coordinates to spatial geographic coordinates for each camera, a precise conversion from two-dimensional image information to a unified three-dimensional physical space is achieved. By processing the video stream of each camera frame by frame and detecting personnel targets, the image pixel coordinates and corresponding facial information of each personnel target can be obtained in real time. Further utilizing the coordinate mapping relationship, the image pixel coordinates are converted into the spatial geographic coordinates of the personnel targets, thereby assigning an actual physical location to each target. Finally, based on the spatial geographic coordinates and facial information at continuous time points, trajectory information of personnel targets containing target identification, spatial geographic coordinates, facial information, and camera device identification is generated and transmitted in real time, providing standardized, complete data input with spatiotemporal and identity information for subsequent multi-camera target fusion.

[0011] Furthermore, it also includes recording the identifier and corresponding facial information of each global human target obtained after fusion into the global target queue.

[0012] The beneficial effects of adopting the above-mentioned further scheme are as follows: By recording the identifier and corresponding facial information of each global person target obtained after fusion into a global target queue, a unified identity mapping library is constructed. This recording ensures a stable binding between each global person target confirmed by fusion and its key biometric features. When processing the trajectory information of new person targets later, the existing global person target identifier can be quickly retrieved from the global target queue based on the facial information, thereby ensuring that the identity identifier of the same person remains consistent and continuous across different times or different camera views, effectively supporting long-term, cross-camera continuous tracking and status maintenance.

[0013] Furthermore, it also includes: when the trajectory information of a person target carrying valid facial information is received again, a search and match is performed from the global target queue based on the facial information.

[0014] The beneficial effects of adopting the above-mentioned further solution are: when trajectory information of a person carrying valid facial information is received again, the existing global person target identifier can be quickly searched and matched from the global target queue based on this facial information. This process achieves continuous identification and association of the same person's identity, effectively avoiding the need to repeatedly create new global person target identifiers due to the target briefly leaving the field of view or switching cameras. By reusing historical identity identifiers, the uniqueness of the person's identity and the continuity of their trajectory are ensured throughout the entire monitoring time and space range, enhancing the stability of long-term tracking and the consistency of data.

[0015] 2) In a second aspect, the present invention also provides a multi-video target fusion system based on facial information, the specific technical solution of which is as follows: The system includes a neighboring camera determination module, a filtering module, a recognition module, a spatial location coordinate acquisition module, and a display module. The neighboring camera determination module identifies cameras with overlapping fields of view from each camera deployed within a preset spatial area as neighboring cameras. The filtering module filters out trajectory information of personnel targets that do not carry valid facial information. The recognition module identifies personnel targets with the same facial information from different neighboring cameras within the trajectory information of the filtered personnel targets carrying valid facial information. The spatial location coordinate acquisition module merges the identified personnel targets with the same facial information into a single global personnel target, using the target identifier in the trajectory information of the personnel target with the longest trajectory length before merging as the identifier of this global personnel target, and using the spatial geographic coordinates in the trajectory information of the personnel target with the longest trajectory length as the spatial location coordinates of this global personnel target. The display module displays the spatial location coordinates of the global personnel targets on a front-end map.

[0016] Based on the above scheme, the multi-video target fusion system based on face information of the present invention can be further improved as follows.

[0017] Furthermore, it also includes a trajectory information acquisition module, which is used to: establish a coordinate mapping relationship between the image pixel coordinates and spatial geographic coordinates of each camera; process the video stream acquired in real time by each camera frame by frame, detect each person target contained in each image frame, and acquire the image pixel coordinates and corresponding face information of each person target in that image frame; for each camera, for each person target in the image frame acquired by the camera at each moment, use the coordinate mapping relationship established for the camera to convert the image pixel coordinates of the person target in the current image frame into the spatial geographic coordinates of the person target at the current moment; for each detected person target, generate and send the trajectory information of the person target in real time based on the spatial geographic coordinates and corresponding face information of the person target at consecutive moments; wherein, the trajectory information of the person target includes: target identifier, spatial geographic coordinates, face information, and the camera device identifier for acquiring the face information.

[0018] Furthermore, it also includes a recording module, which is used to record the identifier and corresponding facial information of each global human target obtained after fusion to the global target queue.

[0019] Furthermore, it also includes a search and matching module, which is used to: when the trajectory information of a person target carrying valid facial information is received again, perform a search and matching from the global target queue based on the facial information.

[0020] 3) In a third aspect, the present invention also provides an electronic device, the electronic device including a processor coupled to a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor, so as to enable the electronic device to implement any of the above-mentioned multi-video target fusion methods based on face information.

[0021] 4) In a fourth aspect, the present invention also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the above-mentioned multi-video target fusion methods based on facial information.

[0022] It should be noted that the beneficial effects of the technical solutions of the second to fourth aspects of the present invention and their corresponding possible implementations can be found in the above description of the technical effects of the first aspect and its corresponding possible implementations, and will not be repeated here. Attached Figure Description

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments of the present invention will be briefly introduced below: Figure 1This is a flowchart illustrating a multi-video target fusion method based on facial information according to an embodiment of the present invention. Figure 2 A spatial plan of the pre-defined spatial area; Figure 3 This is a schematic diagram of the structure of a multi-video target fusion system based on facial information according to an embodiment of the present invention. Detailed Implementation

[0024] The principles and features of the present invention are described below. The examples given are only for explaining the present invention and are not intended to limit the scope of the present invention.

[0025] The technical solution of the present invention and how the technical solution of the present invention solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of the present invention will now be described with reference to the accompanying drawings.

[0026] like Figure 1 As shown in the figure, a multi-video target fusion method based on facial information according to an embodiment of the present invention includes the following steps: S1. Identify cameras with overlapping fields of view from each camera deployed within a preset spatial area as adjacent cameras. The specific implementation process is as follows: S10. Access a configuration database that stores basic information about all cameras. The database records a key piece of data for each camera deployed within a preset spatial area: the camera device identifier. Simultaneously, it needs to retrieve the coordinate mapping parameters that are independently established and stored for each camera. Furthermore, if information exists recording physical parameters such as camera installation location and horizontal and vertical field of view angles, it should also be retrieved. This data is the raw input for calculating the actual monitoring coverage area of ​​each camera.

[0027] S11. For each camera, perform a reverse derivation calculation using its coordinate mapping relationship. Specifically, take the image pixel coordinates of the four corner points of the camera's image frame, such as the top left, top right, bottom right, and bottom left corners. Transform these corner point image pixel coordinates into a spatial geographic coordinate system using the coordinate mapping relationship established for the camera. The four resulting spatial geographic coordinate points roughly define a quadrilateral area that the camera can cover in geographic space. This quadrilateral can be considered as a projected polygon of the camera's effective field of view on the ground plane. For more accurate calculations, more points can be taken from the image, such as adding a center point or dividing points, to generate a polygon with more sides to better match the actual field of view shape.

[0028] S12. Generate a list containing all camera device identifiers, and systematically combine them in pairs to form all possible camera pairs. For each camera pair, extract the field-of-view projection polygons corresponding to the two cameras from the calculation results of S11. Then, use a polygon intersection detection algorithm from computational geometry to analyze the two polygons. The algorithm determines whether the two polygons have a common intersection region, or whether the boundary of at least one polygon is contained within the interior of the other polygon. If the detection algorithm confirms that the two polygons have a common area of ​​any size, regardless of the size of the common area, it determines that the field of view of the two cameras overlaps.

[0029] S13. For each camera pair identified as having overlapping fields of view, record this relationship as a data item. The recorded data item must contain at least the camera device identifiers of both cameras to uniquely identify which two cameras are adjacent. Summarize all such data items to form a global list of adjacent camera relationships. This list can be stored as a configuration file, such as in JSON or XML format, or in a specific table in a relational database. Each record in the list explicitly indicates which two cameras have overlapping fields of view, requiring association processing with detected personnel targets in subsequent fusion services.

[0030] S14. When the fusion service program starts or initializes, it reads the adjacent camera relationship list generated in the above steps. Based on the records in the list, the fusion service program internally establishes a relationship mapping table. When the fusion service program receives trajectory information of people targets sent from different cameras, it queries this relationship mapping table based on the camera device identifier attached to the trajectory information to quickly determine whether the two cameras sending the trajectory information are adjacent cameras. This determination is a direct prerequisite for subsequent execution of the target fusion logic, especially for determining whether face comparison and merging of people targets from different cameras is necessary.

[0031] S15. Considering that camera positions may change due to maintenance, or coordinate mapping relationships may be updated after recalibration, the effective field-of-view projection polygon of the camera may also change. Therefore, a periodic task can be set up, such as daily or weekly, to automatically re-execute the entire process from S10 to S13. By recalculating, the list of adjacent camera relationships can be updated to ensure that the list always reflects the current true state of camera field-of-view overlap relationships, guaranteeing the accuracy of the fusion processing. The entire recognition process is fully automated and requires no manual intervention.

[0032] S2. Filter out the trajectory information of personnel targets that do not carry valid facial information. The specific implementation process is as follows: S20: The fusion service program continuously listens to the network port, receiving data packets transmitted in real time via wireless or wired networks. These data packets are serialized byte streams of trajectory information of personnel targets sent by various cameras. The service program deserializes each received data packet, restoring the byte stream into a structured data object according to a predefined communication protocol format. This data object contains all fields of the personnel target's trajectory information, including the facial information field to be examined.

[0033] S21. Access the internal structure of the trajectory information data object and locate the field named "Face Information". The content of this field may be a high-dimensional floating-point array, representing the feature vector extracted from the face image; it may also be a string or integer, representing the face identifier obtained after comparison; in some cases, this field may also be a composite structure containing subfields such as feature vectors and quality scores.

[0034] S22. A clear set of standards is needed to determine whether a set of facial information is valid. These rules and thresholds are usually pre-configured in the settings file of the fusion service program. Key judgment rules include: the facial feature vector cannot be an array of all zeros or an empty array; the dimension of the feature vector must meet the expected length; if the facial information is accompanied by a quality score, the score must be higher than a preset minimum quality threshold; if the facial information is a facial identifier, the identifier cannot be a null value or a default value representing an unknown identity.

[0035] S23. Based on the rules loaded in S22, perform logical judgments on the face information fields extracted from S21. The judgment is a sequential, multi-condition check process. For example, first, check if the face information field is empty or does not exist; then, if the field is a feature vector, check if its data is complete and its dimensions are correct; then, check if there is a usable quality score and whether the score meets the standard. All preset check conditions must pass sequentially to produce a "valid" judgment result. If any check fails, an "invalid" judgment result is immediately generated.

[0036] S24. Maintain two independent data processing queues or channels. When a person's trajectory information is determined to carry valid facial information, this complete trajectory information data object is added to a buffer called the "Valid Trajectory Queue," awaiting subsequent fusion processing. Conversely, when a person's trajectory information is determined not to carry valid facial information, this data will not be sent to the Valid Trajectory Queue. The data object can be discarded directly, or it can be transferred to an audit log area and its memory released, thus implementing the filtering operation.

[0037] S25. To monitor operational status and data quality, logging is performed during filtering operations. A log entry is generated whenever a person's trajectory information is filtered out because it does not carry valid facial information. The log content may include the target identifier of the filtered trajectory information, the source camera device identifier, a timestamp, and the reason for filtering. Simultaneously, real-time statistical counters can be updated, for example, recording the total number of trajectory information received per unit time and the number filtered out.

[0038] S26. The valid trajectory queue serves as the input source for subsequent fusion logic. An independent fusion processing thread continuously retrieves the trajectory information of personnel targets from the valid trajectory queue, performing target identification and merging operations. The entire filtering process is a continuously running loop. The fusion service program repeatedly executes the operations from S20 to S25, performing real-time judgment and sorting on each newly received trajectory information, thereby ensuring that the data stream input into the core fusion algorithm is always cleaned trajectory information carrying valid facial information.

[0039] Valid facial information refers to facial data that meets the system's preset quality standards and can be reliably used for person identification and comparison. It not only indicates the existence of facial information fields but also emphasizes the usability and credibility of its content. Valid facial information typically means a sufficiently discriminative feature vector successfully extracted from a clear, frontal facial image, which is complete and conforms to the algorithm's required format. In systems that include quality assessment, valid facial information also requires its corresponding image quality score to exceed a certain threshold to ensure the credibility of the features. A set of facial information deemed valid is crucial evidence for determining whether individuals detected from different camera perspectives or at multiple time points belong to the same person, and is a prerequisite for performing subsequent target fusion operations.

[0040] S3. For the trajectory information of the filtered personnel targets carrying valid facial information, identify personnel targets with the same facial information in the trajectory information of personnel targets from different adjacent cameras. S4. Merge identified personnel targets with the same facial information into a single global personnel target. Use the target identifier in the trajectory information of the personnel target with the longest trajectory length before merging as the identifier of the global personnel target, and use the spatial geographic coordinates in the trajectory information of the personnel target with the longest trajectory length as the spatial location coordinates of the global personnel target. The specific implementation process is as follows: S40. The fusion service continuously receives and caches filtered trajectory information of personnel targets carrying valid facial information. Based on the camera device identifiers attached to these trajectory information, it queries the adjacent camera relationship list, filters out trajectory information from adjacent cameras that have been determined to have overlapping fields of view, and puts this information into a batch dataset to be fused. This dataset contains multiple trajectory information reports within similar time windows, each containing target identifier, spatial geographic coordinates, facial information, and camera device identifier.

[0041] S41. Extract the facial information carried by the trajectory information of each person target from the batch dataset to be fused. For the facial information of each trajectory, calculate its similarity with the facial information of other trajectory information in the same batch. Similarity calculation is usually done by comparing the cosine distance or Euclidean distance between two facial feature vectors. A similarity threshold is preset. When the similarity score of two facial information is higher than this threshold, the two facial information are determined to be the same, that is, they correspond to the same person. Based on the comparison results, the trajectory information of all person targets in the batch dataset to be fused is grouped. Each group contains all the trajectory information that is determined to have the same facial information. This means that these trajectory information come from different cameras but point to the same real person.

[0042] S42. For each group generated in S41, the trajectory length of each personnel target within the group needs to be calculated. The trajectory length is calculated based on historical data. The fusion service program maintains a historical trajectory cache database, which records a series of spatial geographic coordinates accumulated over time for each personnel target identifier reported by a single camera. Based on the target identifier and camera device identifier in each trajectory information within the group, this historical trajectory cache database is queried to obtain all historical location points of the target within the most recent period. The trajectory length can be measured by the total number of historical location points, or by the cumulative geographic length of the polyline formed by connecting these location points sequentially. A specific trajectory length value is calculated for each personnel target within the group.

[0043] S43. Within the same group, compare the trajectory length values ​​of all personnel targets. By iterating through the comparisons, find the trajectory information of the personnel target with the largest trajectory length value. If multiple personnel targets have the same trajectory length value and are all the maximum value, additional rules can be used to select one, such as selecting the first personnel target tracked, or selecting the personnel target whose latest reported spatial geographic coordinates are closer to the center of the area. Determining the trajectory information of the unique personnel target with the longest trajectory length is the basis for subsequent operations.

[0044] S44. The target identifier in the trajectory information of the personnel target with the longest trajectory length determined in the previous step is directly adopted as the identifier of the newly generated global personnel target. This identifier is upgraded from a local identifier of a single camera to a globally unique identifier within the system. At the same time, the latest reported spatial geographic coordinates contained in the trajectory information of the personnel target with the longest trajectory length are directly adopted as the spatial location coordinates of that global personnel target at the current moment. These coordinates represent the location reported by the most stable or most reliable tracking source selected from multiple perspectives.

[0045] S45. Create a new data entity, namely the global personnel target. This entity's fields include: global personnel target identifier, current spatial coordinates, and identical facial information used as the basis for fusion. Add or update this global personnel target entity to the global target queue. The global target queue is a data structure maintained in the fusion service program's memory, used to record the status of all currently active, fused global personnel targets. Simultaneously, associate the trajectory information of other personnel targets within the group with this newly created global personnel target, and internally logically treat these local trajectories as part of this global personnel target, thereby achieving merging.

[0046] S46. After completing the merging operation of a group, the identifier and spatial coordinates of the newly generated global personnel target, along with their facial information, are packaged into a new data format and sent to the subsequent server module for front-end display and historical trajectory storage. Then, the temporary computational data related to that group is cleared or reset to release memory resources. The fusion service program continuously monitors new data input. Whenever new trajectory information of a personnel target carrying valid facial information arrives and meets the batch processing conditions, the process starting from S40 is re-triggered, thus achieving uninterrupted real-time fusion processing.

[0047] S5. Display the spatial coordinates of all personnel targets on a front-end map. The specific implementation process is as follows: S50. After the fusion service program completes the creation or update of global personnel targets, the server program needs to proactively push this data to the front end. The server program encapsulates the global personnel target's identifier, latest spatial coordinates, corresponding facial information identifier, and timestamp into a specific data exchange object. This data exchange object is typically organized in JSON format. After encapsulation, the server pushes the data exchange object to all online front-end clients via a real-time communication link, such as a WebSocket connection. The push frequency is high, synchronized with the frame rate of the fusion processing, for example, ten times per second, to ensure the real-time display on the front end.

[0048] S51. A front-end program running in the user's browser or a dedicated client application continuously receives WebSocket messages from the server through a network listening module. When a new message arrives, the front-end program first performs security and format verification on the message content. After successful verification, the front-end program parses the JSON-formatted message body and extracts the global personnel target data list contained within it. For each global personnel target data in the list, the front-end program converts it into a data model that can be processed internally. This model explicitly includes fields such as global personnel target identifier and spatial location coordinates.

[0049] S52. The front-end interface integrates a web map engine, such as the Baidu Maps JavaScript API or the Gaode Maps JS API. When the page loads, the front-end program initializes the map engine, setting the map's center point, zoom level, and display style. Since the spatial coordinates sent from the server may be in latitude and longitude format or in planar meter coordinates specific to the scenario, the front-end program needs to call the coordinate transformation interface provided by the map engine according to a pre-agreed coordinate system standard. The coordinate transformation interface converts the unified spatial coordinates into screen pixel coordinates or geographic latitude and longitude coordinates used internally by the map engine for direct rendering.

[0050] S53. The front-end program maintains a mapping dictionary between global personnel target identifiers and map primitive objects. When it receives global personnel target data, the front-end program first checks whether its global personnel target identifier already exists in the mapping dictionary. If it does not exist, the front-end program needs to create a new map display element for the global personnel target. This element is usually a custom icon, such as an arrow icon representing a person. The front-end program calls the map engine's interface, uses the converted coordinates, adds this icon to the corresponding position on the map, and stores the icon object in the mapping dictionary after associating it with the global personnel target identifier. If the global personnel target identifier already exists in the mapping dictionary, the front-end program calls the map engine's interface based on the new coordinate data to update the position of the existing icon on the map.

[0051] S54. To more clearly demonstrate the movement trends of personnel, the front-end program needs to calculate and display the heading angle of all personnel targets. The front-end program maintains a short-term historical position queue for each global personnel target, storing the spatial coordinates of the most recent few moments in chronological order. When a new position arrives, the front-end program calculates a direction vector based on the current coordinates and the previous coordinate in the historical queue. Based on this direction vector, the front-end program calculates the heading angle. The front-end program calls the map engine's interface to update the rotation angle of the corresponding icon marker, making its arrow point in the direction of movement. Simultaneously, the front-end program can connect the continuous historical position coordinates of this global personnel target and draw a semi-transparent polyline on the map as a historical trajectory line to visually display its movement path.

[0052] S55. The data pushed by the server not only includes active global personnel targets, but may also include information on target status changes. For example, when a global personnel target has not been updated for a long time, the server may push a status marked as offline. After receiving the status change information, the front-end application needs to update the visual status of the corresponding element on the map, such as changing the icon to gray or adding an offline indicator. For global personnel targets that the server has explicitly indicated have disappeared, the front-end application needs to remove their icon markers and trajectory lines from the map, delete the relevant records from the internal mapping dictionary, and release front-end resources.

[0053] S56. The front-end map interface provides basic interactive functions. For example, clicking a person's icon will pop up an information window displaying the identifier of that person, the latest time, and other information. The front-end program implements performance optimization measures, such as visually throttling frequent location updates to avoid excessive map rendering and resulting lag; or dynamically adjusting the number of displayed people and the level of detail in their trajectories based on the current map zoom level. The entire display process is a dynamic loop. The front-end program continuously listens for server pushes and repeatedly executes the operations from S51 to S55, thereby ensuring that the global people display on the map is real-time, smooth, and information-rich.

[0054] Optionally, the above technical solution also includes: S010. Establish the coordinate mapping relationship between the image pixel coordinates and spatial geographic coordinates for each camera. The specific implementation process is as follows: S0100 involves establishing a set of control points with known spatial geographic coordinates within a pre-defined spatial area. These control points need to be easily identifiable by images and have fixed locations, such as using checkerboard patterns or physical markers with specific shapes and colors. The spatial geographic coordinates of the control points need to be measured and recorded beforehand using precision surveying instruments, such as a total station or a high-precision GPS receiver. The number of control points should be sufficient, and they should be evenly distributed across the entire field of view of the camera, especially within the area where mapping is desired, to ensure that the subsequent mapping calculations have sufficient accuracy and robustness.

[0055] S0101 involves using a pre-installed and fixed camera to acquire multiple images containing all or most of the aforementioned control points from different angles and positions. For each acquired image, the image pixel coordinates of each visible control point are accurately determined through automatic identification using image processing algorithms or manual annotation. This step establishes multiple sets of correspondence datasets, each containing the known spatial geographic coordinates of a control point in the real world and its corresponding image pixel coordinates in a specific image.

[0056] S0102. Based on the multiple sets of correspondence data established in the previous step, a transformation model from image pixel coordinates to spatial geographic coordinates is calculated and fitted. This transformation model can typically employ a direct linear transformation model or a more complex perspective transformation model. The calculation process uses the spatial geographic coordinates of control points as the target output and their corresponding image pixel coordinates as the input, employing mathematical optimization methods such as the least squares method to solve for all the parameters required for the transformation model. The solved transformation model characterizes the mapping law determined by the viewpoint, position, and internal optical parameters of a specific camera.

[0057] S0103 standardizes and stores the transformation model parameters calculated independently for each camera. These parameters are stored in configuration files or written to a database and strongly correlated with the corresponding camera's device identifier. In subsequent real-time processing, the corresponding transformation model parameters are invoked based on the camera's device identifier to quickly transform the coordinates of any image pixel detected in the real-time video stream, thereby obtaining its corresponding spatial geographic coordinates. To ensure continuous mapping accuracy, the transformation model can be periodically recalculated or calibrated using control points.

[0058] The pre-defined spatial area refers to the physical environment where multi-video surveillance and personnel target fusion are required. This area is typically an indoor or outdoor space with clearly defined boundaries, such as a lobby, plaza, workshop, or corridor. In this invention, multiple cameras are deployed at specific locations within this area to cover the entire area or key sections within it, ensuring, in particular, that there is necessary overlap in the fields of view between the cameras to facilitate subsequent target fusion processing.

[0059] Image pixel coordinates refer to the two-dimensional position of a target point in a digital image. A digital image is composed of tiny units arranged in rows and columns, called pixels. Image pixel coordinates are usually represented as ordered pairs, where the first value represents the column number of the target point from the left edge of the image, and the second value represents the row number of the target point from the top edge of the image. For example, the coordinates of the first pixel in the top left corner of an image are often defined as follows: Image pixel coordinates are fundamental data in computer vision processing, describing the position of an object in a two-dimensional image plane.

[0060] Spatial geographic coordinates refer to the geographic reference representation of the target point's location in the three-dimensional real world. In the context of this invention, spatial geographic coordinates can be Cartesian coordinates in meters, with its origin defined at a reference point within a preset spatial area; or they can be globally universal latitude and longitude coordinates plus altitude. Spatial geographic coordinates provide absolute or relative positional information of the target within a unified geographic framework, serving as the foundation for multi-camera target fusion and unified display on the front-end map. Through coordinate mapping relationships, image pixel coordinates limited to a single image can be transformed into a globally consistent spatial geographic coordinate system.

[0061] S011. Perform frame-by-frame processing on the video stream captured in real time by each camera. For each image frame, detect each person target contained therein, and obtain the image pixel coordinates and corresponding facial information of each person target in that image frame. The specific implementation process is as follows: S0110. Continuously receive real-time video streams from each deployed camera via network protocol or direct data interface. These video streams are typically transmitted in compressed encoding formats, such as H.264 or H.265. Call the corresponding video decoding library to perform real-time decoding on the received compressed streams. The decoding process restores the continuous streams to a series of uncompressed raw image frames arranged in chronological order, each image frame representing a complete scene captured by the camera at a specific moment.

[0062] S0111. To improve the accuracy and efficiency of subsequent detection algorithms, a series of standardization processes need to be performed on the original image frames. These processes include adjusting the image size to the fixed resolution required by the algorithm model, normalizing pixel values ​​to conform to the model input range, and performing possible image enhancement operations to improve contrast or reduce noise. The preprocessed image frames are transformed into a standardized multidimensional data matrix, waiting to be input into the target detection model.

[0063] S0112. The preprocessed image frame data is input into a pre-trained neural network model for people detection. This model is built on a deep learning framework, such as a convolutional neural network, which can automatically analyze image content and identify all objects belonging to the "people" category. After processing, the model outputs information for one or more detection boxes. Each detection box represents an identified person target, and the detection box information includes the rectangular boundary position of the box in the image, usually described by the image pixel coordinates of the top-left corner of the rectangle, as well as the width and height of the rectangle. Simultaneously, the model outputs a confidence score for each detection box, indicating the degree of certainty that the box actually contains a person target. Unreliable detection boxes with low scores are filtered out based on a preset confidence threshold.

[0064] S0113. For each confirmed person target detection box, the face region is located. For each person target detection box obtained in S0112, the image region within the rectangular bounding box is further analyzed. A dedicated face detection model is used to scan and analyze this local region, which can accurately locate the position of the person's face. The output of face detection is usually a more refined rectangular box, i.e., a face bounding box, which precisely frames the face portion in the image. Similarly, the position of the face bounding box is represented by its image pixel coordinates within the complete image frame.

[0065] S0114. For each face bounding box located in the previous step, extract the facial image region within the bounding box. This facial image region is then fed into a facial feature extraction neural network model. This model performs deep analysis on the input face image, transforming it into a high-dimensional, discriminative feature vector. This feature vector is a sequence of numbers, a digital and abstract representation of facial information, capable of characterizing the uniqueness of a face. In some implementations, this feature vector is also compared with a known facial feature database to assign a unique facial identifier to the current face. Ultimately, the facial information obtained for each person target is mainly composed of this feature vector and possible facial identifiers.

[0066] S0115. It is necessary to accurately associate and encapsulate the information of all personnel targets from the same image frame that belong to the same person target. For each personnel target confirmed in the image frame, the center point coordinates of the overall detection box of the personnel target obtained in S0112 are calculated as the image pixel coordinates of that personnel target in the current image frame. Simultaneously, the facial information generated for that personnel target in S0114, including the feature vector and possible facial identifiers, is bound to this personnel target. At this point, for the currently processed image frame, a complete data packet for each personnel target is obtained, containing the image pixel coordinates of the target in this frame and the corresponding facial information.

[0067] S0116. Temporarily buffer all personnel target data packets obtained after processing the current image frame or directly pass them to the next processing module. After completing all processing of the current image frame, immediately begin repeating the operations from S0111 to S0115 for the next decoded image frame. This process is repeated continuously, forming a real-time processing pipeline, thereby realizing frame-by-frame processing and information extraction of each camera video stream.

[0068] Facial information refers to a set of data extracted from a person's facial image that can be used to characterize that person's identity. Facial information is not the original portrait image, but rather an abstract digital description obtained through specific algorithms. This data typically includes a high-dimensional feature vector, extracted using a deep learning model, which captures and encodes the unique attributes of a face in terms of geometric structure, texture, etc., making the facial information of different individuals distinguishable. In some applications, facial information may also include a unique identifier generated after comparison with a database. Facial information is a key basis for achieving person target fusion in multi-camera scenarios, as it is direct evidence to determine whether targets detected from different camera perspectives belong to the same person.

[0069] S012. For each camera, for each person target in the image frame acquired by the camera at each time moment, using the coordinate mapping relationship established for that camera, the image pixel coordinates of the person target in the current image frame are transformed into the spatial geographic coordinates of the person target at the current time moment. The specific implementation process is as follows: S0120. Before processing the real-time video stream, coordinate mapping relationships need to be established for each camera deployed within a preset spatial area. This process is the same as before, involving calibrating control points and calculating transformation model parameters. These parameters are persistently stored in a configuration file or system database and strongly bound to the unique device identifier of each camera. When a processing thread for a particular camera is started or initialized, all coordinate mapping relationship parameters corresponding to that camera are retrieved from the storage medium and loaded based on the input camera device identifier. These parameters are loaded into a specific data structure in memory for quick access during subsequent frame processing.

[0070] S0121. For a particular camera currently being processed, following the aforementioned frame-by-frame processing procedure, the analysis of the image frame acquired at a specific moment has been completed. For this image frame, each person target has been detected, and a data packet has been generated for each person target. A key field in this data packet is the image pixel coordinates of the person target in the current image frame. These coordinates are usually represented in the form of numerical pairs, such as the x-coordinate and y-coordinate of the center point of the bottom edge of the target bounding box.

[0071] S0122. Extract the image pixel coordinate pairs from the personnel target data packet. To ensure correct input to the transformation model, this coordinate data needs to be normalized. This process may include converting the coordinate values ​​to floating-point numbers, or performing simple linear translation or scaling on the coordinate values ​​according to the requirements of the transformation model, such as adjusting the origin of the coordinate system from the top left corner of the image to the center of the image. Normalization ensures that the input image pixel coordinates conform to the coordinate system standard agreed upon when establishing the coordinate mapping relationship.

[0072] S0123. The system invokes the coordinate mapping model corresponding to the current camera, which has been loaded into memory. Mathematically, this model is represented as a function whose input is the normalized image pixel coordinates, and whose output is the corresponding spatial geographic coordinates. The transformation calculation process inputs the normalized image pixel coordinate pairs into this function. Internally, depending on its model type (e.g., direct linear transformation or perspective transformation model), the function performs a series of mathematical operations using pre-calculated and loaded model parameters. These operations may include matrix multiplication, vector transformations, and homogeneous coordinate normalization. The entire calculation process is efficiently completed by the system's central processing unit or graphics processing unit.

[0073] S0124. After the calculation in the previous step, the coordinate mapping relationship model outputs a result. This result is usually a tuple containing two or three values, representing different components of the spatial geographic coordinates. For example, in a Cartesian coordinate system, the output might be the X-axis and Y-axis coordinates; in a latitude and longitude coordinate system, the output might be the longitude and latitude values. This calculation result tuple is received and parsed into explicit spatial geographic coordinate components.

[0074] S0125. Perform necessary post-processing on the parsed spatial geographic coordinate components. This processing includes attaching explicit units of measurement to the coordinate values, such as meters or degrees; or normalizing the coordinate values ​​to a preset decimal precision. After processing, the final determined spatial geographic coordinate data is backfilled into a field named "Spatial Geographic Coordinates" in the currently processed personnel target data packet. At this point, the spatial location information of the personnel target at the current moment has been transformed from the image coordinate system to a unified geographic coordinate system.

[0075] S0126. For each person target detected in the current image frame, independently and repeatedly execute the operations from S0124 to S0125 to ensure that all targets within the frame have completed coordinate transformation. After processing all targets in the current frame, a personnel target data packet containing complete information such as image pixel coordinates and spatial geographic coordinates is sent to the subsequent modules. Then, the processing of the image frame acquired by the camera at the next moment begins, repeating the entire process starting from S011, thus forming a continuous, real-time pipeline for converting image pixel coordinates to spatial geographic coordinates. For each deployed camera, such an independent processing pipeline runs in parallel.

[0076] S013. For each detected person target, generate and transmit the person target's trajectory information in real time based on the person target's spatial geographic coordinates and corresponding facial information at continuous time intervals. The trajectory information includes: target identifier, spatial geographic coordinates, facial information, and the identifier of the camera device that acquired the facial information. The specific implementation process is as follows: S0130. Maintain a target tracking context table for each independently operating camera. When a camera detects a new person target for the first time in an image frame at a certain moment, a tracking instance needs to be created for it. Generate a unique sequence code within the current camera context; this code serves as the target identifier for the person target within the current camera's field of view. Simultaneously, record the spatial geographic coordinates of the person target at that initial moment, as well as the facial information extracted from that frame image. This initial data constitutes the starting point for this person target tracking instance.

[0077] S0131. When processing the image frame of the next moment of the camera, a set of newly detected personnel targets, their spatial geographic coordinates, and facial information will be obtained. It is necessary to match the newly detected targets with the targets being tracked in the previous moment. The matching process mainly relies on two criteria: first, the proximity of spatial geographic coordinates, i.e., the position of the new target and the position of an existing target in the previous step are within a preset distance threshold; second, the similarity of facial information, determined by calculating the similarity score between the facial feature vector of the new target and the facial feature vector of the existing target. When a match is successful, the newly detected spatial geographic coordinates and facial information are merged into the existing tracking instance, thereby continuing the trajectory of the personnel target.

[0078] S0132. Due to occlusion, rapid movement, or brief disappearance from view, a person target may be lost in several frames. A survival timer is set for each tracking instance. If an existing target is not successfully matched within several consecutive frames, the tracking instance is not immediately deleted, but marked as temporarily lost. In subsequent frames, if a newly detected target appears, its spatial geographic coordinates are within the range of historical trajectory prediction, and its facial information is highly similar to the record of a temporarily lost instance, it is considered a target recovery. At this time, the original target identifier is reused, and the new location information is continued onto the trajectory.

[0079] S0133. For each person target successfully matched or continuously tracked in each frame, a standard format data structure needs to be assembled. This data structure is the trajectory information of the person target. This data structure contains several fixed fields. The target identifier field is filled with the unique identification code assigned to this tracking instance. The spatial geographic coordinate field is filled with the physical location coordinates of the person target at this moment, calculated based on the current frame image. The face information field is filled with the latest face feature vector extracted from the current frame. If the face information for the current frame is unavailable, the face information of the last valid record is filled in. The timestamp field is filled with the precise acquisition time corresponding to the current image frame. The camera device identifier field is filled with the unique device number of the currently processed camera.

[0080] S0134. The assembled trajectory information data structure of the personnel target exists in memory as an object. For transmission over a network, this data structure needs to be serialized into a standard data exchange format, such as JSON or Protocol Buffers. The serialization process converts all fields in the data structure, including target identifiers, spatial geographic coordinate pairs, feature vector arrays of facial information, timestamp strings, and camera device identifier strings, into a continuous byte stream according to predefined key-value pair rules. This byte stream is the message body prepared for transmission.

[0081] S0135. Maintain a network connection to the fusion service program for each camera. This connection is typically established during initialization and maintained as a persistent connection to reduce communication overhead. The connection uses a reliable network transport protocol, such as TCP, to ensure the ordered and reliable delivery of data packets. The network output stream is encapsulated as a send queue, which manages the data to be transmitted.

[0082] S0136. The serialized trajectory information byte stream of the personnel target is immediately placed into the network transmission queue of the corresponding camera. An independent network transmission thread or asynchronous transmission mechanism continuously monitors this transmission queue. Once new data is in the queue, the transmission mechanism immediately retrieves the data from the queue and sends the data packet to the designated network address and port of the fusion service program through the established network connection. This transmission action is performed at a very high frequency, such as ten times per second or synchronized with the video frame rate, thereby achieving real-time transmission of the personnel target's trajectory information.

[0083] S0137. When a tracking instance of a person target is confirmed to have permanently disappeared, for example, if it leaves the camera's field of view and exceeds the maximum allowed interruption time, the memory resources occupied by the target identifier and its tracking context will eventually be released. For all other continuously active person targets, at the end of each frame processing cycle, the latest trajectory information will be generated and sent to them, and then the temporary data of the current frame will be cleared to prepare for the next processing cycle to process newly acquired image frames from the camera, and so on.

[0084] The target identifier is a unique identification code assigned to each person successfully tracked within the field of view of a single camera. This code is typically generated when the target first appears and remains unchanged throughout its tracked lifecycle. The target identifier is used to uniquely identify and distinguish different moving person targets within the context of a single camera. It is a temporary, internally used tracking tag, which can be in the format of an incrementing integer sequence number or a combined string containing camera number and time information. The target identifier is a key field that links the trajectory information of the same target generated at different times.

[0085] The camera device identifier is a unique number or string code assigned to each physical camera deployed within a pre-defined spatial area. This identifier is determined during camera installation and deployment, and is typically associated with the camera's hardware serial number, network MAC address, or registration number in the management system. The camera device identifier is used at the system-wide level to clearly distinguish the data source, indicating which specific camera collected and reported the trajectory information of a particular person target. In subsequent multi-camera target fusion processes, the camera device identifier is one of the important bases for determining the adjacency relationship between cameras and for data association.

[0086] Optionally, the above technical solution also includes: recording the identifier and corresponding facial information of each global human target obtained after fusion into a global target queue.

[0087] Optionally, the above technical solution also includes: when trajectory information of a person carrying valid facial information is received again, a search and match is performed from the global target queue based on the facial information. The specific implementation process is as follows: 1) The fusion service program runs continuously, with its network receiving module constantly acquiring the latest data from various cameras. When a new trajectory of a person arrives, the program parses it to confirm that the facial information it carries has been deemed valid by the pre-filtering logic. From the data structure of this trajectory information, the program extracts two core fields: one is the camera device identifier associated with the trajectory information, and the other is the standardized facial information data, usually a feature vector.

[0088] 2) The global target queue is a core data structure maintained in memory by the fusion service program. It exists as a list or dictionary, storing all global personnel target objects currently being tracked by the system. Each global personnel target object contains several fixed fields, including the global personnel target identifier and the facial information bound to that target, used for initial fusion or the last update. The program is preparing to perform a traversal query operation on this global target queue.

[0089] 3) Retrieve each global person target object sequentially from the global target queue. For each retrieved object, the program accesses its stored face information field. The program calculates the similarity between the face information carried in the trajectory information of the newly received person target and the face information stored in the current global person target object. Similarity calculation typically involves mathematically comparing two facial feature vectors, such as calculating the cosine similarity or inverse cosine distance between the two vectors. This calculation process generates a specific similarity score; a higher score indicates that the two facial images are more likely to belong to the same person.

[0090] 4) A preset similarity threshold for identity matching is established. This threshold, determined through experimentation and experience, balances the accuracy and recall of the matching process. The program compares each calculated similarity score with this preset threshold. If, during the traversal, a similarity score exceeding the preset threshold is found for a global person target, the program immediately determines that the newly received trajectory information has successfully matched that global person target. The program records the global person target identifier of this successfully matched global person target and terminates subsequent traversal to save computational resources. If, after traversing the entire global target queue, none of the calculated similarity scores reach the preset threshold, the program determines that no matching global person target identifier has been found.

[0091] 5) If the result is a successful match, the program outputs the found global personnel target identifier as the return value. This identifier will be used in subsequent steps, such as updating the latest location coordinates of the global personnel target, or associating the new trajectory information with the global target's history. If the result of the previous step is no match found, which usually means that the new trajectory information may correspond to a person appearing for the first time, the program outputs a special identifier value representing "no match," such as a null value or a specific error code. This result will trigger the process of creating a new global personnel target.

[0092] 6) Regardless of whether a match is successful or not, the program completes a query operation. The program records relevant information from this query, such as query time, facial features of the queried person, matching results, and execution time, in the system log for monitoring and auditing. Afterward, the program releases the temporary computing resources used for this query, prepares to receive the trajectory information of the next new person target, and executes the entire search and matching process again. This search and matching process is a high-frequency, fundamental step in the fusion service, ensuring that the system can continuously and correctly aggregate a constant stream of real-time data under the correct global identity.

[0093] The technical solution of the present invention will be further described through another embodiment, which specifically includes the following steps: 1) To achieve full coverage of the area, at least four cameras will be installed within the pre-defined spatial area for facial recognition. Figure 2A spatial plan of a pre-defined spatial area is shown. Cameras are typically installed in the corners of this area for face recognition. A coordinate mapping relationship is established between the image pixel coordinates and spatial geographic coordinates of each camera deployed within the pre-defined spatial area; this process is also called mapping calibration. The video streams acquired in real-time by each camera are processed frame by frame. For each image frame, each person target is detected, and the image pixel coordinates and corresponding facial information of each person target in that image frame are obtained. For each camera, for each person target in the image frames acquired at each time moment, the coordinate mapping relationship established for that camera is used to convert the image pixel coordinates of the person target in the current image frame into the spatial geographic coordinates of the person target at the current time moment. Through a tracking algorithm, the trajectory is basically guaranteed to be output normally even when the person is occluded within a certain period of time. For each detected person target, the trajectory information of the person target is generated and sent to the fusion service program in real time based on the spatial geographic coordinates and corresponding facial information of the person target at continuous time moments. The trajectory information of the person target includes: target identifier, spatial geographic coordinates, facial information, and the identifier of the camera device acquiring this information.

[0094] 2) The fusion service program receives trajectory information of personnel targets from each camera in real time. Cameras with overlapping fields of view are identified as adjacent cameras from each camera deployed within a preset spatial area. This is achieved through the configuration of the fusion service program and camera device identification. Trajectory information of personnel targets without valid facial information is filtered out. For the filtered trajectory information of personnel targets carrying valid facial information, personnel targets with the same facial information from different adjacent cameras are identified. To simplify the fusion process, only targets with the same facial information are merged. The identified personnel targets with the same facial information are merged into a single global personnel target. The target identifier in the trajectory information of the personnel target with the longest trajectory length before merging is used as the identifier of this global personnel target, and the spatial geographic coordinates in the trajectory information of the personnel target with the longest trajectory length are used as the spatial location coordinates of this global personnel target. This global personnel target identifier is bound to the personnel's facial identity information and recorded in the global target queue. Since the trajectory information of the personnel targets sent to the fusion program by the algorithm already takes into account occlusion, trajectory interruption is not considered during fusion. Since the cameras installed in the room basically cover the entire space, the brief absence of facial information does not affect the tracking of personnel positions. When trajectory information of a person target carrying valid facial information is received again, a matching global personnel target identifier is searched from the global target queue based on the facial information. If a match is found, the corresponding identifier is used to continue tracking. For single-camera scenarios, when only one person target's trajectory information carries facial information, no fusion processing is performed, and the output is directly used to display the spatial coordinates of the global personnel targets on the front-end map.

[0095] The fusion service program typically sends global personnel target information to the server at a frequency of more than ten frames per second. The server calculates the heading angle information based on historical trajectory information. Since personnel movement is relatively slow, the heading angle is calculated when the movement distance exceeds a certain number of meters. The geographic information system platform displays the target information in real time, achieving the uniqueness and continuity of personnel target display.

[0096] For personnel targets within a preset spatial range, this method proposes a scheme where the algorithm only needs to detect facial information and image pixel coordinates. It then calculates the spatial geographic coordinates from the image pixel coordinates through a coordinate mapping relationship. The algorithm sends the trajectory information of the personnel targets, which includes both spatial geographic coordinates and facial information, to the fusion service program. By identifying the relationship between adjacent cameras, it performs fusion processing on personnel targets under adjacent cameras, fusing targets with the same facial information. After fusion, the same facial target uses a global personnel target identifier, filtering out targets without valid facial information, and displaying targets with real identity information in real time, while simultaneously reducing database storage requirements.

[0097] In the above embodiments, although the steps are numbered S1, S2, etc., they are only specific embodiments given by the present invention. Those skilled in the art can adjust the execution order of S1, S2, etc. according to the actual situation. The scheme after adjusting the order is also within the protection scope of the present invention. It can be understood that in some embodiments, some or all of the above embodiments may be included.

[0098] like Figure 3 As shown, an embodiment of the present invention provides a multi-video target fusion system 200 based on facial information, which includes an adjacent camera determination module 201, a filtering module 202, a recognition module 203, a spatial location coordinate acquisition module 204, and a display module 205. The adjacent camera determination module 201 is used to: identify cameras with overlapping fields of view from each camera deployed in a preset spatial area as adjacent cameras; The filtering module 202 is used to: filter out the trajectory information of personnel targets that do not carry valid facial information; The recognition module 203 is used to: identify personnel targets with the same facial information in the trajectory information of personnel targets from different adjacent cameras, based on the filtered trajectory information of personnel targets carrying valid facial information; The spatial location coordinate acquisition module 204 is used to: merge the identified personnel targets with the same facial information into the same global personnel target, and use the target identifier in the trajectory information of the personnel target with the longest trajectory length before merging as the identifier of the global personnel target, and use the spatial geographic coordinates in the trajectory information of the personnel target with the longest trajectory length as the spatial location coordinates of the global personnel target. The display module 205 is used to display the spatial coordinates of all personnel targets on a front-end map.

[0099] Optionally, the above technical solution further includes a trajectory information acquisition module, which is used for: Establish the coordinate mapping relationship between image pixel coordinates and spatial geographic coordinates for each camera; The video streams captured in real time by each camera are processed frame by frame. For each image frame, each person target contained therein is detected, and the image pixel coordinates and corresponding face information of each person target in that image frame are obtained. For each camera, for each person target in the image frame captured by the camera at each moment, the image pixel coordinates of the person target in the current image frame are transformed into the spatial geographic coordinates of the person target at the current moment using the coordinate mapping relationship established for the camera. For each detected person target, the trajectory information of the person target is generated and sent in real time based on the spatial geographic coordinates of the person target at continuous time and the corresponding facial information; the trajectory information of the person target includes: target identifier, spatial geographic coordinates, facial information and the identifier of the camera device that collected the facial information.

[0100] Optionally, the above technical solution also includes a recording module, which is used to record the identifier and corresponding facial information of each global human target obtained after fusion to the global target queue.

[0101] Optionally, the above technical solution also includes a search and matching module, which is used to: when the trajectory information of a person target carrying valid facial information is received again, perform a search and matching from the global target queue based on the facial information.

[0102] It should be noted that the beneficial effects of the multi-video target fusion system 200 based on face information provided in the above embodiments are the same as those of the multi-video target fusion method based on face information described above, and will not be repeated here. Furthermore, the system provided in the above embodiments is only illustrated by the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the system can be divided into different functional modules according to the actual situation to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and their specific implementation process is detailed in the method embodiments, and will not be repeated here.

[0103] An electronic device according to an embodiment of the present invention includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements any of the above-mentioned multi-video target fusion methods based on facial information.

[0104] An embodiment of the present invention provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements any of the above-described multi-video target fusion methods based on facial information.

[0105] The above description is merely a preferred embodiment of the present invention and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of disclosure in this invention is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-disclosed concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features with similar functions disclosed in this invention.

[0106] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A multi-video target fusion method based on facial information, characterized in that, include: Identify cameras with overlapping fields of view from each camera deployed within a predefined spatial area as adjacent cameras; Filter out the trajectory information of individuals who do not carry valid facial information; For the trajectory information of the personnel target carrying valid facial information after filtering, identify personnel targets with the same facial information in the trajectory information of the personnel targets from different adjacent cameras; The identified personnel targets with the same facial information are merged into a single global personnel target. The target identifier in the trajectory information of the personnel target with the longest trajectory length before merging is used as the identifier of the global personnel target. The spatial geographic coordinates in the trajectory information of the personnel target with the longest trajectory length are used as the spatial location coordinates of the global personnel target. Display the spatial coordinates of all personnel targets on a front-end map.

2. The multi-video target fusion method based on facial information according to claim 1, characterized in that, Also includes: Establish the coordinate mapping relationship between image pixel coordinates and spatial geographic coordinates for each camera; The video streams captured in real time by each camera are processed frame by frame. For each image frame, each person target contained therein is detected, and the image pixel coordinates and corresponding face information of each person target in the image frame are obtained. For each camera, for each person target in the image frame acquired by the camera at each moment, the image pixel coordinates of the person target in the current image frame are converted into the spatial geographic coordinates of the person target at the current moment using the coordinate mapping relationship established for the camera. For each detected person target, trajectory information of the person target is generated and sent in real time based on the spatial geographic coordinates and corresponding facial information of the person target at continuous time intervals; wherein, the trajectory information of the person target includes: target identifier, spatial geographic coordinates, facial information, and camera device identifier for collecting facial information.

3. The multi-video target fusion method based on facial information according to claim 2, characterized in that, Also includes: The identifier and corresponding facial information of each global human target obtained after fusion are recorded in the global target queue.

4. The multi-video target fusion method based on facial information according to claim 3, characterized in that, Also includes: When trajectory information of a person carrying valid facial information is received again, a search and match is performed from the global target queue based on that facial information.

5. A multi-video target fusion system based on facial information, characterized in that, It includes a neighboring camera determination module, a filtering module, a recognition module, a spatial location coordinate acquisition module, and a display module; The adjacent camera determination module is used to: identify cameras with overlapping fields of view from each camera deployed in a preset spatial area as adjacent cameras; The filtering module is used to: filter out the trajectory information of personnel targets that do not carry valid facial information; The identification module is used to: identify personnel targets with the same facial information from the trajectory information of the personnel targets from different adjacent cameras, based on the filtered trajectory information of the personnel targets carrying valid facial information; The spatial location coordinate acquisition module is used to: merge the identified personnel targets with the same facial information into the same global personnel target, and use the target identifier in the trajectory information of the personnel target with the longest trajectory length before merging as the identifier of the global personnel target, and use the spatial geographic coordinates in the trajectory information of the personnel target with the longest trajectory length as the spatial location coordinates of the global personnel target. The display module is used to display the spatial coordinates of all personnel targets on a front-end map.

6. A multi-video target fusion system based on facial information according to claim 5, characterized in that, It also includes a trajectory information acquisition module, which is used for: Establish the coordinate mapping relationship between image pixel coordinates and spatial geographic coordinates for each camera; The video streams captured in real time by each camera are processed frame by frame. For each image frame, each person target contained therein is detected, and the image pixel coordinates and corresponding face information of each person target in the image frame are obtained. For each camera, for each person target in the image frame acquired by the camera at each moment, the image pixel coordinates of the person target in the current image frame are converted into the spatial geographic coordinates of the person target at the current moment using the coordinate mapping relationship established for the camera. For each detected person target, trajectory information of the person target is generated and sent in real time based on the spatial geographic coordinates and corresponding facial information of the person target at continuous time intervals; wherein, the trajectory information of the person target includes: target identifier, spatial geographic coordinates, facial information, and camera device identifier for collecting facial information.

7. A multi-video target fusion system based on facial information according to claim 6, characterized in that, It also includes a recording module, which is used to record the identifier and corresponding facial information of each global human target obtained after fusion to the global target queue.

8. A multi-video target fusion system based on facial information according to claim 7, characterized in that, It also includes a search and matching module, which is used to: when the trajectory information of a person target carrying valid facial information is received again, perform a search and matching from the global target queue based on the facial information.

9. An electronic device, characterized in that, The method includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the multi-video target fusion method based on facial information as described in any one of claims 1 to 4.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the multi-video target fusion method based on facial information as described in any one of claims 1 to 4.