Camera topology mapping method and device, electronic equipment and readable medium
By acquiring and processing video captured by cameras, performing quality optimization and target tracking, and generating camera topology maps, the problem of high hardware costs in existing technologies is solved, achieving lower-cost and more efficient camera topology map drawing.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- DMALL LIFE (CHINA) NETWORK CO LTD
- Filing Date
- 2026-02-25
- Publication Date
- 2026-05-29
AI Technical Summary
In existing technologies, drawing camera topology maps requires the deployment of additional positioning base stations, resulting in high hardware costs.
By acquiring videos from various cameras, performing quality optimization and coordinate calibration, generating a calibrated video frame sequence, performing target tracking to obtain a set of target person information, and performing cross-camera target tracking based on a preset image edge distance range and camera identifiers, a camera topology map is generated.
It reduces the hardware cost of drawing camera topology maps, improves the drawing effect of camera topology maps, and ensures the accuracy of the relative position information between cameras.
Smart Images

Figure CN122120436A_ABST
Abstract
Description
Technical Field
[0001] The embodiments disclosed herein relate to the field of computer technology, and more specifically to a method, apparatus, electronic device, and readable medium for drawing camera topology maps. Background Technology
[0002] As supermarkets and shopping malls become larger and more intelligent, the number of cameras deployed in a single store typically reaches 50-200 units, used for security monitoring, customer flow analysis, and abnormal behavior early warning. A camera topology map is an image representing the relative positions of camera devices. Currently, camera topology mapping is generally achieved by deploying additional positioning base stations to obtain the relative positional relationships between cameras, thereby creating the camera topology map.
[0003] However, when drawing the camera topology map using the above method, the following technical problems often occur: Additional positioning base stations are deployed to obtain the relative positional relationships between cameras, thereby enabling the creation of a camera topology map. However, the hardware costs are high due to the need to deploy additional positioning base stations to capture the relative positional relationships between camera devices.
[0004] The information disclosed in this background section is only intended to enhance the understanding of the background of the inventive concept, and therefore may contain information that does not form prior art known to those skilled in the art. Summary of the Invention
[0005] The summary portion of this disclosure is intended to provide a brief overview of the concepts, which will be described in detail in the detailed description portion. This summary portion is not intended to identify key or essential features of the claimed technical solutions, nor is it intended to limit the scope of the claimed technical solutions.
[0006] Some embodiments of this disclosure provide methods, apparatus, electronic devices, and computer-readable media for drawing camera topology maps to address one or more of the technical problems mentioned in the background section above.
[0007] In a first aspect, some embodiments of this disclosure provide a method for drawing a camera topology map. The method includes: acquiring videos captured by various cameras, wherein each camera has a corresponding camera identifier, and each video corresponds to one of the cameras; performing quality optimization and coordinate calibration processing on the videos to obtain a sequence of calibrated video frames; performing target tracking processing on each calibrated video frame sequence to obtain a set of target person information; performing cross-camera target tracking processing on the obtained set of target person information based on a preset edge distance range, camera identifiers, and the calibrated video frame sequences to obtain trajectory information of each target person; performing cross-camera association and constraint verification processing on the set of target person information based on the trajectory information of each target person to obtain target association information pairs for each camera; and generating a camera topology map based on the target association information pairs for each camera.
[0008] Secondly, some embodiments of this disclosure provide a camera topology map drawing apparatus, the apparatus comprising: an acquisition unit configured to acquire videos captured by various cameras, wherein each of the cameras has a corresponding camera identifier, and each video corresponds to one of the cameras; a first processing unit configured to perform quality optimization and coordinate calibration processing on the videos to obtain a sequence of calibration video frames; a second processing unit configured to perform target tracking processing on each of the calibration video frame sequences to obtain a set of target person information; a third processing unit configured to perform cross-camera target tracking processing on the obtained set of target person information based on a preset screen edge distance range, the camera identifiers, and the calibration video frame sequences to obtain trajectory information of each target person; a fourth processing unit configured to perform cross-camera association and constraint verification processing on the set of target person information based on the trajectory information of each target person to obtain a pair of target association information for each camera; and a generation unit configured to generate a camera topology map based on the pair of target association information for each camera.
[0009] Thirdly, some embodiments of this disclosure provide an electronic device, including: one or more processors; and a storage device having one or more programs stored thereon, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation of the first aspect above.
[0010] Fourthly, some embodiments of this disclosure provide a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the method described in any of the implementations of the first aspect above.
[0011] The above-described embodiments of this disclosure have the following beneficial effects: the camera topology mapping method of some embodiments of this disclosure reduces the hardware cost required for camera topology mapping. Specifically, the reason for the high hardware cost required for camera topology mapping is that additional positioning base stations are deployed to obtain the relative positional relationships between cameras, thereby drawing the camera topology map. Since additional positioning base stations are needed to capture the relative positional relationships between camera devices, the required hardware cost is high. Based on this, the camera topology mapping method of some embodiments of this disclosure first acquires each video captured by each camera, wherein each of the cameras has a corresponding camera identifier, and each video corresponds to one of the cameras. Thus, each video captured by each camera can be obtained for generating each calibration video frame sequence. Then, each video is subjected to quality optimization and coordinate calibration processing to obtain each calibration video frame sequence. Thus, each calibration video frame sequence used to generate each target person information set can be obtained. Afterwards, for each calibration video frame sequence, target tracking processing is performed to obtain the target person information set. Thus, each target person information set used to generate target person trajectory information can be obtained. Next, based on the preset edge distance range of the screen, the identifiers of each camera, and the aforementioned calibrated video frame sequences, cross-camera target tracking processing is performed on the obtained target person information sets to obtain the trajectory information of each target person. This provides trajectory information for determining the relative positions between cameras. Then, based on the aforementioned target person trajectory information, cross-camera association and constraint verification processing is performed on the aforementioned target person information sets to obtain target association information pairs between cameras. This allows the generation of target association information pairs between cameras that characterize the relative positions between cameras, based on the target person trajectory information. Finally, a camera topology map is generated based on the aforementioned target association information pairs. This method involves target detection in the video footage from each camera during the topology map drawing process, followed by tracking of the detected targets to obtain the relative position information between cameras. The camera topology map is drawn based on the relative position information between cameras. Since this method does not require additional positioning hardware, it reduces the hardware cost of drawing the camera topology map. Attached Figure Description
[0012] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and elements are not necessarily drawn to scale.
[0013] Figure 1 This is a flowchart of some embodiments of the camera topology diagram drawing method according to this disclosure; Figure 2 These are schematic diagrams illustrating the structure of some embodiments of the camera topology drawing apparatus according to this disclosure; Figure 3 This is a schematic diagram of the structure of an electronic device suitable for implementing some embodiments of the present disclosure. Detailed Implementation
[0014] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0015] It should also be noted that, for ease of description, only the parts relevant to the invention are shown in the accompanying drawings. Unless otherwise specified, the embodiments and features described in this disclosure can be combined with each other.
[0016] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0017] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0018] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0019] This disclosure will now be described in detail with reference to the accompanying drawings and embodiments.
[0020] Figure 1 A flow 100 of some embodiments of a camera topology mapping method according to the present disclosure is shown. The camera topology mapping method includes the following steps: Step 101: Acquire the videos captured by each camera.
[0021] In some embodiments, the execution entity (e.g., a computing device) of the camera topology mapping method can acquire individual videos captured by each camera via a wired or wireless connection. Each camera has a corresponding camera identifier, and each video corresponds to one of the cameras. The cameras can be those installed in public places (e.g., shopping malls). Each video can be a video captured by the camera showing pedestrian movement trajectories. The camera identifier can be a number representing the camera. For example, the camera number could be B-1.
[0022] It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G connection, WiFi connection, Bluetooth connection, WiMAX connection, Zigbee connection, UWB (ultra wideband) connection, and other currently known or future wireless connection methods.
[0023] Step 102: Perform quality optimization and coordinate calibration on each video to obtain a sequence of calibrated video frames.
[0024] In some embodiments, the aforementioned execution entity may perform quality optimization and coordinate calibration processing on the aforementioned videos to obtain various calibrated video frame sequences.
[0025] In some optional implementations of certain embodiments, the aforementioned execution entity may perform quality optimization and coordinate calibration processing on the aforementioned videos through the following steps to obtain a sequence of calibrated video frames: The first step is to filter each of the aforementioned videos to obtain a sequence of filtered video frames. In practice, the execution entity can perform Gaussian filtering on each of the aforementioned videos to obtain a sequence of filtered video frames. Each of these filtered video frame sequences can be a sequence of clear images arranged in chronological order after noise removal through filtering.
[0026] The second step involves correcting each of the filtered video frame sequences to obtain a distortion-corrected video frame sequence. In practice, the execution entity can correct lens distortion in each of the filtered video frame sequences using Zhang's calibration method to obtain a distortion-corrected video frame sequence. Each distortion-corrected video frame sequence can be a sequence of corrected, distortion-free images (edges restored to normal, object proportions conforming to physical reality) arranged in chronological order.
[0027] The third step involves generating calibration video frame sequences based on the distortion-corrected video frame sequences. In practice, the executing entity can perform inter-frame difference processing on each distortion-corrected video frame sequence to obtain the calibration video frame sequences. Each calibration video frame sequence can be an image sequence that retains the pixel information of the moving target, while the pixels of the remaining static background (such as shelves) are processed into black or transparent images arranged in chronological order. Each calibration video frame sequence corresponds to one camera identifier among the various camera identifiers.
[0028] Step 103: For each calibration video frame sequence in each calibration video frame sequence, perform target tracking processing on each calibration video frame sequence to obtain a target person information set.
[0029] In some embodiments, the execution entity may perform target tracking processing on each calibration video frame sequence to obtain a target person information set.
[0030] In some optional implementations of certain embodiments, the execution entity can perform target tracking processing on each of the aforementioned calibration video frame sequences through the following steps to obtain a target person information set: The first step, for each of the above calibration video frame sequences, is to perform the following steps: The first sub-step involves performing target detection processing on the aforementioned calibration video frame sequence to obtain target person bounding box information. This target person bounding box information includes the target person's location information, target person identifier, and target candidate box size. In practice, the executing entity can input the aforementioned calibration video frame sequence into a YOLOv8 model to perform target detection processing and obtain target person bounding box information. The target person's location information can be the coordinates of the bounding box's center point. The target person identifier can be a number representing the identified target person. For example, the target person identifier could be A-1. The target candidate box size can represent the size (height and width) of the bounding box output by the YOLOv8 model. For example, the target person's location information could be [center point coordinates: (56.3, 120.5)]. The target person identifier could be A-1. The target candidate box size could be [height: 450.9, width: 340.5].
[0031] The second step is to generate a target person information set based on the target person frame information mentioned above.
[0032] In addressing the aforementioned technical problems in the application scenario of drawing a topology map of cameras in large shopping malls, the following technical issues often arise: Due to differences in installation angles, lighting conditions, and resolutions of different cameras, the features of the same target person can vary significantly visually. Conversely, the features of different target persons (such as customers wearing similar clothing) can be very similar. This can easily lead to identity confusion or trajectory interruption when tracking target persons, resulting in failure to identify the same target person. Furthermore, multiple detection results for each video frame need to be correctly matched with existing historical information, and using a simple greedy matching algorithm is prone to errors. This results in low accuracy in matching the identity of target persons within the cameras, leading to low accuracy in determining the camera direction based on the tracked target person, and consequently, poor quality of the drawn camera topology map. Therefore, this application scenario requires the following characteristics: suitability for high appearance similarity, multi-viewpoint, cross-time period association, and high-concurrency real-time matching, in order to obtain accurate camera directions based on the tracked target person, thereby drawing the camera topology map. Faced with the above technical problems, we decided to adopt the following solution: In some optional implementations of certain embodiments, the aforementioned execution entity may generate a target person information set based on the aforementioned target person frame information through the following steps: The first step is to perform the following steps for each target person frame in the above target person frame information: The first sub-step is to determine the image represented by the target person bounding box information in the above-mentioned calibration video frame as the target image.
[0033] The second sub-step involves extracting human features from the target image to obtain target human feature information. In practice, the executing entity can input the target image into a ResNet-50 model to obtain target human feature information. This target human feature information can be a 512-dimensional feature vector representing human appearance (clothing, body shape, posture, etc.) extracted from the target image region by the ResNet-50 deep neural network model.
[0034] The third sub-step generates comprehensive cost information based on the preset list of historical figures, the target figure bounding box information, and the target figure feature information. In practice, firstly, the executing entity determines the cosine similarity between the target figure feature information and the feature vectors of each historical figure in the preset list of historical figures as a similarity value. Then, the difference between the first preset value and the similarity value is determined as the appearance cost. Next, the Mahalanobis distance between the target figure location information (in the target figure bounding box information) and the predicted location information (in the trajectory information of each historical figure in the preset list of historical figures) is determined as the location distance cost. Each appearance cost corresponds to one of the location distance costs. Then, the product of each appearance cost and the second preset value is determined as the first value. Then, the product of the location distance cost corresponding to each appearance cost and the third preset value is determined as the second value. Finally, the sum of the first and second values is determined as the comprehensive cost information. Each historical figure in the preset list of historical figures includes a historical figure feature vector and historical figure trajectory information. The historical figure trajectory information includes a predicted location. The predicted location can be the coordinates of a historical figure predicted by the Kalman filter. For example, the predicted location could be (158.0, 235.5). The first preset value can be 1. The second and third preset values can be pre-set parameters, and the sum of the second and third preset values can be 1.
[0035] The second step is to generate a cost matrix based on the obtained comprehensive cost information. In practice, the aforementioned implementing entity can create an M×N cost matrix. Each element of the cost matrix can be a comprehensive cost information. M in the M×N cost matrix can be the number of target character feature information. N in the M×N cost matrix can be the number of historical character information. For example, there are two target character feature information (D1, D2) and two historical character information (T1, T2). The comprehensive cost information between target character feature information D1 and historical character information T1 is 0.15. The comprehensive cost information between target character feature information D1 and historical character information T2 is 0.80. The comprehensive cost information between target character feature information D2 and historical character information T1 is 0.60. The comprehensive cost information between target character feature information D2 and historical character information T2 is 0.25. The cost matrix can be... .
[0036] The third step involves performing inter-frame correlation matching on the aforementioned cost matrix to obtain a target person information set, which corresponds to one calibration video frame sequence from each of the calibration video frame sequences. In practice, firstly, the executing entity can use the Hungarian algorithm to match each historical person information and each target person information in the cost matrix, and identify a successfully matched historical person information and a target person information as a person information pair, resulting in each person information pair, each unmatched target person information, and each unmatched historical person information. Then, the target person information in each person information pair is identified as historical person information to update the historical person information. Next, an unused identifier is assigned to the target person identifier included in each unmatched target person information to update the target person information. Finally, the target person information included in each successfully matched person information pair and the updated target person information are identified as the target person information set. The target person information includes target person feature information, target person identifier, and target person location information.
[0037] The above technical solution, combined with steps 104-106 and related content, serves as an inventive point of this disclosure, solving the technical problem of "poor quality of the drawn camera topology map." Factors contributing to this poor quality often include: Drawing topology maps for large shopping mall cameras often presents the following technical problems: During the drawing process, differences in installation angles, lighting conditions, and resolution between different cameras cause significant visual variations in the features of the same target person. Conversely, the features of different target persons (such as customers wearing similar clothing) can be very similar, easily leading to identity confusion or trajectory interruption during target person tracking, resulting in failure to identify the same target person. Furthermore, multiple detection results for each video frame need to be correctly matched with existing historical information; using a simple greedy matching algorithm is prone to errors. This results in low accuracy in matching the identity of the target person within the camera, leading to low accuracy in determining the camera direction based on the tracked target person, thus resulting in poor quality of the drawn camera topology map. Solving these factors can improve the quality of the drawn camera topology map. To achieve this effect, the following steps are first performed on each target person bounding box in the target person bounding box information: First, the image represented by the target person bounding box information in the calibration video frame is determined as the target image. This yields the target image used to generate target person feature information. Then, the target image undergoes person feature extraction processing to obtain target person feature information. This allows for the extraction of deep-level appearance features (body shape, posture, etc.) of the target person in the target image to capture subtle textures, body shapes, and other deep differences, resulting in target person feature information used to generate comprehensive cost information. Finally, based on a preset historical person information list, the target person bounding box information, and the target person feature information, comprehensive cost information is generated. This yields various comprehensive cost information for generating the cost matrix. Next, based on the obtained comprehensive cost information, a cost matrix is generated. This yields a cost matrix used to generate the target person information set. Finally, inter-frame correlation matching processing is performed on the cost matrix to obtain the target person information set, wherein the target person information set corresponds to one calibration video frame sequence in each calibration video frame sequence. Therefore, the Hungarian algorithm can be used to match target people in the current frame with target people in historical frames to obtain a target person information set, thereby achieving accurate tracking of target people within the same camera and maintaining the continuity of their movement trajectory. Furthermore, by determining the target person's location through target detection and extracting deep features (body shape, pose, etc.) using ResNet-50, and tracking each target person in the target person information set by comparing the similarity of their features, the association of target people within the same camera can be achieved.Simultaneously, the Hungarian algorithm combined with a cost matrix ensures the continuity of each customer's trajectory and the stability of target person identification within a single camera. This provides reliable and continuous target person data for subsequent automatic learning of camera topology relationships. Furthermore, combining steps 104-106, cross-camera identity tracking is performed on each target person in the target person information set to obtain accurate camera orientation. This accurate camera orientation is then used to draw the camera topology map, improving its effectiveness.
[0038] Step 104: Based on the preset edge distance range of the screen, the identifiers of each camera and the sequence of each calibrated video frame, perform cross-camera target tracking processing on the obtained target person information set to obtain the trajectory information of each target person.
[0039] In some embodiments, the execution entity may perform cross-camera target tracking processing on the obtained target person information sets based on a preset screen edge distance range, each camera identifier, and each of the above-mentioned calibration video frame sequences, to obtain the trajectory information of each target person.
[0040] In some optional implementations of certain embodiments, the execution entity can perform cross-camera target tracking processing on the obtained target person information sets based on a preset image edge distance range, each camera identifier, and the aforementioned calibration video frame sequences, to obtain the trajectory information of each target person: The first step is to perform the following steps for each target person information in each of the above target person information sets: The first sub-step involves, in response to determining that the target person's location information, including the target person information, falls within the preset edge distance range of the screen, performing edge-triggered detection processing on the calibration video frame sequence corresponding to the target person information to obtain the target person's motion trajectory data. The target person's motion trajectory data corresponds to one of the camera identifiers mentioned above. The preset edge distance range can be a pre-defined screen area (e.g., at a resolution of 1920×1080, the area within 96 pixels from the edge is a pre-defined screen area). In practice, firstly, in response to determining that the target person's location information, including the target person information, falls within the preset edge distance range of the screen, the executing entity can continuously record the target person's various position information and timestamps in the calibration video frame sequence. Then, the camera identifier corresponding to the calibration video frame sequence is identified as the disappearing camera identifier. Finally, the target person's feature information, target person identifier, the target person's various position information in the calibration video frame sequence, various timestamps, and the disappearing camera identifier are identified as the target person's motion trajectory data. Each position information of the target person in the calibration video frame sequence can be the center coordinates of the target person (e.g., the coordinates of the midpoint between the customer's feet).
[0041] The second sub-step involves storing the motion trajectory data of the aforementioned target individuals into a preset database. This preset database stores the motion trajectory data of various historical target individuals within a preset time period. This historical target individual motion trajectory data may include historical target individual characteristic information, historical target individual identifiers, location information within the historical target individual's calibration video frame sequence, and timestamps. The preset time period can be a pre-defined information representing a time interval, for example, 60 seconds.
[0042] The second step involves performing cross-camera identity tracking on the target person's movement trajectory data, the camera identifiers corresponding to the target person's movement trajectory data, and the aforementioned preset database to obtain the target person's trajectory information.
[0043] In some optional implementations of certain embodiments, the execution entity may perform cross-camera identity tracking processing on the target person information based on the target person's motion trajectory data, the camera identifiers corresponding to the target person's motion trajectory data, and the aforementioned preset database, thereby obtaining the target person's trajectory information: The first step, in response to determining that the center coordinates of each target person in the target person's motion trajectory data are not within the preset edge distance range, is to perform similarity matching processing on the target person's feature information included in the target person's motion trajectory data based on the historical target person's feature information included in the preset database, thereby obtaining various similarity matching information. In practice, the execution entity can perform the following steps for each historical target person's feature information included in the preset database: determining the cosine similarity between the target person's feature information and the historical target person's feature information as the similarity matching information.
[0044] The second step involves generating target trajectory information based on the target person's motion trajectory data, the preset database, and the camera identifier corresponding to the target person's motion trajectory data, in response to determining that one of the aforementioned similarity matching information is greater than or equal to a preset similarity threshold. In practice, firstly, the executing entity can extract the target person's position information in the last five calibration video frames from the target person's motion trajectory data. Then, the timestamp of the target person in the last calibration video frame is extracted from the target person's motion trajectory data as the disappearance timestamp. Next, the disappearance camera identifier included in the target person's motion trajectory data is identified as the target camera identifier. Then, the target person's position information in the first three calibration video frames is extracted from the preset database. Then, the camera identifier corresponding to the calibration video frame sequence containing the first three calibration video frames is identified as the source camera identifier. Next, the timestamp of the first video frame in which the target person appears in the calibration video frames is extracted from the preset database as the appearance timestamp. Finally, the difference between the position information of adjacent calibration video frames in the last five calibration video frames is determined as each horizontal and vertical coordinate displacement vector and each vertical coordinate displacement vector. Next, the average value of each horizontal coordinate displacement vector is determined as the horizontal average displacement. Then, the average value of each vertical coordinate displacement vector is determined as the vertical average displacement. Then, the arctangent function of the above vertical and horizontal average displacements is determined as the fourth value. Then, the fourth value is rounded to the nearest multiple of 45° (e.g., 30° is rounded to 45°) to determine the vanishing direction angle. Next, based on the target person's position information in the first three calibration video frames, the appearance direction angle is generated. It should be noted that the method used to generate the appearance direction angle based on the target person's position information in the first three calibration video frames is the same as the method used to generate the vanishing direction angle based on the target person's position information in the last five calibration video frames. Finally, the vanishing direction angle, appearance direction angle, source camera identifier, target camera identifier, vanishing timestamp, and appearance timestamp are determined as the target person's trajectory information. The vanishing and appearance direction angles can correspond to eight directions (e.g., 0° is horizontal to the right, 90° is vertical upward). The preset similarity threshold can be a pre-defined value. For example, the positional information of the target person in the last five calibration video frames could be: (150, 1000) in the first calibration video frame, (135, 1010) in the second, (120, 1020) in the third, (105, 1030) in the fourth, and (90, 1040) in the fifth. The horizontal displacement vectors are: 135-150, 120-135, 105-120, and 90-105, respectively.The displacement vectors for each vertical axis are: 1010-1000, 1020-1010, 1030-1020, and 1040-1030. The average displacement of the above horizontal axis can be (135-150+120-135+105-120+90-105) / 4. The average displacement of the above vertical axis can be (1010-1000+1020-1010+1030-1020+1040-1030) / 4. The fourth value can be arctan2((1010-1000+1020-1010+1030-1020+1040-1030) / 4, (135-150+120-135+105-120+90-105) / 4)≈146.31°. The vanishing direction angle can be 135°.
[0045] Step 105: Based on the trajectory information of each target person, perform cross-camera association and constraint verification processing on each target person information set to obtain target association information pairs for each camera.
[0046] In some embodiments, the execution entity may perform cross-camera association and constraint verification processing on the target person information sets based on the trajectory information of each target person to obtain target association information pairs for each camera.
[0047] In addressing the aforementioned technical problems by employing technical solutions, the following technical issues often arise in the application scenario: drawing a topology map of cameras in large shopping malls. Due to the dense flow of people in and out of these malls, relying solely on the characteristic information of the target person for cross-camera tracking can lead to mismatches in feature similarity due to changes in lighting, similar clothing, and different viewing angles, resulting in two people being incorrectly associated with the same person. Furthermore, the movement behavior of pedestrians recorded by the cameras is not absolutely standardized (they may linger, turn back, or detour), which generates noisy samples (for example, cameras A and B are located on opposite sides of a straight line, connected by a direct passage. Leaving from the right side of camera A, one should enter the left side of camera B. However, the target person entering camera B from camera A does not follow the straight passage but instead detours around to the lower left corner of camera B. This generates a sample that conforms to the logic of cross-camera association but is actually incorrect). Due to the influence of these noisy samples, the accuracy of the determined camera directions is low, resulting in a poorly drawn camera topology map. The following requirements are necessary for this application scenario: accurate camera orientation is needed when drawing the camera topology map of a large shopping mall to ensure the map is drawn correctly. To address these technical challenges, we have decided to adopt the following solution: In some optional embodiments, the execution entity may perform cross-camera association and constraint verification processing on the target person information sets based on the trajectory information of each target person, thereby obtaining target association information pairs for each camera: The first step is to perform the following steps for each target person's trajectory information in the above-mentioned target person trajectory information: The first sub-step involves generating a cross-camera transfer time difference based on the target person's trajectory information, including the disappearance and appearance timestamps. In practice, the executing entity can determine the cross-camera transfer time difference as the difference between the appearance and disappearance timestamps.
[0048] The second sub-step, in response to determining that the aforementioned cross-camera transfer time difference is greater than or equal to a first preset time and less than or equal to a second preset time, performs direction angle matching processing on the appearance direction angle and disappearance direction angle included in the aforementioned target person trajectory information, based on the aforementioned target person trajectory information, to obtain direction angle matching information. The aforementioned direction angle matching information can be one of the following: the appearance direction angle is not within the aforementioned target value range, or the appearance direction angle is within the aforementioned target value range. In practice, firstly, the executing entity can determine the sum of the disappearance direction angle and a fourth preset value as the first target value. Then, the sum of the aforementioned first target value and a fifth preset value is determined as the second target value. Next, the difference between the aforementioned first target value and the fifth preset value is determined as the third target value. Then, the range between the aforementioned second target value and the aforementioned third target value is determined as the target value range. Finally, it is determined whether the appearance direction angle is within the aforementioned target value range to obtain direction angle matching information. The aforementioned fourth preset value can be 180°. The aforementioned fifth preset value can be 30°. The aforementioned first preset time can be 0.5 seconds. The aforementioned second preset time can be 3 seconds.
[0049] The third sub-step involves determining the target person trajectory information as sample target person trajectory information in response to the determination that the aforementioned direction angle matching information meets a preset condition. The preset condition may be that the aforementioned direction angle is within the aforementioned target value range.
[0050] The second step involves generating an initial probability matrix based on the obtained trajectory information of each target person in the sample and each preset direction. Each sample's trajectory information includes the source camera identifier and target camera identifier corresponding to each preset direction. In practice, firstly, the executing entity can set an initial confidence level for each preset direction. Then, the initial confidence level of each preset direction is used as an element of the initial direction probability matrix to obtain the initial probability matrix. Each sample's trajectory information corresponds to an element in the initial probability matrix. The preset directions can be eight directional angles (0° for horizontal to the right, 45° for upper right, 90° for vertical upward, 135° for upper left, 180° for horizontal to the left, 225° for lower left, 270° for vertical downward, and 315° for lower right). The initial confidence level can be 12.5%. The initial probability matrix can be an N×N×8 three-dimensional probability matrix, where N is the total number of cameras and 8 represents the confidence level of the eight directional angles. In the initial probability matrix above, each element P[A][B][d_A] represents the target camera identifier, [B] represents the source camera identifier, and [d_A] represents the initial confidence level of the direction d_A from the target camera to the source camera. For example, if the target camera identifier in the sample target person trajectory information is B-1, the source camera identifier is B-2, the vanishing direction angle is 180°, and the appearing direction angle is 0°, then the direction d_A from the target camera to the source camera could be 180° horizontally to the left.
[0051] The third step involves updating the initial probability matrix based on the trajectory information of each target person in the sample and a preset update rule, resulting in an updated probability matrix. In practice, firstly, for each target person trajectory in the sample trajectory information, the executing entity can extract the appearance and disappearance directions from the trajectory information to obtain a pair of directions. Then, the pair of directions is sequentially searched in a preset direction inverse relation table. If the pair of directions is found in the preset direction inverse relation table, the appearance and disappearance directions of the target person trajectory information match. If the pair of directions is not found in the preset direction inverse relation table, the appearance and disappearance directions of the target person trajectory information do not match. Finally, each element in the initial probability matrix is updated according to the preset update rule to obtain the updated probability matrix. The preset update rule can be that if the appearance and disappearance directions of the target person trajectory information match, the executing entity can first determine the element in the initial probability matrix corresponding to the target person trajectory information as the target element. Then, the difference between the total confidence level and the aforementioned target element is determined as the first target confidence level. Next, the product of the first target confidence level and the preset learning rate is determined as the second target confidence level. Then, the sum of the second target confidence level and the aforementioned target element is determined as the target element, thus updating the target element. If the appearance direction angle and disappearance direction angle included in each sample target person trajectory information do not match, then firstly, the executing entity can determine the element corresponding to the sample target person trajectory information in the aforementioned initial probability matrix as the target element. Then, the product of the aforementioned target element and the fifth preset value is determined as the target element, thus updating the aforementioned target element. The aforementioned preset direction reciprocity table can be a list representing the mapping relationship between a direction and its opposite direction. For example, if the appearance direction angle is 0° and the disappearance direction angle is 180°, then the aforementioned direction angle pair can be (180°, 0°). The aforementioned preset directional reciprocal relationship table can be [(0°——>180°), (90°——>270°), (45°——>225°), (135°——>315°), (180°——>0°), (90°——>270°), (225°——>45°), (315°——>135°)]. The aforementioned preset learning rate can be an adjustable parameter (for example, when the number of target person trajectory information for each sample reaches 1000, the learning rate drops to 1%). The aforementioned fifth preset value can be 0.95. The aforementioned total confidence level can be the sum of the initial confidence levels for each preset direction. For example, if the number of preset directions is 8, and the initial confidence level for each preset direction is set to 12.5%, then the aforementioned total confidence level is 100%.
[0052] The fourth step involves performing local stability determination on the updated probability matrix to obtain a set of stable relationships. In practice, firstly, the executing entity can sequentially traverse each camera identifier pair in the updated probability matrix to obtain the maximum confidence level for each preset direction corresponding to each camera identifier pair. Finally, in response to determining that the maximum confidence level is greater than or equal to a preset threshold, the camera identifier pair, the preset direction corresponding to the maximum confidence level, and the maximum confidence level are determined as stable relationships. Each stable relationship is a quadruple. For example, the camera identifier pair is (B-1, B-2). The maximum confidence level is 0.98. The preset direction corresponding to the maximum confidence level is 180°. The stable relationship can be (B-1, B-2, 180°, 0.98). The preset threshold can be 0.95.
[0053] Fifth, based on the updated probability matrix, the stable relationship set is globally optimized to obtain cross-camera target association information pairs. In practice, firstly, the executing entity can determine the stable relationship proportion as the ratio of the total number of stable relationships in the stable relationship set to all camera identifier pairs (excluding diagonals) in the updated probability matrix. Then, in response to determining that the stable relationship proportion is greater than or equal to a preset proportion, the updated probability matrix is adjusted using the MCMC algorithm to obtain the adjusted probability matrix. Next, each camera identifier pair in the adjusted probability matrix is sequentially traversed to obtain the maximum confidence level for each preset direction corresponding to each camera identifier pair. Then, in response to determining that the maximum confidence level is greater than or equal to a preset threshold, the camera identifier pair and the preset direction corresponding to the maximum confidence level are determined as cross-camera target association information pairs.
[0054] The above technical solution, combined with step 106 and related content, serves as an inventive point of this disclosure, addressing the technical problem of "poor quality of the drawn camera topology map." Factors contributing to this poor quality often include: Due to the dense flow of people in shopping malls, cross-camera tracking based solely on the target person's feature information can easily lead to mismatches due to changes in lighting, similar clothing, and different viewing angles, resulting in incorrect association of two people as the same person. Furthermore, the movement behavior of pedestrians recorded by the cameras is not absolutely standardized (they may wander, turn back, or detour), generating noise samples (for example, cameras A and B are located on opposite sides of a straight line, connected by a direct passage. Leaving from the right side of camera A, one should enter the left side of camera B. However, the target person entering camera B from camera A does not follow the straight passage but instead detours from the lower left corner of camera B. This generates a sample that conforms to the logic of cross-camera association but is actually incorrect). Due to the influence of noise samples, the accuracy of the determined camera direction is low, resulting in a poor quality of the drawn camera topology map. Solving these factors can improve the quality of the drawn camera topology map. To achieve this effect, firstly, for each of the target person trajectory information, the following steps are performed: First, based on the disappearance timestamp and appearance timestamp included in the target person trajectory information, a cross-camera transfer time difference is generated. This provides a cross-camera transfer time difference to ensure the target person trajectory information has a reasonable, non-instantaneous movement time. Then, in response to determining that the cross-camera transfer time is greater than or equal to a first preset time and less than or equal to a second preset time difference, based on the target person trajectory information, directional angle matching processing is performed on the appearance and disappearance direction angles included in the target person trajectory information to obtain directional angle matching information. This provides directional angle matching information to verify whether the target person's movement direction conforms to the physical logic of the relative positions of the two cameras. Finally, in response to determining that the directional angle matching information meets preset conditions, the target person trajectory information is determined as sample target person trajectory information. This provides sample target person trajectory information for generating the initial probability matrix. Then, based on the obtained sample target person trajectory information and each preset direction, an initial probability matrix is generated. This provides an initial probability matrix for generating the updated probability matrix. Next, based on the target trajectory information of each sample and the preset update rules, the initial probability matrix is updated to obtain the updated probability matrix. This yields the updated probability matrix used to generate a stable set of relationships. Then, the updated probability matrix undergoes local stability determination processing to obtain a stable set of relationships. This provides a stable set of relationships used to generate cross-camera target association information pairs.Finally, based on the updated probability matrix, the stable relationship set is globally optimized to obtain cross-camera target association information pairs. Therefore, the updated probability matrix can be adjusted using the MCMC algorithm to eliminate logical conflicts between camera pairs (e.g., if camera A → camera B is 0°, then camera B → camera A cannot be 0°). Also, combined with step 106, the transfer time difference and orientation angle matching of each target trajectory information eliminates mismatched target trajectory information. Simultaneously, the probability matrix is updated using a preset update rule, and the MCMC algorithm is used to adjust the probability matrix to eliminate logical conflicts between camera pairs, reducing the impact of noise samples on camera orientation judgment and obtaining cross-camera target association information pairs. This improves the accuracy of camera orientation judgment. Then, using cross-camera target association information pairs with high camera orientation accuracy, a camera topology map is generated, further improving the quality of the drawn camera topology map.
[0055] Step 106: Generate a camera topology map based on the target association information pairs of each camera.
[0056] In some embodiments, the aforementioned execution entity can generate a camera topology map based on the aforementioned camera target association information pairs. In practice, the aforementioned execution entity can use Gephi software, import the various camera target association information pairs into Gephi software, and use its built-in layout algorithm (such as Force Atlas 2) and rendering engine to generate the camera topology map.
[0057] The above-described embodiments of this disclosure have the following beneficial effects: the camera topology mapping method of some embodiments of this disclosure reduces the hardware cost required for camera topology mapping. Specifically, the reason for the high hardware cost required for camera topology mapping is that additional positioning base stations are deployed to obtain the relative positional relationships between cameras, thereby drawing the camera topology map. Since additional positioning base stations are needed to capture the relative positional relationships between camera devices, the required hardware cost is high. Based on this, the camera topology mapping method of some embodiments of this disclosure first acquires each video captured by each camera, wherein each of the cameras has a corresponding camera identifier, and each video corresponds to one of the cameras. Thus, each video captured by each camera can be obtained for generating each calibration video frame sequence. Then, each video is subjected to quality optimization and coordinate calibration processing to obtain each calibration video frame sequence. Thus, each calibration video frame sequence used to generate each target person information set can be obtained. Afterwards, for each calibration video frame sequence, target tracking processing is performed to obtain the target person information set. Thus, each target person information set used to generate target person trajectory information can be obtained. Next, based on the preset edge distance range of the screen, the identifiers of each camera, and the aforementioned calibrated video frame sequences, cross-camera target tracking processing is performed on the obtained target person information sets to obtain the trajectory information of each target person. This provides trajectory information for determining the relative positions between cameras. Then, based on the aforementioned target person trajectory information, cross-camera association and constraint verification processing is performed on the aforementioned target person information sets to obtain target association information pairs between cameras. This allows the generation of target association information pairs between cameras that characterize the relative positions between cameras, based on the target person trajectory information. Finally, a camera topology map is generated based on the aforementioned target association information pairs. This method involves target detection in the video footage from each camera during the topology map drawing process, followed by tracking of the detected targets to obtain the relative position information between cameras. The camera topology map is drawn based on the relative position information between cameras. Since this method does not require additional positioning hardware, it reduces the hardware cost of drawing the camera topology map.
[0058] Further reference Figure 2 As an implementation of the methods shown in the figures, this disclosure provides some embodiments of a camera topology mapping device, which are similar to... Figure 1 Corresponding to the method embodiments shown, the device can be specifically applied to various electronic devices.
[0059] like Figure 2 As shown, the camera topology mapping apparatus 200 in some embodiments includes: an acquisition unit 201, a first processing unit 202, a second processing unit 203, a third processing unit 204, a fourth processing unit 205, and a generation unit 206. The system comprises the following components: an acquisition unit 201, configured to acquire videos captured by various cameras, wherein each camera has a corresponding camera identifier, and each video corresponds to one of the cameras; a first processing unit 202, configured to perform quality optimization and coordinate calibration processing on the videos to obtain a sequence of calibrated video frames; a second processing unit 203, configured to perform target tracking processing on each of the calibrated video frame sequences to obtain a set of target person information; a third processing unit 204, configured to perform cross-camera target tracking processing on the obtained set of target person information based on a preset edge distance range, camera identifiers, and the calibrated video frame sequences to obtain trajectory information of each target person; a fourth processing unit 205, configured to perform cross-camera association and constraint verification processing on the target person information set based on the trajectory information of each target person to obtain target association information pairs for each camera; and a generation unit 206, configured to generate a camera topology map based on the target association information pairs for each camera.
[0060] It is understandable that the units described in the device 200 are related to the reference. Figure 1 The steps in the method described above correspond to each other. Therefore, the operations, features, and beneficial effects described above for the method also apply to the device 200 and the units contained therein, and will not be repeated here.
[0061] The following is for reference. Figure 3 It shows a schematic diagram of the structure of an electronic device 300 suitable for implementing some embodiments of the present disclosure. Figure 3 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of this disclosure.
[0062] like Figure 3As shown, the electronic device 300 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 301, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 302 or a program loaded from a storage device 308 into a random access memory (RAM) 303. The RAM 303 also stores various programs and data required for the operation of the electronic device 300. The processing unit 301, ROM 302, and RAM 303 are interconnected via a bus 304. An input / output (I / O) interface 305 is also connected to the bus 304.
[0063] Typically, the following devices can be connected to I / O interface 305: input devices 306 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 307 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 308 including, for example, magnetic tapes, hard disks, etc.; and communication devices 309. Communication device 309 allows electronic device 300 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 3 An electronic device 300 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively. Figure 3 Each box shown can represent a device or multiple devices as needed.
[0064] In particular, according to some embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, some embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 309, or installed from storage device 308, or installed from ROM 302. When the computer program is executed by processing device 301, it performs the functions defined in the methods of some embodiments of this disclosure.
[0065] It should be noted that, in some embodiments of this disclosure, the computer-readable medium may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In some embodiments of this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In some embodiments of this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0066] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0067] A computer-readable medium may be contained within an electronic device or may exist independently, not assembled into the electronic device. The computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: acquire videos captured by various cameras, wherein each of the cameras has a corresponding camera identifier, and each video corresponds to one of the cameras; perform quality optimization and coordinate calibration on the videos to obtain calibrated video frame sequences; perform target tracking processing on each calibrated video frame sequence to obtain a set of target person information; perform cross-camera target tracking processing on the obtained target person information sets based on a preset image edge distance range, camera identifiers, and the calibrated video frame sequences to obtain trajectory information for each target person; perform cross-camera association and constraint verification processing on the target person information sets based on the trajectory information to obtain target person association pairs for each camera; and generate a camera topology map based on the target person association pairs for each camera.
[0068] Computer program code for performing operations of some embodiments of this disclosure can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages—such as Java, Smalltalk, and C++—and conventional procedural programming languages—such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0069] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0070] The units described in some embodiments of this disclosure can be implemented in software or hardware. The described units can also be housed in a processor; for example, a processor may be described as including an acquisition unit, a first processing unit, a second processing unit, a third processing unit, a fourth processing unit, and a generation unit. The names of these units do not necessarily limit the specific unit; for example, the generation unit may also be described as "a unit that generates a camera topology map based on the aforementioned pairs of camera target association information."
[0071] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0072] The above description is merely a selection of preferred embodiments of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of the invention involved in the embodiments of this disclosure is not limited to technical solutions formed by specific combinations of technical features, but should also cover other technical solutions formed by arbitrary combinations of technical features or their equivalents without departing from the inventive concept. For example, technical solutions formed by substituting features with (but not limited to) technical features with similar functions disclosed in the embodiments of this disclosure.
Claims
1. A method for drawing a camera topology map, comprising: Acquire each video captured by each camera, wherein each camera has a corresponding camera identifier, and each video corresponds to one of the cameras. Each video is subjected to quality optimization and coordinate calibration to obtain a sequence of calibrated video frames. For each calibration video frame sequence in the various calibration video frame sequences, target tracking processing is performed on each calibration video frame sequence to obtain a target person information set; Based on the preset edge distance range of the screen, the identification of each camera and the sequence of each calibrated video frame, cross-camera target tracking processing is performed on the obtained target person information set to obtain the trajectory information of each target person. Based on the trajectory information of each target person, cross-camera association and constraint verification processing is performed on the information sets of each target person to obtain target association information pairs for each camera; A camera topology map is generated based on the target association information pairs of each camera.
2. The method according to claim 1, wherein, The process of performing quality optimization and coordinate calibration on each video to obtain a sequence of calibrated video frames includes: Each video is filtered to obtain a sequence of filtered video frames. The filtered video frame sequences are corrected to obtain the distortion-corrected video frame sequences. Based on each distortion-corrected video frame sequence, generate each calibrated video frame sequence.
3. The method according to claim 1, wherein, For each calibration video frame sequence in the respective calibration video frame sequences, target tracking processing is performed on each calibration video frame sequence to obtain a target person information set, including: For each of the calibration video frame sequences, perform the following steps: The calibration video frame sequence is subjected to target detection processing to obtain target person bounding box information, wherein the target person bounding box information includes target person location information, target person identifier, and target candidate box size; Based on the target person frame information, a target person information set is generated.
4. The method according to claim 1, wherein, The method involves performing cross-camera target tracking processing on the obtained target person information sets based on a preset image edge distance range, each camera identifier, and each calibrated video frame sequence to obtain the trajectory information of each target person, including: For each piece of information about a target person in each of the aforementioned target person information sets, the following steps are performed: In response to determining that the target person information includes the target person's position information within the preset edge distance range of the screen, edge trigger detection processing is performed on the calibration video frame sequence corresponding to the target person information to obtain target person motion trajectory data, wherein the target person motion trajectory data corresponds to one of the camera identifiers in the various camera identifiers; The movement trajectory data of the target person is stored in a preset database, wherein the preset database stores the movement trajectory data of each historical target person within a preset time period; Based on the target person's movement trajectory data, the camera identifiers corresponding to the target person's movement trajectory data, and the preset database, cross-camera identity tracking processing is performed on the target person's information to obtain the target person's trajectory information.
5. The method according to claim 4, wherein, The target person's motion trajectory data includes the center coordinates of each target person, target person feature information, and the target person's cross-camera identity tracking processing based on the target person's motion trajectory data, the camera identifiers corresponding to the target person's motion trajectory data, and the preset database, to obtain the target person's trajectory information, including: In response to determining that the center coordinates of each target person in the target person's motion trajectory data are not within the preset screen edge distance range, a similarity matching process is performed on the target person's feature information in the target person's motion trajectory data based on the feature information of each historical target person included in the preset database, to obtain each similarity matching information. In response to determining that one of the similarity matching information is greater than or equal to a preset similarity threshold, target person trajectory information is generated based on the target person's motion trajectory data, the preset database, and the camera identifier corresponding to the target person's motion trajectory data.
6. A camera topology mapping device, comprising: The acquisition unit is configured to acquire each video captured by each camera, wherein each camera has a corresponding camera identifier, and each video corresponds to one of the cameras. The first processing unit is configured to perform quality optimization and coordinate calibration processing on each of the videos to obtain each calibrated video frame sequence. The second processing unit is configured to perform target tracking processing on each calibration video frame sequence in the respective calibration video frame sequences to obtain a target person information set. The third processing unit is configured to perform cross-camera target tracking processing on the obtained target person information set based on a preset screen edge distance range, each camera identifier and each calibrated video frame sequence, to obtain the trajectory information of each target person. The fourth processing unit is configured to perform cross-camera association and constraint verification processing on the target person information set based on the trajectory information of each target person, so as to obtain target association information pairs for each camera. The generation unit is configured to generate a camera topology map based on the target association information pairs of each camera.
7. An electronic device, comprising: One or more processors; A storage device on which one or more programs are stored; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1 to 5.
8. A computer-readable medium having a computer program stored thereon, wherein, When the program is executed by the processor, it implements the method as described in any one of claims 1 to 5.