Street View Information Analysis Method, Device, Electronic Device, and Storage Medium

Through mobile devices, video streams and GPS data of street scene information are collected simultaneously, and the alignment and multi-dimensional feature fusion is combined with geospatial information, which solves the problems of limited video range, difficulty in splicing of multiple cameras, and errors in pedestrian misjudgment and geographic location inference in traditional street scene information analysis, achieving higher analysis accuracy.

CN120182900BActive Publication Date: 2025-07-22HARBIN INSTITUTE OF TECHNOLOGY (SHENZHEN) (INSTITUTE OF SCIENCE AND TECHNOLOGY INNOVATION HARBIN INSTITUTE OF TECHNOLOGY SHENZHEN)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510665860.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-22
Publication Date
2025-07-22
Estimated Expiration
2045-05-22

AI Technical Summary

Technical Problem

Traditional street scene information analysis technology has problems such as limited video range, difficult multi-camera splicing, location deviation, pedestrian misjudgment and geographical location inference errors, resulting in inaccurate analysis.

Method used

The mobile device synchronously collects video streams and GPS data of street view information, performs multi-dimensional feature fusion, and aligns with preset geospatial information to realize pedestrian tracking and re-identification, and obtains pedestrian attribute information.

Benefits of technology

It improves the accuracy of pedestrian location information, reduces misjudgment and quantitative statistics errors, and improves the accuracy of street view information analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182900B_ABST
    Figure CN120182900B_ABST
Patent Text Reader

Abstract

The present application relates to a street view information analysis method, apparatus, electronic device, and storage medium. The method includes: synchronously and dynamically collecting a video stream of street view information and GPS data through a mobile device, where the GPS data is used to record the geographical location information corresponding to the street view information in the video stream; aligning the video stream with the GPS data and preset geospatial information in sequence to map the geographical locations in the street view information to the geospatial in the real world; for the video stream after two alignment processes, adopting the method of pedestrian multi-dimensional feature fusion to perform pedestrian tracking and re-identification to obtain the attribute information of each pedestrian, where the attribute information includes the location and behavior category of the pedestrian at different times; analyzing the street view information according to the attribute information of multiple pedestrians. The present application improves the accuracy of street view analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a street view information analysis method, apparatus, electronic device, and storage medium. Background Art

[0002] In today's digital age, fields such as urban planning, traffic management, public safety, and the construction of smart cities are increasingly relying on data. Among them, the ability to obtain relevant information based on the population in a specific area or path is crucial. However, traditional data collection and analysis technologies face many difficulties in practical applications.

[0003] The traditional method generally collects video data at fixed monitoring points (such as traffic intersections, shopping malls, squares, etc.), and then uses feature extraction technology to identify and classify the dynamic behaviors of people in consecutive video frames. However, this brings a series of problems. First, due to the fixed perspective of the fixed camera, the video range it collects is limited. When it is necessary to analyze the pedestrian situation in a large area or complex path, the splicing and calibration of multiple cameras are difficult, and position deviation is likely to occur. Second, due to the installation position, angle, and time synchronization problems of different cameras, when integrating video data for pedestrian analysis, the images of the same pedestrian in different cameras are often misjudged as different individuals, resulting in incorrect statistics of the number of pedestrians, and further affecting the analysis of pedestrian behavior patterns. Finally, when relying on video data for pedestrian analysis, since the video itself does not have accurate geographical location information, researchers can only make rough inferences through the scene features in the video, such as inferring the specific street location where the pedestrian is located and the relative position relationship with surrounding buildings and facilities. However, such inferences often have large errors, making it impossible to deeply carry out the correlation analysis of crowd behavior and geographical space. These situations will all lead to inaccurate street view information analysis. Summary of the Invention

[0004] This application provides a street view information analysis method, apparatus, electronic device, and storage medium to solve the problem of inaccurate street view information analysis.

[0005] In a first aspect, this application provides a street view information analysis method, and the method includes:

[0006] Dynamically collect a video stream of street view information and GPS data through a mobile device, where the GPS data is used to record the geographical location information corresponding to the street view information in the video stream;

[0007] Align the video stream with the GPS data and preset geographical space information in sequence to map the geographical location in the street view information to the real-world geographical space;

[0008] For the video stream after two alignment processes, pedestrian multi-dimensional feature fusion is adopted for pedestrian tracking and re-identification to obtain the attribute information of each pedestrian, where the attribute information includes the location and behavior category of the pedestrian at different times;

[0009] Analyze the street view information according to the attribute information of multiple pedestrians.

[0010] Optionally, aligning the video stream with the GPS data and the preset geospatial information in sequence includes:

[0011] Convert the video stream into continuous image frames, where each image frame carries a time stamp of video shooting;

[0012] Based on the time stamp in the image frame, align each image frame with the GPS data to determine the shooting time and corresponding geographical location of the image frame;

[0013] Determine at least three identical and non-collinear feature points between the image frame and the geospatial information, and record the pixel coordinates of each feature point in the image frame and the geographical coordinates in the geospatial information;

[0014] Determine the conversion parameters for converting from pixel coordinates to geographical coordinates according to at least three of the feature points;

[0015] Map each image frame to the geospatial space of the geospatial information according to the conversion parameters.

[0016] Optionally, for the video stream after two alignment processes, adopting pedestrian multi-dimensional feature fusion for pedestrian tracking and re-identification to obtain the attribute information of each pedestrian includes:

[0017] When a target pedestrian in the current image frame is detected, perform multi-level matching of the multi-dimensional features of the target pedestrian with the pedestrian feature set in the database;

[0018] If the hierarchical feature matching fails, assign a pedestrian ID to the target pedestrian in the database, and store the association relationship between the pedestrian ID, the current image frame ID, and the multi-dimensional features in the database;

[0019] If each hierarchical feature matches successfully, determine the pedestrian ID of the target pedestrian in the database, and update the database according to the current image frame ID and the multi-dimensional features. Repeat the above update process until the target pedestrian no longer appears in the image frame.

[0020] Optionally, determining that each hierarchical feature matches successfully includes:

[0021] Match the biometric features of the target pedestrian with the set of pedestrian features in the database, where the biometric features include gender and age range.

[0022] If the biometric feature matching is successful, determine the first feature set composed of the pedestrian IDs that match successfully in the database, and match the overall features of the target pedestrian with the first feature set, where the overall features include skeletal features and clothing features.

[0023] If the overall feature matching is successful, determine the second feature set composed of the pedestrian IDs that match successfully in the first feature set, and match the accessory features of the target pedestrian with the second feature set, where the accessory features include the items carried by the target pedestrian, the behavior category, and the accompanying group.

[0024] If the accessory feature matching is successful, it is determined that each level of feature matching is successful.

[0025] Optionally, determining that the accessory feature matching is successful includes:

[0026] If the type of the item carried by the target pedestrian is the same as that of the item carried by the pedestrian to be matched in the second feature set, the item binding position is the same, and the color difference value of the main color of the item is within the set color difference range, configure a preset item score for the item carried by the target pedestrian, where each successfully matched item corresponds to an item score.

[0027] If the behavior category of the target pedestrian is the same as that of the pedestrian to be matched, configure a preset behavior score for the behavior category of the target pedestrian.

[0028] If the group ID of the accompanying group where the target pedestrian is located is the same as the group ID of the accompanying group where the pedestrian to be matched is located, configure a preset group score for the accompanying group of the target pedestrian.

[0029] Perform a weighted sum of the total item score, the behavior score, and the group score. If the weighted result exceeds the set score, it is determined that the accessory feature matching is successful.

[0030] Optionally, determining the accompanying group where the target pedestrian is located includes:

[0031] If the distance between the target pedestrian and at least one adjacent pedestrian is less than the set distance threshold in multiple consecutive image frames, determine that the target pedestrian and the adjacent pedestrian form an accompanying group; or,

[0032] If the target pedestrian and at least one adjacent pedestrian have the same motion trajectory, pose change, or have communication and interaction in multiple consecutive image frames, it is determined that the target pedestrian and the adjacent pedestrian form a companion group; or,

[0033] If the carry-on items of the target pedestrian and the carry-on items of at least one adjacent pedestrian are functionally complementary, it is determined that the target pedestrian and the adjacent pedestrian form a companion group.

[0034] Optionally, analyzing the street view information according to the attribute information of multiple pedestrians includes:

[0035] Performing semantic segmentation on the image frames in the video stream to obtain the spatial element distribution information in the image frames, where the spatial element distribution information includes natural landscape distribution information and public facility distribution information;

[0036] Analyzing the spatial element distribution information and the attribute information of the pedestrians to determine the association between the spatial element distribution and the behavior categories and quantities of the pedestrians.

[0037] In a second aspect, the present application provides a street view information analysis device, and the device includes:

[0038] An acquisition module, configured to synchronously acquire the video stream of the street view information and GPS data through a mobile device, where the GPS data is used to record the geographical location information corresponding to the street view information in the video stream;

[0039] An alignment module, configured to align the video stream with the GPS data and the preset geographical space information in sequence to map the geographical location in the street view information to the real-world geographical space;

[0040] A tracking and re-identification module, configured to perform pedestrian tracking and re-identification on the video stream after two alignment processes in a manner of fusing multi-dimensional features of pedestrians to obtain the attribute information of each pedestrian, where the attribute information includes the location and behavior category of the pedestrian at different times;

[0041] An analysis module, configured to analyze the street view information according to the attribute information of multiple pedestrians.

[0042] In a third aspect, the present application provides an electronic device, including: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor connected to the at least one bus; and at least one memory connected to the at least one bus.

[0043] Fourthly, the present application further provides a computer storage medium storing computer-executable instructions for executing the street view information analysis method according to any one of the above-mentioned aspects of the present application.

[0044] The above technical solutions provided by the embodiments of the present application have the following advantages compared with the prior art: by using a mobile device, street view data can be flexibly and comprehensively collected. By aligning the collected video stream with GPS data and preset geospatial information, accurate positioning can be achieved, improving the accuracy of pedestrian location information. Then, by combining the multi-dimensional features of pedestrians for pedestrian tracking and re-identification, pedestrian misjudgment and quantity statistical errors can be reduced. Finally, based on the behavior and quantity of pedestrians, the accuracy of street view information analysis is higher. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application and, together with the specification, are used to explain the principles of the present application.

[0046] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0047] One or more embodiments are exemplarily illustrated by the pictures in the corresponding accompanying drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings represent similar elements, unless otherwise stated, and the drawings in the figures do not constitute a proportional limitation.

[0048] Figure 1 Schematic diagram of the hardware environment for street view information analysis provided by the embodiments of the present application;

[0049] Figure 2 Flowchart of a method for street view information analysis provided by the embodiments of the present application;

[0050] Figure 3 Flowchart of a method for identifying pedestrian behavior categories provided by the embodiments of the present application;

[0051] Figure 4 Schematic diagram of the crowd distribution heat map provided by the embodiments of the present application;

[0052] Figure 5 Schematic diagram of the behavior distribution report provided by the embodiments of the present application;

[0053] Figure 6 Schematic diagram of the green view rate provided by the embodiments of the present application;

[0054] Figure 7 Schematic diagram of bench distribution provided by an embodiment of the present application;

[0055] Figure 8 Schematic diagram of behavior-facility association matrix provided by an embodiment of the present application;

[0056] Figure 9 Schematic diagram of the structure of a street view information analysis device provided by an embodiment of the present application;

[0057] Figure 10 Schematic diagram of the structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0058] To make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0059] The following disclosure provides many different embodiments or examples for implementing different structures of the present application. To simplify the disclosure of the present application, the components and settings of specific examples are described below. Of course, they are only examples and are not intended to limit the present application. In addition, the present application may repeat reference numerals and / or letters in different examples. Such repetition is for the purpose of simplification and clarity and does not itself indicate the relationship between the various embodiments and / or settings discussed.

[0060] To solve the problem of inaccurate street view information analysis mentioned in the background art, in the embodiments of the present application, the continuous video collected by the mobile device is aligned with the real geographical space in terms of position, and then the pedestrians in the video are re-identified based on multi-dimensional features, and the street view information is analyzed based on the obtained pedestrian attribute information, avoiding problems such as deviation in the integration of video and video position and duplicate pedestrian recognition, improving the accuracy of the identified pedestrian positions and quantities, and thus improving the accuracy of street view information analysis.

[0061] The embodiments of the present application can be applied to multiple application scenarios, including but not limited to: traffic management optimization, adding public facilities, and commercial site selection planning, etc.

[0062] Optionally, in the embodiments of the present application, the above-mentioned street view information analysis method can be applied to the hardware environment composed of a terminal 101 and a server 103 as shown in Figure 1 and as shown in Figure 1As shown in the figure, the server 103 is connected to the terminal 101 through a network and can be used to provide services for the terminal or the client installed on the terminal. A database 105 can be set up on the server or independently of the server to provide data storage services for the server 103. The above-mentioned network includes, but is not limited to: wide area network, metropolitan area network or local area network. The terminal 101 includes, but is not limited to: PC, mobile phone, tablet computer, etc.

[0063] Next, in combination with specific embodiments, a street view information analysis method provided by an embodiment of the present application will be described in detail. Taking the application to a server as an example, as Figure 2 shown, the specific steps are as follows:

[0064] Step 201: Dynamically collect the video stream and GPS data of the street view information through a mobile device. Among them, the GPS data is used to record the geographical location information corresponding to the street view information in the video stream;

[0065] Step 202: Align the video stream with the GPS data and the preset geospatial information in sequence to map the geographical location in the street view information to the real-world geospatial;

[0066] Step 203: For the video stream after two alignment processes, adopt the method of pedestrian multi-dimensional feature fusion to perform pedestrian tracking and re-identification to obtain the attribute information of each pedestrian. Among them, the attribute information includes the location and behavior category of the pedestrian at different times;

[0067] Step 204: Analyze the street view information according to the attribute information of multiple pedestrians.

[0068] Next, the nouns appearing in the embodiments of the present application will be explained, specifically including the following content.

[0069] Mobile device: A portable electronic device with data collection and processing capabilities, such as a panoramic camera, action camera, smart phone, handheld or head-mounted camera, etc.

[0070] Video stream: A series of image frame sequences continuously captured by the mobile device camera. These image frames are arranged in chronological order and record the picture information of the street view at different times. The resolution ≥ 1080P, the frame rate ≥ 30fps, and it can clearly present the pedestrians, objects and environmental details in the street view.

[0071] GPS data: Data obtained by the Global Positioning System (GPS), including longitude, latitude, time, ground speed, etc., and is used to accurately record the shooting time and geographical location corresponding to each frame of the street view in the video stream.

[0072] Geospatial information: A pre-prepared dataset containing information such as geographic coordinates, terrain, and distribution of ground features, such as high-precision maps, CAD (Computer-AIDed Designfile) data, or GIS (Geographic Information System) data. By aligning with this information, the geographical location in the street view information can be accurately corresponded to the real-world geospatial.

[0073] In step 201, a mobile device (such as a panoramic camera, smartphone, handheld camera, etc.) synchronously conducts dynamic acquisition of video stream and GPS data during movement and uploads the acquired data to the server. The mobile device has portability and flexibility, can penetrate various scenarios, break through the limitation of fixed perspectives, comprehensively cover all street scenes to be photographed, and obtain rich and comprehensive original street scene data.

[0074] The video stream can accurately capture the detailed information of pedestrians and the surrounding environment, while the GPS data accurately records the geographical location information corresponding to each frame of the video stream, covering key data such as longitude, latitude, time, and ground speed, laying a foundation for the subsequent geographical positioning of street view information. For example, in the daily acquisition scenario of urban streets, a staff member walks along the street with a mobile device equipped with a high-resolution camera. The mobile device continuously acquires the video stream, and at the same time, the GPS module real-time records the location information of the mobile device, ensuring the close association between the street view image and the geographical location.

[0075] In step 202, the video stream itself lacks accurate geographical coordinates and time information, while the GPS data records accurate time and geographical location. Aligning the video stream with the GPS data can determine the accurate shooting time and corresponding geographical location for each frame in the video stream, so that the scene in the video stream can be accurately positioned in the time and space dimensions. For example, when monitoring the crowd activities on a certain street, by aligning with the GPS data, it is possible to accurately know the specific time when the crowd appears in the video and their exact location on the street, providing a reliable spatio-temporal basis for subsequent analysis.

[0076] The geospatial information stores rich geospatial data, including information on terrain, landform, ground features, etc. Aligning the video stream after the initial alignment with the geospatial information can visually map the scene in the video stream to the real geographical environment. Users can clearly see the specific area photographed by the video on the map and understand its relative position relationship with surrounding geographical elements (such as roads, buildings, etc.), making the geographical location of the video stream more intuitive and easy to understand. For example, the ground feature data can be used to identify objects such as buildings and roads in the video and associate them with the corresponding objects in the actual geospatial, enabling the video stream to accurately correspond to the location in the real geographical environment.

[0077] Video streams contain rich visual information, such as crowd behaviors, street scene features, etc.; GPS data provides location and time information; and geospatial information covers various real geospatial information. After aligning the three, the fusion of multi-source data is achieved, which can accurately locate the scenes in the video stream into the real-world geospatial, determine their specific geographical locations, and facilitate the accuracy of subsequent data analysis.

[0078] In step 203, the server combines the multi-dimensional features of pedestrians (including skeletal features, appearance features, carried item features, etc.) for pedestrian tracking and re-identification. In consecutive frames of the video stream, these features are comprehensively used to construct a pedestrian feature database to continuously track pedestrians. When a pedestrian appears in different frames, by comparing the database, it can be accurately judged whether it is the same pedestrian, avoiding misjudgment of pedestrians caused by factors such as perspective changes and similar clothing. At the same time, accurate pedestrian tracking also makes the pedestrian count more accurate, reducing the situation of double counting or missed counting.

[0079] In step 204, the server analyzes the street scene information by combining the attribute information of multiple pedestrians, and can obtain at least one of the following analysis results: a crowd distribution heat map, a behavior distribution report, and the association between spatial elements and pedestrians.

[0080] The server can generate a crowd distribution heat map based on the location information of multiple pedestrians at different times. The crowd distribution heat map is used to show which areas are crowded and which areas are deserted, which helps the managers of commercial areas to reasonably plan the store layout, guide the flow of people, or help relevant departments anticipate in advance the safety risks that may be brought by crowd gathering.

[0081] The server can generate a behavior distribution report by statistically analyzing according to the behavior category information of pedestrians in different regions and time periods. The behavior distribution report details the proportion of pedestrians with different behaviors such as walking, running, and standing in each region. During the urban planning process, these data can be used to reasonably set facilities in the park, such as adding benches in areas where pedestrian rest behaviors are concentrated, and planning fitness areas in areas where sports behaviors are concentrated, to improve the use experience of public spaces.

[0082] The server can combine the spatial elements in the street scene (such as water areas, benches, green spaces, etc.) and pedestrian attribute information to analyze the association between the two. Specifically, it can count the number and behaviors of pedestrians around different public facilities, or generate a relationship graph between the change in the proportion of spatial elements and the change in behavior distribution, which is used to analyze the impact of the change in spatial elements on pedestrian behavior. For example, it is found through analysis that in areas with a relatively high green view rate in a certain park, the number of pedestrians staying (behavior categories are standing or sitting) is relatively large; near bus stops, the number of pedestrians gathering increases significantly, and most of them are in a standing waiting state.

[0083] Exemplarily, in the operation and management scenario of a large scenic area, the scenic area staff use smartphones to walk along the main tourist routes in the scenic area, and the smartphones synchronously collect video streams and GPS data. After the collection is completed, the data is uploaded to the server, and the server aligns the video stream with the GPS data and the geospatial information of the scenic area, so that each section of the street view can be accurately corresponded to the specific location on the scenic area map. Then, a multi-dimensional feature fusion analysis is performed on the pedestrians in the video stream to obtain the location and behavior information of each tourist. Through analysis, it is found that in the lakeside area of the scenic area (where the water view rate is relatively high), most tourists will choose to stay and enjoy the scenery (the behavior category is standing); while on the main roads connecting various scenic spots, tourists mainly walk and ride (there is a bicycle rental service in the scenic area). Based on these analysis results, the scenic area management department can optimize the tourist route signs, add rest facilities and service points in the areas where tourists stay more, and improve the tourist experience.

[0084] In this application, by using a mobile device, street view data can be collected flexibly and comprehensively. By aligning the collected video stream with the GPS data and the preset geospatial information, accurate positioning can be achieved, the accuracy of pedestrian location information can be improved, and pedestrian tracking and re-identification can be carried out by combining the multi-dimensional features of pedestrians, which can reduce pedestrian misjudgment and quantity statistics errors. Finally, based on the behavior and quantity of pedestrians, the accuracy of street view information analysis is higher.

[0085] As an optional implementation manner, in step 202, aligning the video stream with the GPS data and the preset geospatial information in sequence includes the following contents:

[0086] Step S11: Convert the video stream into continuous image frames, where each image frame carries a time stamp of the video shooting;

[0087] Step S12: Based on the time stamp in the image frame, align each image frame with the GPS data to determine the shooting time and the corresponding geographical location of the image frame;

[0088] Step S13: Determine at least three identical and non-collinear feature points between the image frame and the geospatial information, and record the pixel coordinates of each feature point in the image frame and the geographical coordinates in the geospatial information;

[0089] Step S14: Determine the conversion parameters for converting from pixel coordinates to geographical coordinates based on at least three feature points;

[0090] Step S15: Map each image frame to the geospatial of the geospatial information according to the conversion parameters.

[0091] Among them, step S11 includes the following contents.

[0092] Removing invalid video segments: During the street view data collection process, the video stream obtained by the mobile device may contain various invalid video segments. For example, black screen images caused by device failures, occlusions, or other factors, or some special images that do not meet the analysis requirements. Then, it is necessary to clip and remove this invalid video segment to avoid the interference of invalid data on subsequent analysis and improve the efficiency and accuracy of data processing.

[0093] Image frame extraction: After completing the video clip, use a dedicated video processing algorithm to convert the processed video into consecutive image frames frame by frame in chronological order. During the conversion process, ensure that each frame of the image is accurately attached with the timestamp at the time of collection. This timestamp records the specific moment when the image frame was taken and is an important identifier for achieving precise alignment with GPS data. By attaching timestamps to each frame of the image, subsequent operations can correspond the image frames with GPS data one by one based on the key dimension of time, thereby accurately associating the street view images with the geographical location information at the time of shooting.

[0094] GPS data smoothing: Since GPS signals are easily interfered by various factors during transmission, such as building occlusions, signal reflections, etc., the obtained GPS data has noise and abnormal points. By using the Kalman smoothing algorithm to correct GPS noise and abnormal points, a smooth GPS trajectory is generated, so that the GPS data can more truly reflect the actual movement path of the mobile device.

[0095] Among them, the following content is included in step S12.

[0096] GPS interpolation: In street view data processing, the acquisition frequency of video frames is usually high, while the acquisition frequency of GPS data is relatively low, which leads to an out-of-sync problem between the two in terms of time. To solve this problem, GPS interpolation operations are carried out. Based on the timestamps of video frames, linear interpolation alignment is performed on the low-frequency GPS data. Specifically, according to adjacent GPS data points and their corresponding times, by means of linear calculation, appropriate GPS data points are inserted within the time intervals of video frames, so that the GPS data is more matched with the video frames in terms of time. GPS interpolation enables better synchronization of GPS data and video frames in terms of time, solving the problem of time inconsistency caused by different data acquisition frequencies.

[0097] Filtering invalid segments: In the collected street view data, in addition to invalid video segments in the video stream, the corresponding GPS data may also contain invalid segments. To ensure the validity and accuracy of the data, it is necessary to perform the operation of filtering invalid segments. By identifying black screen images or special image features in the video and simultaneously removing the corresponding GPS data points of these invalid images. In addition, duplicate points and blurred images in the data are also filtered. Duplicate points may be caused by GPS signal fluctuations or device errors, and blurred images will affect the analysis of street view information. By removing these invalid data, the quality of the data and the accuracy of the analysis can be improved.

[0098] In step S13, the system supports users to upload open-source maps, CAD files, or GIS files. The server extracts feature points from the image frames and geospatial information (such as high-precision maps). At least three identical and non-collinear feature points are selected through a matching algorithm. These feature points correspond to the same position in reality, and the pixel coordinates (with the upper left corner as the origin) of each feature point in the image frame and the geospatial coordinates (such as longitude and latitude) are recorded.

[0099] In step S14, the server determines the conversion parameters from pixel coordinates to geospatial coordinates based on the affine transformation model. Affine transformation can combine linear and translation transformations in two-dimensional space for coordinate mapping. By substituting the pixel and geospatial coordinates of at least three feature points into the matrix formula of affine transformation, an overdetermined system of equations is constructed, and then algorithms such as the least squares method are used to solve it to obtain the conversion parameters that minimize the sum of the squares of the errors of the system of equations. These parameters accurately describe the conversion relationship between the two types of coordinates.

[0100] In step S15, each pixel of the image frame is traversed, and the pixel coordinates are substituted into the previously determined affine transformation formula to calculate the corresponding geospatial coordinates of the pixel. In this way, all pixels of the entire image frame are converted, enabling elements such as pedestrians and buildings in the image frame to accurately correspond to real geospatial positions, laying a geographical positioning foundation for subsequent street view information analysis.

[0101] The above operations of deleting invalid video segments and selecting feature points support users' interactive operations, and users can adjust and optimize according to the actual situation.

[0102] In this application, the system supports users to upload geospatial files. Through a series of complex and precise coordinate conversion algorithms, the GPS data points in the video stream are automatically matched with the geospatial data. This matching process is not completely automated and also fully supports users' interactive operations, realizing the accurate mapping of street view information to the real-world geospatial, providing a data basis for subsequent street view information analysis that includes accurate geospatial information.

[0103] As an alternative implementation, in step 203, for the video stream after two alignment processes, a pedestrian multi-dimensional feature fusion method is adopted for pedestrian tracking and re-identification, and the attribute information of each pedestrian includes the following:

[0104] Step S21: When the target pedestrian in the current image frame is detected, the multi-dimensional features of the target pedestrian are hierarchically matched with the pedestrian feature set in the database;

[0105] Step S22: If the hierarchical feature matching fails, a pedestrian ID is assigned to the target pedestrian in the database, and the association relationship between the pedestrian ID, the current image frame ID, and the multi-dimensional features is stored in the database;

[0106] Step S23: If each hierarchical feature matches successfully, the pedestrian ID of the target pedestrian in the database is determined, and the database is updated according to the current image frame ID and the multi-dimensional features. The above update process is repeated until the target pedestrian no longer appears in the image frame.

[0107] After completing the two alignment processes of the video stream with GPS data and the preset geospatial information, it enters the pedestrian tracking and re-identification stage. In this stage, with the help of the pedestrian multi-dimensional feature fusion technology, the attribute information of each pedestrian is accurately obtained, covering key data such as the location and behavior category of the pedestrian at different times. The specific process is as follows.

[0108] When the video stream is played frame by frame and a target pedestrian (a random pedestrian) is detected in the current image frame, the server starts a hierarchical matching process. The server comprehensively extracts the multi-dimensional features of the target pedestrian, including biometric features (such as gender, age group), overall features (skeletal features, clothing features), and accessory features (features of carried items, behavior category, accompanying group).

[0109] The server performs hierarchical matching of these multi-dimensional features with the pedestrian feature set stored in the database. First, start matching from the biometric feature level. If the gender and age group of the target pedestrian do not match the corresponding biometric features in the pedestrian feature set in the database, it is determined that the hierarchical feature matching fails. Once a hierarchical feature matching failure occurs, the system assigns a unique pedestrian ID to the target pedestrian in the database to ensure its recognizability throughout the analysis process. At the same time, the association relationship between this pedestrian ID, the current image frame ID, and the multi-dimensional features of the target pedestrian is completely stored in the database for subsequent query and analysis. The structure of the person feature database is shown in Table 1.

[0110] Table 1

[0111]

[0112] If the biometric features of the target pedestrian match successfully with the set of pedestrian features in the database, then enter the next level, i.e., the overall feature matching. Compare the skeletal features and clothing features of the target pedestrian with the corresponding features in the first feature set determined by the successful biometric feature matching in the database. If a match fails at this level, assign a new ID to the target pedestrian and store the relevant information according to the above process. If the overall feature matching is successful, further determine the second feature set composed of the pedestrian IDs that match successfully in the first feature set, and then carry out the accessory feature matching. Compare the accessory features such as the carried items, behavior categories, and accompanying groups of the target pedestrian with the corresponding features in the second feature set in detail.

[0113] Only when the features of each level match successfully can the system determine the pedestrian ID corresponding to the target pedestrian in the database. Subsequently, the system updates the database based on the current image frame ID and the latest multi-dimensional features of the target pedestrian. The updated content includes the position information of the target pedestrian in the current image frame (determined by aligning the image frame with the geospatial information), changes in behavior categories, etc. Since the target pedestrian may be at a relatively far position when first captured, but the last time the pedestrian is captured is the closest point to the target pedestrian, the above update process will be repeated until the target pedestrian completely disappears from subsequent image frames, thus completely and accurately recording all the action trajectories and changes in attribute information of the target pedestrian in the video stream.

[0114] This application can dynamically record the latest position and behavior status of pedestrians by updating the database according to the current image frame ID and the latest multi-dimensional features. As pedestrians move and their behaviors change in the street scene, the database is continuously updated to ensure that no matter when pedestrians appear in different image frames, the system can accurately track based on the latest data, avoiding tracking interruptions or errors caused by information lag, and achieving accurate recording of the entire action trajectory of pedestrians. In addition, the matching of multi-level features increases the accuracy of pedestrian recognition.

[0115] The following describes each feature of pedestrians.

[0116] 1. Biometric features: including gender and age group. Exemplarily, the division of age groups is shown in Table 2.

[0117] Table 2

[0118]

[0119] 2. Overall features: including skeletal features and clothing features.

[0120] (1) Skeletal features.

[0121] Use Yolo-Pose to obtain human key points (such as shoulders, elbows, knees, ankles, etc.), calculate the ratio values of each feature, normalize them to the range [0, 1] (eliminating the influence of height differences), and construct a feature vector .

[0122] , where represents the shoulder width to height ratio, represents the shoulder width, represents the height from the midpoint of the two shoulders to the top of the head.

[0123] , where represents the thigh to calf ratio, represents the length from the hip to the knee, represents the length from the knee to the ankle.

[0124] , where represents the arm length to torso ratio, represents the arm length (from the shoulder to the wrist), represents the torso height (from the neck to the hip).

[0125] (2) Clothing features.

[0126] Based on the human key points obtained by Yolo-Pose, distinguish the head, upper body, and lower body, segment the person image, calculate the LAB color histograms of each part respectively, and normalize them to probability distributions, finally obtaining 3 sub-block feature vectors .

[0127] Among them, , , are the color features of the head, upper body, and lower body respectively.

[0128] 3. Auxiliary features: including items carried by the target pedestrian, behavior categories, and accompanying groups.

[0129] (1) Items carried. Items carried include but are not limited to: accessories (such as backpacks, hats, etc.), bicycles, skateboards, and pets.

[0130] a. For accessories, features can be extracted in the following way: First, use a segmentation model (such as the Yolov8-seg model) to analyze the street view image, detect accessory areas such as backpacks and hats, and generate a pixel-level mask for each accessory. The mask is used to clearly separate the accessory from the complex background. Then, use a human pose estimation algorithm to obtain the positions of human key points. For example, confirm that the backpack area should be within the back area defined by the human key points, and the hat area should overlap with the head position of the human key points to verify the accuracy of accessory detection. Finally, convert the accessory area to the LAB color space, use the K-means algorithm to cluster the pixels, select the cluster center with a proportion greater than a certain ratio (such as 30%) and a color difference ΔE greater than a set difference (such as ΔE>15) from the clustering results, and record this cluster center as the dominant color in LAB format. Eventually, form features including the accessory type, the position of the accessory, and the dominant color.

[0131] b. For bicycles, skateboards, and pets, features can be extracted in the following way: First, use a detection model (such as the Yolo-v7 model) to scan the street view image, detect the associated items of pedestrians (such as bicycles and scooters) or associated animals (such as cats and dogs), accurately identify their positions in the image, and generate corresponding detection boxes. Then, calculate whether the intersection over union (IoU) of the item detection box and the area of the human key points obtained by the human pose estimation algorithm is greater than a set threshold (such as 0.4) to ensure the accuracy and relevance of the detection. For example, judge whether the bicycle pedal is close to the human key point of the foot, and whether the pet leash is connected to the human key point of the hand. After the IoU meets the condition, output the type of the item or animal, determine the binding position (such as handheld or backpacked) according to its relative position to the human body, and extract the dominant color of the item according to the K-means algorithm. Combine this information to form features including the item type, the binding position, and the dominant color.

[0132] (2) Behavior categories. Such as Figure 3As shown in the figure, a street view image is input into the system. First, a detection module (such as the Yolo-v7 model) is used to detect the target and generate a detection frame containing pedestrians. Then, a human key point detection model (such as Yolo-Pose) is used to estimate the posture. Yolo-Pose determines the posture information of pedestrians by analyzing the human key points of pedestrians. Then, the activity database trained based on the image data set is used to compare and match the posture information of pedestrians with the posture features of various activities such as walking, running, standing, cycling, and sitting in the activity database. If the posture information of the pedestrian matches the posture features in the activity database, the activity type of the pedestrian is directly determined; if not, the posture feature is recorded in the single-person behavior list. Then, according to the proportion of each posture in the single-person behavior list, the posture with a higher probability is selected as the behavior category result. If the proportions of each posture are consistent, the last detected posture is selected as the final behavior category result.

[0133] (3) Companion groups. The methods for determining companion groups include at least one of the following three methods.

[0134] a. If the distance between the target pedestrian and at least one adjacent pedestrian is less than the set distance threshold in multiple consecutive image frames, it is determined that the target pedestrian and the adjacent pedestrian form a companion group. In other words, if the distance between pedestrians is close, such as holding hands or crossing arms, and this relative position relationship is maintained in multiple frames of images, it can be inferred that there is a relationship between these pedestrians.

[0135] b. If the target pedestrian and at least one adjacent pedestrian have the same motion trajectory, posture changes, or communication interactions in multiple consecutive image frames, it is determined that the target pedestrian and the adjacent pedestrian form a companion group. Team members may have a certain degree of synchronization in walking, standing, and other behaviors. For example, team members may stop together, turn together, etc., and their behaviors can be determined by analyzing the motion trajectory and posture changes of pedestrians in multiple frames of images. Or observe whether there are signs of communication and interaction between pedestrians, such as eye contact, gesture interaction, etc.

[0136] c. If the object carried by the target pedestrian complements the object carried by at least one neighboring pedestrian, the target pedestrian and the neighboring pedestrian are determined to form a companion group. Although some objects are not in the hands of the same person, they have a complementary relationship of function and can also be considered as common objects. For example, one person holds a camera and the other holds a tripod. In this case, it can be judged that they may be taking photos together, and these objects are considered common objects.

[0137] In addition, the relative position and distance between the carried object and the pedestrians can be observed. If the distance between a carried object and multiple pedestrians is relatively close and this relative position relationship is maintained in multiple frames of images, it can be speculated that these pedestrians may be associated with the carried object. For example, if a large suitcase is detected and the positions of two pedestrians are both close to the suitcase, and their movement trajectories are consistent with the movement trajectory of the suitcase, it can be speculated that these two pedestrians may be carrying the suitcase together.

[0138] The feature matching of pedestrians is described below.

[0139] 1. Biometric matching. The gender matching must be exactly the same, and the age group can float up or down by one age group deviation. If the matching is successful, the overall feature matching is entered; if the matching fails, a pedestrian ID is configured for the target pedestrian in the database.

[0140] 2. Overall feature matching. The process of overall feature matching is as follows: match the skeletal features of the target pedestrian with the skeletal features of the pedestrians to be matched in the first feature set to determine the skeletal similarity; match the clothing features of the target pedestrian with the clothing features of the pedestrians to be matched to determine the clothing similarity; if the weighted sum value of the skeletal similarity and the clothing similarity is greater than the set threshold, it is determined that the overall feature matching is successful.

[0141] (1) Calculate the skeletal feature set of the target pedestrian and the skeletal feature set of the pedestrians to be matched in the database The weighted Euclidean distance, and the formula for the weighted Euclidean distance is as follows.

[0142] , where represents the weighted Euclidean distance, represents the weight of the i-th skeletal feature, represents the i-th skeletal feature of the target pedestrian, represents the i-th skeletal feature of the pedestrian to be matched. Exemplarily, .

[0143] The formula for skeletal feature similarity is: .

[0144] (2) Use the Bhattacharyya distance to calculate the similarity between the color feature set of the target pedestrian and the color feature set of the pedestrians to be matched in the database. The calculation formula for color feature similarity is as follows.

[0145] , where represents the color feature similarity (clothing similarity), Represents the color probability value of the target pedestrian in the i-th interval. Represents the color probability value of the pedestrian to be matched in the i-th interval, and n represents the total number of color intervals.

[0146] (3)For the similarity of skeletal features and the similarity of color features Perform weighting to obtain the comprehensive similarity D. The formula for the comprehensive similarity is as follows: , where is the weight of the skeletal feature, is the weight of the color feature.

[0147] Exemplarily, if D < 0.75, enter the matching of accessory features. If D ≥ 0.75, the matching fails, and a pedestrian ID is configured for the target pedestrian in the database.

[0148] 3. Matching of accessory features.

[0149] The matching process of the accessory features is as follows: If the item type of the item carried by the target pedestrian is the same as that of the item carried by the pedestrian to be matched in the second feature set, the item binding position is the same, and the color difference value of the main color of the item is within the set color difference range, then a preset item score is configured for the item of the target pedestrian, where each successfully matched item corresponds to an item score; if the behavior category of the target pedestrian is the same as that of the pedestrian to be matched, then a preset behavior score is configured for the behavior category of the target pedestrian; if the group ID of the group that the target pedestrian belongs to is the same as the group ID of the group that the pedestrian to be matched belongs to, then a preset group score is configured for the group that the target pedestrian belongs to; perform weighted summation on the total item score, behavior score, and group score. If the weighted result exceeds the set score, it is determined that the matching of the accessory features is successful.

[0150] Exemplarily, the target pedestrian A is carrying two items, a hat and a backpack, and the pedestrian B to be matched in the second feature set is also carrying similar items. After inspection, the color difference value of the hat is 10 (meeting ΔE < 15), and the binding positions are both on the head. The color difference value of the backpack is 12, and the binding positions are both on the back. According to the rules, if there are 2 carried items, 0.5 points are added for successfully matching one item, and 1 point for a complete match. Here, both the hat and the backpack are successfully matched, so 1 point is obtained for the carried items. The target pedestrian A is walking at this moment, and the pedestrian B to be matched also has the behavior of walking in the behavior list, so the behavior pattern is successfully matched, and 1 point is recorded. The group ID of the group that the target pedestrian A belongs to is G1, and the group ID of the group that the pedestrian B to be matched belongs to is also G1, so the group relationship is successfully matched, and 1 point is recorded. Finally, the carried items, behavior, and group relationship are weighted and scored according to 0.5, 0.3, and 0.2. Since 1 point is obtained for the carried items, 1 point for the behavior pattern, and 1 point for the group relationship. The weighted calculation is: 1 * 0.5 + 1 * 0.3 + 1 * 0.2 = 1. Because 1 exceeds the set 0.75 points, the accessory feature matching is successful.

[0151] In this application, through multi-level matching of biometric features, overall features (bones, clothing, etc.), and accessory features (carried items, behavior patterns, group relationships, etc.), the pedestrian identity is gradually refined and accurately located. This progressive method can avoid misidentification caused by misjudgment of a single feature and improve the accuracy of pedestrian identity recognition. The bone features can reflect the structural and postural information of the human body, the clothing features provide significant appearance identifiers, and the accessory features further enrich the description of the pedestrian's features, enabling the system to identify and track pedestrians from multiple perspectives, improving the recognition accuracy both overall and in detail.

[0152] As an optional implementation manner, analyzing the street view information according to the attribute information of multiple pedestrians includes: performing semantic segmentation on the image frames in the video stream to obtain the spatial element distribution information in the image frames, where the spatial element distribution information includes natural landscape distribution information and public facility distribution information; analyzing the spatial element distribution information and the attribute information of the pedestrians to determine the association between the spatial element distribution and the behavior categories and quantities of the pedestrians.

[0153] Analyzing the street view information can obtain at least one of the following analysis results: a crowd distribution heat map, a behavior distribution report, and the association between spatial elements and pedestrians.

[0154] Figure 4It is a schematic diagram of a population distribution heat map. The person re-identification technology can accurately record the information of each pedestrian at different times and locations, and these location information are the basic data for generating the heat map. The heat map represents the frequency and density of pedestrians by the depth of color. The accurate location information provided by person re-identification enables the heat map to more precisely reflect the pedestrian flow distribution. For example, in a large exhibition, by obtaining the movement trajectories of each exhibitor through person re-identification technology and then generating a heat map, it is possible to intuitively see which areas have a large pedestrian flow and which areas are relatively deserted. Combining with the time series information of person re-identification, the changing trend of the heat map over time can also be analyzed. For instance, the changing situation of the pedestrian flow at the entrances of different stores in a shopping mall at different time periods of a day is helpful for merchants to reasonably arrange business hours and employee schedules.

[0155] Figure 5 It is a schematic diagram of a behavior distribution report. During the person re-identification process, the behaviors of pedestrians will be detected and classified (such as walking, running, standing, cycling, etc.), and these behavior information are important bases for generating the behavior distribution report. By statistically analyzing the behaviors of different pedestrians, the behavior distribution in different regions and time periods can be obtained. For example, in a park, by counting the proportion of the number of people running, walking, and resting at different time periods to form a behavior distribution report, it provides a reference for the park management department to optimize the facility layout and activity arrangements. Person re-identification can also track the behavior changes of individual pedestrians. Summarizing the individual behavior data into the behavior distribution report can provide a more comprehensive understanding of the crowd behavior pattern. For example, analyzing the shopping behavior of a certain customer in a shopping mall, including the staying time and purchasing behavior in different stores during the whole process from entering the mall to leaving, provides data support for merchants to carry out precise marketing.

[0156] The server performs semantic segmentation on the image frames in the video stream and can obtain the spatial element distribution information in the image frames. The spatial element distribution information includes natural landscape distribution information and public facility distribution information. The natural landscape distribution information includes but is not limited to: sky view ratio, water view ratio, and green view ratio.

[0157] The sky view ratio refers to the pixel proportion of the spatial element of the sky in the image, the water view ratio refers to the pixel proportion of the spatial element of the water area in the image. As Figure 6 shown, the green view ratio refers to the pixel proportion of the spatial element of green plants in the image, while trash cans, benches, etc. belong to public facilities. The public facility distribution information refers to the pixel proportion of public facilities such as trash cans and benches in the image. Figure 7It is a distribution schematic diagram of benches. If the video stream is a panoramic video, the video frame needs to be cut into four orientations: east, west, south, and north. Calculate the pixel proportion of each spatial element in the image respectively. Finally, take the average value of the four images as the index value of the natural landscape distribution information. The public facility distribution information is calculated by accumulating the number of public facilities in the four images.

[0158] The server analyzes the spatial element distribution information and the attribute information of pedestrians to determine the association between the spatial element distribution and the behavior categories and quantities of pedestrians, which provides a method for studying the relationship between elements such as sky view rate, water view rate, and trash cans and pedestrians. For example, it can be analyzed whether pedestrians are more inclined to behaviors such as staying and resting in areas with a higher sky view rate; whether the behavior patterns of pedestrians (such as walking, viewing, etc.) are different in places with a high water view rate; and what characteristics the aggregation quantity and behaviors of pedestrians (such as discarding garbage, staying briefly, etc.) will show in areas where trash cans are densely distributed.

[0159] The server can also generate a behavior-facility association matrix, a relationship diagram between the change in element proportion and the change in behavior distribution, etc. Among them, the behavior-facility association matrix is to count the quantities of different behaviors of pedestrians around public facilities, and output the association matrix by calculating the Pearson correlation coefficient, as Figure 8 shown. The relationship diagram between the change in element proportion and the change in behavior distribution refers to a double Y-axis chart. Among them, the horizontal axis represents time or spatial position, the left vertical axis represents the change in green view rate, and the right vertical axis represents the change in the number of standing pedestrians. Through such a double Y-axis chart, it can be intuitively observed whether there is a certain association between the change in the proportion of spatial elements and the change in pedestrian behavior distribution. For example, whether the number of standing pedestrians is relatively large in areas with a higher green view rate; or whether the number of standing pedestrians will fluctuate correspondingly during a certain time period when the green view rate changes during the day. Through this kind of analysis, it can help researchers or relevant personnel understand the impact of environmental factors (spatial elements) on pedestrian behavior, so as to provide a reference basis for urban planning, landscape design, etc.

[0160] The embodiment of this application provides an overall process step for street view information analysis, including the following content.

[0161] 1. Data collection.

[0162] Video stream collection: Use mobile devices equipped with high-definition cameras, such as smartphones, panoramic cameras, etc., to continuously collect street view video streams at a resolution of not less than 1080P and a frame rate of 30fps during the movement process, and comprehensively record pedestrian and surrounding environment information.

[0163] GPS data collection: Synchronously obtain accurate longitude, latitude, time, ground speed, etc. information through the device's built-in GPS module or an external high-precision GPS device, and use it to mark the geographical location corresponding to each frame of the video stream.

[0164] 2. Video Alignment.

[0165] Align the video with GPS data: Convert the video stream into continuous image frames and attach timestamps to each frame. Based on the timestamp of the image frame, perform linear interpolation on the low-frequency GPS data to ensure accurate temporal matching between the video frames and GPS location information, providing a basis for subsequent geolocation.

[0166] Align with geospatial information: Identify at least three identical and non-collinear feature points in the image frame and the preset geospatial information (such as high-precision maps, GIS data), and record the pixel coordinates of the feature points in the image frame and the geographic coordinates in the geospatial. Through the affine transformation model, calculate the transformation parameters based on the feature point coordinates to accurately map the image frame to the real-world geospatial.

[0167] 3. Object Detection.

[0168] Use the detection model to perform object detection on each frame of the video stream, identify the items carried by pedestrians (such as bicycles, scooters, etc.) and pets (such as cats, dogs, etc.), and generate detection boxes.

[0169] Use the segmentation model to detect regions of accessories such as backpacks and hats and generate pixel-level masks.

[0170] 4. Feature Extraction.

[0171] Gender feature: Use a dual-branch MobileNetV3 structure to process face and full-body images and output three classifications: male, female, and unknown.

[0172] Age group feature: Use a soft segmentation regression network to output the age interval with the highest probability.

[0173] Skeletal feature: Obtain the human key point information of pedestrians through the human key point detection model for pose estimation and skeletal feature extraction.

[0174] Clothing feature: Extract the color feature of the clothing from the overall appearance of the pedestrian.

[0175] Feature of carried items: For the detected carried items, extract features such as their type, binding position (such as handheld, carried on the back), and main color (select the dominant color by performing K-means clustering on the accessory region in the LAB color space, K = 3, and select the clustering center with a proportion > 30% and color difference ΔE > 15 as the dominant color).

[0176] Behavior feature: Based on the pose information obtained from the human key point detection model, compare it with the activity database to determine the behavior pattern of the pedestrian (walking, running, standing, cycling, sitting posture, etc.).

[0177] Group relationship feature: Determine the group ID of the group the pedestrian belongs to and analyze their group relationship.

[0178] 5. Multi-level matching.

[0179] Biometric matching: Match the biometric features such as the gender and age range of the target pedestrian with the pedestrian feature set in the database. If the match fails, assign a new pedestrian ID to the target pedestrian and store the relevant information (pedestrian ID, image frame ID, multi-dimensional features).

[0180] Overall feature matching: If the biometric matching is successful, further match the skeletal features and clothing features with the corresponding features in the database. If the match fails, also assign a new ID and store the information.

[0181] Ancillary feature matching: If the overall feature matching is successful, perform a comprehensive feature weighted matching on the carried items, behavior patterns, and group relationships. For the carried items, check the color difference ΔE < 15 and position consistency, and add points according to the number of carried items; match the behavior patterns by type, and record 1 point for the existence of the same type of behavior; match the group relationship by the group ID, and record 1 point for a successful match. Finally, perform a weighted score according to 0.5 (carried items), 0.3 (behavior), and 0.2 (group relationship). If the score exceeds 0.75, the match is successful. If the match fails, record the relevant information in the database.

[0182] 6. Result output and update.

[0183] If the multi-level matching is successful, determine the pedestrian ID of the target pedestrian in the database, and update the database according to the current image frame ID and the latest multi-dimensional features, including information such as location and behavior changes.

[0184] If the match fails, record the relevant information in the database.

[0185] 7. Analyze the street view information in combination with the location changes and behavior changes of pedestrians.

[0186] Based on the same technical concept, this application provides a street view information analysis device, as Figure 9 shown, the device includes:

[0187] An acquisition module 901, configured to synchronously acquire the video stream and GPS data of the street view information through a mobile device, where the GPS data is used to record the geographical location information corresponding to the street view information in the video stream;

[0188] An alignment module 902, configured to align the video stream with the GPS data and the preset geographical space information in sequence, so as to map the geographical location in the street view information to the real-world geographical space;

[0189] A tracking and re-identification module 903, which is used to perform pedestrian tracking and re-identification on the video stream after two alignment processes by adopting the method of pedestrian multi-dimensional feature fusion to obtain the attribute information of each pedestrian, where the attribute information includes the position and behavior category of the pedestrian at different times;

[0190] An analysis module 904, which is used to analyze the street view information according to the attribute information of multiple pedestrians.

[0191] Optionally, the alignment module 902 is used to:

[0192] Convert the video stream into continuous image frames, where each image frame carries the time stamp of video shooting;

[0193] Align each image frame with the GPS data based on the time stamp in the image frame to determine the shooting moment and corresponding geographical location of the image frame;

[0194] Determine at least three identical and non-collinear feature points between the image frame and the geospatial information, and record the pixel coordinates of each feature point in the image frame and the geographical coordinates in the geospatial information;

[0195] Determine the conversion parameters for converting from pixel coordinates to geographical coordinates according to at least three feature points;

[0196] Map each image frame into the geospatial space of the geospatial information according to the conversion parameters.

[0197] Optionally, the tracking and re-identification module 903 is used to:

[0198] In the case of detecting the target pedestrian in the current image frame, perform multi-level matching of the multi-dimensional features of the target pedestrian with the pedestrian feature set in the database;

[0199] If the hierarchical feature matching fails, assign a pedestrian ID to the target pedestrian in the database, and store the association relationship between the pedestrian ID, the current image frame ID, and the multi-dimensional features in the database;

[0200] If each hierarchical feature matches successfully, determine the pedestrian ID of the target pedestrian in the database, and update the database according to the current image frame ID and the multi-dimensional features. Repeat the above update process until the target pedestrian no longer appears in the image frame.

[0201] Optionally, the tracking and re-identification module 903 is specifically used to:

[0202] Match the biometric features of the target pedestrian with the pedestrian feature set in the database, where the biometric features include gender and age group;

[0203] If the biometric match is successful, determine the first feature set composed of the pedestrian IDs that match successfully in the database, and match the overall features of the target pedestrian with the first feature set, where the overall features include skeletal features and clothing features;

[0204] If the overall feature match is successful, determine the second feature set composed of the pedestrian IDs that match successfully in the first feature set, and match the accessory features of the target pedestrian with the second feature set, where the accessory features include the items carried by the target pedestrian, the behavior category, and the companion group;

[0205] If the accessory feature match is successful, it is determined that each level of features matches successfully.

[0206] Optionally, the tracking and re-identification module 903 is specifically configured to:

[0207] If the item type of the item carried by the target pedestrian is the same as that of the item carried by the pedestrian to be matched in the second feature set, the item binding positions are the same, and the color difference value of the main color of the item is within the set color difference range, configure a preset item score for the item of the target pedestrian, where each successfully matched item corresponds to an item score;

[0208] If the behavior category of the target pedestrian is the same as that of the pedestrian to be matched, configure a preset behavior score for the behavior category of the target pedestrian;

[0209] If the group ID of the companion group where the target pedestrian is located is the same as the group ID of the companion group where the pedestrian to be matched is located, configure a preset group score for the companion group of the target pedestrian;

[0210] Perform a weighted sum of the total item score, the behavior score, and the group score. If the weighted result exceeds the set score, it is determined that the accessory feature match is successful.

[0211] Optionally, the tracking and re-identification module 903 is specifically configured to:

[0212] If the distance between the target pedestrian and at least one adjacent pedestrian is less than the set distance threshold in multiple consecutive image frames, determine that the target pedestrian and the adjacent pedestrian form a companion group; or,

[0213] If the target pedestrian and at least one adjacent pedestrian have the same motion trajectory, pose change, or have communication and interaction in multiple consecutive image frames, determine that the target pedestrian and the adjacent pedestrian form a companion group; or,

[0214] If the carried item of the target pedestrian is functionally complementary to the carried item of at least one adjacent pedestrian, determine that the target pedestrian and the adjacent pedestrian form a companion group.

[0215] Optionally, the analysis module 904 is used to:

[0216] Semantically segment the image frames in the video stream to obtain the spatial element distribution information in the image frames, where the spatial element distribution information includes natural landscape distribution information and public facility distribution information;

[0217] Analyze the spatial element distribution information and the attribute information of pedestrians to determine the association between the spatial element distribution and the behavior categories and quantities of pedestrians.

[0218] As Figure 10 shown, an embodiment of the present application provides an electronic device, including a processor 1001, a communication interface 1002, a memory 1003, and a communication bus 1004. Among them, the processor 1001, the communication interface 1002, and the memory 1003 complete mutual communication through the communication bus 1004.

[0219] The memory 1003 is used to store computer programs.

[0220] In an embodiment of the present application, when the processor 1001 executes the program stored on the memory 1003, it implements the street view information analysis method provided by any one of the foregoing method embodiments.

[0221] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the street view information analysis method provided by any one of the foregoing method embodiments.

[0222] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0223] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the related technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., including several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0224] It should be understood that the terms used herein are for the purpose of describing particular example embodiments only and are not intended to be limiting. Unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" as used herein may also include the plural forms. The terms "comprising", "including", "containing", and "having" are inclusive and thus specify the presence of the stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the particular order described or illustrated, unless the order of performance is explicitly stated. It should also be understood that additional or alternative steps may be used.

[0225] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.

Claims

1. A street view information analysis method, characterized in that, The method includes: Dynamically collecting a video stream of street view information and GPS data through a mobile device, wherein the GPS data is used to record the geographical location information corresponding to the street view information in the video stream; Aligning the video stream with the GPS data and preset geospatial information in sequence to map the geographical location in the street view information to the geospatial in the real world; For the video stream after two alignment processes, adopting the method of pedestrian multi-dimensional feature fusion for pedestrian tracking and re-identification to obtain the attribute information of each pedestrian, wherein the attribute information includes the location and behavior category of the pedestrian at different times; Analyzing the street view information according to the attribute information of multiple pedestrians; Among them, for the video stream after two alignment processes, adopting the method of pedestrian multi-dimensional feature fusion for pedestrian tracking and re-identification to obtain the attribute information of each pedestrian includes: When detecting a target pedestrian in the current image frame, performing multi-level matching of the multi-dimensional features of the target pedestrian with the pedestrian feature set in the database; If there is a situation where the hierarchical feature matching fails, assigning a pedestrian ID to the target pedestrian in the database and storing the association relationship between the pedestrian ID, the current image frame ID, and the multi-dimensional features in the database; If each hierarchical feature matches successfully, determining the pedestrian ID of the target pedestrian in the database and updating the database according to the current image frame ID and the multi-dimensional features, repeating the above update process until the target pedestrian no longer appears in the image frame; Among them, analyzing the street view information according to the attribute information of multiple pedestrians includes: Performing semantic segmentation on the image frames in the video stream to obtain the spatial element distribution information in the image frames, wherein the spatial element distribution information includes natural landscape distribution information and public facility distribution information; Analyzing the spatial element distribution information and the attribute information of the pedestrians to determine the association between the spatial element distribution and the behavior category and quantity of the pedestrians.

2. The method according to claim 1, wherein Aligning the video stream with the GPS data and preset geospatial information in sequence includes: Converting the video stream into continuous image frames, wherein each image frame carries a time stamp of video shooting; Based on the time stamp in the image frame, aligning each image frame with the GPS data to determine the shooting moment and corresponding geographical location of the image frame; Determining at least three identical and non-collinear feature points between the image frame and the geospatial information, and recording the pixel coordinates of each feature point in the image frame and the geographical coordinates in the geospatial information; Determining the conversion parameters for converting from pixel coordinates to geographical coordinates according to at least three of the feature points; Mapping each image frame to the geospatial in the geospatial information according to the conversion parameters.

3. The method according to claim 1, wherein Determining that each hierarchical feature matches successfully includes: Matching the biometric features of the target pedestrian with the pedestrian feature set in the database, wherein the biometric features include gender and age group; If the biometric match is successful, determine the first feature set composed of the pedestrian IDs that match successfully in the database, and match the overall features of the target pedestrian with the first feature set, where the overall features include skeletal features and clothing features; If the overall feature match is successful, determine the second feature set composed of the pedestrian IDs that match successfully in the first feature set, and match the accessory features of the target pedestrian with the second feature set, where the accessory features include the items carried by the target pedestrian, the behavior category, and the companion group; If the accessory feature match is successful, it is determined that each level of feature match is successful.

4. The method according to claim 3, characterized in that, Determining that the accessory feature match is successful includes: If the type of item carried by the target pedestrian is the same as the type of item carried by the pedestrian to be matched in the second feature set, the item binding position is the same, and the color difference value of the main color of the item is within the set color difference range, configure a preset item score for the item carried by the target pedestrian, where each successfully matched item corresponds to an item score; If the behavior category of the target pedestrian is the same as the behavior category of the pedestrian to be matched, configure a preset behavior score for the behavior category of the target pedestrian; If the group ID of the companion group where the target pedestrian is located is the same as the group ID of the companion group where the pedestrian to be matched is located, configure a preset group score for the companion group of the target pedestrian; Perform a weighted sum of the total item score, the behavior score, and the group score. If the weighted result exceeds the set score, it is determined that the accessory feature match is successful.

5. The method according to claim 4, wherein Determining the companion group where the target pedestrian is located includes: If the distance between the target pedestrian and at least one adjacent pedestrian is less than the set distance threshold in multiple consecutive image frames, determine that the target pedestrian and the adjacent pedestrian form a companion group; or, If the target pedestrian and at least one adjacent pedestrian have the same motion trajectory, pose change, or have communication and interaction in multiple consecutive image frames, determine that the target pedestrian and the adjacent pedestrian form a companion group; or, If the carried item of the target pedestrian is functionally complementary to the carried item of at least one adjacent pedestrian, determine that the target pedestrian and the adjacent pedestrian form a companion group.

6. A street view information analysis device, characterized in that, The device includes: An acquisition module, configured to synchronously acquire a video stream of street view information and GPS data through a mobile device, where the GPS data is used to record the geographical location information corresponding to the street view information in the video stream; An alignment module, configured to align the video stream with the GPS data and the preset geographical space information in sequence, so as to map the geographical location in the street view information to the geographical space in the real world; A tracking and re-identification module, configured to perform pedestrian tracking and re-identification on the video stream after two alignment processes by using the method of pedestrian multi-dimensional feature fusion, and obtain the attribute information of each pedestrian, where the attribute information includes the location and behavior category of the pedestrian at different times; An analysis module, configured to analyze the street view information according to the attribute information of multiple pedestrians; Among them, the tracking and re-identification module is used for: In the case of detecting a target pedestrian in the current image frame, perform multi-level matching between the multi-dimensional features of the target pedestrian and the set of pedestrian features in the database; If there is a failure in hierarchical feature matching, assign a pedestrian ID to the target pedestrian in the database, and store the association relationship between the pedestrian ID, the current image frame ID, and the multi-dimensional features in the database; If each hierarchical feature matches successfully, determine the pedestrian ID of the target pedestrian in the database, and update the database according to the current image frame ID and the multi-dimensional features. Repeat the above update process until the target pedestrian no longer appears in the image frame; Among them, the analysis module is used to: Perform semantic segmentation on the image frames in the video stream to obtain the spatial element distribution information in the image frames, where the spatial element distribution information includes natural landscape distribution information and public facility distribution information; Analyze the spatial element distribution information and the attribute information of the pedestrians to determine the association between the spatial element distribution and the behavior categories and quantities of the pedestrians.

7. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used to store computer programs; The processor is used to implement the method according to any one of claims 1-5 when executing the program stored on the memory.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Cross-video pedestrian positioning tracking method, system and device

    CN111462200A

  • Pedestrian tracking method, system and device and storage medium

    CN114663835A