Streetscape information analysis method and device, electronic equipment and storage medium
Video streams and GPS data are collected through mobile devices, combined with geospatial information for alignment, and tracking and re-identification with multi-dimensional features of pedestrians, solving the problems of pedestrian position deviation and misjudgment in traditional street scene information analysis, and improving the accuracy of the analysis.
Patent Information
- Application Number
- CN202510665860.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-22
AI Technical Summary
When traditional street scene information analysis technology collects and analyzes pedestrian situations in large areas or complex paths, there are problems such as fixed perspective, position deviation, pedestrian misjudgment and statistical errors, resulting in inaccurate analysis.
The video stream and GPS data of street view information are collected dynamically by mobile devices, and the video stream is aligned with GPS data and preset geospatial information. Pedestrian tracking and re-identification are carried out by fusion of pedestrian multi-dimensional features, and the attribute information of each pedestrian is obtained, and street view information is analyzed based on this information.
It improves the accuracy of pedestrian location information, reduces pedestrian misjudgment and quantitative errors, and enhances the accuracy and depth of street scene information analysis.
Smart Images

Figure CN120182900A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a street view information analysis method, apparatus, electronic device, and storage medium. Background Art
[0002] In today's digital age, fields such as urban planning, traffic management, public safety, and the construction of smart cities are increasingly dependent on data, and the ability to obtain relevant information based on people in specific areas or paths is crucial. However, traditional data collection and analysis technologies face many difficulties in practical applications.
[0003] The traditional method generally collects video data at fixed monitoring points (such as traffic intersections, shopping malls, squares, etc.), and then uses feature extraction technology to identify and classify the dynamic behaviors of people in consecutive video frames. However, this brings a series of problems. First, due to the fixed perspective of the fixed camera, the video range it collects is limited. When analyzing the pedestrian situation in a large area or complex path, the splicing and calibration of multiple cameras are difficult, and position deviation is likely to occur. Second, due to the installation position, angle, and time synchronization problems of different cameras, when integrating video data for pedestrian analysis, the images of the same pedestrian in different cameras are often misjudged as different individuals, resulting in incorrect statistics of the number of pedestrians, and thus affecting the analysis of pedestrian behavior patterns. Finally, when relying on video data for pedestrian analysis, since the video itself does not have accurate geographical location information, researchers can only make rough inferences through the scene features in the video, such as inferring the specific street location where the pedestrian is located and the relative position relationship with surrounding buildings and facilities. However, such inferences often have large errors, making it impossible to deeply carry out the correlation analysis between crowd behavior and geographical space. These situations will all lead to inaccurate street view information analysis. Summary of the Invention
[0004] This application provides a street view information analysis method, apparatus, electronic device, and storage medium to solve the problem of inaccurate street view information analysis.
[0005] In a first aspect, this application provides a street view information analysis method, and the method includes: Dynamically collecting a video stream of street view information and GPS data through a mobile device, where the GPS data is used to record the geographical location information corresponding to the street view information in the video stream; Aligning the video stream with the GPS data and preset geographical space information in sequence to map the geographical location in the street view information to the real-world geographical space; For the video stream after two alignment processes, pedestrian multi-dimensional feature fusion is adopted for pedestrian tracking and re-identification to obtain the attribute information of each pedestrian, where the attribute information includes the location and behavior category of the pedestrian at different times; Analyze the street view information according to the attribute information of multiple pedestrians.
[0006] Optionally, aligning the video stream with the GPS data and the preset geospatial information in sequence includes: Convert the video stream into continuous image frames, where each image frame carries a time stamp of video shooting; Based on the time stamp in the image frame, align each image frame with the GPS data to determine the shooting time and corresponding geographical location of the image frame; Determine at least three identical and non-collinear feature points between the image frame and the geospatial information, and record the pixel coordinates of each feature point in the image frame and the geographical coordinates in the geospatial information; Determine the conversion parameters for converting from pixel coordinates to geographical coordinates according to at least three of the feature points; Map each image frame into the geospatial space of the geospatial information according to the conversion parameters.
[0007] Optionally, for the video stream after two alignment processes, adopting pedestrian multi-dimensional feature fusion for pedestrian tracking and re-identification to obtain the attribute information of each pedestrian includes: When a target pedestrian in the current image frame is detected, perform multi-level matching of the multi-dimensional features of the target pedestrian with the pedestrian feature set in the database; If the hierarchical feature matching fails, assign a pedestrian ID to the target pedestrian in the database, and store the association relationship between the pedestrian ID, the current image frame ID, and the multi-dimensional features in the database; If each hierarchical feature matches successfully, determine the pedestrian ID of the target pedestrian in the database, and update the database according to the current image frame ID and the multi-dimensional features. Repeat the above update process until the target pedestrian no longer appears in the image frame.
[0008] Optionally, determining that each hierarchical feature matches successfully includes: Match the biometric features of the target pedestrian with the pedestrian feature set in the database, where the biometric features include gender and age group; If the biometric match is successful, determine the first feature set composed of the pedestrian IDs that match successfully in the database, and match the overall features of the target pedestrian with the first feature set, where the overall features include skeletal features and clothing features; If the overall feature match is successful, determine the second feature set composed of the pedestrian IDs that match successfully in the first feature set, and match the accessory features of the target pedestrian with the second feature set, where the accessory features include the items carried by the target pedestrian, the behavior category, and the companion group; If the accessory feature match is successful, it is determined that each level of feature match is successful.
[0009] Optionally, determining that the accessory feature match is successful includes: If the type of the item carried by the target pedestrian is the same as the type of the item carried by the pedestrian to be matched in the second feature set, the item binding position is the same, and the color difference value of the main color of the item is within the set color difference range, configure a preset item score for the item carried by the target pedestrian, where each successfully matched item corresponds to an item score; If the behavior category of the target pedestrian is the same as the behavior category of the pedestrian to be matched, configure a preset behavior score for the behavior category of the target pedestrian; If the group ID of the companion group where the target pedestrian is located is the same as the group ID of the companion group where the pedestrian to be matched is located, configure a preset group score for the companion group of the target pedestrian; Perform a weighted sum of the total item score, the behavior score, and the group score. If the weighted result exceeds the set score, it is determined that the accessory feature match is successful.
[0010] Optionally, determining the companion group where the target pedestrian is located includes: If the distance between the target pedestrian and at least one adjacent pedestrian is less than the set distance threshold in multiple consecutive image frames, determine that the target pedestrian and the adjacent pedestrian form a companion group; or, If the target pedestrian and at least one adjacent pedestrian have the same motion trajectory, pose change, or have communication and interaction in multiple consecutive image frames, determine that the target pedestrian and the adjacent pedestrian form a companion group; or, If the carried item of the target pedestrian is functionally complementary to the carried item of at least one adjacent pedestrian, determine that the target pedestrian and the adjacent pedestrian form a companion group.
[0011] Optionally, analyzing the street view information based on the attribute information of multiple pedestrians includes: Perform semantic segmentation on the image frames in the video stream to obtain the spatial element distribution information in the image frames, where the spatial element distribution information includes natural landscape distribution information and public facility distribution information; Analyze the spatial element distribution information and the attribute information of the pedestrians to determine the association between the spatial element distribution and the behavior categories and quantities of the pedestrians.
[0012] In a second aspect, the present application provides a street view information analysis device, which includes: An acquisition module, configured to synchronously and dynamically acquire the video stream and GPS data of the street view information through a mobile device, where the GPS data is used to record the geographical location information corresponding to the street view information in the video stream; An alignment module, configured to align the video stream with the GPS data and the preset geographical space information in sequence to map the geographical location in the street view information to the geographical space in the real world; A tracking and re-identification module, configured to perform pedestrian tracking and re-identification on the video stream after two alignment processes in a manner of fusing multi-dimensional features of pedestrians to obtain the attribute information of each pedestrian, where the attribute information includes the locations and behavior categories of the pedestrians at different times; An analysis module, configured to analyze the street view information according to the attribute information of multiple pedestrians.
[0013] In a third aspect, the present application provides an electronic device, including: at least one communication interface; at least one bus connected to the at least one communication interface; at least one processor connected to the at least one bus; and at least one memory connected to the at least one bus.
[0014] In a fourth aspect, the present application further provides a computer storage medium, storing computer-executable instructions, where the computer-executable instructions are used to execute the street view information analysis method described in any one of the above of the present application.
[0015] The above technical solutions provided by the embodiments of the present application have the following advantages compared with the prior art: By using a mobile device, street view data can be collected flexibly and comprehensively. By aligning the collected video stream with GPS data and preset geographical space information, accurate positioning can be achieved, improving the accuracy of pedestrian location information. Then, by combining the multi-dimensional features of pedestrians for pedestrian tracking and re-identification, pedestrian misjudgment and quantity statistical errors can be reduced. Finally, analyzing the street view information based on the behaviors and quantities of pedestrians is more accurate. Description of the Drawings
[0016] The drawings here are incorporated into the specification and form a part of the specification, showing the embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0017] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0018] One or more embodiments are exemplarily illustrated by the pictures in the corresponding drawings. These exemplary illustrations do not limit the embodiments. Elements with the same reference numerals in the drawings are represented as similar elements. Unless otherwise stated, the drawings in the figures do not constitute a proportional limitation.
[0019] Figure 1 Schematic diagram of the hardware environment for street view information analysis provided by the embodiments of the present application; Figure 2 Flowchart of a method for street view information analysis provided by the embodiments of the present application; Figure 3 Flowchart of a method for identifying pedestrian behavior categories provided by the embodiments of the present application; Figure 4 Schematic diagram of the crowd distribution heat map provided by the embodiments of the present application; Figure 5 Schematic diagram of the behavior distribution report provided by the embodiments of the present application; Figure 6 Schematic diagram of the green view rate provided by the embodiments of the present application; Figure 7 Schematic diagram of the bench distribution provided by the embodiments of the present application; Figure 8 Schematic diagram of the behavior - facility association matrix provided by the embodiments of the present application; Figure 9 Schematic diagram of the structure of a street view information analysis device provided by the embodiments of the present application; Figure 10 Schematic diagram of the structure of an electronic device provided by the embodiments of the present application. Detailed implementation manners
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.
[0021] The following disclosure provides many different embodiments or examples for implementing different structures of the present application. To simplify the disclosure of the present application, components and settings of specific examples are described below. Of course, they are merely examples and are not intended to limit the present application. In addition, the present application may repeat reference numerals and / or letters in different examples. Such repetition is for the purpose of simplification and clarity, and does not itself indicate the relationship between various embodiments and / or settings discussed.
[0022] To solve the problem of inaccurate analysis of street view information mentioned in the background art, the embodiments of the present application align the continuous video collected by the mobile device with the real geographical space, and then re-identify the pedestrians in the video based on multi-dimensional features, and analyze the street view information based on the obtained pedestrian attribute information, avoiding problems such as deviation in the integration of video and video location and duplicate pedestrian recognition, improving the accuracy of the identified pedestrian location and quantity, and thus improving the accuracy of street view information analysis.
[0023] The embodiments of the present application can be applied to multiple application scenarios, including but not limited to: traffic management optimization, adding public facilities, and commercial site selection planning, etc.
[0024] Optionally, in the embodiments of the present application, the above-mentioned street view information analysis method can be applied to Figure 1 the hardware environment composed of the terminal 101 and the server 103 as shown. As Figure 1 shown, the server 103 is connected to the terminal 101 through the network, and can be used to provide services for the terminal or the client installed on the terminal. A database 105 can be set on the server or independently of the server, and is used to provide data storage services for the server 103. The above-mentioned network includes but is not limited to: wide area network, metropolitan area network or local area network, and the terminal 101 includes but is not limited to PC, mobile phone, tablet computer, etc.
[0025] Next, the specific embodiments will be combined to detail a street view information analysis method provided by the embodiments of the present application. Taking the application to the server as an example, as Figure 2 shown, the specific steps are as follows: Step 201: Dynamically collect the video stream and GPS data of the street view information through the mobile device synchronously, where the GPS data is used to record the geographical location information corresponding to the street view information in the video stream; Step 202: Align the video stream with the GPS data and the preset geographical space information in sequence to map the geographical location in the street view information to the geographical space in the real world; Step 203: For the video stream after two alignment processes, adopt the method of fusing pedestrian multi-dimensional features to perform pedestrian tracking and re-identification, and obtain the attribute information of each pedestrian, where the attribute information includes the location and behavior category of the pedestrian at different times; Step 204: Analyze the street view information according to the attribute information of multiple pedestrians.
[0026] The following explains the nouns that appear in the embodiments of the present application, specifically including the following content.
[0027] Mobile device: A portable electronic device with data acquisition and processing capabilities, such as a panoramic camera, action camera, smartphone, handheld or head-mounted camera, etc.
[0028] Video stream: A series of image frame sequences continuously captured by the camera of a mobile device. These image frames are arranged in chronological order, recording the picture information of the street view at different moments, with a resolution ≥ 1080P and a frame rate ≥ 30fps, capable of clearly presenting pedestrians, objects, and environmental details in the street view.
[0029] GPS data: Data obtained by the Global Positioning System (GPS), including information such as longitude, latitude, time, and ground speed, used to accurately record the shooting time and geographical location corresponding to each frame of the street view in the video stream.
[0030] Geospatial information: A dataset prepared in advance containing information such as geographical coordinates, terrain, and land feature distribution, such as high-precision maps, CAD (Computer-AIDed Designfile) data, or GIS (Geographic Information System) data. By aligning with this information, an accurate correspondence can be established between the geographical locations in the street view information and the real-world geospatial space.
[0031] In step 201, a mobile device (such as a panoramic camera, smartphone, handheld camera, etc.) synchronously conducts dynamic acquisition of video stream and GPS data during movement and uploads the acquired data to the server. The mobile device has portability and flexibility, can penetrate various scenarios, break through the limitations of fixed perspectives, comprehensively cover all street views that need to be photographed, and obtain rich and comprehensive original street view data.
[0032] The video stream can accurately capture the detailed information of pedestrians and the surrounding environment, while the GPS data accurately records the geographical location information corresponding to each frame of the video stream, covering key data such as longitude, latitude, time, and ground speed, laying a foundation for the subsequent geographical positioning of street view information. For example, in the daily acquisition scenario of an urban street, a staff member holds a mobile device equipped with a high-resolution camera and walks along the street. The mobile device continuously acquires the video stream, and at the same time, the GPS module real-time records the location information of the mobile device, ensuring the close association between the street view image and the geographical location.
[0033] In step 202, the video stream itself lacks accurate geographical coordinates and time information, while the GPS data records accurate time and geographical locations. Aligning the video stream with the GPS data can determine the accurate shooting time and corresponding geographical location for each frame in the video stream, so that the scenes in the video stream can be accurately located in the time and space dimensions. For example, when monitoring the crowd activities on a certain street, by aligning with the GPS data, it is possible to accurately know the specific time when the crowd appears in the video and their exact positions on the street, providing a reliable spatio-temporal basis for subsequent analysis.
[0034] The geospatial information stores rich geospatial data, including information on terrain, landforms, and ground features. Aligning the video stream after the initial alignment with the geospatial information can visually map the scenes in the video stream to the real geographical environment. Users can clearly see the specific area captured by the video on the map and understand its relative positional relationship with surrounding geographical elements (such as roads, buildings, etc.), making the geographical location of the video stream more intuitive and easy to understand. For example, the ground feature data can be used to identify objects such as buildings and roads in the video and associate them with their corresponding counterparts in the actual geospatial space, enabling the video stream to accurately correspond to the location in the real geographical environment.
[0035] The video stream contains rich visual information, such as crowd behaviors and street scene features; the GPS data provides location and time information; and the geospatial information covers various real geospatial information. After aligning the three, the fusion of multi-source data is achieved, which can accurately locate the scenes in the video stream in the geographical space of the real world, determine their specific geographical locations, and facilitate the accuracy of subsequent data analysis.
[0036] In step 203, the server combines the multi-dimensional features of pedestrians (including skeletal features, appearance features, carried object features, etc.) for pedestrian tracking and re-identification. In the consecutive frames of the video stream, these features are comprehensively used to construct a pedestrian feature database to continuously track pedestrians. When a pedestrian appears in different frames, by comparing with the database, it can be accurately judged whether it is the same pedestrian, avoiding misjudgment of pedestrians caused by factors such as perspective changes and similar clothing. At the same time, accurate pedestrian tracking also makes the pedestrian count more accurate, reducing the situation of double counting or missing records.
[0037] In step 204, the server analyzes the street scene information by combining the attribute information of multiple pedestrians, and the following at least one analysis result can be obtained: a crowd distribution heat map, a behavior distribution report, and the association between spatial elements and pedestrians.
[0038] Based on the location information of multiple pedestrians at different times, the server can generate a crowd distribution heat map, which is used to show which areas are crowded and which areas are deserted. This helps the managers of commercial areas to reasonably plan the layout of stores, guide the flow of people, or help relevant departments anticipate in advance the safety risks that may be brought about by crowd gathering.
[0039] The server can conduct statistics according to different regions and time periods based on the behavior category information of pedestrians and generate a behavior distribution report. The behavior distribution report details the proportion of pedestrians with different behaviors such as walking, running, and standing in each region. During the urban planning process, facilities in the park can be reasonably set according to these data. For example, benches can be added in areas where the pedestrian rest behavior is concentrated, and fitness venues can be planned in areas where the sports behavior is concentrated to improve the usage experience of public spaces.
[0040] The server can combine the spatial elements in the street view (such as water areas, benches, green spaces, etc.) and pedestrian attribute information to analyze the correlation between the two. Specifically, it can count the number and behaviors of pedestrians around different public facilities, or generate a relationship diagram between the change in the proportion of spatial elements and the change in behavior distribution, which is used to analyze the impact of the change in spatial elements on pedestrian behavior. For example, it is found through analysis that in areas with a relatively high green view rate in a certain park, the number of pedestrians staying (behavior category is standing or sitting) is relatively large; near bus stops, the number of pedestrians gathering increases significantly, and most of them are in a standing waiting state.
[0041] Exemplarily, in the operation and management scenario of a large scenic area, the scenic area staff use smartphones to walk along the main tour routes in the scenic area, and the smartphones synchronously collect video streams and GPS data. After the collection is completed, the data is uploaded to the server, and the server aligns the video stream with the GPS data and the geographical spatial information of the scenic area, so that each section of the street view can be accurately corresponding to the specific location on the scenic area map. Then, a multi-dimensional feature fusion analysis is performed on the pedestrians in the video stream to obtain the location and behavior information of each tourist. Through analysis, it is found that in the lake area of the scenic area (with a relatively high water view rate), most tourists will choose to stay and enjoy the scenery (behavior category is standing); while on the main roads connecting various scenic spots, tourists mainly walk and ride (there is a bicycle rental service in the scenic area). Based on these analysis results, the scenic area management department can optimize the tour route signs, add rest facilities and service points in areas where tourists stay more, and improve the tourist experience.
[0042] In this application, by using mobile devices, street view data can be flexibly and comprehensively collected. By aligning the collected video stream with GPS data and preset geographical spatial information, accurate positioning can be achieved, improving the accuracy of pedestrian location information. Then, by combining the multi-dimensional features of pedestrians for pedestrian tracking and re-identification, pedestrian misjudgment and quantity statistics errors can be reduced. Finally, based on the behavior and quantity of pedestrians, the accuracy of analyzing street view information is higher.
[0043] As an optional implementation, in step 202, aligning the video stream with the GPS data and the preset geographic space information in sequence includes the following contents: Step S11: converting the video stream into continuous image frames, wherein each image frame carries a timestamp of video shooting; Step S12: using the timestamp in the image frame as a reference, aligning each image frame with the GPS data to determine the shooting time of the image frame and the corresponding geographical location; Step S13: determining at least three identical, non-collinear feature points of the image frame and the geographic space information, and recording the pixel coordinates of each feature point in the image frame and the geographic coordinates in the geographic space information; Step S14: determining conversion parameters from pixel coordinates to geographic coordinates based on at least three feature points; Step S15: Map each image frame into the geographic space of the geographic space information according to the conversion parameters.
[0044] Wherein, step S11 includes the following contents.
[0045] Removing invalid video segments: During the street view data collection process, the video stream obtained by the mobile device may contain a variety of invalid video segments, such as black screen images caused by equipment failure, occlusion or other factors, or some special images that do not meet the analysis requirements. In this case, it is necessary to edit and remove the invalid video segments to avoid interference of invalid data on subsequent analysis and improve the efficiency and accuracy of data processing.
[0046] Image frame extraction: After completing the video editing, a special video processing algorithm is used to convert the processed video into continuous image frames frame by frame in chronological order. During the conversion process, ensure that each frame of the image is accurately accompanied by the timestamp when it was captured. This timestamp records the specific moment when the image frame was taken and is an important identifier for achieving precise alignment with GPS data. By attaching a timestamp to each frame of the image, the image frames can be matched one by one with the GPS data based on the key dimension of time, thereby accurately associating the street view image with the geographic location information when it was taken.
[0047] GPS data smoothing: Since GPS signals are easily interfered by various factors during transmission, such as building obstruction, signal reflection, etc., the acquired GPS data contains noise and abnormal points. By using the Kalman smoothing algorithm to correct GPS noise and abnormal points and generate a smooth GPS trajectory, the GPS data can more truly reflect the actual movement path of the mobile device.
[0048] Among them, step S12 includes the following contents.
[0049] GPS Interpolation: In street view data processing, the acquisition frequency of video frames is usually high, while the acquisition frequency of GPS data is relatively low, which leads to the problem of asynchronous time between the two. To solve this problem, GPS interpolation operation is carried out. Based on the time stamps of video frames, linear interpolation is performed on the low-frequency GPS data for alignment. Specifically, according to adjacent GPS data points and their corresponding times, appropriate GPS data points are inserted within the time intervals of video frames through linear calculation, so that the GPS data is more matched with the video frames in time. GPS interpolation enables better synchronization of GPS data and video frames in time, and solves the problem of time inconsistency caused by different data acquisition frequencies.
[0050] Filtering Invalid Segments: In the collected street view data, in addition to invalid video segments in the video stream, the corresponding GPS data may also have invalid segments. To ensure the validity and accuracy of the data, the operation of filtering invalid segments is required. By identifying black screen images or special image features in the video, and at the same time removing the GPS data points corresponding to these invalid images. In addition, duplicate points and blurred images in the data are also filtered. Duplicate points may be caused by GPS signal fluctuations or device errors, and blurred images will affect the analysis of street view information. By removing these invalid data, the quality of the data and the accuracy of the analysis can be improved.
[0051] In step S13, the system supports users to upload open-source maps, CAD files or GIS files. The server extracts feature points from the image frames and geospatial information (such as high-precision maps). At least three identical and non-collinear feature points are selected through a matching algorithm. These feature points correspond to the same position in reality, and the pixel coordinates (with the upper left corner as the origin) of each feature point in the image frame and the geographical coordinates (such as longitude and latitude) in the geospatial are recorded.
[0052] In step S14, the server determines the conversion parameters from pixel coordinates to geographical coordinates based on the affine transformation model. Affine transformation can combine linear and translation transformations in two-dimensional space for coordinate mapping. By substituting the pixel and geographical coordinates of at least three feature points into the matrix formula of affine transformation, an overdetermined system of equations is constructed, and then algorithms such as the least squares method are used to solve it, obtaining the conversion parameters that minimize the sum of the squares of the errors of the system of equations. These parameters accurately describe the conversion relationship between the two types of coordinates.
[0053] In step S15, each pixel of the image frame is traversed, and the pixel coordinates are substituted into the previously determined affine transformation formula to calculate the corresponding geographical coordinates of the pixel in the geospatial. In this way, all pixels of the entire image frame are converted, so that elements such as pedestrians and buildings in the image frame can be accurately corresponding to the real geographical space position, laying a geographical positioning foundation for subsequent street view information analysis.
[0054] The above deletion of invalid video segments and selection of feature points support the user's interactive operations, and the user can adjust and optimize according to the actual situation.
[0055] In this application, the system supports the user to upload a geospatial file. Through a series of complex and precise coordinate conversion algorithms, the GPS data points in the video stream are automatically matched with the geospatial data. This matching process is not completely automated and fully supports the user's interactive operations, realizing the accurate mapping of street view information to the real-world geospatial, providing a data basis for subsequent analysis of street view information containing accurate geospatial data.
[0056] As an alternative implementation, in step 203, for the video stream after two alignment processes, a pedestrian multi-dimensional feature fusion method is used for pedestrian tracking and re-identification, and the attribute information of each pedestrian is obtained, including the following content: Step S21: When the target pedestrian in the current image frame is detected, the multi-dimensional features of the target pedestrian are hierarchically matched with the pedestrian feature set in the database; Step S22: If the hierarchical feature matching fails, a pedestrian ID is assigned to the target pedestrian in the database, and the association relationship between the pedestrian ID, the current image frame ID, and the multi-dimensional features is stored in the database; Step S23: If each hierarchical feature matches successfully, determine the pedestrian ID of the target pedestrian in the database, and update the database according to the current image frame ID and the multi-dimensional features. Repeat the above update process until the target pedestrian no longer appears in the image frame.
[0057] After completing the two alignment processes of the video stream with the GPS data and the preset geospatial information, it enters the pedestrian tracking and re-identification stage. In this stage, with the help of the pedestrian multi-dimensional feature fusion technology, the attribute information of each pedestrian is accurately obtained, covering key data such as the location and behavior category of the pedestrian at different times. The specific process is as follows.
[0058] When the video stream is played frame by frame and a target pedestrian (a random pedestrian) is detected in the current image frame, the server starts a hierarchical matching process. The server comprehensively extracts the multi-dimensional features of the target pedestrian, including biometric features (such as gender, age group), overall features (skeletal features, clothing features), and accessory features (carried item features, behavior category, accompanying group).
[0059] The server performs multi-level matching of these multi-dimensional features with the set of pedestrian features already stored in the database. First, the matching starts from the biometric level. If the gender and age range of the target pedestrian do not match the corresponding biometric features in the set of pedestrian features in the database, it is determined that the feature matching at this level fails. Once a situation of hierarchical feature matching failure occurs, the system assigns a unique pedestrian ID to the target pedestrian in the database to ensure its recognizability throughout the analysis process. At the same time, the association relationship between this pedestrian ID, the current image frame ID, and the multi-dimensional features of the target pedestrian is completely stored in the database for subsequent query and analysis. The structure of the human feature database is shown in Table 1.
[0060] Table 1
[0061] If the biometric features of the target pedestrian match successfully with the set of pedestrian features in the database, it enters the next level, i.e., the overall feature matching. The skeletal features and clothing features of the target pedestrian are compared with the corresponding features in the first feature set determined by the successful biometric feature matching in the database. If the matching fails at this level, a new ID is also assigned to the target pedestrian and the relevant information is stored according to the above process. If the overall feature matching is successful, a second feature set composed of the pedestrian IDs that match successfully in the first feature set is further determined, and then the accessory feature matching is carried out. The accessory features such as the items carried, behavior categories, and accompanying groups of the target pedestrian are carefully compared with the corresponding features in the second feature set.
[0062] Only when the feature matching at each level is successful can the system determine the corresponding pedestrian ID of the target pedestrian in the database. Subsequently, the system updates the database based on the current image frame ID and the latest multi-dimensional features of the target pedestrian. The updated content includes the position information of the target pedestrian in the current image frame (determined by aligning the image frame with the geospatial information), the change in behavior categories, etc. Since the target pedestrian may be at a relatively far position when first photographed, but the last time the pedestrian is photographed is the closest point to the target pedestrian, the above update process will be repeated until the target pedestrian no longer appears in subsequent image frames, so as to completely and accurately record all the action trajectories and attribute information changes of the target pedestrian in the video stream.
[0063] By updating the database according to the current image frame ID and the latest multi-dimensional features, this application can dynamically record the latest position and behavior status of pedestrians. As pedestrians move and their behaviors change in the street scene, the database is continuously updated to ensure that no matter when pedestrians appear in different image frames, the system can accurately track based on the latest data, avoiding tracking interruption or errors caused by information lag, and achieving accurate recording of the entire action trajectory of pedestrians. In addition, the multi-level feature matching increases the accuracy of pedestrian recognition.
[0064] Each feature of pedestrians is described below.
[0065] 1. Biometric features: including gender and age group. Exemplarily, the division of age groups is shown in Table 2.
[0066] Table 2
[0067] 2. Overall features: including skeletal features and clothing features.
[0068] (1) Skeletal features.
[0069] Use Yolo-Pose to obtain human key points (such as shoulders, elbows, knees, ankles, etc.), calculate the ratio values of each feature, normalize them to the [0, 1] interval (eliminating the influence of height differences), and construct a feature vector .
[0070] , where represents the shoulder width to height ratio, represents the shoulder width, represents the height from the midpoint of the shoulders to the top of the head.
[0071] , where represents the thigh to calf ratio, represents the length from the hip to the knee, represents the length from the knee to the ankle.
[0072] , where represents the arm length to torso ratio, represents the arm length (from the shoulder to the wrist), represents the torso height (from the neck to the hip).
[0073] (2) Clothing features.
[0074] According to the human key points obtained by Yolo-Pose, distinguish the head, upper body, and lower body, segment the person image, calculate the LAB color histogram of each part respectively, and normalize it to a probability distribution, finally obtaining 3 sub-block feature vectors .
[0075] Among them, , , are the color features of the head, upper body, and lower body respectively.
[0076] 3. Accessories features: including items carried by the target pedestrian, behavior categories, and accompanying groups.
[0077] (1) Items carried. Items carried include, but are not limited to: accessories (such as backpacks, hats, etc.), bicycles, skateboards, and pets.
[0078] a. For accessories, features can be extracted in the following way: First, use a segmentation model (such as the Yolov8-seg model) to analyze street view images, detect accessory regions such as backpacks and hats, and generate a pixel-level mask for each accessory. The mask is used to clearly separate the accessory from the complex background. Then, use a human pose estimation algorithm to obtain the positions of human key points. For example, confirm that the backpack region should be within the back region defined by the human key points, and the hat region should overlap with the head position of the human key points to verify the accuracy of accessory detection. Finally, convert the accessory region to the LAB color space, use the K-means algorithm to cluster the pixels, select the cluster center with a proportion greater than a certain ratio (such as 30%) and a color difference ΔE greater than a set difference (such as ΔE>15) from the clustering results, and record this cluster center as the dominant color in LAB format. Eventually, form features including the type of accessory, the location of the accessory, and the dominant color.
[0079] b. For bicycles, skateboards, and pets, features can be extracted in the following way: First, use a detection model (such as the Yolo-v7 model) to scan street view images, detect the associated items of pedestrians (such as bicycles, scooters) or associated animals (such as cats, dogs), accurately identify their positions in the image, and generate corresponding detection boxes. Then, calculate whether the intersection over union (IoU) of the item detection box and the region of the human key points obtained by the human pose estimation algorithm is greater than a set threshold (such as 0.4) to ensure the accuracy and relevance of the detection. For example, judge whether the bicycle pedals are close to the human key points of the feet, and whether the pet leash is connected to the human key points of the hand. After the IoU meets the condition, output the type of the item or animal, determine the binding position (such as handheld or carried on the back) according to its relative position to the human body, and extract the dominant color of the item according to the K-means algorithm. Combine this information to form features including the type of item, the binding position, and the dominant color.
[0080] (2) Behavior categories. Such as Figure 3As shown in the figure, a street view image is input into the system. First, a detection module (such as the Yolo-v7 model) is used to detect the target and generate a detection frame containing pedestrians. Then, a human key point detection model (such as Yolo-Pose) is used to estimate the posture. Yolo-Pose determines the posture information of pedestrians by analyzing the human key points of pedestrians. Then, the activity database trained based on the image data set is used to compare and match the posture information of pedestrians with the posture features of various activities such as walking, running, standing, cycling, and sitting in the activity database. If the posture information of the pedestrian matches the posture features in the activity database, the activity type of the pedestrian is directly determined; if not, the posture feature is recorded in the single-person behavior list. Then, according to the proportion of each posture in the single-person behavior list, the posture with a higher probability is selected as the behavior category result. If the proportions of each posture are consistent, the last detected posture is selected as the final behavior category result.
[0081] (3) Companion groups. The methods for determining companion groups include at least one of the following three methods.
[0082] a. If the distance between the target pedestrian and at least one adjacent pedestrian is less than the set distance threshold in multiple consecutive image frames, it is determined that the target pedestrian and the adjacent pedestrian form a companion group. In other words, if the distance between pedestrians is close, such as holding hands or crossing arms, and this relative position relationship is maintained in multiple frames of images, it can be inferred that there is a relationship between these pedestrians.
[0083] b. If the target pedestrian and at least one adjacent pedestrian have the same motion trajectory, posture changes, or communication interactions in multiple consecutive image frames, it is determined that the target pedestrian and the adjacent pedestrian form a companion group. Team members may have a certain degree of synchronization in walking, standing, and other behaviors. For example, team members may stop together, turn together, etc., and their behaviors can be determined by analyzing the motion trajectory and posture changes of pedestrians in multiple frames of images. Or observe whether there are signs of communication and interaction between pedestrians, such as eye contact, gesture interaction, etc.
[0084] c. If the object carried by the target pedestrian complements the object carried by at least one neighboring pedestrian, the target pedestrian and the neighboring pedestrian are determined to form a companion group. Although some objects are not in the hands of the same person, they have a complementary relationship of function and can also be considered as common objects. For example, one person holds a camera and the other holds a tripod. In this case, it can be judged that they may be taking photos together, and these objects are considered common objects.
[0085] In addition, the relative position and distance between the carried object and the pedestrians can also be observed. If the distance between a carried object and multiple pedestrians is relatively close, and this relative position relationship is maintained in multiple frames of images, then it can be speculated that these pedestrians may be associated with the carried object. For example, if a large suitcase is detected, and the positions of two pedestrians are both close to the suitcase, and their movement trajectories are consistent with the movement trajectory of the suitcase, then it can be speculated that these two pedestrians may be carrying the suitcase together.
[0086] The feature matching of pedestrians is described below.
[0087] 1. Biometric matching. The gender matching must be exactly the same, and the age group can float up or down by one age group deviation. If the matching is successful, then the overall feature matching is entered; if the matching fails, then a pedestrian ID is configured for the target pedestrian in the database.
[0088] 2. Overall feature matching. The overall feature matching process is as follows: Match the skeletal features of the target pedestrian with the skeletal features of the pedestrians to be matched in the first feature set to determine the skeletal similarity; match the clothing features of the target pedestrian with the clothing features of the pedestrians to be matched to determine the clothing similarity; if the weighted sum value of the skeletal similarity and the clothing similarity is greater than the set threshold, then the overall feature matching is determined to be successful.
[0089] (1) Calculate the skeletal feature set of the target pedestrian and the skeletal feature set of the pedestrians to be matched in the database The weighted Euclidean distance, and the formula for the weighted Euclidean distance is shown below.
[0090] , where represents the weighted Euclidean distance, represents the weight of the i-th skeletal feature, represents the i-th skeletal feature of the target pedestrian, represents the i-th skeletal feature of the pedestrian to be matched. Exemplarily, .
[0091] The formula for skeletal feature similarity is: .
[0092] (2) Use the Bhattacharyya distance to calculate the similarity between the color feature set of the target pedestrian and the color feature set of the pedestrians to be matched in the database. The calculation formula for the color feature similarity is shown below.
[0093] , where represents the color feature similarity (clothing similarity), represents the color probability value of the target pedestrian in the i-th interval. represents the color probability value of the pedestrian to be matched in the i-th interval, and n represents the total number of color intervals.
[0094] (3)Weight the similarity of skeletal features and the similarity of color features to obtain the comprehensive similarity D. The formula for the comprehensive similarity is as follows: , where is the weight of the skeletal feature, is the weight of the color feature.
[0095] Exemplarily, if D < 0.75, enter the matching of accessory features. If D ≥ 0.75, the matching fails, and a pedestrian ID is configured for the target pedestrian in the database.
[0096] 3. Matching of accessory features.
[0097] The matching process of the accessory features is as follows: If the item type of the item carried by the target pedestrian is the same as that of the item carried by the pedestrian to be matched in the second feature set, the item binding position is the same, and the color difference value of the main color of the item is within the set color difference range, then a preset item score is configured for the item of the target pedestrian. Among them, each successfully matched item corresponds to an item score; if the behavior category of the target pedestrian is the same as that of the pedestrian to be matched, then a preset behavior score is configured for the behavior category of the target pedestrian; if the group ID of the companion group where the target pedestrian is located is the same as the group ID of the companion group where the pedestrian to be matched is located, then a preset group score is configured for the companion group of the target pedestrian; perform a weighted sum of the total item score, behavior score, and group score. If the weighted result exceeds the set score, it is determined that the accessory feature matching is successful.
[0098] Exemplarily, the target pedestrian A is carrying two items, a hat and a backpack, and the pedestrian B to be matched in the second feature set is also carrying similar items. After inspection, the color difference value of the hat is 10 (meeting ΔE < 15), and the binding positions are both on the head. The color difference value of the backpack is 12, and the binding positions are both on the back. According to the rules, if there are 2 carried items, getting one match increases the score by 0.5 points, and a complete match gets 1 point. Here, both the hat and the backpack are successfully matched, so 1 point is obtained for the carried items. The target pedestrian A is walking at this moment, and the pedestrian B to be matched also has the behavior of walking in the behavior list, so the behavior pattern is successfully matched and 1 point is recorded. The group ID of the companion group where the target pedestrian A is located is G1, and the group ID of the companion group where the pedestrian B to be matched is located is also G1, so the group relationship is successfully matched and 1 point is recorded. Finally, the carried items, behavior, and group relationship are weighted and scored according to 0.5, 0.3, and 0.2. Since the carried items get 1 point, the behavior pattern gets 1 point, and the group relationship gets 1 point. The weighted calculation is: 1 * 0.5 + 1 * 0.3 + 1 * 0.2 = 1. Because 1 exceeds the set 0.75 points, the accessory feature matching is successful.
[0099] In this application, through multi-level matching of biometric features, overall features (bones, clothing, etc.), and accessory features (carried items, behavior patterns, group relationships, etc.), the pedestrian identity is gradually refined and accurately located. This progressive method can avoid misidentifications caused by misjudgments of single features and improve the accuracy of pedestrian identity recognition. Bone features can reflect the structural and postural information of the human body, clothing features provide significant appearance identifiers, and accessory features further enrich the description of pedestrian features, enabling the system to identify and track pedestrians from multiple perspectives and improving the recognition accuracy both overall and in detail.
[0100] As an optional implementation manner, analyzing the street view information according to the attribute information of multiple pedestrians includes: performing semantic segmentation on the image frames in the video stream to obtain the spatial element distribution information in the image frames, where the spatial element distribution information includes natural landscape distribution information and public facility distribution information; analyzing the spatial element distribution information and the attribute information of the pedestrians to determine the association between the spatial element distribution and the behavior categories and quantities of the pedestrians.
[0101] Analyzing the street view information can obtain at least one of the following analysis results: a crowd distribution heat map, a behavior distribution report, and the association between spatial elements and pedestrians.
[0102] Figure 4It is a schematic diagram of a population distribution heat map. The person re-identification technology can accurately record the information of each pedestrian at different times and locations, and these location information are the basic data for generating the heat map. The heat map represents the frequency and density of pedestrian appearances through the shade of colors. The accurate location information provided by person re-identification enables the heat map to more precisely reflect the flow distribution of people. For example, in a large exhibition, by obtaining the movement trajectories of each exhibitor through person re-identification technology and then generating a heat map, it is possible to visually see which areas have a large flow of people and which areas are relatively deserted. Combining with the time series information of person re-identification, the changing trend of the heat map over time can also be analyzed. For instance, the changing situation of the flow of people at the entrances of different stores in a shopping mall at different time periods of a day helps merchants reasonably arrange business hours and staff scheduling.
[0103] Figure 5 It is a schematic diagram of a behavior distribution report. During the person re-identification process, the behaviors of pedestrians will be detected and classified (such as walking, running, standing, cycling, etc.), and this behavior information is an important basis for generating the behavior distribution report. By statistically analyzing the behaviors of different pedestrians, the behavior distribution in different regions and different time periods can be obtained. For example, in a park, by counting the proportion of the number of people running, walking, and resting in different time periods to form a behavior distribution report, it provides a reference for the park management department to optimize the facility layout and activity arrangements. Person re-identification can also track the behavior changes of individual pedestrians. Summarizing the individual behavior data into the behavior distribution report can more comprehensively understand the crowd behavior pattern. For example, analyzing the shopping behavior of a certain customer in a shopping mall, including the staying time and purchase behavior in different stores during the whole process from entering the mall to leaving, provides data support for merchants to conduct precise marketing.
[0104] The server performs semantic segmentation on the image frames in the video stream and can obtain the distribution information of spatial elements in the image frames. The distribution information of spatial elements includes the distribution information of natural landscapes and public facilities. The distribution information of natural landscapes includes, but is not limited to: sky view ratio, water view ratio, and green view ratio.
[0105] The sky view ratio refers to the pixel proportion of the spatial element of the sky in the image. The water view ratio refers to the pixel proportion of the spatial element of water area in the image. As Figure 6 shown, the green view ratio refers to the pixel proportion of the spatial element of green plants in the image. And trash cans, benches, etc. belong to public facilities. The distribution information of public facilities refers to the pixel proportion of public facilities such as trash cans and benches in the image. Figure 7It is a schematic diagram of the bench distribution. If the video stream is a panoramic video, the video frame needs to be cut into four orientations: east, west, south, and north. Calculate the pixel proportion of each spatial element in the image respectively, and finally take the average value of the four images as the index value of the natural landscape distribution information. The public facility distribution information is obtained by accumulating the number of public facilities in the four images.
[0106] The server analyzes the spatial element distribution information and the pedestrian attribute information to determine the correlation between the spatial element distribution and the pedestrian behavior categories and quantities, which provides a method for studying the relationship between elements such as the sky view rate, water view rate, and trash cans and pedestrians. For example, it can be analyzed whether pedestrians are more inclined to behaviors such as staying and resting in areas with a high sky view rate; whether the pedestrian behavior patterns (such as walking, viewing, etc.) are different in places with a high water view rate; and the characteristics of the number of pedestrians gathering and behaviors (such as discarding garbage, staying briefly, etc.) in areas where trash cans are densely distributed.
[0107] The server can also generate a behavior-facility association matrix, a relationship diagram between the change in element proportion and the change in behavior distribution, etc. Among them, the behavior-facility association matrix is to count the quantities of different behaviors of pedestrians around public facilities, and by calculating the Pearson correlation coefficient, output the association matrix, as Figure 8 shown. The relationship diagram between the change in element proportion and the change in behavior distribution refers to a double-Y axis chart. Among them, the horizontal axis represents time or spatial position, the left vertical axis represents the change in the green view rate, and the right vertical axis represents the change in the number of standing pedestrians. Through such a double-Y axis chart, it can be intuitively observed whether there is a certain correlation between the change in the proportion of spatial elements and the change in pedestrian behavior distribution. For example, whether the number of standing pedestrians is relatively large in areas with a high green view rate; or whether the number of standing pedestrians will fluctuate correspondingly during a certain time period when the green view rate changes in a day. Through this kind of analysis, it can help researchers or relevant personnel understand the influence of environmental factors (spatial elements) on pedestrian behavior, so as to provide a reference basis for urban planning, landscape design, etc.
[0108] The embodiment of this application provides an overall process step for street view information analysis, including the following content.
[0109] 1. Data collection.
[0110] Video stream collection: Use mobile devices equipped with high-definition cameras, such as smartphones, panoramic cameras, etc., to continuously collect street view video streams at a resolution of not less than 1080P and a frame rate of 30fps during the movement process, and comprehensively record pedestrian and surrounding environment information.
[0111] GPS data collection: Synchronously obtain accurate longitude, latitude, time, ground speed, etc. information through the device's built-in GPS module or an external high-precision GPS device, and use it to mark the geographical location corresponding to each frame of the video stream.
[0112] 2. Video alignment.
[0113] Align the video with GPS data: Convert the video stream into continuous image frames and attach timestamps to each frame. Based on the timestamp of the image frame, perform linear interpolation on the low-frequency GPS data to ensure accurate temporal matching between the video frames and GPS location information, providing a basis for subsequent geolocation.
[0114] Align with geospatial information: Identify at least three identical and non-collinear feature points in the image frame and preset geospatial information (such as high-precision maps, GIS data), and record the pixel coordinates of the feature points in the image frame and the geographic coordinates in the geospatial. Through the affine transformation model, calculate the transformation parameters based on the feature point coordinates to accurately map the image frame to the real-world geospatial.
[0115] 3. Object detection.
[0116] Use the detection model to perform object detection on each frame of the video stream, identify the items carried by pedestrians (such as bicycles, scooters, etc.) and pets (such as cats, dogs, etc.), and generate detection boxes.
[0117] Use the segmentation model to detect areas of accessories such as backpacks and hats and generate pixel-level masks.
[0118] 4. Feature extraction.
[0119] Gender feature: Use a dual-branch MobileNetV3 structure to process face and full-body images, and output three classifications: male, female, and unknown.
[0120] Age group feature: Use a soft segmentation regression network to output the age range with the highest probability.
[0121] Skeletal feature: Obtain the human key point information of pedestrians through the human key point detection model for pose estimation and skeletal feature extraction. Clothing feature: Extract the color feature of the clothing from the overall appearance of the pedestrian.
[0122] Feature of carried items: For the detected carried items, extract features such as their type, binding position (such as handheld, carried on the back), and main color (select the dominant color by performing K-means clustering on the accessory area in the LAB color space, K = 3, and select the cluster center with a proportion > 30% and color difference ΔE > 15 as the dominant color).
[0123] Behavior feature: Based on the pose information obtained from the human key point detection model, compare it with the activity database to determine the behavior pattern of the pedestrian (walking, running, standing, cycling, sitting, etc.).
[0124] Group relationship feature: Determine the group ID of the group the pedestrian belongs to and analyze their group relationship.
[0125] 5. Multi-level matching.
[0126] Biometric matching: Match the biometric features such as the gender and age range of the target pedestrian with the set of pedestrian features in the database. If the match fails, assign a new pedestrian ID to the target pedestrian and store the relevant information (pedestrian ID, image frame ID, multi-dimensional features).
[0127] Overall feature matching: If the biometric matching is successful, further match the skeletal features and clothing features with the corresponding features in the database. If the match fails, also assign a new ID and store the information.
[0128] Ancillary feature matching: If the overall feature matching is successful, perform a comprehensive feature weighted matching on the carried items, behavior patterns, and group relationships. For carried items, check the color difference ΔE < 15 and position consistency, and add points according to the number of carried items; match the behavior patterns by type, and record 1 point for the existence of the same type of behavior; match the group relationship by group ID, and record 1 point for a successful match. Finally, perform a weighted score according to 0.5 (carried items), 0.3 (behavior), and 0.2 (group relationship). If the score exceeds 0.75, the match is successful. If the match fails, record the relevant information in the database.
[0129] 6. Result output and update.
[0130] If the multi-level matching is successful, determine the pedestrian ID of the target pedestrian in the database and update the database according to the current image frame ID and the latest multi-dimensional features, including information such as position and behavior changes.
[0131] If the match fails, record the relevant information in the database.
[0132] 7. Analyze the street view information in combination with the position changes and behavior changes of the pedestrians.
[0133] Based on the same technical concept, this application provides a street view information analysis device, as Figure 9 shown, the device includes: An acquisition module 901, configured to synchronously acquire the video stream and GPS data of the street view information through a mobile device, where the GPS data is used to record the geographical location information corresponding to the street view information in the video stream; An alignment module 902, configured to align the video stream with the GPS data and the preset geographical space information in sequence to map the geographical location in the street view information to the real-world geographical space; The tracking and re-identification module 903 is used to perform pedestrian tracking and re-identification on the video stream after two alignment processes by using the method of pedestrian multi-dimensional feature fusion, and obtain the attribute information of each pedestrian, where the attribute information includes the location and behavior category of the pedestrian at different times; The analysis module 904 is used to analyze the street view information according to the attribute information of multiple pedestrians.
[0134] Optionally, the alignment module 902 is used to: Convert the video stream into continuous image frames, where each image frame carries the time stamp of the video shooting; Based on the time stamp in the image frame, align each image frame with the GPS data to determine the shooting time and corresponding geographical location of the image frame; Determine at least three identical and non-collinear feature points between the image frame and the geospatial information, and record the pixel coordinates of each feature point in the image frame and the geographical coordinates in the geospatial information; Determine the conversion parameters from pixel coordinates to geographical coordinates according to at least three feature points; Map each image frame into the geospatial space of the geospatial information according to the conversion parameters.
[0135] Optionally, the tracking and re-identification module 903 is used to: In the case of detecting the target pedestrian in the current image frame, perform multi-level matching of the multi-dimensional features of the target pedestrian with the pedestrian feature set in the database; If the hierarchical feature matching fails, assign a pedestrian ID to the target pedestrian in the database, and store the association relationship between the pedestrian ID, the current image frame ID, and the multi-dimensional features in the database; If each hierarchical feature matches successfully, determine the pedestrian ID of the target pedestrian in the database, and update the database according to the current image frame ID and the multi-dimensional features. Repeat the above update process until the target pedestrian no longer appears in the image frame.
[0136] Optionally, the tracking and re-identification module 903 is specifically used to: Match the biometric features of the target pedestrian with the pedestrian feature set in the database, where the biometric features include gender and age group; If the biometric features match successfully, determine the first feature set composed of the pedestrian IDs that match successfully in the database, and match the overall features of the target pedestrian with the first feature set, where the overall features include skeletal features and clothing features; If the overall feature matching is successful, determine the second feature set composed of the pedestrian IDs that match successfully in the first feature set, and match the accessory features of the target pedestrian with the second feature set, where the accessory features include the items carried by the target pedestrian, the behavior category, and the accompanying group; If the accessory feature matching is successful, it is determined that each level of feature matching is successful.
[0137] Optionally, the tracking and re-identification module 903 is specifically configured to: If the type of the item carried by the target pedestrian is the same as the type of the item carried by the pedestrian to be matched in the second feature set, the item binding position is the same, and the color difference value of the main color of the item is within the set color difference range, configure a preset item score for the item of the target pedestrian, where each successfully matched item corresponds to an item score; If the behavior category of the target pedestrian is the same as the behavior category of the pedestrian to be matched, configure a preset behavior score for the behavior category of the target pedestrian; If the group ID of the accompanying group where the target pedestrian is located is the same as the group ID of the accompanying group where the pedestrian to be matched is located, configure a preset group score for the accompanying group of the target pedestrian; Perform a weighted sum of the total item score, the behavior score, and the group score. If the weighted result exceeds the set score, it is determined that the accessory feature matching is successful.
[0138] Optionally, the tracking and re-identification module 903 is specifically configured to: If the distance between the target pedestrian and at least one adjacent pedestrian is less than the set distance threshold in multiple consecutive image frames, determine that the target pedestrian and the adjacent pedestrian form an accompanying group; or, If the target pedestrian and at least one adjacent pedestrian have the same motion trajectory, pose change, or have communication and interaction in multiple consecutive image frames, determine that the target pedestrian and the adjacent pedestrian form an accompanying group; or, If the carried item of the target pedestrian is functionally complementary to the carried item of at least one adjacent pedestrian, determine that the target pedestrian and the adjacent pedestrian form an accompanying group.
[0139] Optionally, the analysis module 904 is used to: Perform semantic segmentation on the image frames in the video stream to obtain the spatial element distribution information in the image frames, where the spatial element distribution information includes the natural landscape distribution information and the public facility distribution information; Analyze the spatial element distribution information and the attribute information of the pedestrians to determine the association between the spatial element distribution and the behavior category and quantity of the pedestrians.
[0140] Such as Figure 10As shown in the figure, an embodiment of the present application provides an electronic device, including a processor 1001, a communication interface 1002, a memory 1003, and a communication bus 1004. Among them, the processor 1001, the communication interface 1002, and the memory 1003 complete mutual communication through the communication bus 1004.
[0141] The memory 1003 is used to store computer programs.
[0142] In an embodiment of the present application, when the processor 1001 is used to execute the program stored on the memory 1003, it implements the street view information analysis method provided in any of the foregoing method embodiments.
[0143] An embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the street view information analysis method provided in any of the foregoing method embodiments.
[0144] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0145] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0146] It should be understood that the terms used herein are for the purpose of describing particular example embodiments only and are not intended to be limiting. Unless the context clearly dictates otherwise, the singular forms "a", "an", and "the" as used herein may also include the plural forms. The terms "comprising", "including", "containing", and "having" are inclusive and thus specify the presence of stated features, steps, operations, elements, and / or components, but do not preclude the presence or addition of one or more other features, steps, operations, elements, components, and / or combinations thereof. The method steps, processes, and operations described herein are not to be construed as necessarily requiring them to be performed in the particular order described or illustrated, unless an execution order is explicitly stated. It should also be understood that additional or alternative steps may be used.
[0147] The above are only specific embodiments of the present application, enabling those skilled in the art to understand or implement the present application. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application will not be limited to these embodiments shown herein, but rather to the broadest scope consistent with the principles and novel features claimed herein.
Claims
1. A street view information analysis method, characterized in that, The method includes: Dynamically collecting a video stream of street view information and GPS data through a mobile device, where the GPS data is used to record the geographical location information corresponding to the street view information in the video stream; Aligning the video stream with the GPS data and preset geographical space information in sequence to map the geographical location in the street view information to the geographical space in the real world; For the video stream after two alignment processes, adopting a pedestrian multi-dimensional feature fusion method for pedestrian tracking and re-identification to obtain the attribute information of each pedestrian, where the attribute information includes the location and behavior category of the pedestrian at different times; Analyzing the street view information according to the attribute information of multiple pedestrians.
2. The method according to claim 1, characterized in that, Aligning the video stream with the GPS data and preset geographical space information in sequence includes: Converting the video stream into continuous image frames, where each image frame carries a time stamp of video shooting; Based on the time stamp in the image frame, aligning each image frame with the GPS data to determine the shooting moment and corresponding geographical location of the image frame; Determining at least three identical and non-collinear feature points between the image frame and the geographical space information, and recording the pixel coordinates of each feature point in the image frame and the geographical coordinates in the geographical space information; Determining the conversion parameters for converting from pixel coordinates to geographical coordinates according to at least three of the feature points; Mapping each image frame to the geographical space of the geographical space information according to the conversion parameters.
3. The method according to claim 1, characterized in that, Adopting a pedestrian multi-dimensional feature fusion method for pedestrian tracking and re-identification for the video stream after two alignment processes to obtain the attribute information of each pedestrian includes: When a target pedestrian in the current image frame is detected, performing multi-level matching of the multi-dimensional features of the target pedestrian with the pedestrian feature set in the database; If the hierarchical feature matching fails, assigning a pedestrian ID to the target pedestrian in the database and storing the association relationship between the pedestrian ID, the current image frame ID, and the multi-dimensional features in the database; If each hierarchical feature matches successfully, determining the pedestrian ID of the target pedestrian in the database and updating the database according to the current image frame ID and the multi-dimensional features, and repeating the above update process until the target pedestrian no longer appears in the image frame.
4. The method according to claim 3, characterized in that, Determining that each hierarchical feature matches successfully includes: Matching the biometric features of the target pedestrian with the pedestrian feature set in the database, where the biometric features include gender and age group; If the biometric features match successfully, determining the first feature set composed of the pedestrian IDs that match successfully in the database, and matching the overall features of the target pedestrian with the first feature set, where the overall features include skeletal features and clothing features; If the overall feature matching is successful, determine a second feature set composed of the pedestrian IDs that match successfully in the first feature set, and match the accessory features of the target pedestrian with the second feature set, where the accessory features include the items carried by the target pedestrian, the behavior category, and the companion group; If the accessory feature matching is successful, determine that each level of feature matching is successful.
5. The method according to claim 4, characterized in that, Determining that the accessory feature matching is successful includes: If the type of the item carried by the target pedestrian is the same as the type of the item carried by the pedestrian to be matched in the second feature set, the item binding position is the same, and the color difference value of the main color of the item is within the set color difference range, configure a preset item score for the item of the target pedestrian, where each successfully matched item corresponds to an item score; If the behavior category of the target pedestrian is the same as the behavior category of the pedestrian to be matched, configure a preset behavior score for the behavior category of the target pedestrian; If the group ID of the companion group where the target pedestrian is located is the same as the group ID of the companion group where the pedestrian to be matched is located, configure a preset group score for the companion group of the target pedestrian; Perform a weighted sum of the total item score, the behavior score, and the group score. If the weighted result exceeds the set score, determine that the accessory feature matching is successful.
6. The method according to claim 5, characterized in that, Determining the companion group where the target pedestrian is located includes: If the distance between the target pedestrian and at least one adjacent pedestrian is less than the set distance threshold in multiple consecutive image frames, determine that the target pedestrian and the adjacent pedestrian form a companion group; or, If the target pedestrian and at least one adjacent pedestrian have the same motion trajectory, pose change, or have communication and interaction in multiple consecutive image frames, determine that the target pedestrian and the adjacent pedestrian form a companion group; or, If the carried item of the target pedestrian is functionally complementary to the carried item of at least one adjacent pedestrian, determine that the target pedestrian and the adjacent pedestrian form a companion group.
7. The method according to claim 1, characterized in that, Analyzing the street view information according to the attribute information of multiple pedestrians includes: Perform semantic segmentation on the image frames in the video stream to obtain the spatial element distribution information in the image frames, where the spatial element distribution information includes natural landscape distribution information and public facility distribution information; Analyze the spatial element distribution information and the attribute information of the pedestrians to determine the association between the spatial element distribution and the behavior category and quantity of the pedestrians.
8. A street view information analysis device, characterized in that,The device includes: An acquisition module, configured to synchronously acquire the video stream of the street view information and the GPS data through a mobile device, where the GPS data is used to record the geographical location information corresponding to the street view information in the video stream; An alignment module, configured to align the video stream with the GPS data and the preset geographical space information in sequence to map the geographical location in the street view information to the geographical space in the real world; A tracking and re-identification module, which is used to perform pedestrian tracking and re-identification on the video stream after two alignment processes by means of pedestrian multi-dimensional feature fusion, and obtain the attribute information of each pedestrian, wherein the attribute information includes the location and behavior category of the pedestrian at different times; An analysis module, which is used to analyze the street view information according to the attribute information of multiple pedestrians.
9. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; The memory is used to store computer programs; The processor is used to implement the method according to any one of claims 1-7 when executing the program stored on the memory.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the method according to any one of claims 1-7 is implemented.
Citation Information
Patent Citations
Cross-video pedestrian positioning tracking method, system and device
CN111462200A
Pedestrian tracking method, system and device and storage medium
CN114663835A
Pedestrian identification and tracking method and device, readable storage medium and equipment
CN115019241A
Pedestrian tracking method and system based on geographic space information
CN115272949A