Traffic intersection flow image processing method based on biological perception time lag
By employing a bio-sensing time-lag method and utilizing masked images and multi-frame processing technology, the requirements for high-performance equipment and complex software in existing technologies have been addressed, enabling low-cost dynamic intersection video display without foreground and improving server computing efficiency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-02
- Publication Date
- 2026-04-14
AI Technical Summary
Existing dynamic intersection foreground removal methods require high-performance equipment and complex video processing software, leading to resource waste and increased costs.
By employing a bio-sensory time-delay-based method, and by capturing masked images and performing frame interpolation, this method uses a simple camera and multi-frame image processing technology to replace professional cameras and computationally intensive video processing software. It extracts and matches the contour features of motor vehicles, non-motor vehicles, and pedestrians to form a continuous dynamic intersection video without a foreground.
It reduces the requirements for equipment performance, reduces the amount of computation, enables real-time display of dynamic intersection video without foreground, reduces equipment and computing costs, and improves server computing efficiency.
Smart Images

Figure CN116597396B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of intelligent traffic control technology and relates to a traffic intersection flow image processing method based on biological perception time delay, which is used for post-control of scene synthesis in intelligent traffic. Background Technology
[0002] Urban intersections are the most complex traffic scenarios in urban roads, involving the most participants and experiencing the most frequent problems. They serve as nodes and hubs in the road traffic system, handling a large volume of traffic. In intelligent traffic management and control at urban intersections, there is an urgent need to present real-time images of the intersection—images showing the actual intersection but without any foreground vehicles, non-motorized vehicles, or pedestrians—on the central control platform of the command center and the display interface of the regional traffic control unit. This allows for the arbitrary placement of pre-set virtual vehicle and pedestrian scenarios from the intelligent traffic system into the intersection for simulated command and control, thereby achieving the goals of intelligent traffic management.
[0003] Current methods for dynamic intersection foreground removal directly remove foregrounds from high-definition videos captured from all four directions of the intersection using edge computing. This not only places high demands on the cameras themselves but also requires computationally intensive video processing software, posing a significant challenge to the capacity of existing servers. It may even necessitate replacing these servers with larger ones to achieve the desired foreground removal, resulting in unnecessary waste of human and material resources. Therefore, there is an urgent need for a technology that reduces the demands on specialized equipment, eliminates the need for complex video processing software, and avoids relying on massive amounts of video data for computation, thereby producing continuous, foreground-free dynamic intersection video. Summary of the Invention
[0004] The proposed method for processing traffic intersection flow images based on biosensing time delay involves removing images containing moving or stationary motor vehicles, non-motor vehicles, and pedestrians from images captured at a certain frequency. Then, frame interpolation is performed, and the resulting video file is equivalent to a continuous display of dynamic intersection video without motor vehicles, non-motor vehicles, or pedestrians. This method uses a simple, inexpensive camera instead of an expensive professional video camera and multi-frame image processing instead of computationally intensive video processing software, ultimately resulting in a visually continuous dynamic intersection video with foreground removed.
[0005] The technical solution of this invention is implemented as follows:
[0006] A method for processing traffic flow images at intersections based on biosensing time delay, including
[0007] S1: Using historical video footage of similar intersections as sample videos, extract the contour features of motor vehicles, non-motor vehicles, and pedestrians from the sample video images to form auxiliary feature sample domains N1, N2, and N3;
[0008] S2: Using the sample video frame as a reference, perform a occlusion operation on the intersection shooting device, and capture occluded images of the target intersection at an average speed of P frames / second. Extract the contour features of motor vehicles, non-motor vehicles, and pedestrians in the images to form target sample domains M1, M2, and M3.
[0009] S3: Create an expanded scaling calibration board;
[0010] S31: Create an expanded scale calibration board based on the same image size as the historical video single frame image of the intersection and the masked image taken at the target intersection;
[0011] S32: Extract sample video segments distributed at different locations in the video, each representing one of the categories: motor vehicles, non-motor vehicles, or pedestrians. Discretize the video segments into consecutive multi-frame images, each corresponding to a region division on an expanded scaling calibration board. Calculate the proportion of a target category displayed at positions 1 to r in the video frame images using the expanded scaling calibration board, which will be ζ1, ζ2, ..., ζ. r ;
[0012] S4: The feature sample domains N1, N2, N3 and the target sample domains M1, M2, M3 are trained and expanded according to the expansion ratio calibration board to obtain the target feature sample domains N1”, N2”, N3”;
[0013] S41: Based on the expansion scale calibration plate, map the display ratio of each category to the feature sample domain and the target sample domain. Expand each contour feature in the feature sample domains N1, N2, N3 and the target sample domains M1, M2, M3 according to the coordinate ratio. Each feature sample domain is divided into r regions according to the expansion scale calibration plate to obtain r feature sample subdomains; that is... , , ;
[0014] S42: Match the contour features in the target sample domains M1, M2, M3 and the feature sample domains N1, N2, N3 according to the expansion ratio calibration board. Place the contour features in the target sample domains M1, M2, M3 with a matching threshold β < 0.8 into the feature sample domains, and then expand and improve them according to step S41 to obtain the target feature sample domains N1”, N2”, N3”.
[0015] S43: For all images of the target intersection that are captured with the masked view, steps S41 and S42 must be performed;
[0016] Furthermore, the sample matching threshold , where a∈{N1'、N2'、N3'}, b∈{M1、M2、M3};
[0017] S5: The feature contours extracted from the captured occluded images are matched with the feature contours contained in the target feature sample domains N1”, N2”, and N3” according to the expanded scaling calibration plate. Occluded images with a matching threshold β ≥ 0.8 are discarded. The remaining occluded images per second are denoted as... ;
[0018] S6: Process the remaining masked images according to the time sequence using biosensing time-lag frame rate to obtain the frame rate F frames / second within the biosensing time-lag time, denoted as... Continuous playback generates a dynamic video file of the intersection with no motor vehicles, non-motor vehicles, or pedestrians.
[0019] Furthermore, step S6 includes, when m≥n, ;
[0020] When m < n
[0021] The working principle and beneficial effects of this invention are as follows:
[0022] The traffic intersection flow image processing method based on bio-sensing time delay uses a simple camera device to replace a professional camera to acquire video files. It also eliminates the need for large video processing software to perform high-speed operations and consume a lot of memory, thus reducing the server's computing load and lowering the requirements for backend equipment. This creates more room for the introduction of future new technologies without changing the cost of the original equipment.
[0023] The overall feature domains of motor vehicles, non-motor vehicles, and pedestrians are trained and expanded proportionally according to the expanded ratio calibration board, resulting in several times the number of features. However, when the server retrieves the feature domains for target feature comparison, it only needs to compare the feature subdomains corresponding to the proportion of the expanded ratio calibration board, which increases the server's computing efficiency to a certain extent.
[0024] By utilizing the principle of biosensing time delay, a smooth dynamic video biosensing experience is formed by superimposing and playing back frame images, achieving the goal of displaying the foreground of intersections without motor vehicles, non-motor vehicles, and pedestrians in real time. Attached Figure Description
[0025] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.
[0026] Figure 1 This is an overall flowchart of a traffic intersection flow image processing method based on biosensing time delay;
[0027] Figure 2 A schematic diagram of a calibration board for a traffic flow image processing method based on biosensing time delay;
[0028] Figure 3The image processing method for traffic flow at intersections based on biosensing time delay is shown in the image at different positions of the same vehicle on the calibration board.
[0029] Figure 4 The image processing method for traffic flow at intersections based on bio-sensing time delay is shown in the image at different positions on the calibration board for the same pedestrian.
[0030] Figure 5 A schematic diagram of a smeared image captured by an intersection camera device using a traffic flow image processing method based on biosensing time delay. Implementation
[0031] The technical solutions of the present invention will be clearly and completely described below with reference to the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.
[0032] As shown in the figure
[0033] A method for processing traffic flow images at intersections based on biosensing time delay, including
[0034] S1: Using historical videos of similar intersections with a 16:9 aspect ratio as sample videos, extract the contour features of motor vehicles, non-motor vehicles, and pedestrians from the sample video images to form auxiliary feature sample domains N1, N2, and N3;
[0035] S2: Perform a stencil operation on the camera with a default shooting aspect ratio of 4:3, and take a stenciled image of the target intersection at an average speed of P frames / second. Extract the contour features of motor vehicles, non-motor vehicles and pedestrians in the image to form the target sample domains M1, M2 and M3.
[0036] The camera's shooting area is masked to match the sample video's area, thus obtaining a masked image of the same size as the sample video without being limited by the camera's technical parameters;
[0037] S3: Create an expanded scaling calibration board;
[0038] S31: Create an expanded scale calibration board based on the same image area as the historical video single frame image of the intersection and the masked image captured at the target intersection, and divide the expanded scale calibration board into r regions;
[0039] S32: Extract sample video segments distributed at different locations in the video, each representing one of the categories: motor vehicles, non-motor vehicles, or pedestrians. Discretize the video segments into consecutive multi-frame images, each corresponding to a region division on an expanded scaling calibration board. Calculate the proportion of a target category displayed at positions 1 to r in the video frame images using the expanded scaling calibration board, which will be ζ1, ζ2, ..., ζ. r ;
[0040] S4: Train and expand the feature sample domains N1, N2, N3 and the target sample domains M1, M2, M3 to obtain the target feature sample domains N1”, N2”, N3”;
[0041] S41: Based on the expansion scale calibration plate, map the display ratio of each category to the feature sample domain and the target sample domain. Expand each contour feature in the feature sample domains N1, N2, N3 and the target sample domains M1, M2, M3 according to the coordinate ratio. Each feature sample domain is divided into r regions according to the expansion scale calibration plate to obtain r feature sample subdomains; that is... , , ;
[0042] S42: Perform sample matching on the contour features in the target sample domains M1, M2, M3 and the feature sample domains N1, N2, N3 according to the expanded scaling calibration board, and set the sample matching threshold. ,in, The contour features in the target sample domains M1, M2, and M3 with a matching threshold β < 0.8 are placed into the feature sample domains, and then expanded and improved according to step S41 to obtain the target feature sample domains N1”, N2”, and N3”.
[0043] S43: For all images of the target intersection that are covered by the camera, steps S41 and S42 must be performed to continuously optimize and improve the target feature sample domain.
[0044] To better explain step S4, as follows Figure 2 , Figure 3 , Figure 4 As shown, Figure 2 An example is provided: an expanded scale calibration board divided into 12 zones. In this embodiment, the intersection camera is positioned above zone 12 and pointing towards zone 1. Figure 3 Examples of the display effects of the same vehicle in different areas of the expanded scale calibration plate in step S32 are provided. By calculation, the display ratio of the vehicle in positions 1 to 12 of the expanded scale calibration plate in the same video frame image can be obtained, which are ζ1, ζ2, ... ζ. 12 ;like Figure 4The following are examples of the display effects of the same pedestrian in different areas of the expanded scale calibration plate in step S32. By calculation, the display ratio of the motor vehicle in the corresponding positions 1 to 12 of the expanded scale calibration plate in the same video frame image can be obtained, which are ζ1, ζ2, ... ζ. 12 Among them, motor vehicles 21∈ζ1N1∈N1, motor vehicles 22∈ζ2N1∈N1', motor vehicles 23∈ζ7N1∈N1', motor vehicles 24∈ζ12N1∈N1'; pedestrians 31∈ζ1N3∈N3', pedestrians 32∈ζ3N3∈N3', pedestrians 33∈ζ4N3∈N3', pedestrians 34∈ζ6N3∈N3', pedestrians 35∈ζ7N3∈N3', pedestrians 36∈ζ9N3∈N3';
[0045] like Figure 5 As shown, contour extraction is performed on the occluded image captured at the target intersection to obtain vehicle 1, vehicle 2, pedestrian 3, and pedestrian 4. Sample matching is then performed on the corresponding expanded scale calibration board for vehicle 1, vehicle 2, pedestrian 3, and pedestrian 4. This yields vehicle 1 located in calibration board region 3, vehicle 2 in calibration board region 7, pedestrian 3 in calibration board region 6, and pedestrian 4 in calibration board region 9. Based on the sample matching threshold calculation, β1=0.5<0.8, β2=0.9>0.8, β3=0.85>0.8, and β4=0.7<0.8. Therefore, vehicle 1 is placed in the feature sample domain ζ3N1∈N1', and pedestrian 4 is placed in the feature sample domain ζ9N3∈N3'. Then, the expanded scale calibration board maps vehicle 1 and pedestrian 4 to the feature sample subdomains of the other 11 expanded scale calibration board regions. After a period of learning and training, the target feature sample domains N1”, N2”, and N3 are obtained.
[0046] S5: The feature contours extracted from the captured occluded images are matched with the feature contours contained in the target feature sample domains N1”, N2”, and N3” according to the expanded scaling calibration plate. Occluded images with a matching threshold β ≥ 0.8 are discarded. The remaining occluded images per second are denoted as... ;
[0047] S6: Process the remaining masked images according to the time sequence using biosensing time-lag frame rate to obtain the frame rate F frames / second within the biosensing time-lag time, denoted as... Continuous playback generates a dynamic video file of the intersection with no motor vehicles, non-motor vehicles, or pedestrians.
[0048] Step S6 includes, when m≥n, When m < n,
[0049] It is worth noting that, such as Figure 2 , Figure 3 , Figure 4The embodiments shown are exaggerated presentation methods used to illustrate the method and highlight the effect. In practice, the results should be obtained using real video, a real-scale calibration board, and calculations by a server or PTZ.
[0050] It is worth noting that, as shown in the figure, this embodiment only illustrates a simple implementation of the traffic intersection flow image processing method based on biosensing time delay. It is only for the purpose of demonstrating the principle of this method. The shooting frequency and occlusion ratio of the occluded image can be set according to actual needs. The pixel accuracy and processing speed in reality are far higher than those in the embodiment. The specific number, position, angle, etc. can be set according to the needs of actual application.
[0051] The classification of motor vehicles, non-motor vehicles, and pedestrians in this method is not mandatory. The classification will be determined according to the specific algorithm, but the classification rules are consistent. For example, it is not mandatory to classify a non-motor vehicle with people into the non-motor vehicle category or the pedestrian category. However, if a non-motor vehicle with people in the auxiliary feature sample domain is classified into the non-motor vehicle category, then a non-motor vehicle with people in the target sample domain will also be classified into the non-motor vehicle category.
[0052] In the description of this invention, it should be understood that the terms "upper", "lower", "front", "rear", "left", "right", etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings, and are only for the convenience of describing this invention, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of this invention.
[0053] The embodiments described above are merely one possible representation of the present invention and are not intended to limit the scope of the present invention. Any modifications and improvements made to the technical solutions of the present invention by those skilled in the art without departing from the spirit of the present invention should fall within the protection scope defined by the claims of the present invention.
Claims
1. A method for processing traffic flow images at intersections based on biological perception time delay, characterized in that, include S1: Using historical videos of similar intersections as sample videos, extract the contour features of motor vehicles, non-motor vehicles, and pedestrians in the images to form auxiliary feature sample domains N1, N2, and N3; S2: Using the sample video frame as a reference, perform a occlusion operation on the intersection shooting device, and capture occluded images of the target intersection at an average speed of P frames / second. Extract the contour features of motor vehicles, non-motor vehicles, and pedestrians in the images to form target sample domains M1, M2, and M3. S3: Create an expanded scaling calibration board; S31: Create an expanded scaling calibration board based on the same image size as a single frame image of a similar historical video of an intersection and a masked image taken at the target intersection; S32: Extract sample video segments distributed at different locations in the video, each belonging to one of the categories of motor vehicles, non-motor vehicles, or pedestrians. Discretize the video segments into continuous multi-frame images, each corresponding to an expanded scale calibration plate region division. Calculate the proportion of a target category displayed at positions 1 to r in the video frame image using the expanded scale calibration plate, where ζ1, ζ2, ..., ζr are given, where r ≥ 1 and r is an integer. S4: The feature sample domains N1, N2, N3 and the target sample domains M1, M2, M3 are trained and expanded according to the expansion ratio calibration board to obtain the target feature sample domains N1”, N2”, N3”; S41: Based on the expansion ratio calibration board, map the display ratio of each category to the feature sample domain and the target sample domain. Expand each contour feature in the feature sample domains N1, N2, N3 and the target sample domains M1, M2, M3 according to the coordinate ratio. Each feature sample domain is divided into r regions according to the expansion ratio calibration board to obtain r feature sample subdomains; that is, N1'={ζ1N1, ζ2N1, ..., ζrN1}, N2'={ζ1N2, ζ2N2, ..., ζrN2}, N3'={ζ1N3, ζ2N3, ..., ζrN3}; S42: Match the contour features in the target sample domains M1, M2, M3 and the feature sample domains N1, N2, N3 according to the expansion ratio calibration board. Place the contour features in the target sample domains M1, M2, M3 with a matching threshold β < 0.8 into the feature sample domains, and then expand and improve them according to step S41 to obtain the target feature sample domains N1”, N2”, N3”. S43: For all images of the target intersection that are captured with the masked view, steps S41 and S42 must be performed; S5: The feature contours extracted from the captured occluded images are matched with the feature contours contained in the target feature sample domains N1”, N2”, and N3” according to the expanded scaling calibration plate. Occluded images with a matching threshold β ≥ 0.8 are discarded. The remaining occluded images per second are denoted as... , where m≥1 and m is an integer; S6: Process the remaining masked images according to the time sequence using biosensing time-lag frame rate to obtain the frame rate F frames / second within the biosensing time-lag time, denoted as... Where n≥1 and n is an integer, continuous playback forms a dynamic continuous video file of the intersection without motor vehicles, non-motor vehicles, and pedestrians.
2. The traffic intersection flow image processing method based on biosensing time delay according to claim 1, characterized in that, Step S32 includes that the scale values ζ1, ζ2, ..., ζr of the r regions divided by the expanded scale calibration plate are calculated by the display size of the same object in different regions.
3. The traffic intersection flow image processing method based on biosensing time delay according to claim 1, characterized in that, Step S42 includes a sample matching threshold. , where a∈{N1'、N2'、N3'}, b∈{M1、M2、M3}.
Citation Information
Patent Citations
Multi-channel image processing method and device and electronic equipment
CN112101305A
Road traffic scene localization generation and expansion method based on generative adversarial network
CN116071457A