Big data-based intelligent transportation information collection method, system, medium and device
By grouping traffic information collection devices and identifying complementary device groups, and dynamically adjusting the collection strategy, the problem of declining data quality in existing technologies has been solved, and the timeliness of traffic information transmission and the efficiency of resource scheduling have been improved.
Patent Information
- Application Number
- CN202510064039.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-15
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2045-01-15
AI Technical Summary
Existing traffic information collection equipment suffers from decreased data quality and insufficient collection capacity under conditions of drastic fluctuations in traffic flow or sudden changes in equipment load, resulting in low transmission timeliness.
By grouping the acquisition devices using big data-based methods, complementary device groups are identified, and data acquisition strategies are dynamically adjusted based on real-time operating status information to achieve collaborative acquisition and resource scheduling among device groups.
Maintaining high data collection quality under conditions of drastic fluctuations in traffic flow or sudden changes in equipment load improves the timeliness of traffic information transmission and enables the rational scheduling of collection resources.
Smart Images

Figure CN119992825B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of data collection, in particular to a smart traffic information collection method and system based on big data, a medium and equipment. BACKGROUND
[0002] With the acceleration of urbanization, urban traffic congestion problems are increasingly prominent. As an important means to alleviate traffic congestion, the core of the smart traffic system lies in accurately and timely collecting traffic information. At present, urban traffic information collection mainly relies on various collection devices distributed in the road network, such as vehicle detectors, video surveillance, electronic tags, etc.
[0003] In the prior art, traffic information collection devices usually adopt fixed collection strategies for data collection. Each collection device works independently and performs collection tasks according to the preset collection frequency and collection parameters. This fixed collection mode has certain limitations in actual application. Especially in the case of dramatic fluctuations in traffic flow or sudden changes in device load, a single collection device may have problems such as reduced data quality and insufficient collection capacity, resulting in low timeliness of traffic information transmission. SUMMARY
[0004] The present application provides a smart traffic information collection method and system based on big data, a medium and equipment, which improves the timeliness of traffic information transmission and realizes the reasonable scheduling of collection resources.
[0005] In a first aspect, the present application provides a smart traffic information collection method based on big data, which comprises:
[0006] Obtaining historical collection records of each collection device in a target traffic network, and grouping each collection device based on the historical collection records to obtain a plurality of collection device groups;
[0007] Obtaining working state information of each collection device in the collection device group;
[0008] Determining a complementary collection device group for each collection device group;
[0009] Based on the working state information of each collection device and the complementary collection device group, determining a data collection strategy for each collection device group.
[0010] By adopting the technical scheme, the devices with similar collection characteristics are divided into the same group based on the historical collection records, so that the subsequent overall scheduling is facilitated; secondly, the complementary collection device groups are determined, so that when an abnormality occurs in a collection device group, the corresponding complementary device can be timely activated for cooperative collection; thirdly, the data collection strategy is dynamically adjusted based on the real-time working state information of each collection device and the complementary collection device group relationship, so that the problem of data quality decline caused by the fixed collection strategy is avoided. Even in the case of severe traffic flow fluctuation or device load mutation, the data collection quality can be maintained at a high level, the transmission timeliness of traffic information is improved, and the reasonable scheduling of collection resources is realized.
[0011] In a second aspect of the present application, a big data-based intelligent traffic information collection system is provided, which comprises:
[0012] A collection device grouping module is configured to obtain historical collection records of each collection device in a target traffic network, and group each collection device based on the historical collection records, to obtain a plurality of collection device groups.
[0013] A device data acquisition module is configured to obtain working state information of each collection device in the collection device group.
[0014] A complementary device determination module is configured to determine a complementary collection device group for each collection device group.
[0015] A collection strategy determination module is configured to determine a data collection strategy for each collection device group based on the working state information of each collection device and the complementary collection device group.
[0016] In a third aspect of the present application, a computer storage medium is provided, which stores a plurality of instructions suitable for being loaded and executed by a processor to perform the above method steps.
[0017] In a fourth aspect of the present application, an electronic device is provided, which comprises a processor and a memory; wherein the memory stores a computer program suitable for being loaded and executed by the processor to perform the above method steps.
[0018] In summary, the one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages:
[0019] The application groups the collection devices based on historical collection records, so that devices with similar collection characteristics are divided into the same group, facilitating subsequent overall scheduling. Secondly, by determining the complementary collection device group, when an abnormality occurs in a certain collection device group, the corresponding complementary device can be activated in time for collaborative collection. Thirdly, based on the real-time working state information of each collection device and the relationship of the complementary collection device group, the data collection strategy is dynamically adjusted, avoiding the problem of data quality decline caused by using a fixed collection strategy. Even in the case of severe traffic flow fluctuations or device load mutations, high data collection quality can be maintained, improving the transmission timeliness of traffic information and realizing the reasonable scheduling of collection resources. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 is a flow diagram of a smart traffic information collection method based on big data provided by an embodiment of the application;
[0021] Figure 2 is a module diagram of a smart traffic information collection system based on big data provided by an embodiment of the application;
[0022] Figure 3 is a structural diagram of an electronic device provided by an embodiment of the application.
[0023] The following items are explained: 300, electronic device; 301, processor; 302, communication bus; 303, user interface; 304, network interface; 305, memory. DETAILED DESCRIPTION
[0024] In order for those skilled in the art to better understand the technical solutions in the specification, the technical solutions in the specification will be described clearly and completely in the following description of the embodiments of the specification in conjunction with the drawings in the embodiments of the specification. Obviously, the described embodiments are only some of the embodiments of the application, not all.
[0025] In the description of the embodiments of the application, the words such as "for example" or "for instance" are used to represent an example, illustration or description. Any embodiment or design scheme described as "for example" or "for instance" in the embodiments of the application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes. Rather, the words such as "for example" or "for instance" are intended to present the relevant concept in a specific way.
[0026] In the description of the embodiments of the present application, the term "a plurality of" means two or more. For example, a plurality of systems means two or more systems, and a plurality of screen terminals means two or more screen terminals. In addition, the terms "first", "second" are only for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the indicated technical features. Therefore, the features defined with "first", "second" can be explicitly or implicitly included one or more features. The terms "include", "contain", "have" and their variants mean "include but not limited to", unless otherwise specifically emphasized.
[0027] The technical solutions in the embodiments of the present application will be described clearly and completely in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all.
[0028] Please refer to Figure 1 , a flowchart of a big data-based intelligent traffic information collection method is proposed, which can be realized by a computer program, can be realized by a single-chip microcomputer, and can run on a big data-based intelligent traffic information collection system. The computer program can be integrated in a computer device or run as an independent tool application. Specifically, the method includes steps 10 to 40, which are as follows:
[0029] Step 10: Obtain historical collection records of each collection device in the target traffic road network, and group each collection device based on the historical collection records to obtain a plurality of collection device groups.
[0030] The target traffic road network in the embodiments of the present application refers to a road network area that needs to collect traffic information, which can be a road network in a certain area of a city, a main road network connecting multiple areas, or a highway network, etc.
[0031] The collection device in the embodiments of the present application refers to a device installed in the target traffic road network for collecting traffic information, including but not limited to a vehicle detector, a video monitoring device, an electronic tag reader, a GPS signal receiver, a traffic flow monitor, and other devices capable of collecting traffic-related data.
[0032] The historical collection record in the embodiments of the present application refers to the traffic data records collected by each collection device in a preset time period, including but not limited to traffic flow data, vehicle speed data, vehicle type data, traffic event data, and device operation records generated during the collection process, such as data collection time, collection frequency, data transmission records, etc.
[0033] The collection device group in the embodiments of the present application refers to a collection device set grouped according to data characteristics and geographical position information in historical collection records. The devices in the same collection device group have similar data collection characteristics and adjacent geographical position distribution.
[0034] Specifically, the system can obtain historical collection records of each collection device in a target traffic network in a past period of time (for example, the last 30 days). Based on these historical collection records, the data completeness of each collection device is calculated, that is, the ratio of the actual amount of valid data obtained in a collection period to the amount of data that should be theoretically obtained. Subsequently, the system divides each collection device into different data completeness levels according to a preset completeness interval (for example, 90%-100% for high completeness, 70%-90% for medium completeness, and below 70% for low completeness). On this basis, the geographical position information of each collection device is further obtained, and the collection devices with the same data completeness level and adjacent geographical position (for example, less than 500 meters in straight-line distance) are divided into the same collection device group. Through this grouping method based on data completeness and geographical position, the collection devices in the same group have similar data collection characteristics and spatial distribution characteristics, providing a basis for subsequent development of collection strategies according to the overall characteristics of the device group, and facilitating quick positioning of possible backup devices when an abnormality occurs in a collection device in a certain area.
[0035] On the basis of the above embodiments, as an optional embodiment, the step of grouping the collection devices based on the historical collection records to obtain a plurality of collection device groups can further include the following steps:
[0036] Step 101: Obtain the data completeness of each collection device in the historical collection records.
[0037] Specifically, before grouping the collection devices, the data collection quality of each collection device needs to be evaluated first. The system can obtain the historical collection records of each collection device in a preset period of time (for example, the last 30 days) and calculate its data completeness. For each collection device, the system first determines the theoretical amount of data N that should be collected, which can be calculated according to the collection frequency of the device (for example, collecting once per minute) and the statistical length of time (for example, 30 days). Subsequently, the actual amount of valid data M obtained by the collection device in this period of time is counted, where the valid data refers to data records meeting the preset quality standards, such as correct data format, values within a reasonable range, complete time stamp, etc. The data completeness is the ratio of the actual amount of valid data to the theoretical amount of data collected, that is, M / N.
[0038] Step 102: Determine the data completeness level of each collection device according to the data completeness.
[0039] Specifically, after obtaining the data completeness of each collection device, the collection device is classified according to a preset completeness interval. For example, when the data completeness is greater than 90%, the collection device is classified as A level; when the data completeness is between 70% and 90%, the collection device is classified as B level; and when the data completeness is less than 70%, the collection device is classified as C level. In this way, the system can quantitatively evaluate the data collection capability of each collection device, provide a reliable quality basis for subsequent device grouping, and help to divide devices with similar collection performance into the same group, so as to facilitate the subsequent development of targeted data collection strategies.
[0040] Step 103: Obtain the geographic location information of each collection device, divide the collection devices with the same data completeness level and adjacent geographic locations into the same collection device group, and obtain a plurality of collection device groups.
[0041] Specifically, the geographic location information of each collection device, including its longitude and latitude coordinates, is obtained. These geographic location information can be fixed coordinates recorded when the device is installed, or can be obtained in real time through the built-in GPS module of the device. Then, the collection devices with the same completeness level are grouped in a spatial clustering manner. Specifically, first, select an ungrouped collection device as the initial device, and calculate the straight-line distance between the initial device and other ungrouped devices with the same level. When the distance between two devices is less than a preset spatial threshold (for example, 500 meters), it is considered that the two devices are adjacent in geographic location. The system divides all devices with the same level and adjacent geographic locations to the initial device into the same collection device group. Repeat the above process until all collection devices are divided into corresponding groups. For example, the system can obtain a plurality of A-level device groups (A1 group, A2 group, etc.), and the devices in each A-level device group have more than 90% data completeness and are adjacent in geographic location. This grouping method based on data completeness level and geographic location enables the collection devices in the same group to have similar data collection quality and spatial correlation. Such grouping result is beneficial to subsequent quick calling of adjacent devices with the same level for cooperative collection when data collection anomaly occurs in a certain area, thereby improving the fault tolerance and reliability of data collection of the system.
[0042] Step 20: Obtain the working state information of each collection device in the collection device group.
[0043] The working state information in the embodiments of the present application refers to the real-time running state parameters of the collection device, including but not limited to: the current data collection frequency of the device, the CPU usage, the memory occupancy, the data cache capacity, the network transmission bandwidth occupancy, the device power (for devices using battery power), the device temperature, the data collection success rate in the recent period (for example, the past 1 hour), and other real-time information reflecting the working condition of the device.
[0044] Specifically, in order to grasp the running status of each collection device in real time, and timely discover and handle possible abnormal situations, it is necessary to regularly obtain the working status information of each collection device in the collection device group. Through the data communication link established with each collection device, a status query instruction is sent to the collection device every preset time interval (for example, every 5 minutes). After receiving the query instruction, the collection device will collect the current working status parameters and return them after packaging. These working status parameters include: CPU usage of the device (used to evaluate processing load), memory occupancy rate (used to evaluate data processing capacity), data cache capacity (used to evaluate data storage status), network transmission bandwidth occupancy rate (used to evaluate data transmission capacity), current data collection frequency (used to confirm whether it meets the preset collection requirements), device temperature (used to monitor the device operating environment), and data collection success rate in the last statistical period (for example, the past 1 hour) (used to evaluate collection stability). For collection devices powered by batteries, the current power information will also be returned. After receiving these working status information, the system stores it in the database and compares it with the preset normal working parameter range. This mechanism of regularly obtaining working status information enables the system to timely grasp the running status of each collection device,
[0045] Step 30: Determine the complementary collection device group of each collection device group.
[0046] The complementary collection device group in the embodiments of the present application refers to the combination relationship of the collection device group with complementary properties in terms of data collection capacity index, spatial distribution range, etc. Specifically, when two collection device groups meet the following conditions, they are complementary collection device groups: the two device groups have a high degree of complementarity (i.e. their data collection capacity indices can match and support each other), and their spatial distribution ranges have an overlapping or adjacent relationship (i.e. the spatial distance is within a preset threshold range). This complementary relationship ensures that when one collection device group appears abnormal, its complementary collection device group can effectively undertake the data collection task of the original device group in the performance guarantee and spatial coverage dimensions.
[0047] Specifically, in order to construct a high-reliability data acquisition system, it is necessary to establish a complementary mechanism between the acquisition device groups, that is, to determine for each acquisition device group a complementary acquisition device group that can support each other in performance and spatial dimensions. First, the data acquisition capability index of each acquisition device group is calculated. The index is obtained by comprehensively evaluating the working state parameters of each acquisition device in the device group, such as CPU usage, memory occupancy, data cache capacity, network transmission bandwidth occupancy, etc., and combining the device quantity and data integrity level, using a weighted calculation method. Subsequently, the system obtains the spatial distribution range of each acquisition device group, and calculates the minimum circumscribed polygon representing the coverage area of the device group by analyzing the geographic location coordinates of all acquisition devices in the device group. On this basis, the capability complementarity between any two acquisition device groups is calculated. This index reflects the matching degree of the data acquisition capability of the two device groups. When the data acquisition capability indices of the two device groups are close, they have a high capability complementarity; otherwise, when the capability indices differ greatly, the complementarity is low. Finally, considering the capability complementarity and spatial distribution characteristics, the complementary acquisition device group is determined. If the capability complementarity of the two device groups exceeds a preset threshold (e.g., 0.8), and their spatial distribution ranges overlap or the distance is less than a preset value (e.g., 500 meters), the two device groups are determined as each other's complementary acquisition device group. This complementary mechanism based on data acquisition capability and spatial distribution ensures that when an abnormality occurs in a certain acquisition device group, its complementary acquisition device group can effectively assume the data acquisition task of the original device group from the performance guarantee and spatial coverage dimensions, thereby maintaining the stable operation of the system and the continuity of data acquisition.
[0048] On the basis of the above embodiment, as another optional embodiment, the step of determining the complementary acquisition device group for each acquisition device group can further include the following steps:
[0049] Step 301: Calculate the data acquisition capability index of each acquisition device group.
[0050] Specifically, in calculating the data acquisition capability index, the system obtains the real-time working state information of each acquisition device in the acquisition device group, including CPU usage, memory occupancy, data cache capacity, network transmission bandwidth occupancy, etc. For each acquisition device group, the system calculates the average CPU usage and average memory occupancy of all devices in the group, and considers the total number of devices and the data integrity level of the device group. The specific calculation formula of the data acquisition capability index is: I = a x (1- average CPU usage) + b x (1- average memory occupancy) + g x device quantity x integrity level coefficient, where a = 0.4, b = 0.3, g = 0.3 are preset weight coefficients.
[0051] Step 302: Obtain the spatial distribution range of each acquisition device group.
[0052] Specifically, the geographic position coordinates of each collection device in each collection device group are read, and a convex hull algorithm is used to calculate the minimum circumscribed polygon containing all device coordinate points. The Graham scan algorithm can be used, which first finds the left lower corner point as the starting point, then sorts the other points according to the polar angle size with the starting point, and connects them in turn to form a convex polygon. This polygon is the spatial distribution range of the collection device group. For example, there are 4 collection devices in a collection device group, with coordinates (1, 1), (1, 4), (3, 2), and (4, 3). Through the convex hull algorithm, a quadrilateral area is obtained as the spatial distribution range of the device group. In this way, the system not only obtains the data collection capability index reflecting the performance level of the device group, but also obtains the spatial distribution range representing the geographic coverage characteristics of the device group, providing an important evaluation basis for subsequent determination of complementary collection device groups.
[0053] Step 303: Calculate the capability complementarity between each collection device group, which is determined according to the data collection capability index.
[0054] Specifically, the capability complementarity between any two collection device groups is calculated, which reflects the matching degree of the two device groups in data collection capability. For collection device groups A and B, assuming their data collection capability indexes are Ia and Ib, the calculation formula of their capability complementarity is: C = 1 - |Ia - Ib| / (Ia + Ib). This calculation method ensures that the value range of the capability complementarity is between 0 and 1, and the closer the data collection capability indexes of the two device groups, the closer the capability complementarity to 1.
[0055] For example, if the data collection capability index of device group A is 1.81 and the data collection capability index of device group B is 1.65, then the capability complementarity C between them is C = 1 - |1.81 - 1.65| / (1.81 + 1.65) = 1 - 0.16 / 3.46 = 0.954.
[0056] Step 304: Determine the complementary collection device group based on the capability complementarity and the spatial distribution range.
[0057] Specifically, first, it is judged whether the complementary degree of the capabilities of the two device groups exceeds a preset threshold (for example, 0.8), and if the condition is met, the relationship between the spatial distribution ranges of the two device groups is further calculated. The system uses a polygon intersection algorithm to determine whether the spatial distribution ranges of the two device groups overlap, and if not, the shortest distance between the two polygons is calculated, and when the distance is less than a preset distance threshold (for example, 100 meters), it is considered that the two device groups have spatial complementarity. Finally, if the two acquisition device groups simultaneously meet the capability complementarity threshold and the spatial distribution requirement, they are determined as each other's complementary acquisition device group. This complementary mechanism based on capability complementarity and spatial characteristics not only ensures that the backup device group has sufficient data acquisition capability to undertake the task of the original device group, but also ensures the proximity of the complementary device group in geographical location, so as to quickly switch and maintain the continuity and integrity of data acquisition when the system abnormally occurs.
[0058] On the basis of the above embodiment, as another optional embodiment, the step of calculating the data acquisition capability index of each acquisition device group can further include the following steps:
[0059] Step 3011: Obtain the data acquisition success rate of each acquisition device group within a preset time period, and count the number of data types supported by each acquisition device group.
[0060] Specifically, the data acquisition success rate of each acquisition device group within a preset time period (for example, the past 24 hours) is obtained, which is specifically calculated by counting the total number of data acquisition requests and the number of successful times of each acquisition device group within the time period. For example, a certain acquisition device group initiated 1000 data acquisition requests within the past 24 hours, of which 950 were successfully completed, and its data acquisition success rate was 95%. At the same time, the system counts the number of data types supported by each acquisition device group, which includes but is not limited to temperature data, humidity data, pressure data, vibration data, etc.
[0061] Step 3012: Calculate the data acquisition amount per unit time based on the acquisition frequency of each acquisition device group.
[0062] Specifically, the acquisition frequency refers to the sampling number of each type of data per unit time, and the specific calculation formula is: data acquisition amount = Σ(acquisition frequency of each type of data x data amount of single acquisition).
[0063] Step 3013: Obtain the data acquisition capability index of each acquisition device group according to the weighted combination of the data acquisition success rate, the number of data types, and the data acquisition amount.
[0064] Specifically, the data collection success rate, the data type quantity, and the data collection quantity are combined by weighting to calculate the data collection capacity index. Assuming that the weight of the data collection success rate is a (for example, 0.4), the weight of the data type quantity is β (for example, 0.3), and the weight of the data collection quantity is γ (for example, 0.3), the calculation formula of the data collection capacity index is I = a x data collection success rate + β x (data type quantity / reference type quantity) + γ x (data collection quantity / reference collection quantity). The reference type quantity and the reference collection quantity are standard values preset by the system. Continuing the above example, assuming that the reference type quantity is 5 and the reference collection quantity is 50 KB / minute, the data collection capacity index I of the collection device group is 0.4 x 0.95 + 0.3 x (3 / 5) + 0.3 x (31 / 50) = 0.38 + 0.18 + 0.186 = 0.746. Through the multi-dimensional weighted calculation, the data collection capacity index obtained by the system not only reflects the collection reliability of the device group, but also reflects the comprehensiveness and efficiency of the data collection, thereby providing a scientific evaluation basis for determining the complementary collection device group.
[0065] On the basis of the above embodiment, as another optional embodiment, the step of determining the complementary collection device group based on the capacity complementarity and the spatial distribution range can further include the following steps:
[0066] Step 3041: calculating the overlapping coverage rate between each collection device group based on the spatial distribution range.
[0067] Specifically, based on the obtained spatial distribution range, the overlapping coverage rate between any two collection device groups is calculated. For the collection device groups A and B, assuming that their spatial distribution ranges are represented by polygons A and B respectively, the system calculates the area S overlap of the overlapping region by using the polygon intersection algorithm, and calculates the areas SA and SB of the polygons A and B respectively, and the overlapping coverage rate between the two device groups is calculated according to the formula: R = S overlap / min(SA, SB). For example, if the area of polygon A is 100 square meters, the area of polygon B is 120 square meters, and the area of the overlapping region is 60 square meters, then the overlapping coverage rate R = 60 / 100 = 0.6. Through this calculation method, the overlapping coverage rate obtained by the system can accurately reflect the degree of overlap of the two device groups in the spatial distribution.
[0068] Step 3042: screening the collection device group pairs with the overlapping coverage rate greater than the overlapping coverage rate threshold.
[0069] Specifically, the calculated overlap coverage is compared with a preset overlap coverage threshold (e.g. 0.3), and the device group pair with an overlap coverage greater than the threshold is screened out. In the above example, since the overlap coverage 0.6 is greater than the threshold 0.3, the device group pair composed of device groups A and B will be retained for subsequent processing.
[0070] Step 3043: The degree of complementarity of each device group pair is sorted, and the device group pair with a degree of complementarity greater than a degree of complementarity threshold is determined as a complementary device group.
[0071] Specifically, the degrees of complementarity of the screened device group pairs are sorted in descending order, and the calculation of the degree of complementarity uses the aforementioned method, i.e. C = 1 - |Ia - Ib| / (Ia + Ib), where Ia and Ib are the data acquisition capability indexes of device groups A and B, respectively. For example, if the data acquisition capability index of device group A is 0.746 and the data acquisition capability index of device group B is 0.682, the degree of complementarity C between them is C = 1 - |0.746 - 0.682| / (0.746 + 0.682) = 1 - 0.064 / 1.428 = 0.955. Finally, the system compares the sorted device group pairs with a preset degree of complementarity threshold (e.g. 0.85), and determines the device group pair with a degree of complementarity greater than the threshold as a complementary device group. In the above example, since the degree of complementarity 0.955 is greater than the threshold 0.85, device groups A and B are finally determined as a complementary device group. This double screening mechanism based on spatial overlap coverage and degree of complementarity ensures that the complementary device group has both sufficient overlap coverage in geographical position and high matching degree in data acquisition capability, so as to effectively assume the data acquisition task of the original device group when the original device group is abnormal, and ensure the reliability of the system and the continuity of data acquisition.
[0072] Step 40: Based on the working state information of each acquisition device and the complementary device group, the data acquisition strategy of each device group is determined.
[0073] Specifically, in order to realize efficient cooperation and resource optimization configuration of the acquisition device group, the system needs to formulate a reasonable data acquisition strategy according to the working state information of each acquisition device and the characteristics of the complementary acquisition device group. First, the real-time working state information of each acquisition device in each acquisition device group is obtained, including CPU usage, memory occupancy, data cache capacity, network transmission bandwidth occupancy, and other basic running parameters. Specifically, the system takes the CPU usage and memory occupancy of the acquisition device as the key indicators for evaluating the device load. For example, when the CPU usage of a certain acquisition device is 75% and the memory occupancy is 80%, it can be determined that the device is in a high load state. At the same time, the system calculates the overall working state indicators of each acquisition device group, specifically by weighted averaging the working state parameters of all acquisition devices in the device group, and the calculation formula is: W=δ×average CPU usage+ε×average memory occupancy+ζ×average bandwidth occupancy, where δ, ε, ζ are weight coefficients and δ+ε+ζ=1. Then, according to the calculated overall working state indicators, combined with the information of the complementary acquisition device group, a data acquisition strategy is formulated for each acquisition device group. When the overall working state indicator of a certain acquisition device group exceeds the preset high load threshold (for example, 0.8), the system allocates part of the data acquisition tasks to its complementary acquisition device group. The specific task allocation method is: first, calculate the amount of data acquisition tasks to be allocated, the calculation formula is: T=(W-0.8)×current total task amount, where W is the overall working state indicator; then, according to the data acquisition capacity index of the complementary acquisition device group, the task amount T is allocated to each complementary acquisition device group in proportion to the capacity index.
[0074] On the basis of the above embodiment, as another optional embodiment, based on the working state information of each acquisition device and the complementary acquisition device group, the step of determining the data acquisition strategy of each acquisition device group can also include the following steps:
[0075] Step 401: Calculate the resource pressure index of each acquisition device based on the resource occupancy rate.
[0076] Specifically, it is realized by monitoring the CPU usage and memory occupancy of the acquisition device. The calculation formula of the resource pressure index is: Ir=λ1×CPU usage+λ2×memory occupancy, where λ1 and λ2 are weight coefficients and λ1+λ2=1. For example, the CPU usage of a certain acquisition device is 85% and the memory occupancy is 75%, and λ1=0.6 and λ2=0.4 are taken, then the resource pressure index Ir=0.6×0.85+0.4×0.75=0.81. When the resource pressure index exceeds the preset threshold (such as 0.8), it indicates that the computational resource load of the device is heavy.
[0077] Step 402: Calculate the storage pressure index of each acquisition device based on the remaining storage capacity.
[0078] Specifically, the system calculates the storage pressure index of each collection device based on the remaining storage capacity, and the calculation formula is: Is=1-(remaining storage capacity / total storage capacity)^μ, where μ is an adjustment coefficient (for example, the value is 1.5), which is used to accelerate the growth rate of the storage pressure index when the storage space is insufficient. For example, the total storage capacity of a certain collection device is 1000GB, and the remaining storage capacity is 200GB, then its storage pressure index Is=1-(200 / 1000)^1.5=0.911. This calculation method ensures that when the remaining storage space is small, the storage pressure index will quickly rise, reminding the system to clean up or dump data in time.
[0079] Step 403: Calculate the transmission pressure index of each collection device based on the length of the transmission queue.
[0080] Specifically, the system calculates the transmission pressure index of each collection device based on the length of the transmission queue, and the calculation formula is: It=Current queue length / (maximum queue length x η), where η is a buffer coefficient (for example, the value is 0.8), which is used to reserve a safety margin for data transmission. For example, the maximum transmission queue length of a certain collection device is 1000 data records, and the current queue length is 600, and η=0.8 is taken, then its transmission pressure index It=600 / (1000 x 0.8)=0.75. This calculation method can reflect the data transmission load of the device, and when the transmission pressure index is close to or exceeds 1, it indicates that the data transmission capacity of the device is close to saturation.
[0081] Step 404: When there is an abnormal collection device group whose resource pressure index and / or storage pressure index and / or transmission pressure index exceeds the corresponding threshold value, determine the cooperative task of the complementary collection device group corresponding to the abnormal collection device group.
[0082] Specifically, the system identifies abnormal states by comparing the three types of pressure indexes of each collection device group with the corresponding threshold values. Set the resource pressure index threshold value to 0.8, the storage pressure index threshold value to 0.9, and the transmission pressure index threshold value to 0.85. When any pressure index exceeds the corresponding threshold value, the system will mark the collection device group as an abnormal state. For example, the resource pressure index of a certain collection device group is 0.81, the storage pressure index is 0.911, and the transmission pressure index is 0.75. Since the resource pressure index and the storage pressure index both exceed their respective threshold values, the device group is determined to be in an abnormal state.
[0083] Further, for the abnormal acquisition device group, corresponding task allocation strategies are formulated based on different types of pressure overrun situations. When the resource pressure index is overrun, the system calculates the amount of computing tasks that need to be allocated, with the formula being Tr= (Ir-0.8) x current total amount of computing tasks, where Ir is the resource pressure index; when the storage pressure index is overrun, the system calculates the amount of data storage that needs to be transferred, with the formula being Ts= (Is-0.9) x current total amount of storage data, where Is is the storage pressure index; when the transmission pressure index is overrun, the system calculates the amount of data transmission that needs to be shunted, with the formula being Tt= (It-0.85) x current length of transmission queue, where It is the transmission pressure index. For example, an abnormal acquisition device group is currently responsible for 1000 acquisition points / minute of computing tasks, and its resource pressure index is 0.85, so the amount of computing tasks that need to be allocated is (0.85-0.8) x 1000 = 50 acquisition points / minute. Next, the system determines a specific collaborative task allocation scheme for the abnormal acquisition device group. The system first checks the current pressure state of the complementary acquisition device groups of the device group, excludes complementary device groups that are also in a high pressure state, and then allocates tasks in proportion according to the data acquisition capability indices of the remaining complementary device groups. Specifically, if an abnormal acquisition device group has two available complementary acquisition device groups A and B, with data acquisition capability indices of 0.746 and 0.682 respectively, then for the 50 acquisition points / minute of computing tasks that need to be allocated, the amount of tasks allocated to A is 50 x 0.746 / (0.746+0.682) ≈ 26 acquisition points / minute, and the amount of tasks allocated to B is 50 x 0.682 / (0.746+0.682) ≈ 24 acquisition points / minute.
[0084] Step 405: generating a data acquisition strategy containing acquisition tasks and / or collaborative tasks of each acquisition device group.
[0085] Specifically, the system finally generates a complete data acquisition strategy containing original acquisition tasks and collaborative tasks of each acquisition device group. For normally operating acquisition device groups, their original acquisition tasks remain unchanged; for abnormal acquisition device groups, their acquisition task amounts are adjusted, and the excess part is allocated as collaborative tasks to corresponding complementary acquisition device groups; for complementary acquisition device groups that undertake collaborative tasks, corresponding collaborative tasks are added to their original acquisition tasks. Through this dynamic task adjustment and allocation mechanism, the system can effectively alleviate the pressure of abnormal acquisition device groups, ensure the continuity and reliability of data acquisition work, and at the same time, through reasonable allocation of collaborative tasks, avoid the generation of new pressure bottlenecks, and improve the operating efficiency of the entire data acquisition system.
[0086] Exemplarily, it is assumed that at a busy intersection: there are two cameras (A1, A2) and a vehicle flow detector (B1), the collection ranges of A1 and A2 have overlaps and can be divided into group A, and the collection range of B1 partially overlaps with group A, which constitutes a complementary area. The system can formulate a strategy: normally A1 works and A2 is on standby; when A1 data is abnormal, A2 is automatically enabled to collect data, and the data of B1 is referred to for verification.
[0087] On the basis of the above-mentioned embodiments, as another optional embodiment, a big data-based intelligent traffic information collection method can further include the following processes:
[0088] Specifically, in order to ensure the reliability and accuracy of the collected data, the system needs to comprehensively evaluate and timely correct the data quality of each collection device group. First, the data quality indicators of each collection device group are obtained, including data completeness rate, data accuracy rate and data timeliness. Among them, the data completeness rate is determined by calculating the ratio of the actual number of data points obtained to the theoretical number of data points that should be obtained, and the calculation formula is: Rc=actual data points / theoretical data points. For example, a certain collection device group should have collected 3600 data points in the past 1 hour, and actually collected 3420 data points, so its data completeness rate is 3420 / 3600=0.95. The data accuracy rate is evaluated by comparing the collected data with the pre-set reasonable value range, and the calculation formula is: Ra=effective data points / total data points, where the effective data points are the data points falling within the reasonable range. For example, 96 of the 100 data points collected by a certain temperature sensor fall within the pre-set reasonable range of -30℃ to 50℃, so its data accuracy rate is 0.96. The data timeliness is measured by calculating the ratio of the time required from data collection to availability to the pre-set maximum allowed delay time, and the calculation formula is: Rt=1-min(actual delay time / maximum allowed delay time, 1). For example, when the actual processing delay of a certain data point is 200 milliseconds and the maximum allowed delay is 500 milliseconds, its timeliness is 1-200 / 500=0.6.
[0089] The system calculates the data quality score of each collection device group based on the three indicators, and the calculation formula is: Q=ω1×Rc+ω2×Ra+ω3×Rt, wherein ω1, ω2, and ω3 are weight coefficients. For example, ω1=0.3, ω2=0.4, and ω3=0.3, and the data quality score of the above collection device group is Q=0.3×0.95+0.4×0.96+0.3×0.6=0.849. When the calculated data quality score is lower than the preset quality threshold (for example, 0.85), the system marks the collection device group as an abnormal quality state. In this case, the system obtains calibration data from the complementary collection device group corresponding to the abnormal quality collection device group, and uses these data to correct the abnormal quality data. The specific correction process includes: first, obtaining the collection data of the complementary collection device group at the same time period and the same location as the calibration data; then, the system calculates the deviation between the calibration data and the abnormal data, including the system deviation and the random deviation; finally, the system corrects the abnormal data according to the identified deviation. For example, the collection data of a certain temperature sensor is continuously 2℃ higher, and presents ±0.5℃ random fluctuation, the system will first eliminate the 2℃ system deviation, and then smooth the random fluctuation through Kalman filtering algorithm, and finally obtain the corrected collection data. Through this mechanism based on data quality evaluation and complementary device calibration, the system can timely find and correct the abnormality in the collection data, ensure the reliability and accuracy of the data, and provide more valuable data support for subsequent data analysis and decision-making.
[0090] See Figure 2 A module schematic diagram of a smart traffic information collection system based on big data is provided for the embodiments of the present application, wherein the system comprises:
[0091] A collection device grouping module is configured to obtain historical collection records of each collection device in a target traffic network, and group each collection device based on the historical collection records to obtain a plurality of collection device groups.
[0092] A device data acquisition module is configured to obtain working state information of each collection device in the collection device group.
[0093] A complementary device determination module is configured to determine a complementary collection device group for each collection device group.
[0094] A collection strategy determination module is configured to determine a data collection strategy for each collection device group based on the working state information of each collection device and the complementary collection device group.
[0095] Optionally, the collection device grouping module is further configured to obtain the data completeness of each collection device in the historical collection records.
[0096] determine a data integrity level of each of the collection devices according to the data integrity;
[0097] obtain geographical position information of each of the collection devices, divide collection devices with the same data integrity level and adjacent geographical positions into a same collection device group, and obtain a plurality of collection device groups.
[0098] Optionally, the complementary device determination module is further configured to calculate a data collection capability index of each of the collection device groups;
[0099] obtain a spatial distribution range of each of the collection device groups;
[0100] calculate a capability complementarity between each of the collection device groups, the capability complementarity being determined according to the data collection capability index;
[0101] determine a complementary collection device group based on the capability complementarity and the spatial distribution range.
[0102] Optionally, the complementary device determination module is further configured to obtain a data collection success rate of each of the collection device groups within a preset time period, and count a number of data types supported by each of the collection device groups;
[0103] calculate a data collection amount per unit time based on a collection frequency of each of the collection device groups;
[0104] obtain a data collection capability index of each of the collection device groups according to a weighted combination of the data collection success rate, the number of data types, and the data collection amount.
[0105] Optionally, the complementary device determination module is further configured to calculate an overlapping coverage rate between each of the collection device groups based on the spatial distribution range;
[0106] screen out a collection device group pair with an overlapping coverage rate greater than an overlapping coverage rate threshold value;
[0107] sort the capability complementarity of each of the collection device group pairs, and determine a collection device group pair with a capability complementarity greater than a capability complementarity threshold value as a complementary collection device group.
[0108] Optionally, the collection strategy determination module is further configured to calculate an overlapping coverage rate between each of the collection device groups based on the spatial distribution range;
[0109] screen out a collection device group pair with an overlapping coverage rate greater than an overlapping coverage rate threshold value;
[0110] sort the capability complementarity of each of the collection device group pairs, and determine a collection device group pair with a capability complementarity greater than a capability complementarity threshold value as a complementary collection device group.
[0111] Optionally, the collection strategy determination module is further configured to acquire data quality indexes of each of the collection device groups, the data quality indexes including data completeness rate, data accuracy rate and data timeliness;
[0112] The data quality indexes are used to calculate data quality scores of the collection device groups;
[0113] When the data quality scores are lower than a preset quality threshold, the collection device group with quality anomaly is determined;
[0114] Calibration data are acquired from a complementary collection device group corresponding to the collection device group with quality anomaly, and the quality abnormal data are corrected based on the calibration data to obtain corrected collection data.
[0115] It should be noted that the system provided in the above embodiments is used to implement its functions, and only the division of the above functional modules is used as an example for illustration. In actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the above described functions. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is described in the method embodiments, which will not be repeated here.
[0116] The computer storage medium provided in the embodiments of the present application can store a plurality of instructions, and the instructions are suitable for being loaded by a processor and executing the above-mentioned embodiments of a big data-based intelligent traffic information collection method. The specific implementation process can be referred to the specific description of the above-mentioned embodiments, which will not be repeated here.
[0117] Please refer to Figure 3 The present application also discloses an electronic device. Figure 3 is a structural schematic diagram of an electronic device disclosed in the embodiments of the present application. The electronic device 300 can include at least one processor 301, at least one network interface 304, a user interface 303, a memory 305, and at least one communication bus 302.
[0118] The communication bus 302 is used to realize the connection and communication between the components.
[0119] The user interface 303 can include a display screen (Display) and a camera (Camera). Optionally, the user interface 303 can further include a standard wired interface and a wireless interface.
[0120] Optionally, the network interface 304 can include a standard wired interface and a wireless interface (such as a WI-FI interface).
[0121] The processor 301 can include one or more processing cores. The processor 301 connects various parts within the server through various interfaces and lines, performs various functions of the server and processes data by running or executing instructions, programs, code sets or instruction sets stored in the memory 305, and calling data stored in the memory 305. Alternatively, the processor 301 can be implemented in at least one of a hardware form of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 301 can integrate a combination of one or more of a central processing unit (CPU), a graphics processor (GPU), and a modem. Among them, the CPU mainly processes operating systems, user interfaces, and application programs; the GPU is responsible for rendering and drawing the content to be displayed on the display screen; and the modem is used for processing wireless communication. It can be understood that the above-mentioned modem can also not be integrated into the processor 301, but can be realized by a separate chip.
[0122] The memory 305 can include a random access memory (RAM) and a read-only memory (ROM). Alternatively, the memory 305 includes a non-transitory computer-readable storage medium. The memory 305 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 305 can include a program storage area and a data storage area, wherein the program storage area can store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playing function, an image playing function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area can store data involved in the above-mentioned various method embodiments, etc. The memory 305 can alternatively be at least one storage device located away from the aforementioned processor 301. Referring to Figure 3 The memory 305 as a computer storage medium can include an operating system, a network communication module, a user interface module, and an application program of a big data-based intelligent traffic information collection method.
[0123] In Figure 3In the electronic device 300 shown, the user interface 303 is mainly used to provide an interface for the user to input, and obtain data input by the user; and the processor 301 can be used to invoke an application program of a big data-based intelligent transportation information collection method stored in the memory 305, and when executed by one or more processors 301, the electronic device 300 performs the method described in one or more of the above embodiments. It should be noted that, for the foregoing method embodiments, in order to simply describe, they are all described as a combination of a series of actions, but those skilled in the art should know that the present application is not limited to the order of the actions described, because according to the present application, some steps can be performed in other order or at the same time. Secondly, those skilled in the art should know that the embodiments described in the specification all belong to preferred embodiments, and the actions and modules involved are not necessarily required by the present application.
[0124] In the above embodiments, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0125] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are only schematic. The division of the units is only a logical function division. There can be another division manner for actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different parts can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical or other forms.
[0126] The units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0127] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware, or in the form of software functional unit.
[0128] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable memory. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a memory and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the embodiments of the present application. The aforementioned memory includes: a U disk, a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0129] The above-described are only exemplary embodiments of the present disclosure, and cannot limit the scope of the present disclosure. That is, any equivalent changes and modifications made in accordance with the teachings of the present disclosure are still within the scope of the present disclosure. Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon considering the specification and practicing the true principles of the present disclosure.
[0130] The present application is intended to cover any variations, uses, or adaptive changes of the present disclosure that follow the general principles of the present disclosure and include common knowledge or conventional technical means in the technical field not recorded in the present disclosure. The specification and examples are only considered as exemplary, and the scope and spirit of the present disclosure are defined by the claims.
Claims
1. A big data-based intelligent transportation information collection method, characterized in that, The method comprises: acquiring historical collection records of each collection device in a target traffic network, and grouping each collection device based on the historical collection records to obtain a plurality of collection device groups; acquiring working state information of each collection device in the collection device groups; acquiring a data collection success rate of each collection device group within a preset time period, and counting a number of data types supported by each collection device group; calculating a data collection amount per unit time based on a collection frequency of each collection device group; obtaining a data collection capability index of each collection device group according to a weighted combination of the data collection success rate, the number of data types, and the data collection amount; acquiring a spatial distribution range of each collection device group; calculating a capability complementarity between each collection device group, the capability complementarity being determined according to the data collection capability index, and the capability complementarity reflecting a matching degree of two device groups in data collection capability; calculating an overlapping coverage rate between each collection device group based on the spatial distribution range; screening out a collection device group pair with an overlapping coverage rate greater than an overlapping coverage rate threshold value; sorting the capability complementarity of each collection device group pair, and determining a collection device group pair with a capability complementarity greater than a capability complementarity threshold value as a complementary collection device group; determining a data collection strategy of each collection device group based on working state information of each collection device and the complementary collection device group. 2.The big data-based intelligent transportation information collection method according to claim 1, characterized in that, The grouping of each collection device based on the historical collection records to obtain a plurality of collection device groups comprises: acquiring a data completeness of each collection device in the historical collection records; determining a data completeness level of each collection device according to the data completeness; acquiring geographical position information of each collection device, and dividing collection devices with the same data completeness level and adjacent geographical positions into a same collection device group to obtain a plurality of collection device groups. 3.The big data-based intelligent transportation information collection method according to claim 1, characterized in that, The working state information comprises a computing resource occupancy rate, a remaining storage capacity, and a transmission queue length, and the determination of the data collection strategy of each collection device group based on the working state information of each collection device and the complementary collection device group comprises: calculating a resource pressure index of each collection device based on the computing resource occupancy rate; calculating a storage pressure index of each collection device based on the remaining storage capacity; calculating a transmission pressure index of each collection device based on the transmission queue length; when there is an abnormal collection device group with a resource pressure index and / or a storage pressure index and / or a transmission pressure index exceeding a corresponding threshold value, determining that the abnormal collection device group corresponds to a cooperative task of the complementary collection device group; generating a data collection strategy comprising collection tasks and / or cooperative tasks of each collection device group. 4.The big data-based intelligent transportation information collection method according to claim 1, characterized in that, The method further comprises: acquiring a data quality index of each collection device group, the data quality index comprising a data completeness rate, a data accuracy rate, and a data timeliness; calculating a data quality score of each collection device group based on the data quality index; when the data quality score is lower than a preset quality threshold value, determining a quality abnormal collection device group; Obtain calibration data from a complementary acquisition device group corresponding to the acquisition device group with quality anomaly, and correct the quality anomaly data based on the calibration data to obtain corrected acquisition data.
5. A big data-based intelligent transportation information collection system, characterized in that, The system for executing the big data-based intelligent traffic information acquisition method of claim 1 comprises: An acquisition device grouping module is configured to obtain historical acquisition records of each acquisition device in a target traffic network, and group each acquisition device based on the historical acquisition records to obtain a plurality of acquisition device groups. A device data obtaining module is configured to obtain working state information of each acquisition device in the acquisition device groups. A complementary device determining module is configured to determine a complementary acquisition device group for each acquisition device group. An acquisition strategy determining module is configured to determine a data acquisition strategy for each acquisition device group based on the working state information of each acquisition device and the complementary acquisition device group.
6. A computer-readable storage medium, characterized in that, The computer readable storage medium stores a plurality of instructions, which are suitable for being loaded and executed by the processor to execute the method of any one of claims 1-4.
7. An electronic device, comprising: The electronic device comprises a processor, a memory, a user interface, and a network interface. The memory is configured to store instructions. The user interface and the network interface are configured to communicate with other devices. The processor is configured to execute the instructions stored in the memory to enable the electronic device to execute the method of any one of claims 1-4.
Citation Information
Patent Citations
Collaborative data processing
CN109696692A
Sensor sharing method and device for intelligent networked automobile
CN116346862A