Traffic service management big data comprehensive analysis, research and judgment system
By designing a comprehensive analysis and judgment system for traffic business management big data, real-time access to multi-source data flows, dynamically loading and optimizing risk rules, real-time risk identification and early warning are achieved, and the shortcomings of existing systems in risk identification and early warning are solved, and the intelligence and efficiency of the traffic management system are improved.
Patent Information
- Application Number
- CN202510466129.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-15
AI Technical Summary
The existing traffic management system has shortcomings in risk identification and early warning, and cannot effectively respond to the rapidly changing traffic environment, and lacks the ability to dynamically load and optimize rules, resulting in slow response and difficulty in real-time monitoring and timely early warning.
A comprehensive analysis and judgment system for traffic business management big data is designed. Through the access module, the multi-source big data stream is connected in real time, the loading module dynamically loads risk rules and optimizes the order of execution of rules, the matching module performs real-time risk identification, and the early warning module generates risk warning signals and pushes them in real time.
Effectively analyze and judge potential risks in traffic business processes, improve the accuracy and timeliness of risk identification, and realize the intelligence and efficiency of traffic management systems.
Smart Images

Figure CN119989003A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of traffic management, and in particular to a comprehensive analysis and judgment system for traffic business management big data. Background Art
[0002] With the rapid development of social economy and the acceleration of urbanization, the demand for transportation is increasing, and the complexity and diversity of transportation business management are increasing. Traditional traffic management models usually rely on manual monitoring and regular data analysis, which cannot effectively cope with the growing amount of traffic data and real-time changing traffic conditions. Faced with massive amounts of traffic data, how to quickly and accurately identify risks and respond has become an important challenge in transportation business management.
[0003] At present, big data technology has been introduced in the field of traffic management in order to improve the intelligent level of traffic management through efficient processing and analysis of traffic data. Although multiple traffic data analysis systems and management platforms have been widely used, they still have certain shortcomings in risk identification and early warning. For example, existing risk identification methods often rely on static rule bases and outdated risk assessment models, which makes it difficult to effectively identify potential risks in a rapidly changing traffic environment. In addition, the lack of the ability to dynamically load and optimize rules makes the existing system slow to respond to complex and sudden traffic events, making it difficult to achieve real-time monitoring and timely early warning. Summary of the invention
[0004] The purpose of the present invention is to provide a comprehensive analysis and judgment system for traffic business management big data to address the deficiencies in the prior art, effectively analyze and judge the potential risks in the traffic business process, and improve the accuracy and timeliness of risk identification.
[0005] An embodiment of the present application provides a traffic business management big data comprehensive analysis and judgment system, the system comprising: The access module is used to access the multi-source big data stream of traffic business management in real time using distributed data acquisition technology according to traffic business needs, generate standardized data streams, and generate cleaned real-time data streams based on the standardized data streams; The loading module is used to load the predefined risk rules into the rule engine according to the risk rule library by using the dynamic rule loading technology, convert the risk rules into executable rule models, and dynamically adjust the execution order and priority of the risk rules based on the rule model by using the rule optimization algorithm to generate an initialized rule engine; The matching module is used to perform real-time risk identification based on the cleaned real-time data stream and the initialized rule engine using rule matching technology, match the rules for each record in the data stream, and calculate the risk score of each record based on the matching results to generate a risk identification result set; The early warning module is used to generate risk warning signals based on the risk identification result set and use early warning trigger technology to trigger warnings for records whose risk scores exceed the preset threshold, and push the corresponding warning information to users in real time, thereby realizing comprehensive analysis and judgment of traffic business management big data.
[0006] Another embodiment of the present application provides a method for comprehensive analysis and judgment of traffic business management big data, the method comprising: According to the needs of transportation business, use distributed data collection technology to access the multi-source big data stream of transportation business management in real time, generate standardized data streams, and generate cleaned real-time data streams based on the standardized data streams; According to the risk rule library, the predefined risk rules are loaded into the rule engine using dynamic rule loading technology, and the risk rules are converted into executable rule models. Based on the rule model, the execution order and priority of the risk rules are dynamically adjusted using the rule optimization algorithm to generate an initialized rule engine. Based on the cleaned real-time data stream and the initialized rule engine, real-time risk identification is performed using rule matching technology. Rules are matched for each record in the data stream, and based on the matching results, the risk score of each record is calculated to generate a risk identification result set. Based on the risk identification result set, early warning trigger technology is used to generate risk warning signals, and warnings are triggered for records whose risk scores exceed the preset threshold. The corresponding warning information is pushed to users in real time, realizing comprehensive analysis and judgment of traffic business management big data.
[0007] Optionally, the method of using distributed data acquisition technology to access multi-source big data streams of traffic business management in real time according to traffic business needs, generating standardized data streams, and generating cleaned real-time data streams based on the standardized data streams includes: According to the traffic business needs, configure the multi-source data access interface, perform parameter configuration and connection test on the multi-source data access interface, and generate the data access status after initialization, where the multi-source data at least includes database log data, API interface data, message queue data, and sensor data; According to the data access status after initialization, the multi-source data stream is captured in real time using streaming computing technology. The multi-source data stream is sharded by time window or data volume through distributed message queue and data sharding mechanism to generate sharded data blocks. Each data block is marked with a data sharding identifier to generate a sharding identification data set. According to the shard identification data set, the multi-source data is formatted uniformly. The raw data from different data sources is converted into a unified standardized format through data mapping technology and format conversion algorithm to generate a standardized data stream. Based on the standardized data stream, the data cleaning algorithm is used to remove noise data and redundant information. Among them, invalid data, duplicate data and erroneous data are identified and filtered through outlier detection technology and regular expression matching to generate a cleaned real-time data stream.
[0008] Optionally, the method of loading predefined risk rules into a rule engine using a dynamic rule loading technology based on a risk rule library, converting the risk rules into executable rule models, and dynamically adjusting the execution order and priority of the risk rules based on the rule models using a rule optimization algorithm to generate an initialized rule engine includes: According to the risk rule base, the predefined risk rules are parsed using rule parsing technology, wherein the rule text in the risk rule base is converted into a structured rule object through syntax analysis technology and rule description language to generate a rule object set; Based on the rule object set, the rule is converted into an executable rule model by using the rule compilation technology, wherein the rule object is compiled into an execution code by using the intermediate code generation technology and the optimization compiler to generate the executable rule model; According to the executable rule model, the rules are loaded into the rule engine using the rule cache mechanism. The LRU cache algorithm and distributed cache technology are used to cache the frequently used rules into the rule engine memory to generate the cache-optimized rule loading state. Based on the rule loading status after cache optimization, the rule execution order is dynamically adjusted using the rule optimization algorithm. The priority weight of each rule is calculated according to the historical execution frequency and business importance of the rule, and the rule execution order after priority adjustment is generated. The rule engine is configured according to the order of rule execution after priority adjustment. The state management technology and concurrency control mechanism are used to ensure that the rule engine can efficiently execute rules while supporting dynamic rule updates and real-time adjustments, thereby generating an initialized rule engine. Based on the initialized rule engine, the running status of the rule engine is verified. Among them, through rule coverage analysis and execution log monitoring, it is ensured that all rules are correctly loaded and executed. According to the verification results, the feedback mechanism is used to dynamically correct the rule loading and running process to generate the final initialized rule engine status.
[0009] Optionally, the real-time risk identification is performed using rule matching technology based on the cleaned real-time data stream and the initialized rule engine, rule matching is performed on each record in the data stream, and based on the matching result, the risk score of each record is calculated to generate a risk identification result set, including: According to the cleaned real-time data stream, the data stream is secondary sharded according to the time window or data volume. The data stream is divided into multiple data shards through the hash sharding algorithm and the time window division mechanism, and each data shard is assigned to a different computing node to generate a shard task set; Based on the sharding task set and the initialized rule engine, a multi-pattern matching algorithm is used to perform rule matching on each data shard. The Aho-Corasick algorithm and parallel computing technology are used to simultaneously execute rule matching tasks on multiple computing nodes to generate a preliminary matching result set. According to the preliminary matching result set, the matching results of each computing node are summarized, wherein the scattered matching results are merged into a global matching result set through a reduction algorithm and distributed aggregation technology, and the risk score of each record is calculated based on the global matching result set, wherein the risk score data set is generated through a weighted summation algorithm and risk level mapping; Based on the risk score data set, a risk identification result set is generated. Through the threshold judgment mechanism and risk classification technology, records with risk scores exceeding the preset threshold are marked as high risk to generate the final risk identification result set.
[0010] Optionally, based on the risk identification result set, a risk warning signal is generated using the warning trigger technology, a warning is triggered for records whose risk scores exceed a preset threshold, and the corresponding warning information is pushed to the user in real time, so as to realize the comprehensive analysis and judgment of the traffic business management big data, including: According to the risk identification result set, the risk score of each record is filtered, wherein the records with risk scores exceeding the preset threshold are screened out through the preset threshold and condition judgment logic to generate a high-risk record set; Based on the high-risk record set, create warning signals. Each high-risk record is converted into a warning event through event-driven technology, and a corresponding warning signal is generated. The warning signals are sorted by priority using the event queue to generate an ordered warning signal queue. According to the ordered warning signal queue, the warning signal is converted into standardized warning information, wherein the key information in the warning signal is filled into the warning template through the template engine and dynamic data filling technology to generate formatted warning information; Based on the formatted warning information, through multi-channel distribution technology and priority scheduling algorithm, the notification channel and push strategy are determined according to user preferences and warning urgency, and real-time push tasks are generated to push the warning information to relevant users.
[0011] Yet another embodiment of the present application provides a storage medium, wherein the storage medium stores a computer program, wherein the computer program is configured to execute any of the above methods when running.
[0012] Yet another embodiment of the present application provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute any of the methods described above.
[0013] Compared with the prior art, the present invention provides a comprehensive analysis and judgment system for traffic business management big data, including an access module, which is used to access multi-source big data streams in real time according to traffic business needs, and generate a cleaned real-time data stream; a loading module, which is used to load risk rules into a rule engine, convert risk rules into executable rule models, and dynamically adjust them to generate an initialized rule engine; a matching module, which is used to perform real-time risk identification and rule matching based on real-time data streams and rule engines, and generate a risk identification result set based on matching results; an early warning module, which is used to trigger early warnings based on the risk identification result set, and push the corresponding early warning information to users in real time, so as to realize comprehensive analysis and judgment of traffic business management big data, thereby being able to effectively analyze and judge potential risks in traffic business processes and improve the accuracy and timeliness of risk identification. BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 A schematic diagram of the structure of a traffic business management big data comprehensive analysis and judgment system provided by an embodiment of the present invention; Figure 2 A hardware structure block diagram of a computer terminal for a method for comprehensive analysis and judgment of traffic business management big data provided by an embodiment of the present invention; Figure 3 A flow chart of a method for comprehensive analysis and judgment of big data for traffic business management provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0015] The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, but should not be construed as limiting the present invention.
[0016] The embodiment of the present invention first provides a traffic business management big data comprehensive analysis and judgment system, see Figure 1 , the system may include: The access module 101 is used to access the multi-source big data stream of traffic business management in real time using distributed data acquisition technology according to traffic business needs, generate standardized data streams, and generate cleaned real-time data streams based on the standardized data streams; The loading module 102 is used to load the predefined risk rules into the rule engine according to the risk rule library by using the dynamic rule loading technology, convert the risk rules into executable rule models, and dynamically adjust the execution order and priority of the risk rules by using the rule optimization algorithm based on the rule model to generate an initialized rule engine; The matching module 103 is used to perform real-time risk identification using rule matching technology based on the cleaned real-time data stream and the initialized rule engine, perform rule matching on each record in the data stream, and calculate the risk score of each record based on the matching result to generate a risk identification result set; The early warning module 104 is used to generate a risk early warning signal based on the risk identification result set and use the early warning trigger technology to trigger an early warning for records whose risk scores exceed a preset threshold, and push the corresponding early warning information to the user in real time, so as to realize comprehensive analysis and judgment of traffic business management big data.
[0017] It can be seen that according to the needs of traffic business, multi-source big data streams of traffic business management are accessed in real time to generate standardized data streams, and based on the standardized data streams, cleaned real-time data streams are generated; according to the risk rule library, predefined risk rules are loaded into the rule engine, and the risk rules are converted into executable rule models, and based on the rule model, the initialized rule engine is generated; according to the real-time data stream and the rule engine, real-time risk identification is performed, and rule matching is performed for each record in the data stream, and based on the matching results, the risk score of each record is calculated to generate a risk identification result set; according to the risk identification result set, the early warning trigger technology is used to generate risk warning signals, and the corresponding early warning information is pushed to the user in real time, so as to realize the comprehensive analysis and judgment of traffic business management big data, so as to effectively analyze and judge the potential risks in the traffic business process and improve the accuracy and timeliness of risk identification. For detailed explanations of the content and steps of each module, see the following one of the comprehensive analysis and judgment methods of traffic business management big data.
[0018] Correspondingly, another embodiment of the present invention provides a method for comprehensive analysis and judgment of traffic business management big data, which can be applied to electronic devices such as computer terminals, specifically ordinary computers, etc.
[0019] The following describes it in detail by taking running on a computer terminal as an example. Figure 2 The hardware structure block diagram of a computer terminal for a method for comprehensive analysis and judgment of traffic business management big data provided by an embodiment of the present invention. Figure 2 As shown, the computer device includes a processor, a memory, and a network interface connected via a system bus, wherein the memory may include a non-volatile storage medium and an internal memory.
[0020] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any one of the comprehensive analysis and judgment methods for traffic business management big data.
[0021] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.
[0022] The internal memory provides an environment for the operation of computer programs in non-volatile storage media. When the computer program is executed by the processor, the processor can execute any comprehensive analysis and judgment method of traffic business management big data.
[0023] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will understand that Figure 2 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.
[0024] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0025] See also Figure 3 The embodiment of the present invention provides a method for comprehensive analysis and judgment of traffic business management big data, which may include the following steps: S301, according to the traffic business needs, using distributed data collection technology to access the multi-source big data stream of traffic business management in real time, generate a standardized data stream, and generate a cleaned real-time data stream based on the standardized data stream; In the first step of the method, distributed data acquisition technology is used to access multi-source large data streams from multiple traffic-related sources in real time according to traffic business needs. Specifically, these data include real-time traffic flow information collected by road sensors, vehicle location and speed data provided by the GPS system, traffic light status information, and image data from surveillance cameras. These data come from different traffic management systems and equipment, and are diverse and real-time. By standardizing these heterogeneous data, the system can ensure the consistency of various types of data in format and representation, thereby generating standardized data streams. Based on the standardized data stream, data cleaning technology is used to denoise, correct and format the data, eliminate redundant and erroneous information, and generate high-quality cleaned real-time data streams. This process not only improves the availability and readability of the data, but also provides a solid foundation for the subsequent risk identification link, ensuring the accuracy and reliability of the analysis results.
[0026] This step has far-reaching significance for traffic business management. First, by accessing traffic big data in real time, the system can obtain instant information on traffic flow, speed, accidents and signal status, which provides the necessary basic data support for traffic dispatch, traffic flow management and emergency response. Secondly, the standardized and cleaned data stream can effectively reduce the analysis errors caused by data quality problems and ensure that the decisions in the risk identification process are based on accurate information. This high-quality data processing improves the real-time monitoring capabilities of traffic congestion, accident risks and abnormal behaviors, which helps traffic management departments to respond quickly, optimize traffic signal control, dispatch vehicles and issue early warning information, thereby ensuring road safety and improving traffic efficiency. Through these means, the traffic management system can better respond to various challenges in a dynamic environment and achieve more intelligent and efficient traffic management.
[0027] Specifically, a multi-source data access interface can be configured according to traffic business needs, and parameter configuration and connection test can be performed on the multi-source data access interface to generate an initialized data access state, wherein the multi-source data includes at least database log data, API interface data, message queue data, and sensor data; In this phase, we first determine the type of data source to be accessed and design the access interface based on the specific transportation business needs. This includes being able to access log data in the database, calling API interfaces to obtain real-time data, reading asynchronous messages from message queues, and collecting sensor data. The access interface is configured through parameter settings, including connection strings, access permissions, etc., and connection tests are performed to ensure that data can be successfully read.
[0028] This step ensures that the data acquisition system can reliably access multiple data sources, laying the foundation for the real-time capture of subsequent data streams. In addition, connection testing can promptly identify and resolve problems in interface configuration, avoiding data loss or errors in subsequent processes.
[0029] When you start configuring a multi-source data access interface, you first need to communicate with the traffic management department to clarify the required data sources and data types. For example, a traffic flow management system may need to access road sensor data (sensor data), traffic light status (API interface data), vehicle GPS data (message queue data), and traffic accident records (database log data). Then, the development team designs the data access interface based on these requirements and uses the appropriate database driver or API request library to connect to each data source.
[0030] Next, configure the parameters. For database log data, engineers need to set the database connection string, including information such as the host address, port, database name, user name and password, and implement dynamic reading of these parameters in the code. When configuring the API interface, you need to define the API URL, request method (such as GET or POST), request header information, and required parameters. In addition, in order to improve the security of the interface, you may also need to configure an authentication mechanism, such as OAuth or API key. After configuration is complete, the system will attempt to connect to all data sources to ensure that the connection parameters are valid.
[0031] Finally, connection testing is a critical step. By writing connection test code, the system will try to connect to each data source and obtain some test data to verify the validity of the interface. After a successful connection, the system will record the connection status and generate a data access status report after initialization, indicating which data sources are successfully connected, which fail, and the reasons for the failure, so that engineers can quickly locate the problem and make adjustments.
[0032] According to the data access status after initialization, the multi-source data stream is captured in real time using streaming computing technology. The multi-source data stream is sharded by time window or data volume through distributed message queue and data sharding mechanism to generate sharded data blocks. Each data block is marked with a data sharding identifier to generate a sharding identification data set. In this phase, the system uses streaming computing technology to capture the data streams transmitted by various data sources in real time. By using distributed message queues (such as Kafka) and data sharding mechanisms, the received data streams are sharded according to the preset time window or data volume. Each generated data block will be marked with a unique shard identifier for subsequent tracking and processing, while forming a complete shard identification data set.
[0033] This process enables data to be processed and analyzed efficiently. By sharding the data stream, not only can data be processed in parallel, but the system's processing capacity can also be improved under high load conditions, performance bottlenecks can be avoided, and data real-time can be ensured.
[0034] After confirming the data access status, the system will start to capture multi-source data streams in real time through streaming computing technology. Using distributed message queues such as Apache Kafka or RabbitMQ, the system can asynchronously read data streams from various data sources. In this process, the system will create multiple consumer instances, which will listen to the message topic or queue of each data source respectively, to achieve parallel processing of data streams.
[0035] After data is captured, the system will shard the received data according to the preset time window (such as every 10 seconds) or data volume (such as every 1,000 records). Through time window processing, the system can aggregate the data generated in each time period into a data block, and through data volume control, it can ensure that each data block is not too large, which is convenient for subsequent processing. For example, if 1,500 data are collected in each 10-second window, these data will be divided into two data blocks, the first 1,000 and the last 500.
[0036] Each generated data block is assigned a unique shard identifier, which is usually a combination of the access timestamp and sequence number, so that it can be easily tracked and managed during subsequent processing. At the same time, the system aggregates all shard data blocks to form a complete shard identification data set, which lays the foundation for the subsequent format unification and cleaning process.
[0037] According to the shard identification data set, the multi-source data is formatted uniformly. The raw data from different data sources is converted into a unified standardized format through data mapping technology and format conversion algorithm to generate a standardized data stream. In this step, the raw data received from different data sources are formatted in a unified manner. Through data mapping technology and format conversion algorithms, data from different sources (such as JSON, XML, CSV, etc.) are converted into a unified standardized structure, which makes subsequent processing easier and more efficient, thereby generating a standardized data stream for further analysis.
[0038] The unification of data formats lays the foundation for further data cleaning and analysis, and can effectively avoid errors caused by inconsistent data formats. In addition, standardized data flows can improve data processing efficiency and simplify rule application and risk identification processes.
[0039] In this phase, the system will unify the format of the raw data received from different data sources. To achieve this goal, using data mapping technology, you first need to define a standard data model, usually based on the data requirements determined by the team during the preliminary analysis process. The standard model includes the necessary field names, data types, and constraints, which enables the subsequent data mapping process to be executed efficiently.
[0040] Taking the JSON format as an example, assuming that the field name in the road sensor data is "sensor_id" and the field name in the traffic light status data is "signal_id", the system will unify these two fields into "device_id". During the data conversion process, the system will traverse each record and map and convert fields from different sources according to predetermined rules. At the same time, for unstructured data such as text data, the system may need to use regular expressions or text parsing algorithms to extract and standardize it into a structured format for subsequent processing.
[0041] After data mapping and format conversion, the system will generate standardized data streams. These data streams are not only consistent in format, but also convenient for subsequent data cleaning and analysis. The output of standardized data streams will provide a basis for various data processing and analysis, which is crucial to ensure the accuracy and authenticity of risk identification.
[0042] Based on the standardized data stream, the data cleaning algorithm is used to remove noise data and redundant information. Among them, invalid data, duplicate data and erroneous data are identified and filtered through outlier detection technology and regular expression matching to generate a cleaned real-time data stream.
[0043] At this stage, the system uses data cleaning algorithms to identify and remove noise and redundant information in the standardized data stream. The data is analyzed through outlier detection technology, and regular expression matching is used to filter out records that do not conform to the format and duplicate data to ensure the data quality of the final generated real-time data stream. Data cleaning is a key link in improving data quality. Removing invalid information, duplicate records and format errors in the data not only improves the accuracy of the data, but also provides a reliable basis for subsequent risk identification, ensuring the credibility and effectiveness of the results.
[0044] After the standardized data stream is generated, the next step is data cleaning. The data cleaning algorithm first performs noise detection on the standardized data stream. This process is usually achieved through outlier detection technology, such as using the Z-score method or the IQR method to identify and remove outliers to ensure the authenticity and reliability of the data. For example, if a record appears in the traffic flow data showing that the traffic volume of a certain section is 100,000 vehicles / hour, and the normal range is between 1,000 and 10,000 vehicles, the system will mark this record as an anomaly and filter it.
[0045] In addition to outlier removal, the system also applies regular expression matching technology to identify records with incorrect data formats. Assuming that a field should be in date format (such as YYYY-MM-DD), the system can quickly locate records that do not conform to this format through regular expressions and further correct or delete them. At the same time, in order to avoid data redundancy, the system will detect duplicate records in the data stream, usually by maintaining a hash table to store the unique identifiers of all processed records to ensure that they are not processed repeatedly.
[0046] Ultimately, through the above steps, the system will only retain valid and clean records, generating a cleaned real-time data stream. This data stream will provide a stable and reliable basis for subsequent risk identification and early warning, which not only improves data quality, but also enhances the efficiency and accuracy of subsequent processing.
[0047] S302, according to the risk rule library, the predefined risk rules are loaded into the rule engine using the dynamic rule loading technology, the risk rules are converted into executable rule models, and based on the rule model, the execution order and priority of the risk rules are dynamically adjusted using the rule optimization algorithm to generate an initialized rule engine; Based on the risk rule base, the predefined risk rules are loaded into the rule engine using dynamic rule loading technology. This process involves parsing, converting and optimizing the risk rules. First, the system extracts the relevant risk rules from the rule base that stores the risk rules, and loads them into the rule engine through dynamic rule loading technology. Then, the system converts the extracted risk rules into executable rule models to make them suitable for the execution environment of the rule engine. Next, based on the generated rule model, the rule optimization algorithm dynamically adjusts the execution order and priority of the risk rules, comprehensively considering the historical execution frequency, business importance and its impact on system performance, and finally forms an initialized rule engine so that the risk rules can be executed efficiently and accurately.
[0048] The core significance of this step is to ensure that the risk rules in the rule engine can flexibly adapt to the ever-changing traffic business needs and environmental conditions. Through dynamic loading and optimization, the rule engine can update the risk rules in real time to ensure that the identification and response to potential risks can remain efficient and accurate. This not only improves the flexibility and intelligence level of risk management, but also provides a solid technical foundation for subsequent risk identification and early warning.
[0049] Specifically, the predefined risk rules can be parsed using rule parsing technology according to the risk rule base, wherein the rule text in the risk rule base is converted into a structured rule object through syntax analysis technology and rule description language to generate a rule object set; At this stage, the system first extracts predefined risk rules from the rule base. These rules are usually stored in a specific text format, such as XML or JSON. Using rule parsing technology, the system will perform grammatical analysis on these texts and convert them into structured rule objects. The core of grammatical analysis is to use the parser to analyze the structure of each component in the text to ensure that the logical relationship and conditional expression of each rule are clear and usable.
[0050] For example, suppose there is a rule in the rule base: "If the traffic volume of a certain road section exceeds 5,000 vehicles / hour and the speed is less than 20 kilometers / hour, it is marked as a high risk of congestion." The system will parse this rule and identify the conditions (traffic volume and speed) and their corresponding logical relationships. After successful parsing, the system will generate a structured rule object, such as an object containing attributes such as "condition", "action" and "priority", and organize these objects into a rule object set for subsequent processing.
[0051] This set of rule objects not only improves the readability and maintainability of the rules, but also lays the foundation for subsequent rule compilation and execution, allowing the rule engine to process and apply these rules in a more effective way.
[0052] Based on the rule object set, the rule is converted into an executable rule model by using the rule compilation technology, wherein the rule object is compiled into an execution code by using the intermediate code generation technology and the optimization compiler to generate the executable rule model; After the rule object set is generated, the system will use rule compilation technology to convert these rule objects into executable rule models. First, the system will use intermediate code generation technology to convert structured rule objects into an intermediate representation form, which is usually not dependent on a specific programming language, making the subsequent compilation process more flexible.
[0053] For example, assuming that a rule object contains a "condition" part and an "action" part, after the intermediate code is generated, the system may generate a pseudocode-like form: "IF (traffic volume>5000) AND (vehicle speed<20) THEN mark as high risk of congestion." Next, the rule optimization compiler will further analyze and optimize this intermediate code and convert it into an executable rule model, usually expressed in the form of specific machine instructions or high-level languages, to prepare for subsequent execution of the rule engine.
[0054] This conversion process can not only improve execution efficiency, but also optimize the execution process of rules, such as by eliminating redundant conditions or simplifying logical judgments, thereby improving the performance and response speed of the overall system, which is especially important for real-time risk warning systems.
[0055] According to the executable rule model, the rules are loaded into the rule engine using the rule cache mechanism. The LRU cache algorithm and distributed cache technology are used to cache the frequently used rules into the rule engine memory to generate the cache-optimized rule loading state. After the executable rule model is generated, the system loads these rules into the rule engine through the rule cache mechanism. To ensure fast access to frequently used rules, the system uses the LRU (Least Recently Used) cache algorithm to save these rules in memory. The LRU algorithm regularly checks the most recently used records and expels rules that have not been used for a long time from the cache, thereby making room for newly loaded rules.
[0056] For example, if a specific risk rule has been called many times in the past hour, while another rule has hardly been used in the same time, the system will choose to keep the former in the cache and remove the latter. In this way, the system can reduce the delay in the rule matching process and improve the response speed of the rule engine. In addition, distributed caching technology can be used to implement cross-node loading and sharing of rules, so that in a distributed environment, different rule engine instances can quickly access the same rule set, ensuring high availability of the rules.
[0057] Through this optimized rule loading state, the rule engine can execute rules quickly and efficiently, significantly improving the performance of the system in complex, large-scale environments to adapt to the growing transportation business needs and risk management challenges.
[0058] Based on the rule loading status after cache optimization, the rule execution order is dynamically adjusted using the rule optimization algorithm. The priority weight of each rule is calculated according to the historical execution frequency and business importance of the rule, and the rule execution order after priority adjustment is generated. Based on the existing cache optimization status, the system will apply the rule optimization algorithm to dynamically adjust the order of rule execution. By analyzing the historical execution frequency, the system can identify which rules were frequently triggered in the past execution process and which were less used. This analysis result will help the system calculate the priority weight of each rule to ensure that more important or more commonly used rules can be executed first.
[0059] For example, if rule A was triggered 90 times in the past 100 executions, while rule B was only triggered 10 times, then the priority weight of rule A will be significantly higher than that of rule B. In this case, the system will put rule A at the front of the execution sequence to ensure that important rules can be processed quickly when risks are identified. In addition, business importance is also a consideration. For example, rules in certain high-risk areas will be given higher weights to ensure that they are prioritized even when resources are tight.
[0060] Through this dynamic adjustment mechanism, the rule engine can continuously optimize its own rule execution process at runtime to adapt to changing business scenarios and risk environments, and improve the system's flexibility and real-time responsiveness.
[0061] The rule engine is configured according to the order of rule execution after priority adjustment. The state management technology and concurrency control mechanism are used to ensure that the rule engine can efficiently execute rules while supporting dynamic rule updates and real-time adjustments, thereby generating an initialized rule engine. After the priority adjustment is completed, the system will rely on state management technology and concurrency control mechanism to make the final configuration of the rule engine. State management technology will help the system track the execution status of each rule in real time during the operation of the rule engine, including successful execution, failure, and abandonment. This information can help administrators evaluate and optimize the execution effect of the rules.
[0062] For example, when a rule is executed abnormally, the system can immediately record the status and may trigger an alarm or make automatic adjustments. The concurrency control mechanism ensures that in a multi-threaded or multi-process environment, the rule engine can efficiently and safely handle multiple requests, avoiding data conflicts and resource competition, thereby ensuring accurate rule execution and data consistency.
[0063] Through these configurations, the final generated initialized rule engine will have high flexibility and scalability, and can support dynamic updates and real-time adjustments of rules during operation. This capability enables enterprises to quickly respond to new risk challenges and maintain accurate control of the risk environment, thereby gaining a competitive advantage in a complex market environment.
[0064] Based on the initialized rule engine, the running status of the rule engine is verified. Among them, through rule coverage analysis and execution log monitoring, it is ensured that all rules are correctly loaded and executed. According to the verification results, the feedback mechanism is used to dynamically correct the rule loading and running process to generate the final initialized rule engine status.
[0065] After the rule engine is initialized, its running status must be verified to ensure that all loaded rules can be executed correctly. This process is achieved through rule coverage analysis and execution log monitoring. Rule coverage analysis will evaluate whether all predefined risk rules are applied in actual operation, while execution log monitoring tracks the execution of each rule, including the number of executions, execution results, and possible error messages. Through comprehensive analysis of this information, the system can identify rules that are not triggered or executed improperly. Based on the verification results, the feedback mechanism is used to dynamically correct the rule loading and running process. The system can timely update and optimize the rule loading strategy to ensure that the final generated rule engine state is efficient and accurate.
[0066] The significance of this step is to ensure the reliability and accuracy of the risk identification system. By verifying the actual running status of the rule engine, problems in rule loading and anomalies in the execution process can be discovered in a timely manner to avoid potential risks that are not captured or warnings are triggered incorrectly. This process not only helps to improve the performance of the overall system, but also enhances users' trust in risk warning results, ensuring that in actual applications, the system can respond to various risk scenarios in an efficient and accurate manner, thereby effectively reducing risk losses in business operations.
[0067] First, the system deploys automated monitoring tools to collect the execution logs of the rule engine in real time. The execution log records the execution status of each rule, including whether it is triggered, the triggering conditions, the execution time, and the execution results. For example, if a rule is set to detect sections of road with high traffic volume, and the rule has never appeared in the log, then this means that the rule has not been triggered in the actual data.
[0068] While collecting execution logs, the system will perform rule coverage analysis. This analysis compares the matching between the loaded rules and the risk scenarios involved in the actual transaction data, and evaluates which rules have not been effectively triggered. If it is found that some key rules have not been triggered, the system will actively generate alerts and extract relevant operation data for analysis.
[0069] Then, based on the results of coverage analysis and execution log monitoring, the system will introduce a feedback mechanism to dynamically modify the rule loading and running process. For example, for those rules that have not been triggered, the system may adjust their parameters or re-evaluate their importance and priority to ensure that future data flows can effectively activate these rules. In addition, the system can also regularly update the risk rule base, add new rules or modify existing rules to adapt to the ever-changing risk environment.
[0070] Ultimately, after continuous analysis and adjustment, the system will generate an optimized rule engine state, ensuring that it can not only maintain high efficiency in actual applications, but also accurately identify and respond to diverse risk scenarios. Such a dynamic correction mechanism ensures the long-term reliability and sustainability of the system.
[0071] S303, based on the cleaned real-time data stream and the initialized rule engine, use rule matching technology to perform real-time risk identification, perform rule matching on each record in the data stream, and calculate the risk score of each record based on the matching result to generate a risk identification result set; In this step, the system uses the cleaned real-time data stream and the initialized rule engine to perform real-time risk identification for each record in the data stream through rule matching technology. This process first involves matching the cleaned data stream with the predefined risk rules one by one. The rule engine will evaluate the characteristics and attributes of each record based on the set risk rules and identify potential risks. For example, if the traffic volume of a certain section of a traffic flow record exceeds the set threshold, the rule engine will regard the record as a high-risk section and calculate its risk score. The risk score calculation for each record combines multiple factors, such as traffic volume, speed, historical accident rate of the section, etc.
[0072] The implementation of this step greatly improves the efficiency and accuracy of real-time risk identification. By dynamically matching the cleaned real-time data stream with the risk rules in the rule engine, the system can quickly respond to potential risks and generate corresponding risk identification result sets based on specific matching results. This mechanism can not only effectively prevent traffic congestion and accidents, but also provide basic data support for subsequent risk assessment and early warning triggering, thereby optimizing the overall risk management process.
[0073] Specifically, the data stream can be secondary sharded by time window or data volume according to the cleaned real-time data stream. The data stream is divided into multiple data shards through the hash sharding algorithm and the time window division mechanism, and each data shard is assigned to a different computing node to generate a sharding task set. In this step, the cleaned real-time data stream needs to be split twice for subsequent parallel processing. By setting a time window or data volume threshold, the data stream is effectively divided into multiple smaller data shards. The application of the hash sharding algorithm can ensure that the data is evenly distributed to different computing nodes, which can improve the parallelism and efficiency of processing. Each shard will be marked for tracking, and eventually form a shard task set, ready to be distributed to each computing node for subsequent processing.
[0074] This sharding process is designed to improve the system's computing efficiency and processing capabilities, so that real-time data streams can take advantage of distributed computing when matching rules and quickly respond to potential risk events. In addition, data sharding helps with load balancing and avoids performance bottlenecks caused by a single node processing too much data, thereby ensuring the stability and efficiency of the entire system.
[0075] In this step, the system first needs to set the time window and data threshold in order to perform secondary sharding on the cleaned real-time data stream. For example, assuming that the cleaned data stream adds 10,000 records per minute, the system can set each shard to contain 500 records, or the time window of each shard is 1 minute. Then, the system processes this data through the hash sharding algorithm to ensure that the data is evenly distributed. Assuming that the unique identifier of a record is ID, by performing a hash operation on the ID, it can be determined which shard the record will be assigned to, which can effectively avoid data skew.
[0076] After sharding, the system will generate multiple data shards, each of which will be marked with a unique identifier for subsequent processing. For example, shard 1 may contain data with record IDs from 1 to 500, and shard 2 may contain data with record IDs from 501 to 1000. These shards are then distributed to different computing nodes for parallel processing. Through this mechanism, the system can improve computing efficiency while ensuring data integrity, thereby achieving real-time risk identification.
[0077] In addition, the system monitors and records the processing status of each data shard. By building a state management module, the system can track the processing progress of each shard in real time to ensure that there are no omissions or delays. For example, if computing node 1 encounters an abnormality when processing shard 1, the system will detect it immediately and reallocate the task of shard 1 to other nodes that are working normally to ensure stable processing and real-time response of the entire data flow.
[0078] Based on the sharding task set and the initialized rule engine, a multi-pattern matching algorithm is used to perform rule matching on each data shard. The Aho-Corasick algorithm and parallel computing technology are used to simultaneously execute rule matching tasks on multiple computing nodes to generate a preliminary matching result set. Using the prepared sharding task set and the initialized rule engine, the system applies a multi-pattern matching algorithm to each data shard for rule matching. Effective string matching algorithms such as the Aho-Corasick algorithm can simultaneously search for multiple patterns in multiple texts, making the risk identification process more efficient. By executing matching tasks in parallel on multiple computing nodes, the processing speed can be greatly improved, achieving real-time risk identification.
[0079] The implementation of this step can ensure that the system can quickly and effectively identify potential risk events. The combination of multi-pattern matching and parallel computing not only improves the efficiency of rule matching, but also allows the system to maintain real-time performance when facing large-scale data flows, ensuring timely response to risks.
[0080] In this step, the system connects each data shard to the initialized rule engine for rule matching. First, the system loads predefined risk rules on each computing node. These rules can be descriptions of various risk scenarios such as abnormal traffic flow patterns and frequent accident sections. By using the Aho-Corasick algorithm, multiple rules are compiled into automata at one time. The system can quickly find records that meet multiple rules when processing data shards. For example, suppose there is a rule for records with a traffic volume of more than 5,000 vehicles / hour on a certain section of road, and another rule for records with a speed of less than 20 kilometers / hour on a certain section of road. The Aho-Corasick algorithm can identify these records at the same time, greatly improving the matching efficiency.
[0081] During the data matching process, the system uses parallel computing technology to assign data shards to different computing nodes for processing. In this way, even if the amount of data is very large, each computing node can efficiently perform matching tasks at the same time. Each node independently matches the rules of the data shard to which it belongs and generates a preliminary matching result set. These result sets will include all matched high-risk records and their related information, such as the matched rule ID and the specific data information of the record.
[0082] Finally, the system will summarize the matching results of each node to form a preliminary and complete matching result set. During this process, the system will also evaluate the processing performance of each node to facilitate subsequent resource allocation and optimization. If a bottleneck is found in the processing of a certain node, you can consider improving its performance by increasing computing resources to ensure the real-time and accuracy of the overall risk control system.
[0083] According to the preliminary matching result set, the matching results of each computing node are summarized, wherein the scattered matching results are merged into a global matching result set through a reduction algorithm and distributed aggregation technology, and the risk score of each record is calculated based on the global matching result set, wherein the risk score data set is generated through a weighted summation algorithm and risk level mapping; In this step, the preliminary matching result sets returned by each computing node will be summarized through the reduction algorithm and distributed aggregation technology. The reduction algorithm will merge the returned results to form a unified global matching result set. Next, the system will calculate the risk score for each record based on the global matching result set, integrate the influence of each risk factor through the weighted sum algorithm, and perform risk level mapping to finally generate a risk score data set. The significance of this process is that it ensures that risk identification is not only a single dimension, but also integrates multiple risk factors to form a comprehensive assessment. By calculating the risk score of each record, the system can provide more accurate information support for real-time risk decision-making, thereby effectively improving the reliability of the risk warning system.
[0084] In this step, the system first needs to aggregate the preliminary matching result sets returned by each computing node to form a unified global matching result set. The system will apply a reduction algorithm, which is an effective way to merge distributed computing results. In this process, the system will merge the matching results according to the record ID, that is, merge the same records returned by all computing nodes to ensure that each record has only one entry in the global matching result set. For example, if both node 1 and node 2 find the same high-risk transaction record, only one merged record will be retained in the end.
[0085] After the summary is completed, the system needs to calculate the risk score for each record. This risk score will be comprehensively evaluated based on multiple factors, such as traffic volume, vehicle speed, historical accident rate of the road section, etc. The system will use a weighted summation algorithm to weight each risk factor and assign different weight values according to business rules to derive the final risk score. For example, in a traffic flow record, if the traffic volume accounts for 70% of the weight and the speed accounts for 30% of the weight, the system will calculate based on these weights and finally derive a specific risk value.
[0086] Through risk level mapping, the system can convert risk scores into corresponding risk levels. For example, risk scores greater than 100 are set as high risk, 50 to 100 are set as medium risk, and less than 50 are set as low risk. Such classification can help the subsequent early warning system to more quickly identify high-risk records and handle them accordingly. Therefore, through this process, the system can effectively provide quantitative data support for subsequent decision-making and risk management.
[0087] Based on the risk score data set, a risk identification result set is generated. Through the threshold judgment mechanism and risk classification technology, records with risk scores exceeding the preset threshold are marked as high risk to generate the final risk identification result set.
[0088] In this step, the system will use the risk score data set to generate a risk identification result set. Using the threshold judgment mechanism, the system will check whether the risk score of each record exceeds the preset threshold. If it exceeds, it will be marked as high risk. In addition, the system will also use risk classification technology to classify risks according to different transportation business needs and generate a final risk identification result set. The implementation of this method can identify high-risk records in a timely and effective manner, and provide basic data for subsequent risk control and response measures. Through clear risk classification, decision makers can quickly focus on high-risk data that may affect system safety, thereby improving the efficiency of overall risk management.
[0089] In this step, the system will use the risk score data set to generate the final risk identification result set. First, the system needs to set one or more preset risk thresholds. For example, if the threshold is set to 80, if the risk score of a record exceeds 80, the record will be marked as high risk. The setting of this threshold can be adjusted based on previous historical data analysis and statistical results to ensure that possible risks can be effectively identified.
[0090] Next, the system uses the threshold judgment mechanism to check each record in the risk score data set one by one. This process will involve simple conditional judgment logic, comparing the risk score of each record with the preset threshold, and the records that meet the conditions will be marked as "high risk". For example, if the risk score of a traffic flow record is 90, the system will mark it as high risk and record it in the risk identification result set.
[0091] Through risk classification technology, the system will further classify high-risk records. For example, high-risk records can be divided into categories such as "high traffic flow sections" and "low speed sections". Different categories of risk records may require different processing strategies. Finally, the system will generate a result set containing all high-risk records and pass this result set to the subsequent warning trigger and notification module to quickly respond to potential risk issues. This mechanism not only improves the accuracy of risk identification, but also provides effective support for subsequent warnings and processing.
[0092] S304, based on the risk identification result set, use the warning trigger technology to generate a risk warning signal, trigger a warning for records whose risk scores exceed the preset threshold, and push the corresponding warning information to the user in real time, so as to realize the comprehensive analysis and judgment of the traffic business management big data.
[0093] This step mainly involves the whole process of generating risk warning signals based on the risk identification result set. First, the system analyzes the risk score of each record and compares it with the preset risk threshold to screen out those records whose risk scores exceed the threshold. This screening process uses conditional judgment logic, which enables the system to quickly identify high-risk behaviors that really need attention. Then, the system uses early warning trigger technology to mark and process these high-risk records and generate corresponding risk warning signals. The warning signal will contain detailed information, such as the specific location of the high-risk section, traffic volume, speed and related accident information. Finally, these warning signals will be pushed to relevant users in real time to ensure that risk information can be delivered in a timely manner.
[0094] The implementation of this step greatly enhances the real-time response capability of the system, enabling traffic management departments to respond immediately to potential risk events. This efficient early warning mechanism ensures the rapid identification and processing of high-risk records, reducing potential losses caused by delays. In addition, by pushing specific early warning information to relevant users, decision makers can understand the risk status in a timely manner and formulate corresponding risk control measures, thereby better ensuring the safety and efficiency of the transportation system.
[0095] Specifically, the risk score of each record may be filtered according to the risk identification result set, wherein the records with risk scores exceeding the preset threshold are screened out through the preset threshold and conditional judgment logic to generate a high-risk record set; In this step, the system first needs to clearly define the preset risk threshold, which is usually based on historical data analysis and risk assessment models. This threshold can be a fixed value, such as 80, or dynamically adjusted according to traffic business needs. The system will traverse the risk identification result set and check each record one by one. Every time the system reads a record, it will extract the risk score of the record and compare it with the set threshold. When the system finds that the risk score of a record exceeds the threshold, it will add the record to the high-risk record set.
[0096] For example, suppose the system's risk identification result set contains five records, whose risk scores are 50, 85, 90, 70, and 100. If the threshold set by the system is 80, the system will first check record 1 (50). Since it does not exceed the threshold, it will continue to check record 2 (85) and find that its score exceeds 80, so it will be added to the high-risk set. Next, record 3 (90) and record 5 (100) also meet the addition conditions, and the final high-risk record set will contain records 2, 3, and 5.
[0097] The key to this filtering process lies in the efficiency and accuracy of the conditional judgment logic. In order to improve processing efficiency, the system can build an optimized traversal algorithm, such as using parallel processing technology to check multiple records at the same time to speed up the generation of high-risk record sets. In this way, the system can analyze a large amount of data in a short period of time to ensure that potential risks are identified in a timely manner.
[0098] Based on the high-risk record set, create warning signals. Each high-risk record is converted into a warning event through event-driven technology, and a corresponding warning signal is generated. The warning signals are sorted by priority using the event queue to generate an ordered warning signal queue. In this step, the system will build warning events record by record from the set of high-risk records. Each high-risk record will be encapsulated into a warning event object, which contains key information such as record ID, risk type, risk score, and related road section information. The system uses event-driven technology, which means that when a high-risk record is identified, the event handler will be triggered immediately. For example, when the risk score of a record is detected to exceed the threshold, the system will automatically create a warning event, such as "The traffic flow on road section X triggered two warnings within 10 minutes, and the traffic flow was 6,000 vehicles / hour, exceeding the set threshold." Doing so not only responds quickly to potential risks, but also ensures timely communication of information.
[0099] These warning events are then added to an event queue. Since different high-risk records may have different levels of urgency, the system needs to prioritize these events. A set of rules can be set, such as setting the priority of records with high risk scores to "high", records with risk scores near the threshold to "medium", and records with relatively the lowest risk scores to "low". In this way, the system can effectively distinguish which events need to be handled immediately and which can be handled later, thereby optimizing resource allocation.
[0100] Once all high-risk records are converted into warning signals and sorted by priority, these ordered warning signal queues will provide a clear logical order for the processing of subsequent steps. The system can store these events through data structures such as first-in-first-out (FIFO) or priority queues to ensure that the system prioritizes the most urgent warning signals, thereby reducing the risks faced by the enterprise.
[0101] According to the ordered warning signal queue, the warning signal is converted into standardized warning information, wherein the key information in the warning signal is filled into the warning template through the template engine and dynamic data filling technology to generate formatted warning information; In this step, the system uses the information in the warning signal queue to generate standardized warning information. First, the system predefines a structured warning information template, which usually includes fields such as "warning type", "related road section", "risk score", "timestamp", etc. By using the template engine, the system can flexibly fill in different warning event information to ensure that the format of each warning message is consistent and easy to understand. For example, the template may be defined as "Warning: There is a risk of traffic flow on road section {road section ID}, the traffic flow is {traffic flow}, and the risk score is {risk score}."
[0102] The system will traverse the ordered warning signal queue, extract the necessary information from each warning signal, and use dynamic data filling technology to fill this information into the warning template. Assuming that the first warning signal in the queue is "The traffic volume of road section A is 6,000 vehicles / hour, and the risk score is 95", the formatted information generated after filling will be "Warning: There is a risk of traffic flow on road section A, the traffic volume is 6,000 vehicles / hour, and the risk score is 95." In this way, the system can generate a large amount of efficient, accurate, and structured warning information.
[0103] The generated formatted warning information will be stored in a new collection and prepared for subsequent distribution. This standardized warning information not only improves the accuracy of data processing, but also ensures that different users can quickly understand the warning content in the same format, thereby enhancing response efficiency.
[0104] Based on the formatted warning information, through multi-channel distribution technology and priority scheduling algorithm, the notification channel and push strategy are determined according to user preferences and warning urgency, and real-time push tasks are generated to push the warning information to relevant users.
[0105] In this step, the system needs to decide how to distribute the warning information to users based on the formatted warning information and the user's preferences. First, the system collects the user's preference information, such as whether they want to receive warnings via email, SMS, APP push, or other methods. This process can be completed through a one-time setting in the user interface or subsequent dynamic updates. Depending on the user's needs, the system will select different channels for pushing the same warning information.
[0106] In order to ensure that high-priority warning messages are delivered first, the system will use a priority scheduling algorithm. This means that when the system has multiple warning messages to send, high-risk events will be processed first. For example, if user A receives a high-risk warning via SMS, and user B wants to receive a relatively low-risk warning via email, the system will schedule according to this setting.
[0107] When the real-time push task is finally generated, the system will create a task queue. Each task includes information such as the content to be sent, the target user, the receiving method and the sending time. When the system processes each task in the task queue one by one, it will distribute it according to the preset sending strategy. By utilizing multi-channel distribution technology, the system can ensure that relevant users can obtain important risk warning information immediately, so that they can take necessary actions in the first time and reduce the impact of potential risks.
[0108] It can be seen that according to the needs of transportation business, multi-source big data streams of transportation business management are accessed in real time to generate standardized data streams, and based on the standardized data streams, cleaned real-time data streams are generated; according to the risk rule library, the predefined risk rules are loaded into the rule engine, the risk rules are converted into executable rule models, and based on the rule model, an initialized rule engine is generated; according to the real-time data stream and the rule engine, real-time risk identification is performed, and rules are matched for each record in the data stream, and based on the matching results, the risk score of each record is calculated to generate a risk identification result set; according to the risk identification result set, risk warning signals are generated using warning trigger technology, and the corresponding warning information is pushed to users in real time, so as to realize comprehensive analysis and judgment of transportation business management big data, so as to effectively analyze and judge the potential risks in the transportation business process, and improve the accuracy and timeliness of risk identification.
[0109] An embodiment of the present invention further provides a storage medium, in which a computer program is stored, wherein the computer program is configured to execute the steps of any of the above method embodiments when running.
[0110] Specifically, in this embodiment, the above storage medium may be configured to store a computer program for performing the following steps: S301, according to the traffic business needs, using distributed data collection technology to access the multi-source big data stream of traffic business management in real time, generate a standardized data stream, and generate a cleaned real-time data stream based on the standardized data stream; S302, according to the risk rule library, the predefined risk rules are loaded into the rule engine using the dynamic rule loading technology, the risk rules are converted into executable rule models, and based on the rule model, the execution order and priority of the risk rules are dynamically adjusted using the rule optimization algorithm to generate an initialized rule engine; S303, based on the cleaned real-time data stream and the initialized rule engine, use rule matching technology to perform real-time risk identification, perform rule matching on each record in the data stream, and calculate the risk score of each record based on the matching result to generate a risk identification result set; S304, based on the risk identification result set, use the warning trigger technology to generate a risk warning signal, trigger a warning for records whose risk scores exceed the preset threshold, and push the corresponding warning information to the user in real time, so as to realize the comprehensive analysis and judgment of the traffic business management big data.
[0111] It can be seen that according to the needs of transportation business, multi-source big data streams of transportation business management are accessed in real time to generate standardized data streams, and based on the standardized data streams, cleaned real-time data streams are generated; according to the risk rule library, the predefined risk rules are loaded into the rule engine, the risk rules are converted into executable rule models, and based on the rule model, an initialized rule engine is generated; according to the real-time data stream and the rule engine, real-time risk identification is performed, and rules are matched for each record in the data stream, and based on the matching results, the risk score of each record is calculated to generate a risk identification result set; according to the risk identification result set, risk warning signals are generated using warning trigger technology, and the corresponding warning information is pushed to users in real time, so as to realize comprehensive analysis and judgment of transportation business management big data, so as to effectively analyze and judge the potential risks in the transportation business process, and improve the accuracy and timeliness of risk identification.
[0112] An embodiment of the present invention further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.
[0113] Specifically, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.
[0114] Specifically, in this embodiment, the processor may be configured to perform the following steps through a computer program: S301, according to the traffic business needs, using distributed data collection technology to access the multi-source big data stream of traffic business management in real time, generate a standardized data stream, and generate a cleaned real-time data stream based on the standardized data stream; S302, according to the risk rule library, the predefined risk rules are loaded into the rule engine using the dynamic rule loading technology, the risk rules are converted into executable rule models, and based on the rule model, the execution order and priority of the risk rules are dynamically adjusted using the rule optimization algorithm to generate an initialized rule engine; S303, based on the cleaned real-time data stream and the initialized rule engine, use rule matching technology to perform real-time risk identification, perform rule matching on each record in the data stream, and calculate the risk score of each record based on the matching result to generate a risk identification result set; S304, based on the risk identification result set, use the warning trigger technology to generate a risk warning signal, trigger a warning for records whose risk scores exceed the preset threshold, and push the corresponding warning information to the user in real time, so as to realize the comprehensive analysis and judgment of the traffic business management big data.
[0115] It can be seen that according to the needs of transportation business, multi-source big data streams of transportation business management are accessed in real time to generate standardized data streams, and based on the standardized data streams, cleaned real-time data streams are generated; according to the risk rule library, the predefined risk rules are loaded into the rule engine, the risk rules are converted into executable rule models, and based on the rule model, an initialized rule engine is generated; according to the real-time data stream and the rule engine, real-time risk identification is performed, and rules are matched for each record in the data stream, and based on the matching results, the risk score of each record is calculated to generate a risk identification result set; according to the risk identification result set, risk warning signals are generated using warning trigger technology, and the corresponding warning information is pushed to users in real time, so as to realize comprehensive analysis and judgment of transportation business management big data, so as to effectively analyze and judge the potential risks in the transportation business process, and improve the accuracy and timeliness of risk identification.
[0116] The above describes in detail the structure, features and effects of the present invention based on the embodiments shown in the drawings. The above is only a preferred embodiment of the present invention, but the present invention is not limited to the scope of implementation shown in the drawings. Any changes made according to the concept of the present invention, or modifications to equivalent embodiments with equivalent changes, which still do not exceed the spirit covered by the description and drawings, should be within the protection scope of the present invention.
Claims
1. A comprehensive analysis and judgment system for traffic business management big data, characterized in that: The system comprises: The access module is used to access the multi-source big data stream of traffic business management in real time using distributed data acquisition technology according to traffic business needs, generate standardized data streams, and generate cleaned real-time data streams based on the standardized data streams; The loading module is used to load the predefined risk rules into the rule engine according to the risk rule library by using the dynamic rule loading technology, convert the risk rules into executable rule models, and dynamically adjust the execution order and priority of the risk rules based on the rule model by using the rule optimization algorithm to generate an initialized rule engine; The matching module is used to perform real-time risk identification based on the cleaned real-time data stream and the initialized rule engine using rule matching technology, match the rules for each record in the data stream, and calculate the risk score of each record based on the matching results to generate a risk identification result set; The early warning module is used to generate risk warning signals based on the risk identification result set and use early warning trigger technology to trigger warnings for records whose risk scores exceed the preset threshold, and push the corresponding warning information to users in real time, thereby realizing comprehensive analysis and judgment of traffic business management big data.
2. The system according to claim 1, characterized in that The access module is specifically used for: According to the traffic business needs, configure the multi-source data access interface, perform parameter configuration and connection test on the multi-source data access interface, and generate the data access status after initialization, where the multi-source data at least includes database log data, API interface data, message queue data, and sensor data; According to the data access status after initialization, the multi-source data stream is captured in real time using streaming computing technology. The multi-source data stream is sharded by time window or data volume through distributed message queue and data sharding mechanism to generate sharded data blocks. Each data block is marked with a data sharding identifier to generate a sharding identification data set. According to the shard identification data set, the multi-source data is formatted uniformly. The raw data from different data sources is converted into a unified standardized format through data mapping technology and format conversion algorithm to generate a standardized data stream. Based on the standardized data stream, the data cleaning algorithm is used to remove noise data and redundant information. Among them, invalid data, duplicate data and erroneous data are identified and filtered through outlier detection technology and regular expression matching to generate a cleaned real-time data stream.
3. The system according to claim 2, characterized in that The loading module is specifically used for: According to the risk rule base, the predefined risk rules are parsed using rule parsing technology, wherein the rule text in the risk rule base is converted into a structured rule object through syntax analysis technology and rule description language to generate a rule object set; Based on the rule object set, the rule is converted into an executable rule model by using the rule compilation technology, wherein the rule object is compiled into an execution code by using the intermediate code generation technology and the optimization compiler to generate the executable rule model; According to the executable rule model, the rules are loaded into the rule engine using the rule cache mechanism. The LRU cache algorithm and distributed cache technology are used to cache the frequently used rules into the rule engine memory to generate the cache-optimized rule loading state. Based on the rule loading status after cache optimization, the rule execution order is dynamically adjusted using the rule optimization algorithm. The priority weight of each rule is calculated according to the historical execution frequency and business importance of the rule, and the rule execution order after priority adjustment is generated. The rule engine is configured according to the order of rule execution after priority adjustment. The state management technology and concurrency control mechanism are used to ensure that the rule engine can efficiently execute rules while supporting dynamic rule updates and real-time adjustments, thereby generating an initialized rule engine. Based on the initialized rule engine, the running status of the rule engine is verified. Among them, through rule coverage analysis and execution log monitoring, it is ensured that all rules are correctly loaded and executed. According to the verification results, the feedback mechanism is used to dynamically correct the rule loading and running process to generate the final initialized rule engine status.
4. A method for comprehensive analysis and judgment of traffic business management big data, characterized in that: The method comprises: According to the needs of transportation business, use distributed data collection technology to access the multi-source big data stream of transportation business management in real time, generate standardized data streams, and generate cleaned real-time data streams based on the standardized data streams; According to the risk rule library, the predefined risk rules are loaded into the rule engine using dynamic rule loading technology, and the risk rules are converted into executable rule models. Based on the rule model, the execution order and priority of the risk rules are dynamically adjusted using the rule optimization algorithm to generate an initialized rule engine. Based on the cleaned real-time data stream and the initialized rule engine, real-time risk identification is performed using rule matching technology. Rules are matched for each record in the data stream, and based on the matching results, the risk score of each record is calculated to generate a risk identification result set. Based on the risk identification result set, early warning trigger technology is used to generate risk warning signals, and warnings are triggered for records whose risk scores exceed the preset threshold. The corresponding warning information is pushed to users in real time, realizing comprehensive analysis and judgment of traffic business management big data.
5. The method according to claim 4, characterized in that The method uses distributed data collection technology to access multi-source big data streams of traffic business management in real time according to traffic business needs, generates standardized data streams, and generates cleaned real-time data streams based on the standardized data streams, including: According to the traffic business needs, configure the multi-source data access interface, perform parameter configuration and connection test on the multi-source data access interface, and generate the data access status after initialization, where the multi-source data at least includes database log data, API interface data, message queue data, and sensor data; According to the data access status after initialization, the multi-source data stream is captured in real time using streaming computing technology. The multi-source data stream is sharded by time window or data volume through distributed message queue and data sharding mechanism to generate sharded data blocks. Each data block is marked with a data sharding identifier to generate a sharding identification data set. According to the shard identification data set, the multi-source data is formatted uniformly. The raw data from different data sources is converted into a unified standardized format through data mapping technology and format conversion algorithm to generate a standardized data stream. Based on the standardized data stream, the data cleaning algorithm is used to remove noise data and redundant information. Among them, invalid data, duplicate data and erroneous data are identified and filtered through outlier detection technology and regular expression matching to generate a cleaned real-time data stream.
6. The method according to claim 5, characterized in that The method comprises: loading the predefined risk rules into the rule engine by using the dynamic rule loading technology according to the risk rule library, converting the risk rules into an executable rule model, and dynamically adjusting the execution order and priority of the risk rules by using the rule optimization algorithm based on the rule model to generate an initialized rule engine, including: According to the risk rule base, the predefined risk rules are parsed using rule parsing technology, wherein the rule text in the risk rule base is converted into a structured rule object through syntax analysis technology and rule description language to generate a rule object set; Based on the rule object set, the rule is converted into an executable rule model by using the rule compilation technology, wherein the rule object is compiled into an execution code by using the intermediate code generation technology and the optimization compiler to generate the executable rule model; According to the executable rule model, the rules are loaded into the rule engine using the rule cache mechanism. The LRU cache algorithm and distributed cache technology are used to cache the frequently used rules into the rule engine memory to generate the cache-optimized rule loading state. Based on the rule loading status after cache optimization, the rule execution order is dynamically adjusted using the rule optimization algorithm. The priority weight of each rule is calculated according to the historical execution frequency and business importance of the rule, and the rule execution order after priority adjustment is generated. The rule engine is configured according to the order of rule execution after priority adjustment. The state management technology and concurrency control mechanism are used to ensure that the rule engine can efficiently execute rules while supporting dynamic rule updates and real-time adjustments, thereby generating an initialized rule engine. Based on the initialized rule engine, the running status of the rule engine is verified. Among them, through rule coverage analysis and execution log monitoring, it is ensured that all rules are correctly loaded and executed. According to the verification results, the feedback mechanism is used to dynamically correct the rule loading and running process to generate the final initialized rule engine status.
7. The method according to claim 6, characterized in that According to the cleaned real-time data stream and the initialized rule engine, the rule matching technology is used to perform real-time risk identification, rule matching is performed on each record in the data stream, and based on the matching results, the risk score of each record is calculated to generate a risk identification result set, including: According to the cleaned real-time data stream, the data stream is secondary sharded according to the time window or data volume. The data stream is divided into multiple data shards through the hash sharding algorithm and the time window division mechanism, and each data shard is assigned to a different computing node to generate a shard task set; Based on the sharding task set and the initialized rule engine, a multi-pattern matching algorithm is used to perform rule matching on each data shard. The Aho-Corasick algorithm and parallel computing technology are used to simultaneously execute rule matching tasks on multiple computing nodes to generate a preliminary matching result set. According to the preliminary matching result set, the matching results of each computing node are summarized, wherein the scattered matching results are merged into a global matching result set through a reduction algorithm and distributed aggregation technology, and the risk score of each record is calculated based on the global matching result set, wherein the risk score data set is generated through a weighted summation algorithm and risk level mapping; Based on the risk score data set, a risk identification result set is generated. Through the threshold judgment mechanism and risk classification technology, records with risk scores exceeding the preset threshold are marked as high risk to generate the final risk identification result set.
8. The method according to claim 7, characterized in that According to the risk identification result set, the risk warning signal is generated by using the warning trigger technology, the record with the risk score exceeding the preset threshold is triggered, and the corresponding warning information is pushed to the user in real time, so as to realize the comprehensive analysis and judgment of the traffic business management big data, including: According to the risk identification result set, the risk score of each record is filtered, wherein the records with risk scores exceeding the preset threshold are screened out through the preset threshold and condition judgment logic to generate a high-risk record set; Based on the high-risk record set, create warning signals. Each high-risk record is converted into a warning event through event-driven technology, and a corresponding warning signal is generated. The warning signals are sorted by priority using the event queue to generate an ordered warning signal queue. According to the ordered warning signal queue, the warning signal is converted into standardized warning information, wherein the key information in the warning signal is filled into the warning template through the template engine and dynamic data filling technology to generate formatted warning information; Based on the formatted warning information, through multi-channel distribution technology and priority scheduling algorithm, the notification channel and push strategy are determined according to user preferences and warning urgency, and real-time push tasks are generated to push the warning information to relevant users.
9. A storage medium, characterized in that: The storage medium stores a computer program, wherein the computer program is configured to execute the method according to any one of claims 4 to 8 when executed.
10. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the method according to any one of claims 4 to 8.
Citation Information
Patent Citations
Real-time risk control system and method based on rule engine
CN116227916A
Urban rail transit operation risk grade determination method, device, equipment and medium
CN117196292A
Heterogeneous terminal level-to-level management method and system of intelligent traffic system
CN118644069A
Cost auditing method and system based on big data
CN119228420A
Road traffic incident identification and early warning method and system
CN119380538A
Cited By
Business financial reconciliation method and system
CN120147042A
Railway four-electrical-interface inspection method and system based on BIM and rule engine
CN120388022A