Multi-source data fusion and intelligent analysis method for intelligent traffic enterprise

By evaluating the quality of multi-source data in real time in smart transportation systems and dynamically adjusting attention weights, the problem of inflexible data quality assessment in the existing technology is solved, and the abnormal event detection capability and system stability are improved, and road traffic safety is ensured.

CN120197127AActive Publication Date: 2025-06-24GUANGZHOU QINGYUN COMPUTER TECH CO LTD

Patent Information

Application Number
CN202510260144.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-24
Estimated Expiration
2045-03-06

AI Technical Summary

Technical Problem

The existing multi-source data fusion and analysis technology in the field of smart transportation has insufficient real-time data quality evaluation and flexibility in cross-modal data processing, and it is impossible to effectively evaluate the quality of multi-source data. Especially in complex environments, it is difficult to quickly adjust the attention to video, radar and voiceprint data, resulting in insufficient abnormal event detection capabilities.

Method used

By collecting video frames, radar trajectories and voiceprint data in real time, evaluating data quality, calculating video clarity scores, radar signal-to-noise ratios and voiceprint environmental interference index, and dynamically adjusting the cross-modal attention weight matrix based on these scores, ensuring that key information from high-quality data sources is always focused on in complex environments.

Benefits of technology

Real-time and accurate assessment of multi-source data quality has been realized, abnormal event detection capabilities have been improved in complex environments, road traffic safety has been ensured, and the adaptability and stability of smart traffic systems have been improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197127A_ABST
    Figure CN120197127A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source data fusion and intelligent analysis method for an intelligent traffic enterprise, and relates to the technical field of intelligent traffic, and the method comprises the following specific steps: S100, collecting video frame data, radar track data and voiceprint data of the traffic enterprise in real time, S200, calculating a quantized video definition score, S300, calculating a radar signal-to-noise ratio, and S500, calculating a video definition score. According to the method, the video, radar and voiceprint data quality can be accurately judged through multi-modal data quality real-time evaluation and credibility quantification, and a cross-modal attention weight matrix is dynamically adjusted according to a quality evaluation result in a complex environment, that is, when the video quality is reduced, the video, radar and voiceprint data quality can be accurately judged; the system automatically increases the radar data weight, ensures that high-quality data source key information is always focused, learns more accurate and comprehensive abnormal characteristics, effectively reduces the traffic accident rate, and guarantees the road traffic safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of smart transportation technology, and specifically to a multi-source data fusion and intelligent analysis method for a smart transportation enterprise. Background Art

[0002] With the acceleration of urbanization and the continuous growth of traffic flow, smart transportation systems have become a key infrastructure for the development of modern cities. Efficient and accurate technical means are urgently needed to cope with increasingly complex traffic conditions. Multi-source data fusion and intelligent analysis technology, as the core support of smart transportation, aims to integrate multiple data sources and explore the potential value behind the data, so as to achieve goals such as traffic flow optimization, accident prevention, and improved travel efficiency. Data sources such as video surveillance, radar detection, and voiceprint collection have been widely used in the field of smart transportation because they can reflect traffic conditions from different dimensions.

[0003] However, existing multi-source data fusion and analysis methods lack real-time and effective means for data quality assessment. Different data sources are affected by factors such as the environment and equipment aging, and the data quality fluctuates greatly. For example, in bad weather, video surveillance images are prone to blur, and existing methods cannot timely and accurately quantify the degree of video quality degradation, resulting in low-quality video data participating in fusion analysis, which seriously interferes with the accuracy of the results; in cross-modal data processing, traditional methods find it difficult to flexibly focus on key information in complex environments. When there is insufficient light at night or encountering bad weather such as heavy rain and fog, traditional cross-modal processing mechanisms cannot quickly adjust their attention to video, radar and voiceprint data, resulting in a significant reduction in the ability to detect abnormal events, and cannot meet the needs of smart transportation for efficient processing of complex scenarios.

[0004] In summary, the current multi-source data fusion and analysis technology in the field of smart transportation has shortcomings in the real-time data quality assessment and the flexibility of cross-modal data processing. There is an urgent need for an innovative method that can accurately evaluate the quality of multi-source data in real time, intelligently focus on key information in complex environments, and improve the ability to detect abnormal events to meet the growing management needs of smart transportation systems. Summary of the invention

[0005] The purpose of the present invention is to make up for the shortcomings of the existing technology and provide a multi-source data fusion and intelligent analysis method for a smart transportation enterprise. It can accurately judge the quality of video, radar, and voiceprint data through real-time evaluation and credibility quantification of multimodal data quality. In complex environments, it can dynamically adjust the cross-modal attention weight matrix based on the quality evaluation results to ensure that the key information of high-quality data sources is always focused, providing strong guarantees for timely response measures and ensuring road traffic safety.

[0006] To solve the above technical problems, the present invention provides the following technical solutions: A multi-source data fusion and intelligent analysis method for intelligent transportation enterprises. The specific steps of the method are as follows:

[0007] S100. Real-time collect video frame data, radar trajectory data and voiceprint data of transportation enterprises, and synchronize the time stamps and spatial coordinates;

[0008] S200. For video data, by analyzing the characteristics of the edge sharpness and contrast of video images, calculate the quantified video sharpness score q v ;

[0009] S300. For radar data, analyze the ratio of the radar echo signal power to the noise power, and calculate the radar signal-to-noise ratio q r ;

[0010] S400. For voiceprint data, comprehensively consider the environmental noise intensity, frequency distribution and the degree of overlap with the voiceprint signal frequency, and calculate the voiceprint environmental interference index q a ;

[0011] S500. Use the video sharpness score, radar signal-to-noise ratio, and voiceprint environmental interference index as initialization parameters to obtain the initial weights, and fuse the multi-modal data to complete multi-source data fusion and intelligent analysis.

[0012] Furthermore, the video sharpness score in S200 where α and β are adjustment coefficients and α + β = 1, α = 0.6, β = 0.4, represents the gradient value of the pixel at coordinates (x i , y i ) in the video frame, which is used to measure the edge sharpness of the image. n is the total number of pixels participating in the gradient calculation in the video frame. I(x j , y j ) is the pixel gray value at coordinates (x j , y j , is the average gray value of the entire video frame, and m is the total number of pixels in the video frame. The video sharpness score is determined by comprehensively considering the image edge details and overall contrast.

[0013] Even further, the radar signal-to-noise ratio in S300 where P s is the power of the target reflection signal received by the radar, and P n is the sum of the internal noise of the radar and the environmental noise power, which is used to quantify the ratio of the effective signal received by the radar to the noise, so as to evaluate the quality of radar trajectory data.

[0014] Further, in S400, the voiceprint interference index is obtained by transforming the collected voiceprint signal to obtain the time-frequency domain representation S(t,f), and the interference index is defined where T is the total duration of the voiceprint signal, and F n is the noise frequency range, and F t is the full frequency range of the voiceprint signal. This index reflects the proportion of noise components in the voiceprint signal and is used to evaluate the degree of environmental interference on the voiceprint data.

[0015] Further, the specific steps of S500 are as follows:

[0016] Obtain the video clarity score q v , the radar signal-to-noise ratio q r and the voiceprint interference index q a ;

[0017] Perform normalization on q v , q r , q a to obtain the weight coefficients that can be used for fusion where q j =q v +q r +q a , and q i represents the input of q v , q r , q a . This makes the sum of the weights equal to 1, thereby obtaining the initial weight w v of the video data, the initial weight w r of the radar data, and the initial weight w a of the voiceprint data;

[0018] Use the weighted fusion algorithm to fuse the multi-modal data.

[0019] Further, during the process of fusing the multi-modal data using the attention-based weighted fusion algorithm, the video data V, radar data R, and voiceprint data A are obtained through S100. The fused data F = w v V + w r R + w a A. Among them, F represents the fused multi-source data, which is the comprehensive data finally used for intelligent analysis. w v , w r , w aThey are the weights corresponding to video data, radar data, and voiceprint data respectively, reflecting the importance of each modality data in the fusion process. V, R, and A are the original video data, radar data, and voiceprint data respectively. Through the weighted fusion method, high-quality data can play a role in the fusion result, and low-quality data can be controlled, thus optimizing the multi-source data fusion effect.

[0020] Furthermore, the S500 introduces an attention mechanism to calculate attention scores using data features. That is, the video data features after feature extraction of multi-modal data are V f , the radar data features are R f , and the voiceprint data features are A f . By using a scoring function to calculate the attention scores of each modality data, calculate the attention score of the video data the attention score of the radar data the attention score of the voiceprint data where θ is a learnable parameter vector, which can explore the importance of different data features for the final result according to the fusion task. They respectively represent the transposes of the video data feature vector, radar data feature vector, and voiceprint data feature vector. After obtaining the attention scores, convert the scores into weight scores, that is Fuse the weight scores with the initial weights to obtain

[0021] λ is a balance coefficient with a value range of [0, 1], which is used to adjust the influence degree of the initial weights and the weights adjusted by the attention mechanism, and optimize the weight allocation.

[0022] Furthermore, when the quality of the modality data mutates, that is, at night and in bad weather conditions, the S500 immediately triggers an emergency reallocation of the attention weights. Through a dynamic calibration mechanism, it ensures that the cross-modal attention mechanism can always detect abnormal events based on high-quality data.

[0023] Compared with the prior art, the multi-source data fusion and intelligent analysis method of this intelligent transportation enterprise has the following beneficial effects:

[0024] First, through real-time evaluation of multi-modal data quality and credibility quantification, the present invention can accurately judge the quality of video, radar, and voiceprint data. In a complex environment, according to the quality evaluation results, it dynamically adjusts the cross-modal attention weight matrix. That is, when the video quality decreases, the system automatically increases the weight of the radar data, ensuring that it always focuses on the key information of high-quality data sources, learning more accurate and comprehensive abnormal features, effectively reducing the incidence of traffic accidents, and ensuring road traffic safety.

[0025] Second, the pre - support and dynamic calibration mechanism of the present invention provides dual guarantees for traffic operation. The pre - support sets initial weights based on prior knowledge of data quality, allocates attention to different data sources, and reduces judgment errors in the initial stage. The dynamic calibration monitors real - time mutations in data quality and triggers re - allocation of weights. This mechanism ensures that the system can quickly adjust data - processing strategies under various emergency situations, always analyze based on high - quality data, greatly improves the adaptability and stability of intelligent transportation, and reduces the risk of failure caused by environmental changes or equipment failures.

[0026] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. Brief Description of the Drawings

[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or in the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following - described drawings are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.

[0028] Figure 1 It is an operation diagram of a multi - source data fusion and intelligent analysis method for an intelligent transportation enterprise;

[0029] Figure 2 It is a step - flow diagram for user classification in Embodiment 2. Detailed Embodiments

[0030] To further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention objective, the following, in combination with the attached drawings and preferred embodiments, details the specific embodiments, structures, features, and their effects of the present invention as follows.

[0031] Embodiment 1

[0032] As Figure 1 shown, this embodiment focuses on a multi - source data fusion and intelligent analysis method for an intelligent transportation enterprise, and details how to collect video frames, radar trajectories, and voiceprint data in real - time, evaluate data quality, perform data fusion and weight adjustment to achieve efficient abnormal event detection in complex environments, and thus improve the intelligent transportation management level and provide strong support for the safe and efficient operation of urban traffic.

[0033] In a specific implementation, when deploying data acquisition devices, including high-definition cameras for acquiring video frame data, radars for collecting radar trajectory data, and microphone arrays for collecting voiceprint data. These devices work together to collect traffic data in real time. The high-definition camera continuously captures traffic scenes at a certain frame rate, generating a series of video frames. The camera is precisely calibrated to ensure that the shooting angle can cover the traffic area. The radar emits millimeter-wave signals and determines information such as the position, speed, and movement direction of targets by receiving the echo signals reflected by target objects, thereby forming radar trajectory data to ensure the accuracy and stability of the data. The microphone arrays are distributed at intersections to collect various sounds generated during vehicle driving, including engine sounds, braking sounds, horn sounds, etc., and convert these sound signals into digital signals for subsequent processing as voiceprint data. To ensure the spatio-temporal consistency of multi-source data, accurate timestamps and spatial coordinates are added to the collected video frame data, radar trajectory data, and voiceprint data. The timestamp marks the specific moment of data acquisition, and the spatial coordinates are based on the radar geographic coordinate system to determine the position of the data acquisition point on the earth, enabling accurate matching of data from different data sources in the time and space dimensions.

[0034] The collected data is affected by various factors, resulting in uneven quality. Therefore, it is necessary to perform quality assessments on video, radar, and voiceprint data respectively to provide a reliable basis for subsequent data fusion, where:

[0035] Calculation of video clarity score: A certain number of pixel points are selected from the collected video frames for calculating the video clarity score. For each selected pixel point (x i , y i ), the gradient value is calculated to measure the clarity of the image edge. The gradient value reflects the change degree of the image brightness around the pixel point. The greater the change, the clearer the edge. At the same time, the gray value I(x j , y j ) of all pixel points in the entire video frame is calculated, and the average gray value is obtained. According to the formula The video clarity score q v is determined by comprehensively considering the image edge details and overall contrast, where α and β are adjustment coefficients, and α + β = 1. In this embodiment, they are used to balance the influence degrees of edge clarity and overall contrast on the video clarity score. n is the total number of pixel points participating in the gradient calculation, and m is the number of all pixel points in the video frame;

[0036] Calculation of radar signal-to-noise ratio: Analyze the radar echo signal. Denote the power of the received target reflection signal as P s , and denote the sum of the radar internal noise and environmental noise power as P n, according to calculate the radar signal-to-noise ratio q r , which is used to quantify the ratio of the effective signal received by the radar to the noise. The higher the signal-to-noise ratio, the stronger the effective signal received by the radar and the higher the quality of the radar trajectory data;

[0037] Calculation of the voiceprint environmental interference index: Transform the collected voiceprint signal to obtain its representation S(t,f) in the time-frequency domain, determine the total duration T of the voiceprint signal, and the noise frequency range F n and the full frequency range F of the voiceprint signal t , according to calculate the voiceprint environmental interference index q a , which reflects the proportion of the noise component in the voiceprint signal. The higher the proportion, the greater the environmental interference to the voiceprint data and the lower the data quality.

[0038] Based on the data quality assessment results, perform multi-modal data fusion and calculate the weights of each modal data during the fusion process. The process is as follows: Obtain the video clarity score q v , the radar signal-to-noise ratio q r and the voiceprint environmental interference index q a . After that, to make the sum of the weights equal to 1 for subsequent fusion calculations, perform normalization on them. According to the formula calculate the initial weights w of the video data v0 , the initial weights w of the radar data r0 , and the initial weights w of the voiceprint data a0 , where q j =q v +q r +q a , q i represents any one of q v , q r , q a . For the obtained video data V, radar data R, and voiceprint data A, in the weighted fusion algorithm based on the attention mechanism, extract the features of the multi-modal data to obtain the video data features V f , the radar data features R f and the voiceprint data features A f . Use the scoring function to calculate the attention scores of each modal data. For the video data, the attention score For the radar data, the attention score For the voiceprint data, the attention score where θ is a learnable parameter vector, and according to the requirements of the fusion task, continuously adjust its value through training to explore the importance of different data features to the final fusion result. They respectively represent the transposes of the video data feature vector, the radar data feature vector, and the voiceprint data feature vector. After obtaining the attention scores, they are converted into weight scores, and the formula is Finally, the weight scores are fused with the initial weights to obtain the final weights for data fusion. Among them, λ is a balance coefficient with a value range of [0, 1], which is used to adjust the influence degrees of the initial weights and the weights adjusted by the attention mechanism. When λ is close to 1, the influence of the initial weights is greater; when λ is close to 0, the influence of the weights adjusted by the attention mechanism is greater, so as to achieve the best data fusion effect. According to the calculated weights, the multi-modal data are fused, and the fused data F = w v V + w r R + w a A. This fused data integrates the information of video, radar, and voiceprint data and serves as the comprehensive data for final intelligent analysis, enabling high-quality data to play a greater role in the fusion result and effectively controlling the influence of low-quality data, thus optimizing the multi-source data fusion effect.

[0039] The fused data is input into the abnormal event detection model to determine whether an abnormal event occurs, and corresponding handling measures are taken according to the event type and location. In this embodiment, a neural network model based on deep learning is used as the abnormal event detection model. This model is trained with a large amount of historical traffic data and learns the characteristic patterns of normal traffic states and various abnormal events. During the training process, the fused multi-source data is used as the input, and at the same time, the corresponding traffic states (normal or abnormal, and the specific types of abnormal events) of the data are labeled. The model continuously adjusts its internal parameters to make the prediction results close to the labeled results. After the training is completed, the model has the ability to detect abnormal events for newly input data. The fused data is input into the trained abnormal event detection model, and the model analyzes and judges the data. If the model output result indicates that the current traffic state is abnormal, the type of the abnormal event is further determined, such as traffic accidents, traffic jams, vehicle violations, etc., and the location where the event occurs is determined through the spatial coordinate information in the data. Once an abnormal event is detected, the corresponding processing process is immediately started. For traffic accidents, alarm information is automatically sent to the traffic management department and the emergency center. For traffic jams, real-time traffic condition information is released through the traffic guidance system to guide vehicles to avoid the congested sections and optimize the traffic flow. For vehicle violation behaviors, relevant evidence is recorded, such as video screenshots, radar trajectory data, etc., and the violation information is stored and transmitted for subsequent processing.

[0040] In the actual traffic environment, the data quality may change suddenly. For example, insufficient light at night may lead to a decline in video quality, or bad weather may affect the quality of radar and voiceprint data. To ensure that abnormal event detection can always be based on high-quality data, a dynamic calibration mechanism is set up in this embodiment. It monitors the changes in video clarity score, radar signal-to-noise ratio, and voiceprint environmental interference index in real time. Once it is found that the quality of one modality data changes suddenly, such as a significant decrease in the video clarity score at night, or a decrease in the radar signal-to-noise ratio and an increase in the voiceprint environmental interference index due to heavy rain, the dynamic calibration mechanism is immediately triggered. The processing flow of the dynamic calibration mechanism is as follows: when a sudden change in data quality is detected, according to the current data quality assessment result, the weights of each modality data are recalculated, that is, the weight coefficient calculation, weight adjustment based on the attention mechanism, and weight fusion are performed again to obtain new weights. The new weights are used to re-fuse the multi-modal data to ensure that the fused data can more accurately reflect the current traffic conditions and provide a reliable basis for abnormal event detection. Through this dynamic calibration mechanism, in various complex environments, it can always focus on the key information of high-quality data sources, improving the accuracy and reliability of abnormal event detection.

[0041] In summary, this embodiment details the complete implementation process of a multi-source data fusion and intelligent analysis method for a smart transportation enterprise. By acquiring video frames, radar trajectories, and voiceprint data, synchronizing time stamps and spatial coordinates, conducting quantitative evaluation, accurately judging data quality, calculating weights based on the evaluation results and performing weighted fusion to optimize the fusion effect, detecting abnormal events and taking corresponding measures, and timely adjusting weights when the data quality changes suddenly to ensure stability and accuracy, it realizes the efficient fusion and intelligent analysis of multi-source data, effectively solving the deficiencies of the prior art in the real-time data quality assessment and the flexibility of cross-modal data processing. In a complex traffic environment, it can accurately focus on the key information of high-quality data sources, improving the ability to detect abnormal events and providing a strong guarantee for the efficient operation of smart transportation and road traffic safety.

[0042] Embodiment 2

[0043] As Figure 2 shown, based on Embodiment 1, this embodiment focuses on the multi-source data fusion and intelligent analysis method of a smart transportation enterprise, and details how to achieve precise monitoring of traffic conditions and efficient handling of abnormal events in an urban traffic scenario by collecting video frames, radar trajectories, and voiceprint data, evaluating data quality, and using fusion strategies and intelligent analysis means, thereby improving the intelligent level of traffic management and road safety.

[0044] First, enter the data collection stage (S100). Install high-definition cameras, radars, and voiceprint collection points to collect video frame data, radar trajectory data, and voiceprint data of transportation enterprises in real time. To ensure the consistency of multi-source data in terms of time and space, precise timestamps and spatial coordinates are added to each set of collected data, enabling the data from different data sources to be processed and analyzed within a unified spatio-temporal framework.

[0045] Then, enter the data quality assessment stage (S200). Conduct a quality assessment on the collected data. For video data, analyze the contrast of the images. For radar data, evaluate the stability and accuracy of its signals. For voiceprint data, focus on the degree of environmental noise interference.

[0046] Subsequently, enter the data fusion and weight determination stage (S300). According to the results of the data quality assessment, fuse different types of data and determine their weights during the fusion process. During the data fusion process, weighted fusion is performed on video data, radar trajectory data, and voiceprint data based on the weights of the data. In this way, high-quality data can play a greater role in the fusion result, while the impact of low-quality data on the final result is weakened, thereby improving the effect of data fusion.

[0047] Next, enter the abnormal event detection and handling stage (S400). Train the fused data using historical traffic data to train a traffic anomaly model, which can accurately identify various normal and abnormal traffic patterns. When an abnormal change in traffic flow is detected, it is determined that there are abnormal events such as traffic congestion or accidents. Once an abnormal event is detected, the corresponding handling process is immediately initiated.

[0048] Finally, enter the dynamic calibration and optimization stage (S500). During the actual operation process, the traffic environment is constantly changing, and the data quality will also fluctuate accordingly. To ensure that the system can always accurately analyze the traffic conditions, monitor the changes in data quality in real time. Once a mutation in data quality is detected, re-evaluate the data quality and readjust the weights of different data according to the new evaluation results. This dynamic calibration mechanism can always maintain accurate monitoring of traffic conditions and efficient detection of abnormal events in a complex and changing traffic environment.

[0049] In summary, this embodiment details the application process of the multi-source data fusion and intelligent analysis method for intelligent transportation enterprises in urban transportation hubs. Through comprehensive data collection, strict data quality assessment, reasonable data fusion and weight determination, efficient abnormal event detection and handling, as well as dynamic calibration and optimization mechanisms, accurate monitoring and effective management of complex traffic conditions are achieved. This method can timely detect and handle traffic abnormal events, improve road traffic efficiency, and ensure traffic safety.

[0050] As described above, it is only the preferred embodiment of the present invention and does not impose any formal restrictions on the present invention. Although the present invention has been disclosed above in the preferred embodiment, it is not intended to limit the present invention. Any person skilled in the art can make some changes or modifications to equivalent embodiments of equivalent changes within the scope of the technical solution of the present invention by using the above-disclosed technical content. However, as long as it does not depart from the content of the technical solution of the present invention, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention still fall within the scope of the technical solution of the present invention.

Claims

1. A multi-source data fusion and intelligent analysis method for a smart transportation enterprise, characterized in that: The specific steps of this method are: S100, collect video frame data, radar trajectory data and voiceprint data of transportation enterprises in real time, and synchronize timestamps and spatial coordinates; S200, for the video data, by analyzing the edge clarity and contrast characteristics of the video image, a quantitative video clarity score q is calculated v ; S300: Analyze the ratio of radar echo signal power to noise power for radar data, and calculate the radar signal-to-noise ratio q r ; S400: For the voiceprint data, comprehensively consider the environmental noise intensity, frequency distribution and the overlap degree with the voiceprint signal frequency to calculate the voiceprint environmental interference index q a ; S500, the video clarity score, radar signal-to-noise ratio, and voiceprint environment interference index are used as initialization parameters to obtain the initial weights, introduce the attention mechanism and fuse the multimodal data to complete multi-source data fusion and intelligent analysis.

2. The multi-source data fusion and intelligent analysis method of a smart transportation enterprise according to claim 1 is characterized in that: The S200 video clarity rating Among them, α and β are adjustment coefficients and α+β=1, α=0.6, β=0.4, Indicates that the coordinates in the video frame are (x i ,y i ) pixel gradient value, which is used to measure the edge clarity of the image. n is the total number of pixels involved in the gradient calculation in the video frame. j ,y j ) is the coordinate (x j ,y j ), is the average grayscale value of the entire video frame, m is the number of all pixels in the video frame, and the video clarity score is determined by comprehensively considering the image edge details and overall contrast.

3. The multi-source data fusion and intelligent analysis method for a smart transportation enterprise according to claim 1 is characterized in that: The S300 radar signal-to-noise ratio Among them, P s is the target reflected signal power received by the radar, P n It is the sum of the radar internal noise and the environmental noise power. It is used to quantify the ratio of the effective signal to noise received by the radar to evaluate the quality of radar trajectory data.

4. The multi-source data fusion and intelligent analysis method for a smart transportation enterprise according to claim 1 is characterized in that: The voiceprint interference index in S400 is obtained by transforming the collected voiceprint signal to obtain a time-frequency domain representation S(t,f), and defining the interference index Among them, T is the total duration of the voiceprint signal, F n is the noise frequency range, F t It is the full frequency range of the voiceprint signal. This index reflects the proportion of noise components in the voiceprint signal and is used to evaluate the degree of environmental interference on the voiceprint data.

5. The multi-source data fusion and intelligent analysis method for a smart transportation enterprise according to claim 1 is characterized in that: The specific steps of S500 are: Get the video clarity score v , radar signal-to-noise ratio q r and voiceprint interference index q a ; Q v ,q r ,q a Normalize to get the weight coefficient that can be used for fusion where q j =q v +q r +q a ,q i Indicates q v ,q r ,q a The input of the video data is used to make the sum of the weights equal to 1, thus obtaining the initial weight w of the video data. v0 , the initial weight w of the radar data r0 , the initial weight w of the voiceprint data a0 ; A weighted fusion algorithm based on attention mechanism is used to fuse multimodal data.

6. The multi-source data fusion and intelligent analysis method of a smart transportation enterprise according to claim 5 is characterized in that: In the process of fusing multimodal data using the weighted fusion algorithm based on the attention mechanism, video data V, radar data R, and voiceprint data A are obtained through S100, and the fused data F=w v V+w r R+w a A, where F represents the fused multi-source data, which is the comprehensive data ultimately used for intelligent analysis, and w v 、w r 、w a are the weights corresponding to video data, radar data, and voiceprint data, reflecting the importance of each modality in the fusion process. V, R, and A are the original video data, radar data, and voiceprint data, respectively. Through weighted fusion, high-quality data can play a role in the fusion result, and low-quality data can be controlled, thereby optimizing the multi-source data fusion effect.

7. The multi-source data fusion and intelligent analysis method for a smart transportation enterprise according to claim 6 is characterized in that: S500 introduces the attention mechanism and uses data features to calculate the attention score, that is, the video data feature after feature extraction of multimodal data is V f , the radar data feature is R f , the voiceprint data feature is A f , calculate the attention score of each modal data through the scoring function, and calculate the attention score of the video data Attention scores for radar data Attention score of voiceprint data Where θ is a learnable parameter vector, which mines the importance of different data features to the final result according to the fusion task. Represent the transposition of the video data feature vector, radar data feature vector, and voiceprint data feature vector respectively. After obtaining the attention score, the score is converted into a weighted score, that is, The weight score is combined with the initialization weight to obtain λ is the balance coefficient, which takes a value of [0, 1] and is used to adjust the influence of the initialization weights and the weights adjusted by the attention mechanism to optimize the weight distribution.

8. The multi-source data fusion and intelligent analysis method for a smart transportation enterprise according to claim 1 is characterized in that: When the quality of modal data changes suddenly, i.e., at night or in bad weather conditions, the S500 immediately triggers an emergency redistribution of attention weights, and through a dynamic calibration mechanism, ensures that the cross-modal attention mechanism can always detect abnormal events based on high-quality data.

Citation Information

Patent Citations

  • Channel selection method and device

    CN111031609A

  • Smart park multi-source data dynamic monitoring and real-time analysis system and method

    CN116665001A

  • Industrial intelligent detection method and system based on multi-modal large model

    CN118503832A

  • Road traffic accident detection method and system based on roadside radar perception

    CN119091616A

  • Road network safety early warning method and device based on hologram and storage medium

    CN119296322A

Cited By

  • Self-adaptive cross validation thundersight data fusion method and system

    CN120508999A

  • Tunnel traffic flow prediction and control method and system

    CN120564429A

  • Tunnel traffic flow prediction and control method and system

    CN120564429B