A security check method and system based on voice and smell fusion data recognition
Through the collaborative processing of AIoT edge devices and cloud-based LLM large models, multimodal fusion of odor and voice signals is achieved, solving the problem of insufficient cross-domain data integration in existing security inspection technologies and improving the accuracy and timeliness of identifying potential threats.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HANGZHOU MINGGUANG MICROELECTRONICS TECH CO LTD
- Filing Date
- 2026-01-23
- Publication Date
- 2026-05-12
AI Technical Summary
Existing security inspection technologies struggle to effectively combine multiple sensory information for comprehensive judgment in complex scenarios, resulting in limited efficiency and accuracy in identifying potential threats. They are particularly prone to missed detections when odor signals are weak or affected by environmental interference, and it is difficult to effectively correlate voice emotional fluctuations with odor abnormalities.
Odor and speech signals are collected and filtered by AIoT edge devices, odor pattern vectors and speech emotion vectors are extracted, and the feature vectors are weighted and fused using an attention mechanism. Combined with time series analysis and data grouping algorithms, a fusion evaluation report is generated to identify high-risk anomalies.
It enables real-time collaborative processing of multimodal information from odor and voice, improving the accuracy and timeliness of identifying potential threats. It can accurately pinpoint risks in complex environments and generate efficient threat level labels.
Smart Images

Figure CN121598267B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of information technology, and in particular to a security inspection method and system based on AIoT edge devices and cloud-based LLM large models to achieve voice and odor fusion data recognition. Background Technology
[0002] In the field of public safety, innovation in security inspection technology is of paramount importance for ensuring social stability and the safety of personnel.
[0003] With the continuous development of intelligent technology, security inspection methods based on the Internet of Things and artificial intelligence have gradually become a research hotspot. Their importance lies in their ability to quickly and accurately identify potential threats and protect public places from dangerous factors.
[0004] However, existing technologies often face problems such as insufficient cross-domain data integration and limited real-time analysis capabilities when dealing with complex security inspection scenarios, especially the shortcomings in multi-dimensional information fusion, which limits the efficiency and accuracy of potential risk identification.
[0005] The current limitations of security inspection methods are mainly reflected in the isolated processing of different types of information.
[0006] Many systems focus only on data from a single sensing channel and are unable to effectively combine information from multiple sensing sources for comprehensive judgment.
[0007] This limitation makes it difficult for the system to capture abnormal signals hidden behind multiple pieces of information when facing complex environments, thus affecting the comprehensive assessment of potential threats.
[0008] Against this backdrop, the core technical challenges facing the research field lie in how to achieve collaborative processing and in-depth analysis of multi-source information.
[0009] First, the accurate capture of odor characteristics is a key issue, because odor signals are often weak and easily affected by the environment. If they cannot be accurately extracted and identified, it may lead to the missed detection of hazardous substances.
[0010] Secondly, due to the complexity of odor data, it becomes extremely difficult to effectively correlate and integrate it with other information such as speech features. This lack of integration of cross-modal information will directly affect the comprehensive judgment of the behavior and intentions of the examinee.
[0011] For example, in real-world security check scenarios, someone carrying suspicious items might try to mask their nervousness through speech. If the system cannot combine abnormal odors with emotional fluctuations in the voice for analysis, it may miss crucial clues. For instance, the person carrying the item might deliberately use a calm tone to respond to security inquiries, but their voice signal would still contain subtle emotional fluctuations such as a slightly faster speaking speed and a trembling tone. If the system only detects odor, it might fail to identify it due to a weak odor signal; if it only analyzes voice or behavior, it might attribute the nervousness to ordinary security anxiety. Only by integrating the odor characteristics of suspicious items with emotional fluctuations in the voice can the risk be accurately identified.
[0012] Therefore, how to achieve real-time collection and preliminary processing of multimodal information such as smell and voice on edge devices, and how to deeply integrate and comprehensively evaluate this information through cloud technology, has become a key issue in improving the efficiency and accuracy of security checks. Summary of the Invention
[0013] This invention provides a method for speech and odor fusion data recognition based on AIoT edge devices and cloud-based LLM large models, mainly including:
[0014] Odor signals and speech signals are acquired by a signal acquisition device. After removing environmental interference noise by a filtering method, a preliminary odor molecule concentration sequence and speech waveform sequence are obtained.
[0015] Based on the preliminary odor molecule concentration sequence and the speech waveform sequence, an odor pattern vector and a speech emotion vector are extracted using a feature extraction model, thereby determining a preliminary multimodal feature set;
[0016] If the similarity between the odor pattern vector and the speech emotion vector in the multimodal preliminary feature set is lower than a threshold, then the two are weighted and fused through an attention mechanism to obtain a comprehensive anomaly index vector; otherwise, the original vector is retained as the comprehensive anomaly index vector.
[0017] After obtaining the comprehensive anomaly index vector, a time series analysis model is used to analyze its time series changes in order to determine potential anomaly pattern sequences.
[0018] For the potential abnormal pattern sequence, similar patterns are grouped using a data grouping algorithm to obtain classified abnormal groups;
[0019] Based on the aforementioned anomaly groups, a high-risk anomaly subset is determined;
[0020] Based on the high-risk anomaly subset, a fusion assessment report is generated to obtain the final threat level label.
[0021] This invention provides a system for recognizing speech and odor fusion data based on AIoT edge devices and cloud-based LLM large models, mainly comprising:
[0022] The signal acquisition and filtering module is used to acquire odor signals and speech signals through a signal acquisition device, and to obtain preliminary odor molecule concentration sequences and speech waveform sequences after removing environmental interference noise using a filtering method.
[0023] The feature extraction module is used to extract odor pattern vectors and speech emotion vectors based on the preliminary odor molecule concentration sequence and the speech waveform sequence using a feature extraction model, thereby determining a preliminary multimodal feature set;
[0024] The similarity judgment and fusion module is used to obtain a comprehensive anomaly index vector by weighted fusion of the odor pattern vector and the voice emotion vector in the multimodal preliminary feature set if the similarity is lower than a threshold; otherwise, the original vector is retained as the comprehensive anomaly index vector.
[0025] The time series analysis module is used to obtain the comprehensive anomaly index vector and then use a time series analysis model to analyze its time series changes to determine potential anomaly pattern sequences.
[0026] The data grouping module is used to group similar patterns into classification anomaly groups based on the potential anomaly pattern sequence using a data grouping algorithm.
[0027] The high-risk determination module is used to determine a high-risk anomaly subset based on the classified anomaly groups;
[0028] The report generation module is used to generate a fusion assessment report based on the high-risk anomaly subset to obtain the final threat level label.
[0029] The technical solutions provided by the embodiments of the present invention may include the following beneficial effects:
[0030] This invention discloses an anomaly threat assessment method based on multimodal signals, addressing the challenge of accurately identifying potential threats in complex environments, such as the difficulty in correlating and assessing abnormal odors (e.g., chemical leaks) with emotional fluctuations in speech (e.g., panic signals) in the public safety field. This invention acquires odor and speech signals using a signal acquisition device, filters and removes noise, then extracts odor pattern vectors and speech emotion vectors to form a preliminary multimodal feature set. If the similarity between the two is below a threshold, an attention mechanism is used for weighted fusion to obtain a comprehensive anomaly index vector; otherwise, the original vectors are retained. Subsequently, a time-series analysis model is used to identify potential anomaly pattern sequences, and a data grouping algorithm is used to classify anomaly groups, thereby determining a high-risk subset and generating a fusion assessment report to output the final threat level label. This method effectively integrates odor and speech multimodal information, improving the accuracy and timeliness of anomaly detection, and achieving precise assessment and early warning of potential threats. Attached Figure Description
[0031] The above and other objects, features and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which:
[0032] Figure 1 This is a flowchart of a method for recognizing speech and odor data based on AIoT edge devices and cloud-based LLM large models, according to the present invention.
[0033] Figure 2 This is a flowchart of a method for recognizing speech and odor data based on AIoT edge devices and cloud-based LLM large models, according to the present invention.
[0034] Figure 3 This is a schematic diagram of the edge-cloud collaboration architecture of the present invention.
[0035] Figure 4 This is a schematic diagram of the system structure of the present invention.
[0036] Figure 5 This is a schematic diagram comparing the preprocessing of odor signals and speech signals in this invention.
[0037] Figure 6 This is a schematic diagram illustrating the extraction of odor pattern vectors and speech emotion vectors according to the present invention.
[0038] Figure 7 This is a schematic diagram of the weighted fusion process of the attention mechanism of the present invention.
[0039] Figure 8 This is a schematic diagram illustrating the temporal analysis and classification grouping of potential abnormal modes in this invention.
[0040] Figure 9 This is a structural diagram of the AIoT edge device of the present invention.
[0041] Figure 10 This is a schematic diagram of the spatial layout of the odor sensor array of the present invention.
[0042] Figures 11(a)-11(b) This is a waveform comparison diagram of the odor-speech fusion signal under normal and abnormal states of the present invention.
[0043] Figure 12 This is a comparison chart of the detection accuracy of the method of this invention and the traditional method.
[0044] Figures 13(a)-13(b) This is a graph showing the system response time and processing delay of the present invention. Detailed Implementation
[0045] The technical solutions of the embodiments of the present invention will be clearly and thoroughly described below with reference to the accompanying drawings. The described embodiments are merely some embodiments of the present invention.
[0046] like Figure 1 As shown, the method of this invention includes seven main steps, forming a complete speech and odor fusion data recognition process through signal acquisition, feature extraction, fusion analysis, temporal judgment, grouping and classification, risk assessment, and final threat level determination. Step S101 involves acquiring odor and speech signals using a signal acquisition device, removing environmental interference noise using a filtering method to obtain a preliminary odor molecule concentration sequence and a speech waveform sequence, providing clean raw data for subsequent analysis; Step S102 uses a feature extraction model to extract odor pattern vectors and speech emotion vectors based on the preliminary odor molecule concentration sequence and speech waveform sequence, determining a preliminary multimodal feature set; Step S103 if the similarity between the odor pattern vector and the speech emotion vector in the preliminary multimodal feature set is lower than a preset threshold, then the two are weighted and fused using an attention mechanism to obtain a comprehensive anomaly index vector, achieving effective integration of cross-modal information; Step... After obtaining the comprehensive anomaly index vector in step S104, a time-series analysis model is used to analyze its temporal changes to determine potential anomaly pattern sequences and capture the dynamic evolution characteristics of anomaly events. In step S105, for the potential anomaly pattern sequences, similar patterns are grouped using a data grouping algorithm to obtain classified anomaly groups, and similar anomalies are clustered for subsequent accurate processing. In step S106, based on the classified anomaly groups, a high-risk anomaly subset is determined, and anomalies that require key attention are selected. In step S107, based on the high-risk anomaly subset, a fusion assessment report is generated to obtain the final threat level label, completing the entire identification process and outputting actionable threat assessment results.
[0047] like Figure 2As shown, the speech and odor fusion data recognition method of this invention adopts a hierarchical processing architecture, including five functional layers. The signal acquisition layer deploys two types of acquisition units: odor sensors and microphones. The odor sensors acquire odor signals from the environment, and the microphones acquire speech signals, achieving simultaneous acquisition of dual-modal raw data. The preprocessing layer performs filtering and noise reduction on both signals, converting the odor signal into an odor molecule concentration sequence and the speech signal into a speech waveform sequence, eliminating environmental noise and equipment interference. The feature extraction layer extracts features from the preprocessed data. The odor feature extraction module extracts odor pattern vectors from the concentration sequence, and the speech feature extraction module extracts speech emotion vectors from the waveform sequence, converting the raw signals into a feature space representation. The fusion analysis layer first uses an attention mechanism fusion technology to perform multimodal feature fusion of the odor pattern vector and the speech emotion vector. Then, the temporal analysis module analyzes the evolution of the fused features over time to uncover the correlation patterns between the two types of data. The output layer generates two types of outputs based on the temporal analysis results: anomaly detection results and threat level labels, enabling intelligent identification and graded early warning of potential threats. Each functional layer is identified by a dashed box, and the output layer is identified by a solid double box to indicate its final output characteristics. The data flow is connected by arrows and the data type is labeled, forming a complete processing link from signal acquisition to result output.
[0048] like Figure 3 As shown, this invention employs an edge-cloud collaborative architecture to achieve efficient recognition of fused speech and odor data. The entire system is divided into three main layers: the AIoT edge device layer, the cloud-based LLM large model layer, and the real-time response feedback layer. At the AIoT edge device layer, the system sequentially executes three steps: signal acquisition, filtering, and feature extraction. The signal acquisition module acquires raw data from sensors, the filtering module eliminates environmental noise and interference signals, the feature extraction module extracts key features and reduces data dimensionality, and the processed data is uploaded to the cloud via a bidirectional communication channel. At the cloud-based LLM large model layer, the system sequentially executes three steps: deep analysis, pattern recognition, and anomaly detection. The deep analysis module uses multi-layer neural networks to extract high-order semantic features, the pattern recognition module matches known speech and odor patterns based on a pre-trained model, and the anomaly detection module identifies abnormal events through confidence thresholds. The cloud processing results are returned to the edge device via a bidirectional communication channel. At the real-time response feedback layer, the system sends the recognition results, confidence levels, and suggested measures to the end user or control system, achieving closed-loop control from data acquisition to response execution. This collaborative architecture ensures both real-time performance and recognition accuracy, providing an efficient and reliable technical solution for speech and odor fusion data recognition.
[0049] like Figure 4As shown, the system of this invention adopts a modular architecture, with data flow processed sequentially from top to bottom. The signal acquisition and filtering module is responsible for acquiring odor and speech signals and filtering and denoising them to obtain preliminary sensor data; the feature extraction module extracts odor pattern vectors and speech emotion vectors from the two types of signals respectively; the fusion judgment module achieves feature fusion through similarity calculation and attention mechanism to generate a comprehensive anomaly index vector; the time series analysis module performs time series change analysis on the comprehensive anomaly index vector; the anomaly grouping module uses a data grouping algorithm to classify similar anomaly patterns; the risk assessment module determines a high-risk anomaly subset; and the report generation module finally outputs a threat level label.
[0050] like Figure 9 As shown, the AIoT edge device features a compact design with a rectangular casing measuring approximately 200mm × 150mm × 40mm. An odor sensor array containing 12 independent sensing units is located on the top of the device. It samples ambient gases through an air intake, and the sensor array can simultaneously detect changes in the concentration of multiple volatile organic compounds (VOCs). A 7-inch touchscreen display is positioned in the center of the front of the device to display real-time sensor data and system operating status. Below the display are three status indicator lights: a red power indicator, a green data transmission indicator, and a blue fault alarm indicator. A MEMS microphone is located on the left side of the device to collect ambient acoustic signals, with a frequency response range of 20Hz to 20kHz and a sensitivity of -38dB ± 2dB. A full-range speaker is located on the right side of the device to play voice prompts and alarm signals, with an output power of 3W and a frequency response range of 100Hz to 18kHz. The device features two interfaces on its bottom: a USB 3.0 data interface on the left for high-speed data transfer and firmware upgrades, with a maximum transfer rate of 5Gbps; and a DC 12V power interface on the right, supporting external power adapters with a power consumption of less than 15W. This structural design enables the integrated layout of multi-source heterogeneous sensors, providing a hardware foundation for the collaborative acquisition of multimodal signals.
[0051] like Figure 10As shown, the odor sensor array adopts a 3×3 spatial layout structure, with nine gas sensor units S1 to S9 evenly distributed on the PCB substrate. The sensors maintain a regular spacing to avoid gas cross-interference. Sensors S1, S2, and S3 in the first row are used to detect volatile organic compounds, sulfides, and ammonia gases, respectively; sensors S4, S5, and S6 in the second row correspond to alcohols, ketones, and aldehydes, respectively; and sensors S7, S8, and S9 in the third row are for esters, aromatic hydrocarbons, and alkanes, respectively. A signal processing chip is located at the center of the bottom of the PCB substrate to receive the electrical signals output from each sensor and perform analog-to-digital conversion, signal amplification, and preliminary filtering. This array-based distribution allows for the independent detection of nine different types of gas components within the same time window, avoiding the time delay and gas mixing problems caused by traditional single-sensor serial detection, ensuring the accuracy and real-time performance of multi-component synchronous detection.
[0052] Specifically, an embodiment of the present invention provides a method for recognizing speech and odor fusion data based on AIoT edge devices and cloud-based LLM large models, which may include:
[0053] Step S101: Odor signal and speech signal are acquired by signal acquisition device, and environmental interference noise is removed by filtering method to obtain preliminary odor molecule concentration sequence and speech waveform sequence.
[0054] Specifically, odor and speech signals are simultaneously acquired using signal acquisition equipment. Filtering techniques are used to process the acquired raw data, resulting in denoised odor molecule concentration sequences and speech waveform sequences. For these denoised sequences, time-domain analysis is employed to extract significant feature points, identifying key nodes in odor concentration changes and the main amplitude distribution of the speech waveform. Based on these key nodes and amplitude distributions, temporal features of odor concentration changes and temporal distribution features of the speech waveform are constructed, establishing their correspondence on the time axis. If the correspondence between these two temporal features and the speech waveform's temporal distribution features is below a preset threshold, they are time-aligned to obtain aligned temporal data. Using this aligned temporal data, a support vector machine algorithm is used to classify odor concentration changes and speech waveform distributions, identifying potential correlation patterns. Based on these classification patterns, further joint feature extraction is performed on the odor molecule concentration sequences and speech waveform sequences to obtain a comprehensive signal feature set. This comprehensive signal feature set is then stored in a preset database using data storage techniques, determining the foundational data source for subsequent analysis.
[0055] like Figure 5As shown, the preprocessing processes for odor and speech signals correspond to the left and right processing flows, respectively. In the left-hand odor signal processing flow, the original odor signal contains time-series data with concentration values between 0 and 1, mixed with random noise and environmental interference introduced by the acquisition equipment. Low-pass filtering is used to process the original signal, filtering out high-frequency noise components while retaining the main trend of odor molecule concentration changes. The denoised odor molecule concentration sequence exhibits a smooth curve characteristic, eliminating rapid fluctuations caused by noise, with concentration variation controlled within the range of 0.2 to 0.8, suitable for subsequent time-series analysis. In the right-hand speech signal processing flow, the original speech waveform contains high-frequency oscillation signals with amplitudes between -0.5 and 1.5, superimposed with circuit noise and background noise. Band-pass filtering is used to process the speech signal, retaining the main frequency components of human voice and removing interference signals beyond the speech frequency range. The denoised speech waveform sequence maintains clear periodic oscillation characteristics, with amplitude controlled within the range of -0.3 to 1.2, effectively improving the signal-to-noise ratio of the speech signal. By comparing the preprocessing processes of the two signals, we can see the adaptability of filtering techniques in different types of signal processing, providing a high-quality data foundation for subsequent feature extraction and correlation analysis.
[0056] Step S102: Based on the preliminary odor molecule concentration sequence and the speech waveform sequence, an odor pattern vector and a speech emotion vector are extracted using a feature extraction model to determine the preliminary multimodal feature set.
[0057] Specifically, odor molecule concentration sequences and speech waveform sequences are acquired. The raw signals are standardized using a pre-established data acquisition module to obtain normalized concentration and waveform data. For the normalized concentration and waveform data, a feature extraction model is used to perform pattern analysis on the odor molecule concentration sequences, and simultaneously, emotion vectorization is performed on the speech waveform sequences to determine odor pattern vectors and speech emotion vectors. Based on the odor pattern vectors and speech emotion vectors, a preliminary multimodal feature set is constructed. The two types of vectors are then uniformly formatted using a data integration tool to obtain a fused multimodal feature dataset. For the fused multimodal feature dataset, if the feature dimension exceeds a preset threshold, dimensionality reduction is used to compress the dataset, identifying an optimized feature subset. Based on the optimized feature subset, a classification model is used to categorize the multimodal features, resulting in categorized feature groupings. For the categorized feature groupings, if the distribution of the groupings does not meet preset balance conditions, data resampling techniques are used to adjust the groupings, determining the final multimodal feature classification set. Based on the final multimodal feature classification set, it is saved to a preset database through the data storage module to obtain a structured feature library for subsequent analysis.
[0058] like Figure 6As shown, the multimodal feature vector extraction process includes a dual-channel parallel processing mechanism. In the input layer, the odor molecule concentration sequence and the speech waveform sequence are standardized by the data acquisition module, converting the raw signals into standardized concentration and waveform data. The standardized data are then input into the corresponding feature extraction models. The odor feature extraction model employs a three-layer neural network structure, including an input layer, a hidden layer, and an output layer. It extracts pattern information from the odor molecule concentration sequence through layer-by-layer feature abstraction, generating an odor pattern vector. The speech feature extraction model also uses a three-layer neural network structure to analyze and process the speech waveform sequence, extracting emotional features from the speech and generating a speech emotion vector. The odor pattern vector and the speech emotion vector are each represented in eight-dimensional vector form, with each dimension ranging from 0.3 to 0.7, reflecting the activation intensity of different feature dimensions. Finally, the two feature vectors are merged to form a preliminary multimodal feature set, which integrates odor pattern information and speech emotion information, providing basic data support for subsequent multimodal data fusion and feature association analysis. The entire extraction process is completed through four stages: ① standardization, ② feature extraction and modeling, ③ vector generation, and ④ feature aggregation, realizing the transformation from the original multimodal signal to the normalized feature vector.
[0059] Step S103: If the similarity between the odor pattern vector and the voice emotion vector in the multimodal preliminary feature set is lower than a threshold, then the two are weighted and fused through an attention mechanism to obtain a comprehensive anomaly index vector; otherwise, the original vector is retained as the comprehensive anomaly index vector.
[0060] By extracting odor pattern vectors and speech emotion vectors from the multimodal feature set and calculating their similarity value, preliminary comparison results are obtained. If the calculated similarity value is lower than a preset threshold, an attention-weighted method is used to fuse the odor pattern vector and speech emotion vector to obtain a comprehensive anomaly index. If the calculated similarity value is not lower than the preset threshold, the original vector data is directly extracted and identified as the anomaly index vector. For the anomaly index vector formed by the comprehensive anomaly index or the original vector data, pattern recognition is performed using a pre-established classification model to determine its anomaly category. Based on the pattern recognition results, key feature dimensions are extracted from the anomaly index vector to obtain the corresponding anomaly distribution information. Through further comparative analysis of the anomaly distribution information, the location of the anomaly index vector in the multimodal feature set is determined, obtaining the final anomaly judgment criteria. The preliminary feature set is updated using the anomaly judgment criteria to generate optimized feature data, completing the multimodal feature anomaly detection process.
[0061] like Figure 7As shown, the attention mechanism weighted fusion process achieves adaptive integration of multimodal features through similarity judgment. In the input layer, odor pattern vector V1 and voice emotion vector V2 are extracted from the multimodal feature set. In the calculation layer, the similarity calculation module uses the cosine similarity formula sim(V1,V2)=V1·V2 / (|V1|·|V2|) to calculate the similarity between the two. In the judgment layer, the calculated similarity is compared with a preset threshold T, which ranges from 0.6 to 0.8. In the processing layer, when the similarity is lower than the threshold T, it indicates a significant difference between the two modal features. At this point, the attention fusion branch is entered, and weight coefficients a and b are calculated through an adaptive weight allocation mechanism (satisfying a+b=1). The two vectors are then weighted and fused to obtain a*V1+b*V2. When the similarity is higher than or equal to the threshold T, the original vector (V1,V2) is directly retained as the feature representation. In the output layer, a comprehensive anomaly index vector V_combined is output for subsequent anomaly detection analysis. This process ensures adaptive fusion when there are significant differences in features across different modalities, while preserving the original information when features are similar, thereby improving the accuracy and robustness of anomaly detection.
[0062] Step S104: After obtaining the comprehensive anomaly index vector, use a time series analysis model to analyze its time series changes to determine potential anomaly pattern sequences.
[0063] After acquiring the comprehensive vector data, the abnormal indicators are initially organized, and irrelevant noise is removed using preset filtering rules to obtain a processed indicator set. Using this processed indicator set, a time-series analysis model is employed to continuously track trends and identify key fluctuation points. Based on these key fluctuation points and the criteria for potential anomalies, if fluctuations exceed preset thresholds, they are marked as anomaly candidates, resulting in a preliminary anomaly set. For this preliminary anomaly set, a pattern sequence construction process is executed, grouping the anomaly candidates by time window division to obtain serialized anomaly patterns. After obtaining the serialized anomaly patterns, anomaly identification is performed, comparing and analyzing historical data to determine if the current pattern matches known anomaly characteristics. Based on the anomaly identification results, the trend judgment is updated; if the pattern characteristics are consistent with historical anomalies, a final anomaly confirmation record is generated. Using the final anomaly confirmation record and the contextual information from indicator analysis, a detailed description document of the anomaly pattern is generated, completing the entire analysis process.
[0064] Step S105: For the potential abnormal pattern sequence, similar patterns are grouped using a data grouping algorithm to obtain classified abnormal groups.
[0065] A preliminary set of anomalous patterns is obtained by processing latent sequences. For these latent sequences, a pre-established feature extraction tool is used to separate key pattern fragments, resulting in a candidate set of anomalous patterns. Based on this candidate set, cluster analysis is used to group similar patterns. For the pattern fragments in the candidate set, distance indices between patterns are calculated to determine the clustering results for similar patterns. The criteria for classifying these groups are then determined based on the clustering results. If the distance between patterns in the clustering results is below a preset threshold, the corresponding patterns are grouped into the same group, thus identifying the preliminary classification groups. From these preliminary classification groups, a refined structure for anomalous classification is obtained. For patterns within each group, the distribution patterns of their sequence features are analyzed to determine the subcategories of the anomalous classification. Based on these subcategories, the focus of anomalous detection is identified. For the pattern features within the subcategories, a preset anomalous rule base is matched to determine the set of high-risk anomalous patterns. Using the high-risk anomalous pattern set, the final anomalous group division results are obtained. For high-risk patterns, the final classification anomalous groups are determined by combining the contextual information from sequence analysis. Based on the final classification of abnormal groups, obtain a complete mapping of abnormal patterns. For the pattern distribution within each group, generate corresponding anomaly detection labels to determine the classification affiliation of the abnormal patterns.
[0066] like Figure 8 As shown, the abnormal pattern analysis and screening process of this invention is divided into three levels. In the time series analysis stage, the comprehensive abnormal index shows obvious fluctuation characteristics with the change of the time series. The abnormal points 20, 40, 60, 80 and 95 are marked as abnormal points 1 to 5, respectively. The abnormal index values corresponding to these abnormal points reach 0.95, 0.55, 0.65, 0.70 and 0.65, respectively, which are significantly higher than the normal range of 0.3 to 0.4. By extracting features from these abnormal fluctuations through the time series analysis model, potential abnormal pattern sequences can be identified. In the clustering and grouping stage, the extracted abnormal pattern sequences are input into the data grouping algorithm, and they are divided into three classification abnormal groups: group A, group B and group C, according to the similarity of abnormal features. Group A contains 8 similar pattern points, group B contains 10 similar pattern points, and group C contains 7 similar pattern points. The distribution distance of the pattern points in each group in the abnormal feature space is less than a set threshold of 0.5. In the high-risk screening phase, a risk assessment model was applied to quantitatively score the categorized anomalous groups. Groups B and C, due to their anomalous intensity and duration exceeding the risk threshold, were marked as high-risk groups and indicated with dark fill and bold borders, while groups A and D were low-risk groups, represented by light-colored dashed borders. The final high-risk anomalous subset included all anomalous patterns from groups B and C, totaling 17 anomalous areas requiring close attention, providing crucial information for subsequent frost damage identification and decision-making.
[0067] Step S106: Determine a high-risk anomaly subset based on the classified anomaly groups.
[0068] Initial data is obtained from anomalous groups, and the groups are decomposed into multiple subsets based on the group segmentation logic. Each subset is matched with the classification criteria using a data splitting tool to obtain preliminary segmentation results. For these preliminary results, a feature selection method is used to extract key features from each anomalous subset. By comparing these features with a pre-defined feature library, it is determined whether the key features meet high-risk criteria; if so, they are marked as high-risk subsets. Based on these high-risk subsets, relevant information on potential threats is obtained. Using a threat level assessment model combined with risk assessment indicators, each high-risk subset is ranked by threat level to determine threat priority. For threat priority, detailed data on anomaly identification is obtained. Using a data comparison tool, key features are correlated with anomaly identification records to obtain specific anomaly patterns within the anomalous subsets. Based on these anomaly patterns, a logistic regression model is used to analyze the correlation between potential threats and anomalous subsets. The model output is used to determine whether further segmentation of the anomalous subsets is needed, resulting in refined subset units. For these refined subset units, the criteria for priority ranking are obtained. A ranking algorithm is used to comprehensively evaluate threat levels and anomaly patterns to determine the final list of high-risk anomalous subsets.
[0069] Step S107: Based on the high-risk anomaly subset, generate a fusion assessment report to obtain the final threat level label.
[0070] By initially screening the data of a high-risk anomaly subset, key field information within the anomaly set is obtained to determine the initial risk value range. Based on the risk value range obtained from the initial screening, a preset threshold standard is used for comparison. If the risk value exceeds the threshold range, it is marked as a high-risk subset, resulting in a high-risk labeling result. For the high-risk labeling result, the association data of each element within the subset is obtained, and a random forest algorithm is used to classify the association data to determine if potential threat patterns exist. Through the classified threat pattern data, feature information related to the threat level is obtained, and the mapping relationship between feature information and level labels is determined. Based on the mapping relationship, the basis for generating a comprehensive assessment report is obtained, and it is determined whether the basis meets the conditions for report generation. Based on the generation basis that meets the conditions, the final result judgment data is obtained, determining the final correspondence between threat level and label. For the data of the final correspondence, an information integration tool is used to generate a comprehensive assessment report, obtaining the final level labeling result.
[0071] like Figures 11(a)-11(b)As shown, waveform comparison analysis was performed on the odor-speech fusion signal under normal and abnormal states. Figure 11(a) shows the signal waveform under normal state. The odor concentration signal fluctuates steadily within the range of 0.25 to 0.40, the speech energy signal maintains a normal rhythmic range of 0.40 to 0.60, and the comprehensive abnormal index changes steadily between 0.30 and 0.50. None of the three curves exceed the set safety threshold of 0.70, and the signal as a whole exhibits stable periodic characteristics. Figure 11(b) shows the signal waveform under abnormal state. Within a time period of 40 to 60 seconds, the odor concentration signal shows an abnormal peak, reaching a maximum of 0.90, significantly exceeding the normal range; the speech energy signal shows high-frequency fluctuations during this period, with the amplitude increasing to about 0.70; the comprehensive abnormal index rises rapidly in the abnormal area, with a peak value exceeding 0.90, clearly breaking through the safety threshold of 0.70, triggering a threat alarm. Through comparative analysis, abnormal pattern sequences can be clearly identified, providing a reliable basis for subsequent threat level judgment. Abnormal areas are marked with gray shading for easy visual identification of potential threat periods.
[0072] like Figure 12 As shown, to verify the effectiveness of the method of the present invention, comparative experiments were conducted between the multimodal fusion detection method of the present invention and the traditional single-modal detection method in five different test scenarios. The test scenarios included: scenario 1 for chemical leak detection, scenario 2 for emotional abnormality detection, scenario 3 for comprehensive abnormality detection (simultaneous odor and speech abnormalities), scenario 4 for weak signal detection (trace amounts of gas or low sound pressure speech), and scenario 5 for strong interference environment detection (background noise or mixed odors). The experimental results show that the method of the present invention is significantly superior to the traditional single-modal method in all test scenarios. Specifically, in scenario 1 (chemical detection), the accuracy of the method of this invention reached 94.5%, which is 6.3 percentage points higher than the odor-only detection method and 29.2 percentage points higher than the speech-only detection method. In scenario 2 (emotional anomaly detection), the accuracy of the method of this invention was 92.8%, which is 7.2 percentage points higher than the speech-only detection method and 20.3 percentage points higher than the odor-only detection method. In scenario 3 (comprehensive anomaly detection), the accuracy of the method of this invention reached a maximum of 95.2%, while the accuracy of the two single-modal methods was below 79%, fully demonstrating the advantages of multimodal fusion. In scenario 4 (weak signal detection) and scenario 5 (strong interference environment), the accuracy of the method of this invention was 89.3% and 91.7% respectively, maintaining a high level, while the accuracy of the single-modal methods dropped to below 72%. The above comparative experimental results fully verify that the present invention, by fusing odor and speech modal information, can achieve complementary enhancement, significantly improving the accuracy and robustness of anomaly detection, especially in complex scenarios and under weak signal conditions, where the advantages of the multimodal fusion method are more obvious.
[0073] This invention provides a system for recognizing speech and odor fusion data based on AIoT edge devices and cloud-based LLM large models, mainly comprising:
[0074] The signal acquisition and filtering module is used to acquire odor signals and speech signals through a signal acquisition device, and to obtain preliminary odor molecule concentration sequences and speech waveform sequences after removing environmental interference noise using a filtering method.
[0075] The feature extraction module is used to extract odor pattern vectors and speech emotion vectors based on the preliminary odor molecule concentration sequence and the speech waveform sequence using a feature extraction model, thereby determining a preliminary multimodal feature set;
[0076] The similarity judgment and fusion module is used to obtain a comprehensive anomaly index vector by weighted fusion of the odor pattern vector and the voice emotion vector in the multimodal preliminary feature set if the similarity is lower than a threshold; otherwise, the original vector is retained as the comprehensive anomaly index vector.
[0077] The time series analysis module is used to obtain the comprehensive anomaly index vector and then use a time series analysis model to analyze its time series changes to determine potential anomaly pattern sequences.
[0078] The data grouping module is used to group similar patterns into classification anomaly groups based on the potential anomaly pattern sequence using a data grouping algorithm.
[0079] The high-risk determination module is used to determine a high-risk anomaly subset based on the classified anomaly groups;
[0080] The report generation module is used to generate a fusion assessment report based on the high-risk anomaly subset to obtain the final threat level label.
[0081] like Figures 13(a)-13(b)As shown, this invention provides a detailed analysis of the system response time and processing latency. Figure 13(a) illustrates the time distribution of each processing stage. The similarity calculation and fusion stage accounts for the highest time consumption, reaching 25.0%, because this stage requires comprehensive calculation and weighted fusion of multiple feature dimensions. The feature extraction stage accounts for 20.8% of the time consumption, mainly involving the parallel extraction of features in the time domain, frequency domain, and wavelet domain. The time series analysis stage accounts for 16.7%, used to detect the dynamic evolution characteristics of the signal. The classification and grouping stage accounts for 15.0%, responsible for clustering signals based on feature similarity. The signal acquisition and filtering stage accounts for 12.5%, completing the preprocessing of the original data. The risk assessment stage accounts for 10.0%, determining the risk level based on the classification results. Figure 13(b) shows the trend of the system's end-to-end response time with the amount of data. When the amount of data increases from 100 samples to 5000 samples, the edge processing time increases from 50ms to 280ms, the cloud processing time increases from 80ms to 340ms, and the total response time increases from 130ms to 620ms. It can be seen that the edge computing architecture has a significant latency advantage with small data volumes. However, as the data volume increases, the parallel computing capabilities of cloud processing gradually become apparent, keeping the overall response time increase within a reasonable range. This analysis verifies the rationality of the cloud-edge collaborative architecture adopted in this invention, which can achieve low system latency while ensuring processing accuracy.
[0082] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to any specific implementation. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention.
Claims
1. A security inspection method based on AIoT edge devices and cloud-based LLM large-scale models to achieve voice and odor fusion data recognition, characterized in that, The method includes: Odor and speech signals are collected by AIoT edge devices. After removing environmental interference noise by filtering methods, preliminary odor molecule concentration sequences and speech waveform sequences are obtained and uploaded to the cloud-based LLM large model. The cloud-based LLM model extracts odor pattern vectors and speech emotion vectors based on the preliminary odor molecule concentration sequence and the speech waveform sequence, thereby determining the preliminary multimodal feature set. If the similarity between the odor pattern vector and the speech emotion vector in the multimodal preliminary feature set is lower than a threshold, then the two are weighted and fused through an attention mechanism to obtain a comprehensive anomaly index vector; otherwise, the original vector is retained as the comprehensive anomaly index vector. After obtaining the comprehensive anomaly index vector, a time series analysis model is used to analyze its time series changes in order to determine potential anomaly pattern sequences. For the potential abnormal pattern sequence, similar patterns are grouped using a data grouping algorithm to obtain classified abnormal groups; Based on the aforementioned anomaly groups, a high-risk anomaly subset is determined; Based on the high-risk anomaly subset, a fusion assessment report is generated to obtain the final threat level label; The step of determining a high-risk anomaly subset based on the classified anomaly groups includes: Initial data is obtained from abnormal groups, and multiple subset units are decomposed based on the logic of group division; By using data splitting tools, each subset unit is matched with the classification criteria to obtain preliminary division results; Based on the preliminary segmentation results, a feature selection method was used to extract key features from each abnormal subset; By comparing with a pre-defined feature library, it is determined whether key features meet the high-risk criteria. If they do, they are marked as a high-risk subset. Based on the marked high-risk subset, obtain relevant information about potential threats; By using a threat level assessment model and combining it with risk assessment indicators, the threat levels of each high-risk subset are ranked to determine the threat priority. Obtain detailed data on anomaly identification based on threat priority; By using data comparison tools, key features are associated with anomaly identification records to obtain specific anomaly patterns of anomaly subsets; Based on the anomaly patterns, analyze the correlation between potential threats and anomalous subsets; Based on the model output, determine whether the abnormal subset needs to be further split, and obtain the refined subset units; For the refined subset units, obtain the basis for priority sorting; By using a sorting algorithm, the threat level and anomaly pattern are comprehensively evaluated to determine the final list of high-risk anomaly subsets.
2. The method according to claim 1, characterized in that, The edge device includes an odor sensor and a microphone.
3. The method according to claim 1, characterized in that, After obtaining the comprehensive anomaly index vector, the time-series analysis model is used to analyze its temporal changes to determine potential anomaly pattern sequences, including: Irrelevant noise is removed by pre-defined filtering rules to obtain the processed set of indicators; Based on the processed set of indicators, identify the key fluctuation points in the trend; Based on key fluctuation points and the criteria for identifying potential anomalies, a preliminary set of anomalies is obtained. For the initial set of anomalies, the process of constructing a pattern sequence is executed. Anomaly candidates are grouped by dividing the time window to obtain serialized anomaly patterns. After obtaining the serialized abnormal pattern, perform an anomaly identification operation to determine whether the current pattern matches known anomaly characteristics; Based on the results of anomaly identification, if the pattern characteristics are consistent with historical anomalies, a final anomaly confirmation record is generated. By combining the final anomaly confirmation record with the contextual information from the indicator analysis, a detailed description document of the anomaly pattern is generated, completing the entire analysis process.
4. The method according to claim 1, characterized in that, The step of grouping similar patterns into classification anomaly groups using a data grouping algorithm for the potential anomaly pattern sequence includes: By processing the potential sequences, a preliminary set of abnormal patterns is obtained; For the potential sequence, key pattern fragments are separated to obtain a candidate set of anomalous patterns; Group similar patterns based on the candidate set of anomalous patterns; For pattern fragments in the candidate set, the distance index between patterns is calculated, and then the clustering result of similar patterns is determined; Based on the clustering results of similar patterns, obtain the criteria for dividing the classification groups; If the distance between patterns in the clustering results is lower than a preset threshold, the corresponding patterns will be grouped into the same group to determine the preliminary classification group. From the initial classification groups, obtain the refined structure of the anomaly classification; For each group's pattern, determine the subcategories for anomaly classification; Based on the subcategories of anomalies, identify the key areas for anomaly detection; Based on the pattern features in the subcategories, a pre-defined anomaly rule base is matched to identify a set of high-risk anomaly patterns; The final anomaly group classification results are obtained by using a set of high-risk anomaly patterns; For high-risk patterns, the final classification of anomalous groups is determined by combining contextual information from sequence analysis; Based on the final classification of abnormal groups, obtain the complete mapping of abnormal patterns; Based on the pattern distribution within the group, corresponding anomaly detection labels are generated to determine the classification of the anomaly patterns.
5. The method according to claim 1, characterized in that, The step of generating a fusion assessment report based on the high-risk anomaly subset to obtain the final threat level label includes: By initially screening the data of the high-risk anomaly subset, key field information in the anomaly set is obtained, and the initial risk value range is determined. Based on the risk value range obtained from the initial screening, a preset threshold standard is used for comparison. If the risk value exceeds the threshold range, it is marked as a high-risk subset, and the high-risk marking result is obtained. For the high-risk marking results, the association data of each element in the subset group is obtained, and the random forest algorithm is used to classify the association data to determine whether there are potential threat patterns. By using the classified threat pattern data, we can obtain feature information related to the threat level and determine the mapping relationship between the feature information and the level label. Based on the mapping results, obtain the basis for generating the comprehensive evaluation report and determine whether the basis meets the conditions for report generation; By obtaining the final judgment data based on the conditions met, the final correspondence between threat level and tag is determined. Based on the data of the final correspondence, an information integration tool is used to generate a comprehensive assessment report, and the final level labeling result is obtained.
6. A security inspection system based on AIoT edge devices and cloud-based LLM large-scale models to achieve voice and odor data fusion recognition, characterized in that, The system includes: The signal acquisition and filtering module is used to acquire odor signals and speech signals through AIoT edge devices, and to obtain preliminary odor molecule concentration sequences and speech waveform sequences after removing environmental interference noise using filtering methods. The feature extraction module is used to extract odor pattern vectors and speech emotion vectors based on the preliminary odor molecule concentration sequence and the speech waveform sequence using a feature extraction model, thereby determining a preliminary multimodal feature set; The similarity judgment and fusion module is used to obtain a comprehensive anomaly index vector by weighted fusion of the odor pattern vector and the voice emotion vector in the multimodal preliminary feature set if the similarity is lower than a threshold; otherwise, the original vector is retained as the comprehensive anomaly index vector. The time series analysis module is used to obtain the comprehensive anomaly index vector and then use a time series analysis model to analyze its time series changes to determine potential anomaly pattern sequences. The data grouping module is used to group similar patterns into classification anomaly groups based on the potential anomaly pattern sequence using a data grouping algorithm. The high-risk determination module is used to determine a high-risk anomaly subset based on the classified anomaly groups; The report generation module is used to generate a fusion assessment report based on the high-risk anomaly subset to obtain the final threat level label; The step of determining a high-risk anomaly subset based on the classified anomaly groups includes: Initial data is obtained from abnormal groups, and multiple subset units are decomposed based on the logic of group division; By using data splitting tools, each subset unit is matched with the classification criteria to obtain preliminary division results; Based on the preliminary segmentation results, a feature selection method was used to extract key features from each abnormal subset; By comparing with a pre-defined feature library, it is determined whether key features meet the high-risk criteria. If they do, they are marked as a high-risk subset. Based on the marked high-risk subset, obtain relevant information about potential threats; By using a threat level assessment model and combining it with risk assessment indicators, the threat levels of each high-risk subset are ranked to determine the threat priority. Obtain detailed data on anomaly identification based on threat priority; By using data comparison tools, key features are associated with anomaly identification records to obtain specific anomaly patterns of anomaly subsets; Based on the anomaly patterns, analyze the correlation between potential threats and anomalous subsets; Based on the model output, determine whether the abnormal subset needs to be further split, and obtain the refined subset units; For the refined subset units, obtain the basis for priority sorting; By using a sorting algorithm, the threat level and anomaly pattern are comprehensively evaluated to determine the final list of high-risk anomaly subsets.
7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the method as described in any one of claims 1-5.