In-vehicle person hiding detection method based on multi-modal data and related equipment
By acquiring and decomposing the vehicle surface and overall vibration signals and combining with the classification network for detection, the problem of poor reliability of existing vehicle personnel hiding detection methods is solved, and higher detection accuracy and reliability are achieved.
Patent Information
- Application Number
- CN202510051359.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-13
- Publication Date
- 2025-05-13
AI Technical Summary
The existing detection methods for hiding personnel in the vehicle are poorly reliable, making it difficult to effectively detect slight changes in the human body inside the vehicle body, and are greatly affected by external factors.
The in-vehicle personnel hiding detection method based on multimodal data is adopted. By obtaining the surface vibration signal and overall vibration signal of the target vehicle, the respiratory signal component is decomposed and classified using a classification network to obtain the personnel hiding detection results.
The reliability of personnel hiding detection is improved. Through the combination of multimodal information, the risk of singlemodal information being disturbed by noise is reduced, and the accuracy of detection is enhanced.
Smart Images

Figure CN119989078A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of vehicle personnel detection, and in particular to a method for detecting hidden personnel in a vehicle based on multimodal data and related equipment. Background Art
[0002] As the global security situation becomes increasingly severe, vehicles are increasingly used as a covert means of illegal activities. Among them, concealed carrying of people is particularly difficult to detect. Existing vehicle inspection technologies mainly rely on methods such as X-ray scanning and metal detectors. Although these methods are effective to a certain extent, they have some inevitable defects, such as large equipment, high cost, easy to cause harm to the human body, complex operation or low detection efficiency.
[0003] Although magnetic induction, thermal imaging and acoustic wave detection technologies have certain application value as supplements in specific situations, they are usually sensitive to environmental conditions and are easily affected by external factors such as weather, the material and structure of the vehicle itself, resulting in relatively high false alarm and missed alarm rates. In addition, these technologies are often unable to effectively detect small changes in the vehicle body caused by physiological activities (such as breathing, heartbeat, etc.), which is crucial for detecting whether there are people hiding in a confined space. Laser vibrometer, as a non-contact detection method, can effectively detect tiny breathing vibrations of the human body, but noise interference in complex environments will affect its accuracy. It can be seen that the existing methods for detecting hidden people in the car have the problem of poor reliability in detecting hidden people in the car. Summary of the invention
[0004] The present application provides a method and related equipment for detecting hidden persons in a vehicle based on multimodal data, which can solve the problem of poor reliability of hidden persons in the vehicle.
[0005] In a first aspect, an embodiment of the present application provides a method for detecting a hidden person in a vehicle based on multimodal data, the method comprising:
[0006] Acquire surface vibration signals and overall vibration signals of the target vehicle;
[0007] Decomposing the surface vibration signal to obtain a first breathing signal component, and decomposing the overall vibration signal to obtain a second breathing signal component; the first breathing signal component is a signal component in the surface vibration signal related to human breathing, and the second breathing signal component is a signal component in the overall vibration signal related to human breathing;
[0008] Based on the first breathing signal component and the second breathing signal component, classification is performed using a classification network to obtain a hidden person detection result of the target vehicle; the hidden person detection result is used to describe whether there is a hidden person in the target vehicle.
[0009] Optionally, decomposing the surface vibration signal to obtain a first breathing signal component includes:
[0010] The overall vibration signal is decomposed and identified, and a second breathing signal component is identified from the overall vibration signal.
[0011] Optionally, decomposing the overall vibration signal to obtain a second breathing signal component includes:
[0012] The overall vibration signal is decomposed and identified, and a second breathing signal component is identified from the overall vibration signal.
[0013] Optionally, based on the first breathing signal component and the second breathing signal component, a classification network is used to perform classification to obtain a person hiding detection result of the target vehicle, including:
[0014] The first breathing signal component and the second breathing signal component are fused by using a classification network to obtain a fused signal, and classification detection is performed based on the fused signal to obtain a hidden person detection result of the target vehicle.
[0015] Optionally, the classification network includes a convolution unit, a first pooling unit, a first activation function unit, a graph neural unit, a dynamic weight learning unit, a second pooling unit, a second activation function unit, a multi-scale fusion unit and an attention mechanism unit connected in sequence;
[0016] The input of the convolution unit is the input of the classification network, and the output of the attention mechanism unit is the output of the classification network.
[0017] Optionally, the convolution unit includes a first convolution layer, a second convolution layer, and a third convolution layer connected in sequence;
[0018] The input end of the first convolutional layer is the input end of the convolutional unit, and the output end of the third convolutional layer is the output end of the convolutional unit.
[0019] Optionally, the second pooling unit includes a first pooling layer and a second pooling layer connected in sequence;
[0020] The input end of the first pooling layer is the input end of the second pooling unit, and the output end of the second pooling layer is the output end of the second pooling unit;
[0021] The second activation function unit includes a first activation function layer and a second activation function layer connected in sequence;
[0022] The input end of the first activation function layer is the input end of the second activation function unit, and the output end of the second activation function layer is the output end of the second activation function unit.
[0023] In a second aspect, an embodiment of the present application provides a vehicle interior person hiding detection device based on multimodal data, comprising:
[0024] An acquisition module, for acquiring a surface vibration signal and an overall vibration signal of a target vehicle;
[0025] A decomposition module decomposes the surface vibration signal to obtain a first breathing signal component, and decomposes the overall vibration signal to obtain a second breathing signal component; the first breathing signal component is a signal component related to human breathing in the surface vibration signal, and the second breathing signal component is a signal component related to human breathing in the overall vibration signal;
[0026] The classification module uses a classification network to perform classification based on the first breathing signal component and the second breathing signal component to obtain a hidden person detection result of the target vehicle; the hidden person detection result is used to describe whether there is a hidden person in the target vehicle.
[0027] In a third aspect, an embodiment of the present application provides a terminal device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the above-mentioned method for detecting hidden persons in a vehicle based on multimodal data when executing the above-mentioned computer program.
[0028] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the above-mentioned method for detecting hidden people in a vehicle based on multimodal data.
[0029] The above solution of the present application has the following beneficial effects:
[0030] In an embodiment of the present application, the surface vibration signal and the overall vibration signal of the target vehicle are obtained, and then the surface vibration signal is decomposed to obtain the first breathing signal component, and the overall vibration signal is decomposed to obtain the second breathing signal component, and finally based on the first breathing signal component and the second breathing signal component, the classification network is used for classification to obtain the target vehicle's hidden personnel detection result. Among them, the surface vibration signal and the overall vibration signal of the target vehicle are analyzed, and the information of the surface vibration and the overall vibration of the target vehicle is considered to improve the richness of information, and multi-modal information is used for hiding detection to avoid the situation where the single-modal information is interfered by noise and causes poor detection accuracy. Based on the rich information, the classification network is used for classification to improve the accuracy of classification, thereby improving the reliability of the hidden personnel detection result.
[0031] Other beneficial effects of the present application will be described in detail in the subsequent specific implementation section. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0033] Figure 1 A flowchart of a method for detecting a hidden person in a vehicle based on multimodal data provided in an embodiment of the present application;
[0034] Figure 2 A schematic diagram of the structure of a classification network provided in an embodiment of the present application;
[0035] Figure 3 A schematic diagram of the structure of a vehicle occupant hiding detection device based on multimodal data provided in one embodiment of the present application;
[0036] Figure 4 A schematic diagram of the structure of a terminal device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0037] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.
[0038] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, wholes, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components and / or combinations thereof.
[0039] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.
[0040] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.
[0041] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.
[0042] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.
[0043] In view of the problem of poor reliability of existing detection of hidden persons in vehicles, an embodiment of the present application provides a method for detecting hidden persons in vehicles based on multimodal data. The method for detecting hidden persons in vehicles obtains the surface vibration signal and the overall vibration signal of the target vehicle, then decomposes the surface vibration signal to obtain a first breathing signal component, and decomposes the overall vibration signal to obtain a second breathing signal component. Finally, based on the first breathing signal component and the second breathing signal component, a classification network is used for classification to obtain the hidden person detection result of the target vehicle. The surface vibration signal and the overall vibration signal of the target vehicle are analyzed, and the information of the surface vibration and the overall vibration of the target vehicle is considered to improve the richness of information. Multimodal information is used for hiding detection to avoid the situation where the single-modal information is interfered by noise and leads to poor detection accuracy. The classification network is used for classification based on the rich information to improve the accuracy of classification, thereby improving the reliability of the hidden person detection result.
[0044] Next, an exemplary description is given of the method for detecting hidden persons in a vehicle based on multimodal data provided by the present application.
[0045] like Figure 1 As shown, the method for detecting hidden persons in a vehicle based on multimodal data provided by the present application comprises the following steps:
[0046] Step 11, obtaining the surface vibration signal and the overall vibration signal of the target vehicle.
[0047] The target vehicle is a vehicle that needs to be detected for hidden personnel inside. The surface vibration signal is used to describe the vibration information of the surface of the target vehicle (such as the vibration of the vehicle shell, hood, etc.), and the overall vibration signal is used to describe the vibration information of multiple parts of the target vehicle (such as the vibration of the vehicle interior space, engine, etc.).
[0048] In some embodiments of the present application, a surface vibration signal of a target vehicle may be acquired by using a laser vibrometer or other equipment, and an overall vibration signal of the target vehicle may be acquired by installing a vibration sensor or other equipment on the outside of the target vehicle.
[0049] It is worth mentioning that the collection of surface vibration signals and overall vibration signals does not require reliance on X-ray and other equipment, which improves the safety and efficiency of detection of hidden people in the vehicle and reduces the cost of detection of hidden people in the vehicle.
[0050] Step 12: decompose the surface vibration signal to obtain a first breathing signal component, and decompose the overall vibration signal to obtain a second breathing signal component.
[0051] The first breathing signal component is a signal component in the surface vibration signal related to human breathing, and the second breathing signal component is a signal component in the overall vibration signal related to human breathing.
[0052] In some embodiments of the present application, the steps of decomposing the surface vibration signal to obtain the first breathing signal component and decomposing the overall vibration signal to obtain the second breathing signal component are specifically:
[0053] In the first step, the surface vibration signal is decomposed to obtain the first breathing signal component.
[0054] Perform empirical mode decomposition on the surface vibration signal to obtain the first breathing signal component.
[0055] Exemplarily, the surface vibration signal is decomposed using an empirical mode decomposition (EMD) algorithm to obtain multiple first signal components, which respectively correspond to different features in the surface vibration signal. Then, the first signal component corresponding to the vibration of human breathing is used as the first breathing signal component.
[0056] In the second step, the overall vibration signal is decomposed to obtain the second breathing signal component.
[0057] The overall vibration signal is decomposed and identified, and a second breathing signal component is identified from the overall vibration signal.
[0058] Exemplarily, the overall vibration signal can be decomposed using an empirical mode decomposition algorithm to obtain multiple second signal components. Then, all the second signal components can be identified through spectral analysis and other methods to identify the vibration mode corresponding to each second signal component (such as the vibration of the vehicle's engine, the vibration of the suspension system, the vibration of human breathing, etc.), and the second signal component corresponding to the vibration of human breathing can be used as the second breathing signal component.
[0059] It is worth mentioning that by decomposing and identifying the surface vibration signal and the overall vibration signal respectively, the first breathing signal component and the second breathing signal component are determined, which can describe the vibration signal generated by the breathing of people in the target vehicle in terms of surface vibration and overall vibration respectively, providing rich information for subsequent steps.
[0060] Step 13: Based on the first breathing signal component and the second breathing signal component, classification is performed using a classification network to obtain a hidden person detection result of the target vehicle.
[0061] The above-mentioned hidden person detection result is used to describe whether there is a hidden person in the target vehicle.
[0062] Specifically, the first breathing signal component and the second breathing signal component are fused using a classification network to obtain a fused signal, and classification detection is performed based on the fused signal to obtain a hidden person detection result of the target vehicle. After the classification result is obtained by using the classification network for classification, it is necessary to analyze the classification result and the known number of people in the target vehicle to obtain a hidden person detection result. For example, the classification result is a category label, which expresses the number of people in the target vehicle detected according to the vibration information (such as 1 person, 2 people, 3 people, etc.). It is known that there is a driver and a passenger in the target vehicle, but the classification result of the classification network is 3 people, indicating that there is one person hiding in the vehicle, and the hidden person detection result is one person hiding.
[0063] In some embodiments of the present application, the structure of the above classification network is as follows: Figure 2 As shown, it includes a convolution unit, a first pooling unit, a first activation function unit, a graph neural unit, a dynamic weight learning unit, a second pooling unit, a second activation function unit, a multi-scale fusion unit and an attention mechanism unit which are connected in sequence; the input end of the convolution unit is the input end of the classification network, and the output end of the attention mechanism unit is the output end of the classification network.
[0064] The convolution unit includes a first convolution layer, a second convolution layer and a third convolution layer which are connected in sequence; an input end of the first convolution layer is an input end of the convolution unit, and an output end of the third convolution layer is an output end of the convolution unit.
[0065] The second pooling unit includes a first pooling layer and a second pooling layer connected in sequence; the input end of the first pooling layer is the input end of the second pooling unit, and the output end of the second pooling layer is the output end of the second pooling unit.
[0066] The second activation function unit includes a first activation function layer and a second activation function layer connected in sequence; the input end of the first activation function layer is the input end of the second activation function unit, and the output end of the second activation function layer is the output end of the second activation function unit.
[0067] It should be noted that the operations in the above-mentioned graph neural unit are operations of graph neural networks, the operations in the above-mentioned dynamic weight learning unit are dynamic weight learning algorithms, the operations in the first activation function unit and the first activation function layer and the second activation function layer in the second activation function unit can both be ReLU functions, and the operations in the above-mentioned multi-scale fusion unit can be a multi-scale feature fusion algorithm based on a convolutional neural network. The above-mentioned convolution unit is used to perform convolution operations on two input data (the first respiratory signal component and the second respiratory signal component) at the same time, the first pooling unit and the second pooling unit are both used to perform pooling operations on the input data, the first activation function unit and the second activation function unit are both used to perform activation function operations on the input data, the graph neural unit is used to calculate the graph neural unit on the input data, the dynamic weight learning unit is used to dynamically adjust the weight of the input data, the multi-scale fusion unit is used to perform multi-scale fusion of the input data to obtain fused data, and the attention mechanism unit is used to perform attention mechanism operations on the input data to obtain classification results; the first convolution layer, the second convolution layer and the third convolution layer are all used to perform convolution operations on the input data, the first pooling layer and the second pooling layer are both used to perform pooling operations on the input data, and the first activation function layer and the second activation function layer are both used to perform activation function operations on the input data.
[0068] The above classification network combines the powerful feature extraction capability of convolutional neural networks with the ability of graph neural networks to understand complex structured data. By embedding graph neural networks in traditional deep learning architectures, it can simultaneously learn the node features and edge relationships of input data. Multi-scale fusion is designed in the classification network to integrate local features and global context information at different levels of abstraction, thereby improving the accuracy and robustness of classification. The attention mechanism is introduced to achieve adaptive selection of input features, so that the classification network can focus on the most critical features for the classification task during training and ignore noise or irrelevant information.
[0069] Exemplarily, after obtaining the hidden person detection result, the hidden person detection result can be displayed on the user interface for user viewing.
[0070] It is worth mentioning that the surface vibration signal and the overall vibration signal of the target vehicle are analyzed, and the surface vibration and overall vibration information of the target vehicle are taken into consideration to improve the richness of information. Multimodal information is used for hiding detection to avoid the situation where the single modal information is interfered by noise and leads to poor detection accuracy. Based on the rich information, classification is performed using a classification network to improve the accuracy of classification, thereby improving the reliability of the results of hidden person detection.
[0071] In addition, through the combination of laser vibrometer and vibration sensor, high-precision detection of people hidden in the vehicle is achieved, overcoming the limitations of a single sensor in complex environments and improving the reliability and practicality of detection.
[0072] The following is an exemplary description of the vehicle interior person hiding detection device based on multimodal data provided by the present application.
[0073] like Figure 3 As shown, an embodiment of the present application provides a vehicle interior person hiding detection device based on multimodal data, and the vehicle interior person hiding detection device 300 based on multimodal data includes:
[0074] An acquisition module 301 is used to acquire a surface vibration signal and an overall vibration signal of a target vehicle;
[0075] The decomposition module 302 decomposes the surface vibration signal to obtain a first breathing signal component, and decomposes the overall vibration signal to obtain a second breathing signal component; the first breathing signal component is a signal component related to human breathing in the surface vibration signal, and the second breathing signal component is a signal component related to human breathing in the overall vibration signal;
[0076] The classification module 303 performs classification using a classification network based on the first breathing signal component and the second breathing signal component to obtain a hidden person detection result of the target vehicle; the hidden person detection result is used to describe whether there is a hidden person in the target vehicle.
[0077] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.
[0078] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.
[0079] like Figure 4 As shown, an embodiment of the present application provides a terminal device. The terminal device D10 of this embodiment includes: at least one processor D100 ( Figure 4 Only one processor is shown in the figure), a memory D101, and a computer program D102 stored in the memory D101 and executable on the at least one processor D100, wherein the processor D100 implements the steps of any of the above-mentioned method embodiments when executing the computer program D102.
[0080] Specifically, when the processor D100 executes the computer program D102, the surface vibration signal and the overall vibration signal of the target vehicle are obtained, and then the surface vibration signal is decomposed to obtain the first breathing signal component, and the overall vibration signal is decomposed to obtain the second breathing signal component, and finally based on the first breathing signal component and the second breathing signal component, the classification network is used for classification to obtain the personnel hiding detection result of the target vehicle. Among them, the surface vibration signal and the overall vibration signal of the target vehicle are analyzed, and the information of the surface vibration and the overall vibration of the target vehicle is considered to improve the richness of information, and multi-modal information is used for hiding detection to avoid the situation where the single-modal information is interfered by noise and causes poor detection accuracy. The classification network is used for classification based on the rich information to improve the accuracy of classification, thereby improving the reliability of the personnel hiding detection result.
[0081] The processor D100 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.
[0082] In some embodiments, the memory D101 may be an internal storage unit of the terminal device D10, such as a hard disk or memory of the terminal device D10. In other embodiments, the memory D101 may also be an external storage device of the terminal device D10, such as a plug-in hard disk, a smart memory card (SMC, SmartMedia Card), a secure digital (SD, Secure Digital) card, a flash card (Flash Card), etc. equipped on the terminal device D10. Further, the memory D101 may also include both an internal storage unit of the terminal device D10 and an external storage device. The memory D101 is used to store an operating system, an application program, a boot loader (BootLoader), data and other programs, such as the program code of the computer program, etc. The memory D101 may also be used to temporarily store data that has been output or is to be output.
[0083] An embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps in the above-mentioned method embodiments can be implemented.
[0084] An embodiment of the present application provides a computer program product. When the computer program product runs on a terminal device, the terminal device can implement the steps in the above-mentioned method embodiments when executing the computer program product.
[0085] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the process in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes a computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable medium may at least include: any entity or device, recording medium, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium that can carry the computer program code to the vehicle personnel hiding detection method device / terminal device based on multimodal data. For example, a USB flash drive, a mobile hard disk, a disk or an optical disk.
[0086] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0087] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0088] The above is a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles described in the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for detecting hidden persons in a vehicle based on multimodal data, characterized in that: include: Acquire surface vibration signals and overall vibration signals of the target vehicle; Decomposing the surface vibration signal to obtain a first breathing signal component, and decomposing the overall vibration signal to obtain a second breathing signal component; the first breathing signal component is a signal component related to human breathing in the surface vibration signal, and the second breathing signal component is a signal component related to human breathing in the overall vibration signal; Based on the first breathing signal component and the second breathing signal component, classification is performed using a classification network to obtain a hidden person detection result of the target vehicle; the hidden person detection result is used to describe whether there is a hidden person in the target vehicle.
2. The method for detecting a person hiding in a vehicle according to claim 1, characterized in that: Decomposing the surface vibration signal to obtain a first breathing signal component includes: Performing empirical mode decomposition on the surface vibration signal to obtain a first breathing signal component.
3. The method for detecting a person hiding in a vehicle according to claim 1, characterized in that: Decomposing the overall vibration signal to obtain a second breathing signal component includes: The overall vibration signal is decomposed and identified, and a second breathing signal component is identified from the overall vibration signal.
4. The method for detecting a person hiding in a vehicle according to claim 1, characterized in that: The method of performing classification based on the first breathing signal component and the second breathing signal component using a classification network to obtain a person hiding detection result of the target vehicle includes: The first breathing signal component and the second breathing signal component are fused using the classification network to obtain a fused signal, and classification detection is performed based on the fused signal to obtain a person hiding detection result of the target vehicle.
5. The method for detecting a person hiding in a vehicle according to claim 1, characterized in that: The classification network includes a convolution unit, a first pooling unit, a first activation function unit, a graph neural unit, a dynamic weight learning unit, a second pooling unit, a second activation function unit, a multi-scale fusion unit and an attention mechanism unit connected in sequence; The input end of the convolution unit is the input end of the classification network, and the output end of the attention mechanism unit is the output end of the classification network.
6. The method for detecting a person hiding in a vehicle according to claim 5, characterized in that: The convolution unit includes a first convolution layer, a second convolution layer and a third convolution layer connected in sequence; The input end of the first convolution layer is the input end of the convolution unit, and the output end of the third convolution layer is the output end of the convolution unit.
7. The method for detecting a hidden person in a vehicle according to claim 5, characterized in that: The second pooling unit includes a first pooling layer and a second pooling layer connected in sequence; An input end of the first pooling layer is an input end of the second pooling unit, and an output end of the second pooling layer is an output end of the second pooling unit; The second activation function unit includes a first activation function layer and a second activation function layer connected in sequence; The input end of the first activation function layer is the input end of the second activation function unit, and the output end of the second activation function layer is the output end of the second activation function unit.
8. A device for detecting hidden persons in a vehicle based on multimodal data, comprising: An acquisition module, for acquiring a surface vibration signal and an overall vibration signal of a target vehicle; a decomposition module, decomposing the surface vibration signal to obtain a first breathing signal component, and decomposing the overall vibration signal to obtain a second breathing signal component; the first breathing signal component is a signal component related to human breathing in the surface vibration signal, and the second breathing signal component is a signal component related to human breathing in the overall vibration signal; The classification module uses a classification network to perform classification based on the first breathing signal component and the second breathing signal component to obtain a hidden person detection result of the target vehicle; the hidden person detection result is used to describe whether there is a hidden person in the target vehicle.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the method for detecting hidden persons in a vehicle based on multimodal data as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for detecting hidden persons in a vehicle based on multimodal data as described in any one of claims 1 to 7 is implemented.