Method and apparatus for detecting presence of occupant in seat

By combining audio detection methods using a speaker and microphone, along with spectral analysis and neural networks, the problems of poor accuracy and high cost in existing seat detection technologies have been solved, achieving efficient and accurate seat detection.

WO2026152684A1PCT designated stage Publication Date: 2026-07-23HUAWEI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HUAWEI TECH CO LTD
Filing Date
2025-08-08
Publication Date
2026-07-23

AI Technical Summary

Technical Problem

Existing seat detection technologies suffer from poor accuracy and high cost due to factors such as lighting and device limitations, making it difficult to accurately determine whether a seat is occupied.

Method used

Audio is played through the vehicle's speakers and picked up by corresponding microphones. The system uses spectrum analysis and signal change characteristics to determine whether someone is in the seat. This is combined with neural networks for detection, reducing hardware costs.

Benefits of technology

It improves the accuracy of detection, avoids the influence of light and false detections, and reduces hardware costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025113601_23072026_PF_FP_ABST
    Figure CN2025113601_23072026_PF_FP_ABST
Patent Text Reader

Abstract

A method for detecting the presence of an occupant in a seat. The method comprises: playing first audio by means of a first speaker (501), the first speaker being at least one among a plurality of speakers of a vehicle; acquiring a first sound by means of a first microphone to obtain second audio (502), the first sound comprising a sound generated by playing the first audio, the first microphone being a microphone corresponding to a first seat among a plurality of microphones of the vehicle, and the first seat being one of the seats of the vehicle; and, on the basis of the second audio, acquiring the condition of occupant presence of the first seat (503). An apparatus for detecting the presence of an occupant in a seat is further provided. The method can improve the efficiency of detecting the presence of an occupant in a seat and reduce hardware costs.
Need to check novelty before this filing date? Find Prior Art

Description

Personnel on-site testing methods and devices

[0001] This application claims priority to Chinese Patent Application No. 202510068047.X, filed on January 15, 2025, entitled “Method and Apparatus for Detecting Personnel Seated,” the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to smart cockpit technology, and more particularly to a method and apparatus for detecting occupants. Background Technology

[0003] Seat detection technology can accurately determine whether there are passengers in a seat. This is an important basic function of smart cockpits. Based on this function, more and richer functions can be triggered, such as seat belt fastening detection and reminders, personalized cockpit configuration, and audio pickup by position.

[0004] According to classifications by some official testing organizations (such as the China New Car Assessment Program (CNCAP) and the European New Car Assessment Program (ENCAP)), commonly used seat detection technologies can be divided into direct detection and indirect detection. Direct detection refers to determining whether there is a passenger in a seat by directly detecting the characteristics of a person through technical means, such as using optical or infrared cameras, wireless sensing sensors (such as millimeter-wave radar) or ultra-wideband radar for human body detection. Indirect detection refers to inferring whether there is a passenger in a seat by indirectly detecting some measurements, such as seat pressure sensing and capacitance detection.

[0005] However, due to factors such as light and devices, the aforementioned on-site detection technology still suffers from problems such as poor recognition accuracy and high cost. Summary of the Invention

[0006] This application provides a method and apparatus for detecting the presence of people, in order to improve detection efficiency and reduce hardware costs.

[0007] In a first aspect, this application provides a method for detecting the presence of people, comprising: playing a first audio through a first speaker, the first speaker being at least one of a plurality of speakers in a vehicle; acquiring a second audio by acquiring a first sound through a first microphone, the first sound including the sound emitted by playing the first audio, the first microphone being a microphone corresponding to a first seat among a plurality of microphones in the vehicle, the first seat being one of the seats in the vehicle; and obtaining the presence of people in the first seat based on the second audio.

[0008] In this embodiment, an audio signal is played through a speaker in the vehicle, and another audio signal is picked up by a microphone. The presence of people in the seats corresponding to the microphone is then determined based on the picked-up audio. This method is not affected by light, does not cause detection errors due to obstruction, and avoids situations like pressure detection schemes that mistakenly detect non-human objects as people sitting, thereby improving detection efficiency. Furthermore, reusing the design concept of the vehicle's acoustic system can also reduce hardware costs.

[0009] In one possible implementation, the occupant presence detection in this application can be event-triggered. For example, when a user gets out of the car and closes the door, the vehicle engine is turned off, or the vehicle's motor stops, initiating occupant presence detection at these times can prevent elderly people or children from being left behind in a stopped vehicle, thus avoiding safety accidents. Alternatively, when a user opens the car door and gets in, the vehicle engine is started, or the vehicle's motor starts, initiating occupant presence detection at these times can activate personalized services for the occupants, such as seatbelt reminders, personalized seating configurations, and targeted audio pickup, enhancing the user experience. It should be noted that the above events are merely illustrative examples and do not constitute a limitation on the triggering events; this application does not impose any specific limitations in this regard.

[0010] In one possible implementation, the occupant detection in this application can be user-triggered. For example, a user interface can be displayed on the vehicle's central control screen, showing the vehicle's six seats. The user can click on one of these six seats to trigger the occupant detection for that seat. It should be noted that the above example of user-triggered detection can also include voice command, interface selection, physical button selection, etc., and this application does not specifically limit the methods used.

[0011] Typically, in-vehicle speakers consist of three types of sound-producing units: tweeters, midrange speakers, and woofers. Optionally, to support in-seat detection functionality, the cabin speaker system can be optimized to the greatest extent possible during cabin design. For example, adding tweeters (e.g., with a sampling rate of 44.1 kHz or higher) and distributing them throughout the entire cabin. The first speaker in this application can be at least one of multiple speakers installed in the vehicle, and can be a tweeter, a midrange speaker, or a woofer, or a combination of the aforementioned tweeters, midrange speakers, and woofers; no specific limitation is made in this regard.

[0012] The first audio played through the first speaker may include a pre-set test audio segment, which may be ultrasonic audio (ultrasonic audio with a sampling frequency greater than 20 kHz, for example, audio with a sampling frequency of 44 kHz, or frequency-modulated continuous wave (FMCW) with a sampling frequency of 19 kHz to 22 kHz). Since ultrasonic audio is inaudible to the human ear, playing ultrasonic audio while the operator is seated will not interfere with the occupants; alternatively, the test audio may be non-ultrasonic audio, for example, an FMCW modulated signal with a sampling frequency of 8 kHz to 10 kHz, which sounds like periodic noise to the human ear.

[0013] Optionally, the first audio may also include other sounds, such as music, video audio, or call audio. That is, the audio played through the first speaker may be just test audio, or it may include test audio plus ordinary audio, so that the audiovisual needs of the driver and passengers can be met simultaneously when the personnel are present for testing.

[0014] In most cases, in-vehicle microphones are only used for voice interaction and may not necessarily support ultrasonic frequency sampling. Optionally, to support seat detection functionality, the in-cabin microphone system can be optimized to the greatest extent during cabin design. For example, high-frequency microphones (e.g., sampling rate above 44.1 kHz) can be deployed and distributed throughout the entire vehicle cabin. The first microphone in this application can be the microphone corresponding to the first seat among multiple microphones installed in the vehicle, that is, the first microphone can be located near the first seat, such as in front of the first seat or on the seat back. The first microphone can be either a high sampling rate (e.g., sampling rate greater than 20 kHz) microphone or a normal sound microphone (e.g., sampling rate less than 20 kHz), without specific limitations.

[0015] In one possible implementation, when there is one first speaker and one first microphone, the first seat is located between the first speaker and the first microphone; that is, the first seat is located on the line connecting the first speaker and the first microphone. This allows the acoustic path between the first speaker and the first microphone to cover the first seat, thereby improving detection accuracy.

[0016] Optionally, the first seat is a designated seat. For example, a seat selected by the user. Or, for example, a seat selected by the vehicle's control system.

[0017] Optionally, the first seat is the current seat polled in the vehicle in a preset order. For example, multiple seats in the cabin are polled one by one in the order of front to back and left to right, and each seat processed is the first seat.

[0018] Optionally, the first seat can be one of multiple seats detected simultaneously in the vehicle. For example, multiple seats can be detected simultaneously, meaning each seat's corresponding microphone can act as a first microphone to pick up a second audio signal, and then determine whether the seat is occupied based on the respective second audio signal. In this case, each seat can be designated as the first seat, and occupancy detection can be performed separately.

[0019] Based on the first audio, the second audio picked up by the first microphone includes the following two cases:

[0020] (1) The first audio only includes the test audio.

[0021] The first audio signal is played by the first speaker, propagates within the cockpit, and is picked up by the first microphone to obtain the second audio signal. At this point, the second audio signal only contains the audio signal picked up based on the test audio signal, and can be subjected to spectral analysis. For example, the second audio signal can be subjected to a Fourier transform to convert it to the frequency domain before analysis.

[0022] (2) The first audio includes the test audio plus the normal sound.

[0023] The first audio signal is played by the first speaker, propagates within the cockpit, and is picked up by the first microphone to obtain the second audio signal. This second audio signal contains not only the audio signal picked up based on the test audio signal but also the audio signal picked up based on ordinary sound. It can be bandpass filtered and subjected to spectral analysis. For example, the audio signal picked up based on ordinary sound can be filtered out by bandpass filtering, and then a Fourier transform can be performed to convert it to the frequency domain to obtain the processed second audio signal, which can then be analyzed.

[0024] The seating status of the first seat includes either someone sitting in the first seat or no one sitting in the first seat.

[0025] In one possible implementation, when the second audio signal meets the signal change characteristics, it is determined that someone is sitting in the first seat. The signal change characteristics include: the attenuation in the energy domain is greater than a preset threshold, and / or, the signal change in the frequency domain conforms to the pattern of breathing or heartbeat.

[0026] In other words, the basis for determining that someone is sitting in the first seat is that the second audio frequency meets the signal change characteristics. The human body (mainly composed of water and muscle) has two significant effects on audio signals. Firstly, the human acoustic impedance is between 1.45 and 1.70 kgm. -2 s -1 ×10 6 Around 0.00429 kgm, much higher than the atmospheric pressure of 0.00429 kgm. -2 s -1 ×10 6Therefore, in the energy domain, the presence of a person will significantly absorb the energy of the audio signal; on the other hand, due to regular micro-movements of the body such as breathing or heartbeat, the audio signal will also show regular phase changes. Therefore, it is possible to analyze whether the audio signal has regular changes similar to breathing or heartbeat in the frequency domain.

[0027] Based on this, signal change characteristics can include changes in the second audio frequency in the energy domain and / or spectral domain. By analyzing the audio changes from playback to pickup, combined with the influence of the human body on the audio signal, the system can accurately determine whether someone is sitting in the seat. This method is unaffected by light, avoids detection errors due to obstructions, and prevents situations like pressure detection schemes that mistakenly identify non-human objects as occupied. Furthermore, reusing the design principles of a vehicle's acoustic system can reduce hardware costs.

[0028] In this application, obtaining the seating status of the people in the first seat based on the second audio can include the following three methods:

[0029] 1. Compare the second audio with the first audio to obtain the seating information of the person in the first seat.

[0030] The first audio includes a pre-set test audio. Therefore, after obtaining the second audio, the second audio can be compared with the first audio to obtain the change of the second audio compared with the first audio. Then, the components of the change in the energy domain and / or spectrum domain are analyzed to determine whether the second audio meets the signal change characteristics, and thus obtain the detection result of whether someone is sitting in the first seat.

[0031] 2. Perform autocorrelation processing on the second audio to obtain the autocorrelation of the second audio; obtain the seating status of the people in the first seat based on the autocorrelation.

[0032] Correlation analysis between the second audio signal and its own delayed signal yields the corresponding autocorrelation coefficient. The first peak has a value of 1 (indicating a correlation coefficient of 1 between the original signal and itself). The value of the second peak is related to whether the seat is occupied; it is lower when someone is sitting (due to human absorption, the subsequent signal differs significantly from the original signal), and higher when no one is sitting (the chirping waves are more consistent). Therefore, the value of the second peak can be used to determine whether the first seat is occupied. It should be noted that this application can also use other peaks, or even other correlation attributes, for this determination; no specific limitations are imposed.

[0033] 3. Input the second audio signal into a pre-trained neural network to obtain the seating situation of the people in the first seat.

[0034] The neural network of this application can be trained to take a second audio input and output the seating arrangement of the people in the first seat.

[0035] As mentioned above, if the first audio includes both test audio and normal sound, then the second audio contains not only the audio picked up based on the test audio but also the audio picked up based on the normal sound. Therefore, bandpass filtering and spectral analysis must be performed on the second audio first. For example, bandpass filtering can be used to filter out the audio picked up based on the normal sound, and then a Fourier transform can be performed to convert it to the frequency domain to obtain the processed second audio. Afterward, any of the three methods mentioned above can be used to determine whether someone is sitting in the first seat based on the processed second audio. Optionally, even if the first audio only contains test audio, bandpass filtering and spectral analysis can still be performed on the second audio to filter out interfering sounds in the second audio, making the detection results more accurate.

[0036] Secondly, this application provides a person-occupancy detection device, comprising: a playback module for playing a first audio through a first speaker, the first speaker being at least one of a plurality of speakers in a vehicle; a sound receiving module for acquiring a second audio through a first microphone, the first sound including the sound emitted by playing the first audio, the first microphone being a microphone corresponding to a first seat among a plurality of microphones in the vehicle, the first seat being one of the seats in the vehicle; and a detection module for obtaining the person-occupancy status of the first seat based on the second audio.

[0037] In one possible implementation, the detection module is specifically used to determine that someone is sitting in the first seat when the second audio meets the signal change characteristics, wherein the signal change characteristics include: attenuation in the energy domain is greater than a preset threshold, and / or, the frequency domain conforms to the change pattern of breathing or heartbeat.

[0038] In one possible implementation, the detection module is specifically used to compare the second audio with the first audio to obtain information on the occupancy of the first seat.

[0039] In one possible implementation, the detection module is specifically used to perform autocorrelation processing on the second audio to obtain the autocorrelation of the second audio; and to obtain the seating status of the people in the first seat based on the autocorrelation of the second audio.

[0040] In one possible implementation, the detection module is specifically used to input the second audio into a neural network to obtain the seating status of the people in the first seat.

[0041] In one possible implementation, the detection module is further configured to perform bandpass filtering and spectrum analysis on the second audio to obtain a processed second audio; and to obtain the seating status of the people in the first seat based on the processed second audio.

[0042] In one possible implementation, the first audio is an ultrasonic frequency band audio with a sampling frequency greater than 20 kHz.

[0043] In one possible implementation, there is one first speaker and one first microphone, and the first seat is located between the first speaker and the first microphone.

[0044] In one possible implementation, the first seat is a designated seat; or, the first seat is the current seat polled in a preset order within the vehicle; or, the first seat is one of a plurality of seats simultaneously detected in the vehicle.

[0045] In one possible implementation, the first microphone is located in front of the first seat, to the side of the first seat, or on the back of the first seat.

[0046] Thirdly, this application provides a vehicle comprising: one or more speakers for playing a first audio; one or more microphones for capturing the sound of the first audio being played to obtain a second audio; one or more processors; and a memory for storing one or more programs; wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of the first aspects above.

[0047] Fourthly, this application provides a vehicle control system, comprising: one or more processors; a memory for storing one or more programs; and, when the one or more programs are executed by the one or more processors, causing the one or more processors to implement the method as described in any one of the first aspects above.

[0048] Fifthly, this application provides a computer-readable storage medium including a computer program that, when executed on a computer, causes the computer to perform the method described in any one of the first aspects.

[0049] Sixthly, this application provides a computer program product comprising computer program code, which, when run on a computer, causes the computer to perform the method described in any one of the first aspects above. Attached Figure Description

[0050] Figure 1 is an exemplary functional block diagram of the vehicle 100 of this application;

[0051] Figure 2 is an exemplary structural block diagram of the vehicle infotainment system of this application;

[0052] Figure 3 is a schematic diagram of the architecture of the target object acquisition system of this application;

[0053] Figure 4 is a schematic diagram of the CNN structure of this application;

[0054] Figure 5 is a flowchart of the process 500 of the personnel presence detection method provided in this application;

[0055] Figure 6 is a schematic diagram of the vehicle's cabin according to this application;

[0056] Figure 7a is a schematic diagram of the autocorrelation coefficient of the present application when no one is sitting;

[0057] Figure 7b is a schematic diagram of the autocorrelation coefficient when someone is sitting in the present application;

[0058] Figure 8 is a flowchart of the personnel-in-present detection method of this application;

[0059] Figure 9 shows the spectrum of FMCW in the ultrasonic band;

[0060] Figure 10a is a schematic diagram of the original signal received by the loudspeaker in the time domain;

[0061] Figure 10b is a schematic diagram of the ultrasonic frequency band signal obtained after bandpass filtering;

[0062] Figure 10c is a schematic diagram of the signal after Fourier transform, which is converted from the time domain to the frequency domain.

[0063] Figure 11 is a structural schematic diagram of the personnel-in-charge detection device 1100 of this application. Detailed Implementation

[0064] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0065] The terms "first," "second," etc., used in the specification, embodiments, claims, and drawings of this application are for distinguishing purposes only and should not be construed as indicating or implying relative importance or order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion, such as including a series of steps or units. A method, system, product, or apparatus is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes, methods, products, or apparatuses.

[0066] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0067] First, the terminology used in this application will be explained.

[0068] Passenger detection: A technology used to detect and track users' eye movements, mainly used to identify the user's gaze direction and fixation point. It is often used in scenarios such as human-computer interaction, psychological research, and user behavior analysis.

[0069] Child Presence Detection: An improved residual network architecture that batch normalizes and activates the input before performing convolution operations, accelerating model convergence and improving training stability. This architecture is used in deep learning to enhance network training efficiency and model generalization ability.

[0070] Figure 1 is an exemplary functional block diagram of a vehicle 100 according to this application. As shown in Figure 1, components coupled to or included in the vehicle 100 may include a propulsion system 110, a sensor system 120, a control system 130, peripheral devices 140, a power supply 150, a computing device 160, and a user interface 170. The components of the vehicle 100 may be configured to operate in a manner interconnected with each other and / or with other components coupled to the respective systems. For example, the power supply 150 may provide power to all components of the vehicle 100. The computing device 160 may be configured to receive data from and control the propulsion system 110, sensor system 120, control system 130, and peripheral devices 140. The computing device 160 may also be configured to generate an image display on the user interface 170 and receive input from the user interface 170.

[0071] It should be noted that in other examples, vehicle 100 may include more, fewer, or different systems, and each system may include more, fewer, or different components. Furthermore, the systems and components shown can be combined or divided in any manner, and this application does not impose any specific limitations on this.

[0072] The computing device 160 may include a processor 161, a transceiver 162, and a memory 163. The computing device 160 may be a controller of the vehicle 100 or part of a controller. The memory 163 may store instructions 1631 executed by the processor 161, and may also store map data 1632. The processor 161 included in the computing device 160 may include one or more general-purpose processors and / or one or more special-purpose processors (e.g., image processors, digital signal processors, etc.). Where the processor 161 includes more than one processor, such processors may operate individually or in combination. The computing device 160 can implement functions that control the vehicle 100 based on input received through the user interface 170. The transceiver 162 is used for communication between the computing device 160 and various systems. The memory 163 may further include one or more volatile storage components and / or one or more non-volatile storage components, such as optical, magnetic, and / or organic storage devices, and the memory 163 may be wholly or partially integrated with the processor 161. Memory 163 may contain instructions 1631 (e.g., program logic) that can be executed by processor 161 to perform various vehicle functions, including any of the functions or methods described herein.

[0073] The propulsion system 110 can provide powered movement for the vehicle 100. As shown in FIG1, the propulsion system 110 may include an engine / motor 114, an energy source 113, a transmission 112, and wheels / tires 111. In addition, the propulsion system 110 may additionally or alternatively include other components besides those shown in FIG1. ​​This application does not specifically limit this.

[0074] Sensor system 120 may include several sensors for sensing information about the environment in which vehicle 100 is located. As shown in FIG1, the sensors of sensor system 120 include a Global Positioning System (GPS) 126, an Inertial Measurement Unit (IMU) 125, a lidar sensor 124, a camera sensor 123, a millimeter-wave radar sensor 122, and a brake 121 for modifying the position and / or orientation of the sensors. GPS 126 may be any sensor used to estimate the geographic location of vehicle 100. For this purpose, GPS 126 may include a transceiver to estimate the position of vehicle 100 relative to the Earth based on satellite positioning data. In an example, computing device 160 may be used to combine map data 1632 with GPS 126 to estimate the road on which vehicle 100 is traveling. IMU 125 may be used to sense changes in the position and orientation of vehicle 100 based on inertial acceleration and any combination thereof. In some examples, the combination of sensors in IMU 125 may include, for example, an accelerometer and a gyroscope. Other combinations of sensors in IMU 125 are also possible. The lidar sensor 124 can be viewed as an object detection system that uses light sensing to detect objects in the environment in which the vehicle 100 is located. Typically, the lidar sensor 124 can utilize optical remote sensing techniques to measure the distance to a target or other properties of the target by illuminating it with light. As an example, the lidar sensor 124 may include a laser source and / or laser scanner configured to emit laser pulses, and a detector for receiving reflections of the laser pulses. For example, the lidar sensor 124 may include a laser rangefinder reflected by a rotating mirror and scans the laser around a digitized scene in one or two dimensions to acquire distance measurements at specified angular intervals. In this example, the lidar sensor 124 may include components such as a light (e.g., laser) source, scanner and optical system, light detector and receiver electronics, and a positioning and navigation system. By scanning the laser reflected back from an object, the lidar sensor 124 can determine the distance to the object, forming a 3D environmental map with accuracy up to the centimeter level. The camera sensor 123 may include any camera (e.g., a still camera, video camera, etc.) for acquiring images of the environment in which the vehicle 100 is located. For this purpose, camera sensor 123 can be configured to detect visible light, or it can be configured to detect light from other parts of the spectrum, such as infrared or ultraviolet light. Other types of camera sensors 123 are also possible. Camera sensor 123 can be a two-dimensional detector, or it can have three-dimensional spatial range detection capabilities. In some examples, camera sensor 123 can be, for example, a distance detector configured to generate a two-dimensional image indicating the distance from camera sensor 123 to several points in the environment.For this purpose, camera sensor 123 can use one or more distance detection techniques. For example, camera sensor 123 can be configured to use structured light technology, in which vehicle 100 illuminates objects in the environment using a predetermined light pattern, such as a grid or checkerboard pattern, and uses camera sensor 123 to detect reflections from the predetermined light pattern on the objects. Based on the distortion in the reflected light pattern, vehicle 100 can be configured to detect the distance to points on the object. The predetermined light pattern can include infrared light or light of other wavelengths. Millimeter-wave radar sensor 122 typically refers to an object detection sensor with a wavelength of 1 to 10 mm and a frequency range of approximately 10 GHz to 200 GHz. The measurements of millimeter-wave radar sensor 122 contain depth information, which can provide the distance to the target; secondly, because millimeter-wave radar sensor 122 has a significant Doppler effect, it is very sensitive to velocity and can directly obtain the velocity of the target. The velocity of the target can be extracted by detecting its Doppler frequency shift. Currently, the two mainstream automotive millimeter-wave radar application frequency bands are 24GHz and 77GHz, respectively. The former has a wavelength of about 1.25cm and is mainly used for short-range perception, such as the vehicle's surrounding environment, blind spots, parking assistance, lane change assistance, etc.; the latter has a wavelength of about 4mm and is used for medium and long-range measurement, such as automatic following, adaptive cruise control (ACC), emergency braking (AEB), etc.

[0075] Sensor system 120 may also include additional sensors, including, for example, sensors that monitor the internal systems of vehicle 100 (e.g., O2 monitor, fuel gauge, oil temperature, etc.). Sensor system 120 may also include other sensors. This application does not specifically limit this.

[0076] The control system 130 can be configured to control the operation of the vehicle 100 and its components. For this purpose, the control system 130 may include a steering unit 136, a throttle 135, a braking unit 134, a sensor fusion algorithm 133, a computer vision system 132, and a navigation / route control system 131. The control system 130 may additionally or alternatively include components other than those shown in FIG. 1. This application does not specifically limit its scope in this regard.

[0077] Peripheral device 140 can be configured to allow vehicle 100 to interact with external sensors, other vehicles, and / or users. For this purpose, peripheral device 140 may include, for example, a wireless communication system 144, a touchscreen 143, a microphone 142, and / or a speaker 141. Peripheral device 140 may additionally or alternatively include components other than those shown in FIG. 1. This application does not specifically limit its scope in this regard.

[0078] Power source 150 can be configured to provide power to some or all of the components of vehicle 100. For this purpose, power source 150 may include, for example, rechargeable lithium-ion or lead-acid batteries. In some examples, one or more battery packs may be configured to provide power. Other power materials and configurations are also possible. In some examples, power source 150 and energy source 113 may be implemented together, as in some fully electric vehicles.

[0079] Speaker 141, also known as a "horn," converts audio electrical signals into sound signals. Vehicle 100 can listen to music, videos, and other sounds, or make hands-free calls through speaker 141. Typically, in-vehicle speakers consist of three types of sound-producing units: tweeters, mid-range speakers, and woofers. Optionally, to support the presence detection function, the cabin speaker system can be optimized to the greatest extent possible during cabin design; for example, by adding tweeters (e.g., with a sampling rate of 44.1 kHz or higher) and distributing them throughout the entire cabin.

[0080] Microphone 142, also known as a "microphone" or "voice transducer," is used to convert sound signals into electrical signals. When making a phone call or sending a voice message, the user can speak by bringing their mouth close to microphone 142, inputting the sound signal into microphone 142. Vehicle 100 may be equipped with at least one microphone 142. In some embodiments, vehicle 100 may be equipped with two microphones 142, which, in addition to collecting sound signals, can also perform noise reduction. In other embodiments, vehicle 100 may also be equipped with three, four, or more microphones 142, which can collect sound signals, reduce noise, identify sound sources, and perform directional recording, etc. In most cases, in-vehicle microphones are only used for voice interaction and may not necessarily support ultrasonic frequency sampling. Optionally, to support the seat detection function, the in-cabin microphone system can be optimized to the maximum extent during cabin design, for example, by deploying high-frequency microphones (e.g., sampling rate above 44.1 kHz) and distributing them throughout the entire vehicle cabin.

[0081] The components of vehicle 100 can be configured to operate in a manner that interconnects with other components within and / or outside their respective systems. For this purpose, the components and systems of vehicle 100 can be communicatively linked together via system buses, networks, and / or other connection mechanisms.

[0082] The vehicle infotainment system of vehicle 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. This application embodiment uses a layered architecture as an example to illustrate the vehicle infotainment system of vehicle 100.

[0083] Figure 2 is an exemplary structural block diagram of the vehicle infotainment system of this application. As shown in Figure 2, the vehicle infotainment system is divided into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the vehicle infotainment system is divided into four layers, from top to bottom: the application layer, the application framework layer, the runtime and system library, and the kernel layer.

[0084] The application layer can include a series of application packages.

[0085] As shown in Figure 2, the application package may include applications such as gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, and video.

[0086] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0087] As shown in Figure 2, the application framework layer may include a window manager, content provider, view system, phone manager, resource manager, notification manager, etc.

[0088] The window manager is used to manage windowed applications. It can retrieve screen size, determine the presence of a status bar, lock the screen, and capture screenshots, among other things.

[0089] Content providers store and retrieve data, making that data accessible to applications. This data may include video, images, audio, phone calls made and received, etc.

[0090] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including notification icons can include views for displaying text and views for displaying images.

[0091] The phone manager is used to provide communication functions for the vehicle, such as managing call status (including connection and disconnection).

[0092] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.

[0093] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of download completion or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, and flashing indicator lights.

[0094] Runtime includes core libraries and a virtual machine. Runtime is responsible for the scheduling and management of the vehicle's infotainment system.

[0095] The core library consists of two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core library of the vehicle infotainment system.

[0096] The application layer and application framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0097] System libraries can include multiple functional modules. For example: surface manager, media libraries, 3D graphics processing libraries (e.g., OpenGL ES), 2D graphics engines (e.g., SGL), etc.

[0098] The Surface Manager is used to manage the display subsystem and provides the blending of 2D and 3D layers for multiple applications.

[0099] The media library supports playback and recording of various common audio and video formats, as well as still image files. It supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG.

[0100] The 3D graphics processing library is used to implement 3D graphics drawing, image rendering, compositing, and layer processing.

[0101] A 2D graphics engine is a graphics engine for 2D drawing.

[0102] The kernel layer is the layer between hardware and software. The kernel layer contains at least the display driver, camera driver, audio driver, and sensor driver.

[0103] It is understood that the components included in the system framework layer, system library, and runtime layer shown in Figure 2 do not constitute a specific limitation on the vehicle. In other embodiments of this application, the vehicle may include more or fewer components than shown, or combine some components, or split some components, or have different component arrangements.

[0104] Since this application involves the application of AI, for ease of understanding, some AI-related terms or concepts used in this application will be explained below, and these terms or concepts are also part of the content of the invention.

[0105] (1) Neural Network

[0106] Neural Networks (NNs) are machine learning models. A neural network can be composed of neural units, which are computational units that take xs and an intercept of 1 as input. The output of such a computational unit can be:

[0107] Where s = 1, 2, ..., n, where n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neural unit. f is the activation function of the neural unit, used to introduce nonlinear characteristics into the neural network to convert the input signal in the neural unit into the output signal. The output signal of this activation function can be used as the input of the next convolutional layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting many of the above-mentioned individual neural units together, that is, the output of one neural unit can be the input of another neural unit. The input of each neural unit can be connected to the local receptive field of the previous layer to extract the features of the local receptive field, which can be a region composed of several neural units.

[0108] (2) Deep Neural Networks

[0109] Deep neural networks (DNNs), also known as multilayer neural networks, can be understood as neural networks with many hidden layers, though there's no specific metric for "many." DNNs can be categorized into three layers based on their position: input layers, hidden layers, and output layers. Generally, the first layer is the input layer, the last layer is the output layer, and the layers in between are hidden layers. All layers are fully connected, meaning that any neuron in the i-th layer is connected to any neuron in the (i+1)-th layer. Although DNNs appear complex, the operation of each layer is actually quite simple, resembling a linear relationship as follows: in, It is the input vector. It is the output vector. α is the offset vector, W is the weight matrix (also called coefficients), and α() is the activation function. Each layer is simply an adjustment of the input vector. The output vector is obtained through such a simple operation. Because DNNs have many layers, the coefficients W and the offset vector... The number of these parameters is therefore quite large. The definitions of these parameters in a DNN are as follows: Taking the coefficient W as an example: Assuming a three-layer DNN, the linear coefficient from the 4th neuron in the second layer to the 2nd neuron in the third layer is defined as... The superscript 3 represents the layer number where coefficient W resides, while the subscript corresponds to the output third layer index 2 and the input second layer index 4. In summary, the coefficients from the k-th neuron in layer L-1 to the j-th neuron in layer L are defined as follows: It's important to note that the input layer does not have a W parameter. In deep neural networks, more hidden layers allow the network to better represent complex real-world situations. Theoretically, the more parameters a model has, the higher its complexity and "capacity," meaning it can perform more complex learning tasks. Training a deep neural network is essentially the process of learning the weight matrix, with the ultimate goal of obtaining the weight matrix of all layers in the trained deep neural network (a weight matrix formed by the vectors W from many layers).

[0110] (3) Convolutional Neural Network

[0111] A convolutional neural network (CNN) is a deep neural network with convolutional structures. It is a deep learning architecture, which refers to learning at multiple levels of abstraction using machine learning algorithms. As a deep learning architecture, CNN is a feed-forward artificial neural network, where each neuron responds to an input image. A CNN contains a feature extractor consisting of convolutional layers and pooling layers. This feature extractor can be viewed as a filter, and the convolution process can be seen as performing convolution with a trainable filter and an input image or a convolutional feature map.

[0112] A convolutional layer is a layer of neurons in a convolutional neural network that performs convolution processing on the input signal. A convolutional layer can contain multiple convolution operators, also called kernels. In image processing, these operators act as filters, extracting specific information from the input image matrix. Essentially, a convolution operator can be a weight matrix, which is usually predefined. During the convolution operation, the weight matrix typically processes the input image pixel by pixel (or two pixels by two pixels, depending on the stride) along the horizontal direction, thus extracting specific features from the image. The size of the weight matrix should be related to the image size. It's important to note that the depth dimension of the weight matrix is ​​the same as the depth dimension of the input image; during convolution, the weight matrix extends to the entire depth of the input image. Therefore, convolution with a single weight matrix produces a single-depth convolutional output. However, in most cases, multiple weight matrices of the same size (rows × columns) are used instead of a single weight matrix. The outputs of each weight matrix are stacked to form the depth dimension of the convolutional image. This dimension can be understood as being determined by the "multiple" factors mentioned above. Different weight matrices can be used to extract different features from the image. For example, one weight matrix can be used to extract edge information, another to extract specific colors, and yet another to blur unwanted noise. These multiple weight matrices have the same size (rows × columns), and the feature maps extracted by these weight matrices also have the same size. These extracted feature maps are then merged to form the output of the convolution operation. The weight values ​​in these weight matrices need to be obtained through extensive training in practical applications. The weight matrices formed by these trained weight values ​​can be used to extract information from the input image, enabling the convolutional neural network to make correct predictions. When a convolutional neural network has multiple convolutional layers, the initial convolutional layers often extract more general features, which can also be called low-level features. As the depth of the convolutional neural network increases, the features extracted by later convolutional layers become increasingly complex, such as high-level semantic features. Features with higher semantic levels are more suitable for the problem being solved.

[0113] Because it's often necessary to reduce the number of training parameters, pooling layers are frequently introduced periodically after convolutional layers. This can be a single convolutional layer followed by a pooling layer, or multiple convolutional layers followed by one or more pooling layers. In image processing, the sole purpose of pooling layers is to reduce the spatial size of the image. Pooling layers can include average pooling and / or max pooling operators to sample the input image to obtain a smaller image size. Average pooling calculates the average value of pixel values ​​within a specific range as the result of average pooling. Max pooling takes the pixel with the largest value within a specific range as the result of max pooling. Furthermore, just as the size of the weight matrix in a convolutional layer should be related to the image size, the operators in a pooling layer should also be related to the image size. The size of the output image after pooling can be smaller than the size of the input image of the pooling layer. Each pixel in the output image represents the average or maximum value of the corresponding sub-region of the input image of the pooling layer.

[0114] After processing by convolutional / pooling layers, a convolutional neural network (CNN) is still insufficient to output the required information. As mentioned earlier, convolutional / pooling layers only extract features and reduce the parameters introduced by the input image. However, to generate the final output information (the required class information or other relevant information), the CNN needs to utilize neural network layers to generate one or a set of desired class numbers of output. Therefore, the neural network can include multiple hidden layers, the parameters of which can be pre-trained based on training data relevant to a specific task type, such as image recognition, image classification, image super-resolution reconstruction, etc.

[0115] Optionally, after the multiple hidden layers in the neural network, there is also an output layer of the entire convolutional neural network. This output layer has a loss function similar to the classification cross-entropy, which is specifically used to calculate the prediction error. Once the forward propagation of the entire convolutional neural network is completed, the backpropagation will begin to update the weight values ​​and biases of the aforementioned layers to reduce the loss of the convolutional neural network and the error between the result output by the convolutional neural network through the output layer and the ideal result.

[0116] (4) Recurrent Neural Network

[0117] Recurrent neural networks (RNNs) are used to process sequential data. In traditional neural network models, the layers from the input layer to the hidden layer and then to the output layer are fully connected, but the nodes within each layer are unconnected. While this type of neural network has solved many difficult problems, it remains inadequate for many others. For example, predicting the next word in a sentence generally requires using the preceding words because words in a sentence are not independent. RNNs are called recurrent neural networks because the current output of a sequence is related to the outputs of previous sequences. Specifically, the network memorizes previous information and applies it to the calculation of the current output; that is, nodes within the same hidden layer are no longer unconnected but connected, and the input to a hidden layer includes not only the output of the input layer but also the output of the hidden layer at the previous time step. Theoretically, RNNs can process sequential data of any length. Training an RNN is similar to training a traditional CNN or DNN. This algorithm also uses the backpropagation algorithm, but with one key difference: when an RNN is expanded, its parameters, such as W, are shared; however, this is not the case with traditional neural networks as illustrated above. Furthermore, in gradient descent, the output at each step depends not only on the network at the current step but also on the states of the network in previous steps. This learning algorithm is called Backpropagation Through Time (BPTT).

[0118] Since we already have convolutional neural networks (CNNs), why do we need recurrent neural networks (RNNs)? The reason is simple. CNNs rely on the fundamental assumption that elements are independent of each other, and that input and output are also independent—like a cat and a dog. However, in the real world, many elements are interconnected. For example, stock prices fluctuate over time. Or, imagine someone saying, "I love traveling, and my favorite place is Yunnan. I definitely want to go there someday." Humans know the answer to this question is "Yunnan." Humans can infer from context, but how can machines do the same? This is where RNNs come in. RNNs aim to give machines the ability to remember, just like humans. Therefore, the output of an RNN depends on both the current input information and historical memory information.

[0119] (5) Loss Function

[0120] In training a deep neural network, to ensure the output closely approximates the desired predicted value, we compare the network's prediction with the target value. Based on the difference, we update the weight vector of each layer (usually pre-configuring parameters before the initial update). For example, if the prediction is too high, the weight vector is adjusted to predict a lower value. This adjustment continues until the deep neural network predicts the target value or a value very close to it. Therefore, we need to predefine "how to compare the difference between the predicted and target values," which is the loss function or objective function. These are important equations used to measure the difference between the predicted and target values. Taking the loss function as an example, a higher output value (loss) indicates a greater difference, and training the deep neural network becomes a process of minimizing this loss.

[0121] (6) Backpropagation algorithm

[0122] Convolutional neural networks can employ backpropagation (BP) to correct the parameters in the initial super-resolution model during training, thereby reducing the reconstruction error loss. Specifically, forward propagation of the input signal to the output generates an error loss; this error loss information is then propagated back to update the parameters in the initial super-resolution model, leading to convergence of the error loss. The backpropagation algorithm is an error-loss-driven backpropagation process aimed at obtaining the optimal parameters of the super-resolution model, such as the weight matrix.

[0123] (7) Generative Adversarial Networks

[0124] Generative adversarial networks (GANs) are a type of deep learning model. This model comprises at least two modules: a generative model and a discriminative model. These two modules learn from each other through a game-like interaction, resulting in better outputs. Both the generative and discriminative models can be neural networks, specifically deep neural networks or convolutional neural networks. The basic principle of GANs is as follows: Taking an image-generating GAN as an example, suppose there are two networks, G (Generator) and D (Discriminator). G is a network that generates images by receiving random noise z and using this noise, denoted as G(z). D is a discriminative network used to determine whether an image is "real." Its input parameter is x, representing an image, and its output D(x) represents the probability that x is a real image. A value of 1 indicates that the image is 100% real, while a value of 0 indicates that the image is impossible to be real. During the training of this generative adversarial network (GAN), the goal of the generative network G is to generate realistic images to deceive the discriminator network D, while the goal of the discriminator network D is to distinguish the images generated by G from real images as much as possible. Thus, G and D constitute a dynamic "game," which is the "adversarial" aspect of the GAN. Ideally, the game will result in G generating images G(z) that are sufficiently realistic, while D struggles to determine whether the images generated by G are real or not, i.e., D(G(z)) = 0.5. This yields a superior generative model G that can be used to generate images.

[0125] Regardless of which AI technology is used, this application can pre-train the model. For example, Figure 3 is a schematic diagram of the target object acquisition system of this application. As shown in Figure 3, the data acquisition device 360 ​​is used to collect data and store it in the database 330. The training device 320 generates the target model / rule 301 based on the data maintained in the database 330. The following will describe in more detail how the training device 320 obtains the target model / rule 301 based on the data. The target model / rule 301 can realize the function of a multimodal model and output the target image.

[0126] The function of each layer in a deep neural network can be expressed mathematically. To describe it: From a physical perspective, the work of each layer in a deep neural network can be understood as transforming the input space (the set of input vectors) to the output space (i.e., from the row space to the column space of a matrix) through five operations on the input space. These five operations include: 1. Dimensionality increase / decrease; 2. Magnification / scaling; 3. Rotation; 4. Translation; 5. "Bending". Operations 1, 2, and 3 are... Completed, operation 4 is performed by Complete, operation 5 is implemented by α(). The term "space" is used here because the object being classified is not a single thing, but a class of things; space refers to the set of all individuals of this class of things. Here, W is the weight vector, where each value represents the weight of a neuron in that layer of the neural network. This vector W determines the spatial transformation from the input space to the output space, as described above; that is, the weights W of each layer control how the space is transformed. The purpose of training a deep neural network is to ultimately obtain the weight matrix of all layers of the trained neural network (a weight matrix formed by the vectors W of many layers). Therefore, the training process of a neural network is essentially learning how to control spatial transformation, more specifically, learning the weight matrix.

[0127] Because the goal is for the output of a deep neural network to be as close as possible to the actual predicted value, we can compare the current network's prediction with the desired target value and update the weight vector of each layer based on the difference. (Of course, there's usually an initialization process before the first update, where parameters are pre-configured for each layer in the deep neural network). For example, if the network's prediction is too high, the weight vector is adjusted to predict a lower value, and this adjustment continues until the neural network can predict the actual target value. Therefore, it's necessary to predefine "how to compare the difference between the predicted value and the target value," which is the loss function or objective function. These are important equations used to measure the difference between the predicted and target values. Taking the loss function as an example, a higher output value (loss) indicates a greater difference, so training a deep neural network becomes a process of minimizing this loss as much as possible.

[0128] The target model / rule 301 obtained from training device 320 can be applied to different systems or devices.

[0129] The execution device 310 is equipped with an I / O interface 312 for data interaction with external devices. The "user" can input data to the I / O interface 312 through the terminal device 340.

[0130] The execution device 310 can call data, code, etc. in the data storage system 350, and can also store data, instructions, etc. in the data storage system 350.

[0131] The calculation module 311 uses the target model / rule 301 to process the input data in order to realize the function of the multimodal model in the embodiments of this application.

[0132] The associated function module 313 and associated function module 314 can respectively implement related functions in the training process, such as preprocessing and filtering.

[0133] Finally, the I / O interface 312 returns the processing result to the terminal device 340 for the user.

[0134] At a deeper level, the training device 320 can generate corresponding target models / rules 301 based on different data for different objectives, in order to provide users with better results.

[0135] In the scenario shown in Figure 3, the user can manually specify the data to be input into the execution device 310, for example, by operating through the interface provided by the I / O interface 312. Alternatively, the terminal device 340 can automatically input data into the I / O interface 312 and obtain the results. If the terminal device 340 requires user authorization to automatically input data, the user can set the corresponding permissions in the terminal device 340. The user can view the results output by the execution device 310 on the terminal device 340, which can be presented in various ways such as display, sound, or animation. The terminal device 340 can also act as a data acquisition terminal, storing the acquired data into the database 330.

[0136] It is worth noting that Figure 3 is merely a schematic diagram of a target image acquisition system provided in an embodiment of this application. The positional relationships between the devices, components, modules, etc. shown in the figure do not constitute any limitation. For example, in Figure 3, the data storage system 350 is an external memory relative to the execution device 310. In other cases, the data storage system 350 can also be placed in the execution device 310. As another example, in Figure 3, the terminal device 340 and the execution device 310 are two devices. In other cases, the terminal device 340 and the execution device 310 can also be integrated into one device.

[0137] A convolutional neural network (CNN) is a deep neural network with convolutional structures. It is a deep learning architecture, which refers to learning at multiple levels of abstraction using machine learning algorithms. As a deep learning architecture, CNN is a feed-forward artificial neural network where each neuron responds to overlapping regions in the input image.

[0138] For example, Figure 4 is a schematic diagram of the structure of the CNN of this application. As shown in Figure 4, the CNN400 may include an input layer 410, a convolutional layer / pooling layer 420, wherein the pooling layer is optional, and a neural network layer 430.

[0139] Convolutional / pooling layers 420:

[0140] Convolutional layers:

[0141] As shown in Figure 4, the convolutional / pooling layer 420 may include layers 421-426 as in Examples 421. In one implementation, layer 421 is a convolutional layer, layer 422 is a pooling layer, layer 423 is a convolutional layer, layer 424 is a pooling layer, layer 425 is a convolutional layer, and layer 426 is a pooling layer. In another implementation, layers 421 and 422 are convolutional layers, layer 423 is a pooling layer, layers 424 and 425 are convolutional layers, and layer 426 is a pooling layer. That is, the output of the convolutional layer can be used as the input of a subsequent pooling layer, or as the input of another convolutional layer to continue the convolution operation.

[0142] Taking convolutional layer 421 as an example, it can include multiple convolution operators, also known as kernels. In image processing, a convolution operator acts as a filter, extracting specific information from the input image matrix. Essentially, a convolution operator can be a weight matrix, which is usually predefined. During the convolution operation, the weight matrix processes the input image pixel by pixel (or two pixels by two pixels, depending on the stride) along the horizontal direction, thus extracting specific features. The size of the weight matrix should be related to the image size. It's important to note that the depth dimension of the weight matrix is ​​the same as the depth dimension of the input image; during convolution, the weight matrix extends to the entire depth of the input image. Therefore, convolution with a single weight matrix produces a single-depth convolutional output. However, in most cases, multiple weight matrices of the same dimension are applied instead of a single weight matrix. The outputs of each weight matrix are stacked to form the depth dimension of the convolutional image. Different weight matrices can be used to extract different features from an image. For example, one weight matrix can be used to extract image edge information, another weight matrix can be used to extract specific colors from the image, and yet another weight matrix can be used to blur unwanted noise in the image. These multiple weight matrices have the same dimension, and the feature maps extracted by these multiple weight matrices with the same dimension also have the same dimension. The extracted feature maps with the same dimension are then merged to form the output of the convolution operation.

[0143] The weight values ​​in these weight matrices need to be obtained through extensive training in practical applications. The weight matrices formed by the weight values ​​obtained through training can extract information from the input image, thereby helping CNN400 to make correct predictions.

[0144] When a CNN400 has multiple convolutional layers, the initial convolutional layers (e.g., 421) tend to extract more general features, which can also be called low-level features. As the depth of the CNN400 increases, the features extracted by later convolutional layers (e.g., 426) become more and more complex, such as high-level semantic features. Features with higher semantic levels are more suitable for the problem to be solved.

[0145] Pooling layer:

[0146] Because it's often necessary to reduce the number of training parameters, pooling layers are frequently introduced periodically after convolutional layers, as illustrated in layers 421-426 of Figure 4 (420). This can be a convolutional layer followed by a pooling layer, or multiple convolutional layers followed by one or more pooling layers. In image processing, the sole purpose of pooling layers is to reduce the spatial size of the image. Pooling layers can include average pooling and / or max pooling operators to sample the input image to obtain a smaller image size. Average pooling calculates the average value of pixel values ​​within a specific range. Max pooling takes the pixel with the largest value within a specific range as the result of max pooling. Furthermore, just as the size of the weight matrix in a convolutional layer should be related to the image size, the operators in a pooling layer should also be related to the image size. The size of the output image after pooling can be smaller than the size of the input image of the pooling layer. Each pixel in the output image represents the average or maximum value of the corresponding sub-region of the input image of the pooling layer.

[0147] Neural network layer 430:

[0148] After processing by the convolutional / pooling layers 420, the CNN400 is still insufficient to output the required information. As mentioned earlier, the convolutional / pooling layers 420 only extract features and reduce the parameters introduced by the input image. However, to generate the final output information (the required class information or other relevant information), the CNN400 needs to utilize neural network layers 430 to generate one or a set of required class numbers of output. Therefore, neural network layers 430 can include multiple hidden layers (431, 432 to 43n as shown in Figure 4) and an output layer 440. The parameters contained in these hidden layers can be pre-trained based on relevant data for a specific task type, such as image recognition, image classification, or image super-resolution reconstruction.

[0149] After the multiple hidden layers in neural network layer 430, which is the last layer of the entire CNN400, is the output layer 440. The output layer 440 has a loss function similar to the classification cross-entropy, which is used to calculate the prediction error. Once the forward propagation of the entire CNN400 is completed (as shown in Figure 4, the propagation from 410 to 440 is the forward propagation), the back propagation (as shown in Figure 4, the propagation from 440 to 410 is the back propagation) will begin to update the weight values ​​and biases of the aforementioned layers, in order to reduce the loss of CNN400 and the error between the result output by CNN400 through the output layer and the ideal result.

[0150] It should be noted that the CNN400 shown in Figure 4 is only an example of a convolutional neural network. In specific applications, convolutional neural networks can also exist in the form of other network models. For example, as shown in Figure 4, multiple convolutional / pooling layers can be parallelized, and the extracted features can be input into the full neural network layer 430 for processing.

[0151] Based on the above, the technical solution of the personnel presence detection method provided in this application will be described below.

[0152] Figure 5 is a flowchart of process 500 of the personnel presence detection method provided in this application. Process 500 can be performed by the vehicle 100 described above. Process 500 is described as a series of steps or operations, and it should be understood that process 500 can be performed in various orders and / or occur simultaneously, and is not limited to the execution order shown in Figure 5. Process 500 may include:

[0153] Step 501: Play the first audio through the first speaker.

[0154] In one possible implementation, the occupant presence detection in this application can be event-triggered. For example, when a user gets out of the car and closes the door, the vehicle engine is turned off, or the vehicle's motor stops, initiating occupant presence detection at these times can prevent elderly people or children from being left behind in a stopped vehicle, thus avoiding safety accidents. Alternatively, when a user opens the car door and gets in, the vehicle engine is started, or the vehicle's motor starts, initiating occupant presence detection at these times can activate personalized services for the occupants, such as seatbelt reminders, personalized seating configurations, and targeted audio pickup, enhancing the user experience. It should be noted that the above events are merely illustrative examples and do not constitute a limitation on the triggering events; this application does not impose any specific limitations in this regard.

[0155] In one possible implementation, the occupant detection in this application can be user-triggered. For example, a user interface can be displayed on the vehicle's central control screen, as shown in Figure 6 (Figure 6 is a schematic diagram of the vehicle's cabin according to this application). The vehicle's cabin contains six seats, and the user can click on one of the six seats to trigger the occupant detection for that seat. It should be noted that the above is an example of user-triggered detection; other user-triggered methods may include voice command, interface selection, physical button selection, etc., and this application does not specifically limit these methods.

[0156] Vehicle speakers typically consist of three types of sound-emitting units: tweeters, midrange speakers, and woofers. Optionally, to support in-seat detection functionality, the vehicle's speaker system can be optimized to the greatest extent possible during cabin design. For example, additional tweeters (e.g., with a sampling rate of 44.1 kHz or higher) can be distributed throughout the entire vehicle cabin. The first speaker in this application can be at least one of multiple speakers in the vehicle (including inward-facing speakers (typically used for playing music, video, conversations, etc.) and outward-facing speakers (typically used for playing horns, human voices, etc.)). It can be a tweeter, a midrange speaker, or a woofer, or a combination of the aforementioned tweeter, midrange speaker, and woofer; no specific limitation is made in this regard.

[0157] The first audio played through the first speaker can be a pre-set audio segment (also known as test audio). This test audio can be ultrasonic audio (ultrasonic audio with a sampling frequency greater than 20 kHz, for example, audio with a sampling frequency of 44 kHz, or frequency-modulated continuous wave (FMCW) with a sampling frequency of 19 kHz to 22 kHz). Since ultrasonic audio is inaudible to the human ear, playing ultrasonic audio while personnel are seated will not interfere with the occupants. Alternatively, the test audio can be non-ultrasonic audio, for example, an FMCW modulated signal with a sampling frequency of 8 kHz to 10 kHz, which sounds like periodic noise to the human ear.

[0158] Optionally, the vehicle's speakers (including the first speaker or other speakers besides the first speaker) can play other sounds, such as music, video sounds, or call sounds, while playing the first audio. That is, the audio played by the vehicle's speakers can be only the first audio, or it can include the first audio plus ordinary sounds (the aforementioned other sounds), so that the visual and auditory needs of the personnel can be met simultaneously when they are conducting the inspection.

[0159] Step 502: Obtain the second audio by acquiring the first sound through the first microphone.

[0160] In most cases, in-vehicle microphones are only used for voice interaction and may not necessarily support ultrasonic frequency sampling. Optionally, to support seat detection, the vehicle's microphone system can be optimized to the greatest extent during cabin design. For example, high-frequency microphones (e.g., sampling rate above 44.1 kHz) can be deployed and distributed throughout the entire vehicle cabin. The first microphone in this application can be the microphone corresponding to the first seat (the first seat is one of the seats in the vehicle, for example, any one of the seats in Figure 6) among multiple microphones in the vehicle. That is, the first microphone can be located near the first seat, for example, in front of the first seat, to the side of the first seat, or on the seat back. The first microphone can be either a high sampling rate (e.g., sampling rate greater than 20 kHz) microphone or a normal sound microphone (e.g., sampling rate less than 20 kHz), without specific limitations.

[0161] In one possible implementation, when there is one first speaker and one first microphone, the first seat is located between the first speaker and the first microphone; that is, the first seat is located on the line connecting the first speaker and the first microphone. This allows the acoustic path between the first speaker and the first microphone to cover the first seat, thereby improving detection accuracy.

[0162] Optionally, the first seat is a designated seat. For example, a seat selected via the user-triggered method described above. Or, for example, a seat selected by the vehicle's control system.

[0163] Optionally, the first seat is the current seat polled in the vehicle in a preset order. For example, in Figure 6, the six seats in the cabin are polled one by one in the order from front to back and from left to right. Each time a seat is processed, the seat being processed is the first seat.

[0164] Optionally, the first seat can be one of multiple seats in the vehicle that are simultaneously detected. Simultaneous detection can refer to the simultaneous detection of occupancy status across multiple seats in the vehicle. For example, the six seats in Figure 6 can be detected simultaneously, meaning that each corresponding microphone can act as a first microphone, simultaneously collecting (also known as picking up, recording, etc.) the sound inside the vehicle to obtain a second audio signal. Then, based on the respective second audio signals, it can be determined whether the seat is occupied. In this case, each seat can be designated as the first seat, and occupancy detection can be performed separately.

[0165] The second audio picked up by the first microphone includes the following two cases:

[0166] (1) The second audio only includes the audio obtained by collecting the sound from the first audio.

[0167] The first audio signal is played by the first speaker, propagates within the cockpit, and is picked up by the first microphone to obtain the second audio signal. At this point, the second audio signal only contains the audio signal picked up based on the test audio signal, and can be subjected to spectral analysis. For example, the second audio signal can be subjected to a Fourier transform to convert it to the frequency domain before analysis.

[0168] (2) The second audio includes the audio of the first audio plus the audio of ordinary sound.

[0169] The first audio signal plus a normal sound is played by the first speaker, propagates within the cockpit, and is picked up by the first microphone to obtain the second audio signal. At this point, the second audio signal contains not only the audio signal picked up based on the test audio signal but also the audio signal picked up based on the normal sound signal. Bandpass filtering and spectral analysis can be performed on it. For example, the audio signal picked up based on the normal sound signal can be filtered out by bandpass filtering, and then a Fourier transform can be performed to convert it to the frequency domain to obtain the processed second audio signal, which can then be analyzed.

[0170] Step 503: Obtain the seating situation of the people in the first seat based on the second audio.

[0171] The seating status of the first seat includes either someone sitting in the first seat or no one sitting in the first seat.

[0172] In one possible implementation, when the second audio signal meets the signal change characteristics, it is determined that someone is sitting in the first seat. The signal change characteristics include: the attenuation in the energy domain is greater than a preset threshold, and / or, the signal change in the frequency domain conforms to the pattern of breathing or heartbeat.

[0173] In other words, the basis for determining that someone is sitting in the first seat is that the second audio frequency meets the signal change characteristics. The human body (mainly composed of water and muscle) has two significant effects on audio signals. Firstly, the human acoustic impedance is between 1.45 and 1.70 kgm. -2 s -1 ×10 6 Around 0.00429 kgm, much higher than the atmospheric pressure of 0.00429 kgm. -2 s -1 ×10 6 Therefore, in the energy domain, the presence of a person will significantly absorb the energy of the audio signal; on the other hand, due to regular micro-movements of the body such as breathing or heartbeat, the audio signal will also show regular phase changes. Therefore, it is possible to analyze whether the audio signal has regular changes similar to breathing or heartbeat in the frequency domain.

[0174] Based on this, signal change characteristics can include changes in the second audio frequency in the energy domain and / or spectral domain. By analyzing the audio changes from playback to pickup, combined with the influence of the human body on the audio signal, the system can accurately determine whether someone is sitting in the seat. This method is unaffected by light, avoids detection errors due to obstructions, and prevents situations like pressure detection schemes that mistakenly identify non-human objects as occupied. Furthermore, reusing the design principles of a vehicle's acoustic system can reduce hardware costs.

[0175] In this application, obtaining the seating status of the people in the first seat based on the second audio can include the following three methods:

[0176] 1. Compare the second audio with the first audio to obtain the seating information of the person in the first seat.

[0177] The first audio includes a pre-set test audio. Therefore, after obtaining the second audio, the second audio can be compared with the first audio to obtain the change of the second audio compared with the first audio. Then, the components of the change in the energy domain and / or spectrum domain are analyzed to determine whether the second audio meets the signal change characteristics, and thus obtain the detection result of whether someone is sitting in the first seat.

[0178] 2. Perform autocorrelation processing on the second audio to obtain the autocorrelation of the second audio; obtain the seating status of the people in the first seat based on the autocorrelation.

[0179] By performing correlation analysis between the second audio signal and its own delayed signal, the autocorrelation of the second audio signal can be obtained. That is, the autocorrelation of the second audio signal represents the correlation between the second audio signal and its own delayed signal. Autocorrelation can be represented by the autocorrelation coefficient. For example, Figure 7a shows a schematic diagram of the autocorrelation coefficient when no one is sitting, and Figure 7b shows a schematic diagram of the autocorrelation coefficient when someone is sitting. As shown in Figures 7a and 7b, the value of the first peak is 1 (indicating that the correlation coefficient between the original signal and itself is 1). The value of the second peak is related to whether someone is sitting. When someone is sitting, the value is lower (because of human absorption, the subsequent signal differs greatly from the original signal), and when no one is sitting, the value is higher (the chirping waves are more consistent). Therefore, the value of the second peak can be used to determine whether the first seat is occupied. It should be noted that this application can also use other peaks, or even other correlation attribute values, for judgment, without specific limitations.

[0180] 3. Input the second audio signal into a pre-trained neural network to obtain the seating situation of the people in the first seat.

[0181] The neural network in this application can be trained to take a second audio input and output the seating situation of the people in the first seat. Its principle and structure can be referred to the description of AI technology above, and will not be repeated here.

[0182] As mentioned above, if the vehicle's speakers play both the first audio and ordinary sounds, then the second audio will contain both the audio picked up based on the test audio and the audio picked up based on the ordinary sounds. Therefore, bandpass filtering and spectral analysis must be performed on the second audio first. For example, bandpass filtering can be used to filter out the audio picked up based on ordinary sounds, and then a Fourier transform can be performed to convert it to the frequency domain to obtain the processed second audio. Afterward, any of the three methods mentioned above can be used to determine whether someone is sitting in the first seat based on the processed second audio. Optionally, even if the first audio only contains the test audio, bandpass filtering and spectral analysis can still be performed on the second audio to filter out interfering sounds in the second audio, making the detection results more accurate.

[0183] In this embodiment, an audio signal is played through a speaker in the vehicle, and another audio signal is picked up by a microphone. The presence of people in the seats corresponding to the microphone is then determined based on the picked-up audio. This method is not affected by light, does not cause detection errors due to obstruction, and avoids situations like pressure detection schemes that mistakenly detect non-human objects as people sitting, thereby improving detection efficiency. Furthermore, reusing the design concept of the vehicle's acoustic system can also reduce hardware costs.

[0184] The technical solutions of the method embodiments shown in Figure 5 will be described in detail below using several specific examples.

[0185] Figure 8 is a flowchart of the personnel presence detection method of this application. As shown in Figure 8, the method includes the following steps:

[0186] 1) Configure transceiver pairs: For seats in different locations within the cabin, based on the vehicle's acoustic system (including microphones and speakers) deployment, some or all of the high-frequency microphones and high-frequency speakers can be pre-configured as transceiver pairs. Optionally, when the seat to be detected is on the line connecting the transceiver pairs, detecting the occupant's presence using ultrasonic signals is more effective. This step is performed in advance, meaning the transceiver pairs for each seat can be set up during the cabin design phase.

[0187] 2) Speaker plays ultrasonic frequency band signal (first audio): When the presence detection of a person in a certain seat (first seat) is triggered, the selected speaker emits a pre-set ultrasonic frequency band signal s. Optionally, this ultrasonic frequency band signal can also be mixed with the music or sound being played before playback.

[0188] 3) Microphone receives ultrasonic frequency signals (second audio) propagating in the cockpit: The microphone receives ultrasonic frequency signals s' after propagation in the cockpit.

[0189] 4) Obtaining passenger seating information: The ultrasonic signal s' is a delayed, weakened version of the ultrasonic signal s, or even a version mixed with other sounds in the cockpit. Analyze the attenuation of the ultrasonic signal s' in the energy domain and whether there are regular changes similar to breathing or heartbeats in the spectral domain. Combined with the actual cockpit conditions, the passenger seating information can be obtained based on the changes in the energy domain and / or spectral domain.

[0190] Taking the vehicle shown in Figure 6 as an example, the method for detecting whether people are seated includes the following steps:

[0191] 1) Configure send and receive pairs

[0192] This embodiment takes the left seat in the middle row as an example. The loudspeaker at the rear left and the microphone array at the left position of the middle row can be used as a transceiver pair. It should be noted that the solution of this application can also be implemented using other transceiver pairs, but compared with the seat being located on the line connecting the transceiver pairs, the ultrasonic frequency signal may not directly pass through the human body to reach the microphone.

[0193] 2) The speaker plays an ultrasonic frequency signal (first audio).

[0194] This embodiment uses FMCW as an example. The sampling frequency of the ultrasonic band signal played by the speaker is 19kHz to 22kHz (refer to Figure 9, which is a spectrum diagram of FMCW in the ultrasonic band, with the vertical axis representing frequency (Hz) and the horizontal axis representing time (s)). FMCW has obvious characteristics in the frequency domain, making it relatively easy to distinguish from other non-ultrasonic band signals, and it is also less susceptible to interference from full-band signals such as loudly slamming car doors.

[0195] 3) The microphone receives ultrasonic frequency signals (second audio) propagating inside the cockpit.

[0196] 4) Obtain information on the number of people present

[0197] The microphone picks up the signal, which can first be bandpass filtered (preserving the signal in the 19kHz to 22kHz frequency band) to obtain the ultrasonic frequency band signal. Then, a Fourier transform is performed on the ultrasonic frequency band signal to convert the signal from the time domain to the frequency domain. Referring to Figures 10a, 10b, and 10c, Figure 10a is a schematic diagram of the original signal received by the speaker in the time domain, Figure 10b is a schematic diagram of the ultrasonic frequency band signal obtained after bandpass filtering, and Figure 10c is a schematic diagram of the signal converted from the time domain to the frequency domain after Fourier transform.

[0198] For the aforementioned frequency domain signal, autocorrelation processing can be performed to obtain the corresponding autocorrelation coefficient, as shown in Figures 7a and 7b. Since FMCW is a periodic signal, the time interval between each chirp wave is fixed. Therefore, the autocorrelation coefficient shows several peaks with equal spacing but progressively decreasing peak values. The first peak's value is always 1 (the original signal and its own correlation coefficient is 1). The correlation coefficient of the second peak is related to whether or not someone is present. When someone is present, the value is lower (due to human absorption, subsequent signals differ significantly from the original signal), and higher when no one is present (the chirps are more consistent). By comparing the peak value of the second peak in the autocorrelation coefficient with a threshold, it can be determined whether the seat is occupied. The typical value for the left seat in the middle row is around 0.93 when unoccupied and around 0.80 when occupied, but it may be lower. As shown in Figure 7a, the peak value of the second peak in the autocorrelation spectrum is 0.93, indicating that the seat is unoccupied; as shown in Figure 7b, the peak value of the second peak in the autocorrelation spectrum is 0.75, indicating that the seat is occupied.

[0199] Figure 11 is a structural schematic diagram of the personnel presence detection device 1100 of this application. As shown in Figure 11, the personnel presence detection device 1100 of this embodiment can be applied to the vehicle mentioned above. The personnel presence detection device 1100 may include: a playback module 1101, a sound receiving module 1102, and a detection module 1103.

[0200] The playback module 1001 is used to play a first audio through a first speaker, which is at least one of a plurality of speakers in the vehicle; the sound receiving module 1002 is used to collect a first sound through a first microphone to obtain a second audio, the first sound including the sound emitted when the first audio is played, the first microphone being a microphone corresponding to a first seat among a plurality of microphones in the vehicle, the first seat being one of the seats in the vehicle; the detection module 1003 is used to obtain the occupancy status of the first seat based on the second audio.

[0201] In one possible implementation, the detection module 1003 is specifically used to determine that someone is sitting in the first seat when the second audio meets the signal change characteristics, wherein the signal change characteristics include: attenuation in the energy domain is greater than a preset threshold, and / or, the frequency domain conforms to the change pattern of breathing or heartbeat.

[0202] In one possible implementation, the detection module 1003 is specifically used to compare the second audio with the first audio to obtain the seating status of the people in the first seat.

[0203] In one possible implementation, the detection module 1003 is specifically used to perform autocorrelation processing on the second audio to obtain the autocorrelation of the second audio; and to obtain the seating status of the people in the first seat based on the autocorrelation of the second audio.

[0204] In one possible implementation, the detection module 1003 is specifically used to input the second audio into a neural network to obtain the seating status of the people in the first seat.

[0205] In one possible implementation, the detection module 1003 is further configured to perform bandpass filtering and spectrum analysis on the second audio to obtain a processed second audio; and to obtain the seating status of the people in the first seat based on the processed second audio.

[0206] In one possible implementation, the first audio is an ultrasonic frequency band audio with a sampling frequency greater than 20 kHz.

[0207] In one possible implementation, there is one first speaker and one first microphone, and the first seat is located between the first speaker and the first microphone.

[0208] In one possible implementation, the first seat is a designated seat; or, the first seat is the current seat polled in a preset order within the vehicle; or, the first seat is one of a plurality of seats simultaneously detected in the vehicle.

[0209] In one possible implementation, the first microphone is located in front of the first seat, to the side of the first seat, or on the back of the first seat.

[0210] The apparatus in this embodiment can be used to execute the technical solution of the method embodiment shown in FIG5. Its implementation principle and technical effect are similar, and will not be described again here.

[0211] In implementation, each step of the above method embodiments can be completed by integrated logic circuits in the processor hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this application can be directly implemented by a hardware encoding processor, or by a combination of hardware and software modules in the encoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method.

[0212] The memory mentioned in the above embodiments can be volatile memory or non-volatile memory, or may include both. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory used in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0213] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0214] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0215] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0216] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0217] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0218] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0219] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

A method for detecting the presence of people, characterized in that, include: A first audio signal is played through a first speaker, which is at least one of a plurality of speakers in the vehicle; A second audio is obtained by acquiring a first sound through a first microphone. The first sound includes the sound emitted when the first audio is played. The first microphone is the microphone corresponding to the first seat among a plurality of microphones in the vehicle. The first seat is one of the seats in the vehicle. The seating arrangement of the people in the first seat is obtained based on the second audio. The method according to claim 1, characterized in that, The step of obtaining the seating status of the first seat based on the second audio includes: If the second audio signal meets the signal change characteristics, it is determined that someone is sitting in the first seat. The signal change characteristics include: attenuation in the energy domain is greater than a preset threshold, and / or, the frequency domain conforms to the change pattern of breathing or heartbeat. The method according to claim 1 or 2, characterized in that, The step of obtaining the seating status of the first seat based on the second audio includes: The second audio is compared with the first audio to obtain the seating situation of the people in the first seat. The method according to claim 1 or 2, characterized in that, The step of obtaining the seating status of the first seat based on the second audio includes: The second audio is subjected to autocorrelation processing to obtain the autocorrelation of the second audio. The seating status of the first seat is obtained based on the autocorrelation of the second audio. The method according to claim 1 or 2, characterized in that, The step of obtaining the seating status of the first seat based on the second audio includes: The second audio is input into the neural network to obtain the seating situation of the people in the first seat. The method according to any one of claims 1-5 is characterized in that, Before obtaining the seating status of the first seat based on the second audio, the method further includes: The second audio is subjected to bandpass filtering and spectral analysis to obtain the processed second audio; The step of obtaining the seating status of the first seat based on the second audio includes: The seating arrangement of the people in the first seat is obtained based on the processed second audio. The method according to any one of claims 1-6, characterized in that, The first audio is an ultrasonic frequency band audio with a sampling frequency greater than 20kHz. The method according to any one of claims 1-7 is characterized in that, There is one first speaker and one first microphone, and the first seat is located between the first speaker and the first microphone. The method according to any one of claims 1-8, characterized in that, The first seat is a designated seat; or, the first seat is the current seat polled in the vehicle in a preset order; or, the first seat is one of a plurality of seats detected synchronously in the vehicle. The method according to any one of claims 1-9 is characterized in that, The first microphone is located in front of the first seat, to the side of the first seat, or on the back of the first seat. A personnel presence detection device, characterized in that, include: A playback module for playing a first audio signal through a first speaker, wherein the first speaker is at least one of a plurality of speakers in the vehicle; The sound receiving module is used to acquire a first sound through a first microphone to obtain a second audio, the first sound including the sound emitted when the first audio is played, the first microphone being the microphone corresponding to the first seat among a plurality of microphones in the vehicle, and the first seat being one of the seats in the vehicle; The detection module is used to obtain the seating status of the people in the first seat based on the second audio. The apparatus according to claim 11 is characterized in that, The detection module is specifically used to determine that someone is sitting in the first seat when the second audio meets the signal change characteristics. The signal change characteristics include: attenuation in the energy domain is greater than a preset threshold, and / or, the frequency domain conforms to the change pattern of breathing or heartbeat. The apparatus according to claim 11 or 12 is characterized in that, The detection module is specifically used to compare the second audio with the first audio to obtain information on the occupancy of the first seat. The apparatus according to claim 11 or 12 is characterized in that, The detection module is specifically used to perform autocorrelation processing on the second audio to obtain the autocorrelation of the second audio; and to obtain the seating status of the people in the first seat based on the autocorrelation of the second audio. The apparatus according to claim 11 or 12 is characterized in that, The detection module is specifically used to input the second audio into a neural network to obtain information about the occupants of the first seat. The apparatus according to any one of claims 11-15 is characterized in that, The detection module is further configured to perform bandpass filtering and spectrum analysis on the second audio to obtain a processed second audio; and to obtain the seating status of the people in the first seat based on the processed second audio. The apparatus according to any one of claims 11-16 is characterized in that, The first audio is an ultrasonic frequency band audio with a sampling frequency greater than 20kHz. The apparatus according to any one of claims 11-17 is characterized in that, There is one first speaker and one first microphone, and the first seat is located between the first speaker and the first microphone. The apparatus according to any one of claims 11-18 is characterized in that, The first seat is a designated seat; or, the first seat is the current seat polled in the vehicle in a preset order; or, the first seat is one of a plurality of seats detected synchronously in the vehicle. The method according to any one of claims 11-19 is characterized in that, The first microphone is located in front of the first seat, to the side of the first seat, or on the back of the first seat. A vehicle characterized in that, include: One or more speakers are used to play the first audio. One or more microphones are used to capture the sound of the first audio being played to obtain the second audio; One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-10. A vehicle control system, characterized in that, include: One or more processors; Memory, used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in any one of claims 1-10. A computer-readable storage medium, characterized in that, It includes a computer program that, when executed on a computer, causes the computer to perform the method of any one of claims 1-10. A computer program product, characterized in that, The computer program product includes computer program code that, when run on a computer, causes the computer to perform the method of any one of claims 1-10.