Indoor person detection method based on ultrasonic array and convolutional neural network and related equipment

By calculating the relative arrival time difference of the receiving elements in the ultrasonic array and using a convolutional neural network for pattern recognition, the problem of clock synchronization dependence in ultrasonic indoor personnel detection was solved, achieving higher detection reliability and adaptability.

CN122362286APending Publication Date: 2026-07-10XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
XI'AN UNIVERSITY OF ARCHITECTURE AND TECHNOLOGY
Filing Date
2026-03-27
Publication Date
2026-07-10

Smart Images

  • Figure CN122362286A_ABST
    Figure CN122362286A_ABST
Patent Text Reader

Abstract

This invention discloses an indoor personnel detection method and related equipment based on an ultrasonic array and a convolutional neural network, comprising: acquiring multi-channel ultrasonic received signals collected by all receiving elements in the ultrasonic array within the same time period; calculating the arrival time difference between signals received by any two different receiving elements in the ultrasonic array based on the multi-channel ultrasonic received signals, to construct a time delay feature matrix characterizing the relative delay relationship of received signals between receiving elements; inputting the time delay feature matrix into a pre-trained personnel detection classification model, and having the personnel detection classification model output a judgment result indicating whether there are personnel indoors; wherein, the personnel detection classification model is obtained by training a convolutional neural network using a training sample set containing the time delay feature matrix. The purpose of this invention is to avoid or eliminate the dependence on high-precision cross-device clock synchronization and improve the reliability of ultrasonic indoor personnel detection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of array signal processing technology, specifically relating to an indoor personnel detection method and related equipment based on an ultrasonic array and a convolutional neural network. Background Technology

[0002] The use of ultrasound to sense or detect people indoors is primarily based on the reflection and propagation characteristics of sound waves. A typical existing technology involves deploying a network of ultrasonic transmitters and receivers within the indoor space. The transmitters periodically emit ultrasonic signals, while the receivers receive the echo signals or direct signals reflected from people or objects. By measuring the flight time of the signal from transmission to reception, calculating the time difference of arrival between different receivers, or analyzing the pattern characteristics of the reflected signals, the presence of people can be inferred.

[0003] The effectiveness of such methods relies on the ability to accurately measure the absolute propagation time or absolute arrival time of ultrasonic signals. Therefore, the system must ensure strict and high-precision time synchronization between the transmitter and all receivers. In practical hardware implementations, each device typically relies on its own crystal oscillator as a local time reference. However, crystal oscillators inherently exhibit frequency drift, and even minute drifts introduce tiny clock skews between devices that are difficult to completely eliminate. Since the propagation time of ultrasound in indoor environments is on the order of milliseconds, these microsecond or even nanosecond-level clock skews directly translate into distance measurement errors, thus limiting the reliability of the final human perception. Summary of the Invention

[0004] To address the problems existing in the prior art, this invention provides an indoor personnel detection method and related equipment based on ultrasonic arrays and convolutional neural networks. Its purpose is to avoid or eliminate the dependence on high-precision cross-device clock synchronization and improve the reliability of ultrasonic indoor personnel detection.

[0005] To solve the above-mentioned technical problems, the present invention is achieved through the following technical solution: According to a first aspect of the present invention, an indoor people detection method based on an ultrasonic array and a convolutional neural network is provided, comprising: Acquire multi-channel ultrasonic received signals collected by all receiving elements in the ultrasonic array within the same time period; Based on the multi-channel ultrasonic received signal, the arrival time difference between the signals received by any two different receiving elements in the ultrasonic array is calculated to form a time delay feature matrix that characterizes the relative delay relationship between the received signals of the receiving elements. The time delay feature matrix is ​​input into a pre-trained personnel detection and classification model, which outputs a judgment result on whether there are personnel indoors; wherein, the personnel detection and classification model is obtained by training a convolutional neural network using a training sample set containing the time delay feature matrix.

[0006] In one possible implementation of the first aspect, the calculation of the time difference of arrival between signals received by any two different receiving elements in the ultrasonic array is specifically performed using a generalized cross-correlation method.

[0007] In one possible implementation of the first aspect, the generalized cross-correlation method employs phase transformation weighting.

[0008] In one possible implementation of the first aspect, the time delay feature matrix is ​​an N×N square matrix, where N is the number of receiving elements in the ultrasonic array, and the matrix contains the N×N square matrix. i Line 1 j The elements of the column represent the first... i The receiving array element and the first j The time difference of arrival of the received signal between each receiving array element.

[0009] In one possible implementation of the first aspect, the convolutional neural network model sequentially comprises: an input layer, at least one convolutional layer, at least one pooling layer, at least one fully connected layer, and an output layer; the input layer is used to receive the time delay feature matrix.

[0010] In one possible implementation of the first aspect, the personnel detection classification model is trained in the following manner: Obtain a training dataset, the training sample set including a first type of time delay feature matrix sample collected when there are only static obstacles indoors, and a second type of time delay feature matrix sample collected when there are people in the same indoor environment; The convolutional neural network is trained using the training dataset until its loss function converges, thus obtaining the personnel detection and classification model.

[0011] In one possible implementation of the first aspect, the receiving elements in the ultrasonic array are arranged in a spiral divergent pattern.

[0012] In one possible implementation of the first aspect, prior to acquiring the multi-channel ultrasonic received signals collected by all receiving elements in the ultrasonic array within the same time period, the method further includes: The detection signal is controlled to emit an ultrasonic source at a frequency within a preset ultrasonic frequency band. The acquired multi-channel ultrasonic received signal includes the echo signal after the detection signal is reflected by the indoor environment.

[0013] In one possible implementation of the first aspect, after acquiring the multi-channel ultrasonic received signal, the received signal is further preprocessed. The preprocessing includes: performing a main frequency analysis on the received signal and performing bandpass filtering based on the analysis results. The passband range of the bandpass filter matches the main frequency component of the detected signal.

[0014] According to a second aspect of the present invention, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the computer program, implements the aforementioned indoor personnel detection method based on an ultrasonic array and a convolutional neural network.

[0015] According to a third aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing a computer program, which, when executed by a processor, implements the aforementioned indoor personnel detection method based on an ultrasonic array and a convolutional neural network.

[0016] According to a fourth aspect of the present invention, a computer program product is provided that, when executed by a processor, implements the aforementioned indoor personnel detection method based on an ultrasonic array and a convolutional neural network.

[0017] Compared with the prior art, the present invention has at least the following beneficial effects: This invention provides an indoor personnel detection method based on an ultrasonic array and a convolutional neural network. It fundamentally avoids the reliance on high-precision cross-device clock synchronization. By using the arrival time difference between any two receiving elements within the ultrasonic array as the core feature, these time differences represent the relative delay relationships between channels within the array itself. Their calculation does not depend on whether the absolute time reference between the transmitter and receiver array, or between different receiving devices, is strictly synchronized. Therefore, it effectively avoids the problem of small time errors introduced by the inherent drift of the crystal oscillator being amplified into distance errors, thus eliminating, in principle, the main error source limiting the accuracy of traditional ultrasonic detection methods. The calculated time differences between all array element pairs are organized into a time delay feature matrix. This matrix can characterize the relative delay patterns of sound waves arriving at various spatial locations within the array. When people are active indoors, their reflection, scattering, and blocking of sound waves alter the spatial distribution of the sound field, thereby inducing characteristic pattern changes in the time delay feature matrix. This feature based on relative relationships is more sensitive to the entry and activity of people. The time delay feature matrix is ​​input into a pre-trained convolutional neural network model for processing. The model can automatically learn and extract discriminative features that distinguish between the presence and absence of people from a large number of time delay feature matrix samples, and directly output the detection results. This avoids the difficulty of manually designing complex thresholds or rule algorithms, reduces the dependence on prior knowledge and parameter tuning, and enhances the adaptability to different indoor scenarios. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the specific embodiments of the present invention, the drawings used in the description of the specific embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0019] Figure 1 This is a flowchart of an indoor personnel detection method based on an ultrasonic array and a convolutional neural network according to the present invention.

[0020] Figure 2 This is a flowchart of the main process for feature processing and fusion in convolutional neural networks.

[0021] Figure 3 This diagram illustrates how model performance is evaluated using a multi-dimensional metric based on a confusion matrix on the test set.

[0022] Figure 4 This is a diagram showing the spiral divergent arrangement of the receiving elements in an ultrasonic array.

[0023] Figure 5 This is a schematic diagram of an ultrasound array and sound source sensing the human body.

[0024] Figure 6 This is a frequency distribution diagram of the data received by the ultrasonic array. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0026] The core of this invention lies in extracting and analyzing the relative arrival time difference between the signals of each receiving element within the ultrasonic array as a feature, and using a convolutional neural network to perform pattern recognition on this feature, thereby realizing the detection of whether people are present indoors, effectively avoiding the dependence of traditional methods on high-precision clock synchronization across devices.

[0027] like Figure 1 As shown, this invention provides an indoor people detection method based on an ultrasonic array and a convolutional neural network, specifically including the following steps: Step 1: Acquire multi-channel ultrasonic received signals collected by all receiving elements in the ultrasonic array within the same time period.

[0028] Specifically, the system is activated, and all receiving elements in the ultrasonic array synchronously acquire ultrasonic signals over a period of time. In this embodiment, 5 seconds of audio data are saved for each data collection. The ultrasonic array contains 64 receiving channels, and its data is stored in two WAV files, corresponding to the 32 channels of audio data collected from channels 1-32 and channels 33-64, respectively. The data processing steps are as follows: first, the two 32-channel audio data are merged into 64-channel audio data; then, the audio data is segmented into short audio files of 0.1 seconds each. These short audio files constitute the multi-channel ultrasonic receiving signal.

[0029] Step 2: Based on the multi-channel ultrasonic received signal, calculate the arrival time difference between the signals received by any two different receiving elements in the ultrasonic array, so as to form a time delay feature matrix that characterizes the relative delay relationship of the received signals between the receiving elements.

[0030] Specifically, the purpose of this step is to extract key features that characterize the presence or absence of a person from the aforementioned multi-channel ultrasonic received signals. This invention no longer relies on calculating the absolute flight time of the signal from transmission to reception, but rather focuses on the relative relationships between different receiving elements within the ultrasonic array. Specifically, for a short audio file of 0.1 seconds, the arrival time difference between the signals received by any two different receiving elements in the ultrasonic array is calculated. The time differences of all receiving element pairs are organized in a specific order to form a matrix that clearly characterizes the relative delay relationships between all receiving elements; this is called the time delay feature matrix. In indoor environments, the close proximity of the equipment results in typically small collected time delay feature data values; in this embodiment, it is considered to simultaneously amplify these values ​​by a factor of 1000.

[0031] Step 3: Input the time delay feature matrix into the pre-trained personnel detection classification model, and have the personnel detection classification model output the judgment result of whether there are personnel indoors; wherein, the personnel detection classification model is obtained by training a convolutional neural network using a training sample set containing the time delay feature matrix.

[0032] Specifically, the time delay feature matrix obtained in step 2 is input into a pre-trained people detection and classification model. It should be noted that the people detection and classification model is essentially a convolutional neural network that has been trained on a large number of labeled time delay feature matrix samples, learning to identify different time delay patterns corresponding to the presence or absence of people. After receiving the input, the people detection and classification model performs a series of nonlinear transformations and calculations internally, ultimately providing a binary judgment result at the output layer: whether there are people or not indoors at the current time.

[0033] In one possible implementation, the calculation of the time difference of arrival between signals received by any two different receiving elements in the ultrasonic array is specifically performed using a generalized cross-correlation method.

[0034] In detail, this embodiment employs a generalized cross-correlation method when calculating the time difference of arrival between any two received array element signals. The generalized cross-correlation method effectively suppresses the influence of noise and improves the accuracy of time delay estimation. Specifically, for signals from the first... i The receiving array element and the first j Two received signals from each receiving array element and First, perform Fourier transforms on each to obtain its spectrum. and ,Right now Then, calculate the cross-power spectrum of these two spectra. ,in yes The complex conjugate of the cross-power spectrum. The phase information of the cross-power spectrum contains the two received signals. The time delay between and.

[0035] Preferably, to further improve the robustness of the generalized cross-correlation method in complex indoor acoustic environments, this embodiment applies a specific weighting function to the cross-power spectrum during calculation. Specifically, phase transform weighting is used, and the weighted cross-power spectrum... .in, Indicates the amplitude of the cross-power spectrum. It is a very small constant used to prevent the denominator from being zero. The weighted cross-power spectrum applies the weighting function to the cross-power spectrum. It mainly reflects the phase difference of the signal.

[0036] Weighted cross-power spectrum Performing an inverse Fourier transform yields the generalized cross-correlation function, the peak position of which corresponds to the time delay between the two signals. (i.e., the time difference of arrival), specifically: .

[0037] In one possible implementation, the time delay feature matrix is ​​an N×N square matrix, where N is the number of receiving elements in the ultrasonic array, and the matrix contains the N-th element. i Line 1 j The elements of the column represent the first... i The receiving array element and the first j The time difference of arrival of the received signal between each receiving array element.

[0038] In other words, the time differences between all receiver array element pairs calculated by the above process are organized into a matrix. In this embodiment, this time delay feature matrix is ​​constructed as an N×N square matrix, where N equals the number of receiver array elements in the ultrasonic array; in this example, N=64. In this matrix, the... i Line 1 j Column elements That is to say, the first i The receiving array element and the first j The arrival time difference of the received signals between each receiving element is represented by a 64×64 matrix structure, which characterizes the spatial relative delay relationship of the sound waves arriving at each receiving element of the ultrasonic array at the current moment.

[0039] In one implementation, the convolutional neural network model sequentially comprises: an input layer, at least one convolutional layer, at least one pooling layer, at least one fully connected layer, and an output layer; the input layer is used to receive the time delay feature matrix.

[0040] The main flowchart of the feature processing and fusion of the convolutional neural network used in this embodiment for the personnel detection and classification model is as follows: Figure 2As shown, the network's input layer is designed to receive the aforementioned 64×64 time delay feature matrix.

[0041] The first convolutional layer uses 32 3×3 convolutional kernels, performs convolution with a stride of 1, and uses the ReLU activation function. By introducing nonlinearity, the output feature map size is maintained at 64×64×32. Subsequently, a 2×2 max pooling layer with a stride of 2 is applied to downsample the feature map to 32×32×32.

[0042] The second convolutional layer uses 64 3×3 convolutional kernels, also with a stride of 1, and ReLU activation, outputting a feature map of size 32×32×64. This is followed by a 2×2 max pooling layer with a stride of 2, resulting in a 16×16×64 feature map.

[0043] The third convolutional layer uses 128 3×3 convolutional kernels with a stride of 1 and ReLU activation, outputting a 16×16×128 feature map. After passing through a 2×2 max pooling layer with a stride of 2, the feature map scale becomes 8×8×128.

[0044] Each convolutional kernel is initialized using a He normal distribution.

[0045] Next, the network is fed into a flattening layer, which flattens the 8×8×128 three-dimensional feature map into an 8192-dimensional vector. This 8192-dimensional vector is then input into the first fully connected layer, which contains 256 neurons and uses the ReLU activation function. Its output is then fed into a second fully connected layer (containing 128 neurons) for computation. To prevent overfitting, a Dropout layer is placed after the fully connected layers, randomly ignoring 30% of the neuron outputs during training.

[0046] Finally, the network passes through an output layer and uses a sigmoid activation function to generate binary classification probabilities. The training objective of the entire network is achieved through a binary cross-entropy loss function. Optimize, where y is the true label and p is the predicted probability.

[0047] In one possible implementation, the personnel detection classification model is trained by: acquiring a training dataset, the training sample set including first-class time delay feature matrix samples collected when there are only static obstacles indoors, and second-class time delay feature matrix samples collected when there are personnel in the same indoor environment; training a convolutional neural network with the training dataset until its loss function converges to obtain the personnel detection classification model.

[0048] In detail, the construction of training data for the personnel detection and classification model determines whether the model can effectively distinguish between personnel and static environments. First, in the target indoor environment, data is pre-collected under two scenarios: the first is a scenario without personnel, where the room contains only static obstacles such as tables, chairs, and cabinets. Ultrasonic signals are collected in this scenario and processed to obtain a large number of Type I time delay feature matrix samples, which encode the reflection characteristics of the static environment. The second scenario involves personnel. Under the same static environment setup, personnel enter in different postures, such as standing, sitting, and walking. Signals are collected and processed to obtain a large number of Type II time delay feature matrix samples. All samples are labeled with their corresponding category label, i.e., no one or someone.

[0049] Subsequently, the entire dataset was divided into training and test sets in a 6:4 ratio, with a random seed of 42 to ensure the reproducibility of experimental results. During training, the convolutional neural network with the above structure was iteratively trained using the training set data. The Adam optimizer was used to adjust the network parameters, with a maximum training epoch of 100 and a batch size of 32. A further 20% of the training set was allocated as a validation set to monitor the training process.

[0050] Two callback strategies were implemented during training to prevent overfitting and optimize efficiency: First, early stopping, which monitors the validation set loss and automatically terminates training and restores the optimal weights after 15 consecutive training epochs without improvement. Second, an adaptive learning rate adjustment strategy, which, based on validation set loss performance, reduces the learning rate by 0.3 after 8 consecutive epochs without improvement, with a minimum learning rate limit of 1×10⁻⁶. -6 Training continues until the model's loss function converges and its performance stabilizes. The resulting model then possesses the ability to identify patterns of human disturbance from the time-delay feature matrix, thereby eliminating interference from fixed, static obstacles. Figure 3 As shown, the model performance is evaluated using a multi-dimensional metric, the confusion matrix on the test set. By cross-referencing the prediction results with the true labels, the limitations of a single accuracy metric are overcome.

[0051] In one possible implementation, the receiving elements in the ultrasonic array are arranged in a spiral divergent pattern.

[0052] In other words, in this embodiment, the 64 receiving array elements used are SPH0641LU4H-1 miniature microphone sensors, which are not arranged in a regular grid, but rather in a spiral divergent arrangement, as shown in the attached diagram. Figure 4 As shown. This arrangement allows for better anisotropy in the spatial distribution of the receiving array elements, ensuring that reflected sound waves arriving at the array from different directions generate more discriminative time difference combinations among the receiving array elements.

[0053] In one possible implementation, before acquiring the multi-channel ultrasonic received signals collected by all receiving elements in the ultrasonic array within the same time period, the method further includes: controlling the emission frequency of the ultrasonic emission source to be within a preset ultrasonic frequency band; the acquired multi-channel ultrasonic received signals include echo signals reflected by the detection signals after passing through the indoor environment.

[0054] In other words, before the signal acquisition step begins, the system controls an independent ultrasonic transmitter to emit a detection signal within a preset ultrasonic frequency band; in this embodiment, a Hesent EC ultrasonic transducer is used. Subsequently, the 64 receiving elements (SPH0641LU4H-1) in the ultrasonic array begin synchronous signal acquisition. The multi-channel ultrasonic received signal acquired at this time mainly includes the echo signal formed by the reflection of the detection signal through the indoor environment (including walls, furniture, and human bodies), such as... Figure 5 The diagram shows the array and sound source used to sense the human body.

[0055] In one possible implementation, after acquiring the multi-channel ultrasonic received signal, the method further includes preprocessing the received signal. The preprocessing includes performing a main frequency analysis on the received signal and performing bandpass filtering based on the analysis results. The passband range of the bandpass filter matches the main frequency component of the detected signal.

[0056] In other words, after merging channels and segmenting signal frames, before performing the actual delay calculation, preprocessing of the 64-channel time-domain signal in each frame is required to improve the signal-to-noise ratio and feature quality. First, the input 0.1-second 64-channel time-domain signal undergoes main frequency analysis, such as... Figure 6 The frequency distribution diagram is used to determine the frequency components where the energy is mainly concentrated in the received signal. Based on the results of the main frequency analysis, bandpass filtering is performed. A bandpass filter with a passband range matching the main frequency component is designed to filter the received signal of each channel. This effectively removes low-frequency noise from the environment and high-frequency noise introduced by the circuit, outputting a filtered time-domain signal while retaining useful ultrasonic echo signals.

[0057] This invention extracts the phase feature without using the conventional echo time. When people are present indoors, the propagation path of the ultrasonic signal emitted by the ultrasonic transducer changes, which in turn causes a change in the arrival time of the signal collected by the array. The time difference between the received signals of different array elements is extracted as the phase feature based on the arrangement of the ultrasonic array elements. In this way, it is only necessary to calculate the time difference when there are people present, and there is no need to consider the clock synchronization between ultrasonic devices.

[0058] In another embodiment of the present invention, a computer device is provided, comprising a processor and a memory. The memory stores a computer program, which includes program instructions. The processor executes the program instructions stored in the computer storage medium. The processor may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing and control core of the terminal, suitable for implementing one or more instructions, specifically suitable for loading and executing one or more instructions in the computer storage medium to achieve a corresponding method flow or corresponding function. The processor described in this embodiment of the present invention can be used in the operation of an indoor personnel detection method based on an ultrasonic array and a convolutional neural network.

[0059] In another embodiment of the present invention, a storage medium is provided, specifically a computer-readable storage medium (Memory), which is a memory device in a computer device used to store programs and data. It is understood that the computer-readable storage medium here can include both the built-in storage medium in the computer device and extended storage media supported by the computer device. The computer-readable storage medium provides storage space that stores the operating system of the terminal. Furthermore, the storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more computer programs (including program code). It should be noted that the computer-readable storage medium here can be Random Access Memory (RAM) or non-volatile memory, such as at least one disk storage device. The processor can load and execute one or more instructions stored in the computer-readable storage medium to implement the corresponding steps of the indoor personnel detection method based on an ultrasonic array and a convolutional neural network in the above embodiments.

[0060] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, optical storage, etc.) containing computer-usable program code.

[0061] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0062] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0063] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0064] This invention also provides a computer program product for executing any of the aforementioned indoor personnel detection methods based on ultrasonic arrays and convolutional neural networks. Since the computer program product provided by this invention belongs to the same inventive concept as the aforementioned indoor personnel detection method based on ultrasonic arrays and convolutional neural networks, it possesses all the advantages of the aforementioned method. Therefore, the beneficial effects of the computer program product provided by this invention will not be elaborated upon here.

[0065] In this invention, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to a specific feature, structure, material, or characteristic described in connection with that embodiment or example, which is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0066] Finally, it should be noted that the above-described embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit them. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the scope of the technology disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention.

Claims

1. An indoor personnel detection method based on ultrasonic array and convolutional neural network, characterized in that, include: Acquire multi-channel ultrasonic received signals collected by all receiving elements in the ultrasonic array within the same time period; Based on the multi-channel ultrasonic received signal, the arrival time difference between the signals received by any two different receiving elements in the ultrasonic array is calculated to form a time delay feature matrix that characterizes the relative delay relationship between the received signals of the receiving elements. The time delay feature matrix is ​​input into a pre-trained personnel detection and classification model, which outputs a judgment result on whether there are personnel indoors; wherein, the personnel detection and classification model is obtained by training a convolutional neural network using a training sample set containing the time delay feature matrix.

2. The indoor personnel detection method based on ultrasonic array and convolutional neural network according to claim 1, characterized in that, The calculation of the arrival time difference between signals received by any two different receiving elements in the ultrasonic array is specifically performed using the generalized cross-correlation method.

3. The indoor personnel detection method based on ultrasonic array and convolutional neural network according to claim 2, characterized in that, The generalized cross-correlation method employs phase transformation weighting.

4. The indoor personnel detection method based on ultrasonic array and convolutional neural network according to claim 1, characterized in that, The time delay feature matrix is ​​an N×N square matrix, where N is the number of receiving elements in the ultrasonic array, and the matrix is ​​represented by the N×N square matrix. i Line number j The elements of the column represent the first... i The receiving array element and the first j The time difference of arrival of the received signal between each receiving array element.

5. The indoor personnel detection method based on ultrasonic array and convolutional neural network according to claim 1, characterized in that, The convolutional neural network model sequentially comprises: an input layer, at least one convolutional layer, at least one pooling layer, at least one fully connected layer, and an output layer; the input layer is used to receive the time delay feature matrix.

6. The indoor personnel detection method based on ultrasonic array and convolutional neural network according to claim 5, characterized in that, The personnel detection and classification model was trained in the following way: Obtain a training dataset, the training sample set including a first type of time delay feature matrix sample collected when there are only static obstacles indoors, and a second type of time delay feature matrix sample collected when there are people in the same indoor environment; The convolutional neural network is trained using the training dataset until its loss function converges, thus obtaining the personnel detection and classification model.

7. The indoor personnel detection method based on ultrasonic array and convolutional neural network according to claim 1, characterized in that, The receiving elements in the ultrasonic array are arranged in a spiral divergent pattern.

8. The indoor personnel detection method based on ultrasonic array and convolutional neural network according to claim 1, characterized in that, Before acquiring the multi-channel ultrasonic received signals collected by all receiving elements in the ultrasonic array within the same time period, the method further includes: The detection signal is controlled to emit an ultrasonic source at a frequency within a preset ultrasonic frequency band. The acquired multi-channel ultrasonic received signal includes the echo signal after the detection signal is reflected by the indoor environment.

9. The indoor personnel detection method based on ultrasonic array and convolutional neural network according to claim 8, characterized in that, After acquiring the multi-channel ultrasonic received signal, the method further includes preprocessing the received signal. The preprocessing includes performing main frequency analysis on the received signal and performing bandpass filtering based on the analysis results. The passband range of the bandpass filter matches the main frequency component of the detected signal.

10. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements an indoor personnel detection method based on an ultrasonic array and a convolutional neural network as described in any one of claims 1 to 9.