Object recognition method and apparatus, object recognition model training method and apparatus, electronic device, storage medium and program product
By sending sound signals of different frequency bands to obtain echo signals for object recognition, the problem of low accuracy caused by the dependence on image recognition is solved, and higher object recognition accuracy and stability are achieved.
Patent Information
- Application Number
- PCT/CN2024/116117
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-05-23
- Filing Date
- 2024-08-30
- Publication Date
- 2025-11-27
AI Technical Summary
In existing technologies, the accuracy of object recognition depends on image recognition, which means that images containing sensitive information may result in lower accuracy in object recognition.
By sending sound signals of different frequency bands at different time periods, corresponding echo signals are obtained, and object recognition is performed based on these echo signals. The first frequency band signal is used to locate the object's position, and the second frequency band signal is used to detect the object's movement. The signal processing is combined with short-time Fourier transform to improve the recognition accuracy.
It improves the accuracy and stability of object recognition, reduces the impact of environmental interference and signal distortion, and is suitable for application scenarios that require precise control and positioning.
Smart Images

Figure CN2024116117_27112025_PF_FP_ABST
Abstract
Description
Object recognition method, object recognition model training method, device, electronic equipment, storage medium and program product
[0001] The present application claims priority to the Chinese patent application No. 202410649636.2, filed on May 23, 2024, and entitled "Object recognition method, object recognition model training method, device, electronic equipment, storage medium and program product", the content of which is incorporated herein by reference in its entirety. TECHNICAL FIELD
[0002] The present application relates to the technical field of computer, and in particular to an object recognition method, an object recognition model training method, a device, an electronic equipment, a storage medium and a program product. BACKGROUND
[0003] Artificial intelligence (AI) is a comprehensive discipline, involving a wide range of fields, both hardware and software technologies. Artificial intelligence basic technologies generally include, such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, pre-training model technology, operation / interaction system, mechatronics, etc. Among them, the pre-training model is also called large model, basic model, which can be widely applied to downstream tasks in various directions of artificial intelligence after fine-tuning. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning, etc.
[0004] SUMMARY
[0005] The embodiments of the present application provide an object recognition method, device, electronic equipment, computer readable storage medium and computer program product, which can effectively improve the accuracy of object recognition.
[0006] The technical solutions of the embodiments of the present application are implemented as follows:
[0007] The embodiments of the present application provide an object recognition method, comprising: sending a sound signal within a certain period, the period comprising a first period and a second period; wherein the sound signal in the first period is a sound signal of a first frequency band, and the sound signal in the second period is a sound signal of a second frequency band; obtaining a first echo signal corresponding to the sound signal of the first frequency band and a second echo signal corresponding to the sound signal of the second frequency band; and obtaining an object recognition result based on the first echo signal and the second echo signal.
[0008] The embodiment of the present application provides a kind of object identification model training method, comprising: obtaining the echo signal sample of sample object, the echo signal sample includes first echo signal sample and second echo signal sample;Call initial object identification model, based on the first echo signal sample and the second echo signal sample, obtain the object identification result of the sample object;
[0009] Based on the object identification result of the sample object and the sample label carried by the echo signal sample, the initial object identification model is trained to obtain the object identification model;Wherein, the object identification model is used to obtain the object identification result based on the first echo signal and the second echo signal of the object to be identified.
[0010] The embodiment of the present application provides an object identification device, comprising: a sending module for sending sound signals within a certain period, the period includes first period and second period;Wherein, the sound signal in the first period is the sound signal of the first frequency band, and the sound signal in the second period is the sound signal of the second frequency band;Division module is used to obtain the first echo signal corresponding to the sound signal of the first frequency band and the second echo signal corresponding to the sound signal of the second frequency band;Object identification module is used to obtain object identification result based on the first echo signal and the second echo signal.
[0011] The embodiment of the present application provides a kind of object identification model training device, comprising: obtaining module, obtains the echo signal sample of sample object, the echo signal sample includes first echo signal sample and second echo signal sample;Sample identification module is used to call initial object identification model, based on the first echo signal sample and the second echo signal sample, obtain the object identification result of the sample object;
[0012] Training module is used to train initial object identification model based on the object identification result of the sample object and the sample label carried by the echo signal sample, to obtain the object identification model;Wherein, the object identification model is used to obtain the object identification result based on the first echo signal and the second echo signal of the object to be identified.
[0013] The embodiment of the present application provides an electronic device, comprising: memory, for storing computer executable instructions or computer programs;Processor is used to execute the computer executable instructions or computer program stored in the memory, when the object identification method or the training method of object identification model provided by the embodiment of the present application is realized.
[0014] The embodiment of the present application provides a kind of computer readable storage medium, stores computer executable instructions, for causing processor to execute when, realize the object identification method or the training method of object identification model provided by the embodiment of the present application.
[0015] An embodiment of the present application provides a computer program product, which comprises a computer program or computer executable instructions stored in a computer readable storage medium. A processor of an electronic device reads the computer executable instructions from the computer readable storage medium, and the processor executes the computer executable instructions, so that the electronic device executes the object recognition method or the training method of the object recognition model described above in the embodiments of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0016] FIG. 1 is a schematic diagram of an architecture of an object recognition system according to an embodiment of the present application;
[0017] FIG. 2 is a schematic diagram of a structure of an electronic device for object recognition according to an embodiment of the present application;
[0018] FIG. 3 is a schematic diagram of a structure of an electronic device for training an object recognition model according to an embodiment of the present application;
[0019] FIG. 4 is a schematic diagram of a flow of an object recognition method according to an embodiment of the present application;
[0020] FIG. 5 is a schematic diagram of a flow of a training method of an object recognition model according to an embodiment of the present application;
[0021] FIG. 6 is a schematic diagram of a principle of an object recognition model according to an embodiment of the present application;
[0022] FIG. 7 is a schematic diagram of a principle of an object recognition method according to an embodiment of the present application;
[0023] FIG. 8 is a schematic diagram of a principle of sound signal processing according to an embodiment of the present application. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical scheme and advantages of the present application clearer, the present application will be described in further detail below with reference to the drawings, and the described embodiments should not be regarded as limiting the present application. All other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0025] In the following description, “some embodiments” are related to a subset of all possible embodiments, but it can be understood that “some embodiments” can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0026] In the following description, the terms "first, second, third" are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that "first, second, third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0027] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0028] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0029] 1) Artificial Intelligence (AI): This is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Pre-trained models, also known as large models or foundational models, can be widely applied to downstream tasks across various AI fields after fine-tuning. AI software technologies mainly include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0030] 2) Machine Learning (ML): This is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and pre-trained learning. Pre-trained models are the latest development in deep learning, integrating the above techniques.
[0031] 3) Convolutional Neural Networks (CNN): A type of feed-forward neural network (FNN) that contains convolutional computations and has a deep structure, it is one of the representative algorithms of deep learning (Deep Learning). Convolutional neural networks have representation learning capabilities and can perform shift-invariant classification on input images according to their hierarchical structure.
[0032] 4) In response to: used to represent the conditions or states that the executed operation depends on, when the dependent conditions or states are met, the executed one or more operations can be real-time or have a set delay; in the absence of special instructions, there is no restriction on the execution order of multiple operations.
[0033] 5) Object Recognition: In the field of computer vision and artificial intelligence, it refers to the process of identifying and classifying objects in images or videos through computer algorithms and models. Object recognition is the foundation of computer vision and artificial intelligence, and is one of the key technologies in many application fields such as autonomous driving, medical diagnosis, intelligent monitoring, etc. Object recognition usually involves knowledge of image processing, computer vision, machine learning, etc. Common algorithms include feature extraction-based algorithms and deep learning-based algorithms.
[0034] 6) Short-Time Fourier Transform (STFT): A commonly used technique in signal processing, it is used to decompose signals in the time-frequency plane. It can decompose a signal into a series of weighted, overlapping windows, each corresponding to a frequency component of the signal. Fourier transform converts the signal in these windows into its representation in both time and frequency, providing time-frequency analysis, which is very important for signal processing and noise suppression tasks.
[0035] In the implementation process of the embodiments of the present application, the applicant found that the related art has the following problems:
[0036] In the related art, the image recognition method is usually used to identify the to-be-identified object. Since the accuracy of object recognition depends on the accuracy of image recognition, the to-be-identified image containing sensitive information will result in low accuracy of object recognition.
[0037] The object recognition method and device, the electronic device, the computer readable storage medium and the computer program product provided in the embodiments of the present application can effectively improve the accuracy of object recognition. The exemplary application of the object recognition system provided in the embodiments of the present application is described below.
[0038] Referring to FIG. 1, FIG. 1 is a schematic diagram of an architecture of an object recognition system 100 provided in the embodiments of the present application. A terminal (exemplarily shown as terminal 400) is connected to a server 200 through a network 300. The network 300 can be a wide area network or a local area network, or a combination of the two.
[0039] The terminal 400 is configured to be used by a user to use a client 410 to display an object recognition result on a graphical interface 410-1 (exemplarily shown as graphical interface 410-1). The terminal 400 and the server 200 are connected to each other through a wired or wireless network.
[0040] In some embodiments, the server 200 can be a stand-alone physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and basic cloud computing services such as big data and artificial intelligence platforms. The terminal 400 can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart speaker, a smart television, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto. The electronic device provided in the embodiments of the present application can be implemented as a terminal or a server. The terminal and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in the embodiments of the present application.
[0041] In some embodiments, the terminal 400 sends a sound signal in a certain period of time, the server 200 obtains a first echo signal corresponding to the sound signal of a first frequency band and a second echo signal corresponding to the sound signal of a second frequency band, obtains an object recognition result based on the first echo signal and the second echo signal, and sends the object recognition result to the terminal 400.
[0042] In some other embodiments, the terminal 400 sends a sound signal in a certain period of time, obtains a first echo signal corresponding to the sound signal of a first frequency band and a second echo signal corresponding to the sound signal of a second frequency band, obtains an object recognition result based on the first echo signal and the second echo signal, and sends the object recognition result to the server 200.
[0043] In some other embodiments, the embodiments of the present application can be implemented by means of cloud technology. Cloud technology refers to a kind of hosting technology that unifies a series of resources such as hardware, software, network, etc. in a wide area network or a local area network to realize data calculation, storage, processing and sharing.
[0044] Cloud technology is a general term for network technology, information technology, integration technology, management platform technology and application technology applied based on cloud computing business model, can form a resource pool, and is used on demand, flexible and convenient. Cloud computing technology will become an important support. The background service of the technical network system needs a large amount of calculation and storage resources.
[0045] Referring to FIG. 2, FIG. 2 is a structural schematic diagram of an electronic device 500 for object recognition provided by the embodiments of the present application, wherein the electronic device 500 shown in FIG. 2 can be the server 200 or the terminal 400 in FIG. 1, and the electronic device 500 shown in FIG. 2 includes at least one processor 430, a memory 450, and at least one network interface 420. The various components in the electronic device 500 are coupled together by a bus system 440. It can be understood that the bus system 440 is used to realize the connection and communication between the components. In addition to including a data bus, the bus system 440 also includes a power bus, a control bus, and a status signal bus. However, for the purpose of clear illustration, all kinds of buses are marked as the bus system 440 in FIG. 2.
[0046] The processor 430 can be an integrated circuit chip with signal processing capability, such as a general purpose processor, a digital signal processor (DSP), or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, etc., wherein the general purpose processor can be a microprocessor or any conventional processor.
[0047] The memory 450 can be removable, non-removable or a combination thereof. Exemplary hardware devices include solid state memory, hard disk drive, optical disk drive, etc. The memory 450 can optionally include one or more storage devices that are physically located away from the processor 430.
[0048] The memory 450 includes volatile memory or non-volatile memory, and can also include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), and the volatile memory can be random access memory (RAM). The memory 450 described in the embodiments of the present application is intended to include any suitable type of memory.
[0049] In some embodiments, the memory 450 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or a subset or superset thereof, which are exemplarily illustrated below.
[0050] The operating system 451 includes system programs for processing various basic system services and performing hardware-related tasks, such as a framework layer, a core library layer, a driver layer, and the like, for implementing various basic services and processing hardware-based tasks.
[0051] The network communication module 452 is used to reach other electronic devices via one or more (wired or wireless) network interfaces 420, exemplary network interfaces 420 including Bluetooth, Wireless Fidelity (WiFi), and Universal Serial Bus (USB), and the like.
[0052] In some embodiments, the object recognition apparatus provided by the embodiments of the present application can be implemented in a software manner, and FIG. 2 shows an object recognition apparatus 455 stored in the memory 450, which can be software in the form of programs and plug-ins, including the following software modules: a sending module 4551, a division module 4552, and an object recognition module 4553. These modules are logical, and thus can be combined or further split according to the functions implemented. The functions of each module will be described below.
[0053] Referring to FIG. 3, FIG. 3 is a structural schematic diagram of an electronic device for training an object recognition model provided by the embodiments of the present application, wherein the electronic device 600 shown in FIG. 3 can be the server 200 or the terminal 400 in FIG. 1, and the electronic device 600 shown in FIG. 3 includes at least one processor 530, a memory 550, and at least one network interface 520. The various components in the electronic device 600 are coupled together through a bus system 540. It can be understood that the bus system 540 is used to realize the connection communication between the components. In addition to including a data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for the purpose of clarity and conciseness, all kinds of buses are marked as the bus system 540 in FIG. 3.
[0054] The processor 530 can be an integrated circuit chip having a signal processing capability, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, and the like, wherein the general-purpose processor can be a microprocessor or any conventional processor.
[0055] The memory 550 can be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state memory, hard drives, optical drives, and the like. The memory 550 optionally includes one or more storage devices remotely located from the processor 530.
[0056] The memory 550 includes volatile memory or nonvolatile memory, and can also include both volatile and nonvolatile memory. Nonvolatile memory can be read only memory (ROM), and volatile memory can be random access memory (RAM). The memory 550 described in embodiments of the present application is intended to include any suitable type of memory.
[0057] In some embodiments, the memory 550 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or a subset or superset thereof, which are exemplarily illustrated below.
[0058] The operating system 551 includes systems programs for processing various basic system services and performing hardware-dependent tasks, such as a framework layer, a core library layer, a driver layer, and the like, for implementing various basic services and processing hardware-based tasks.
[0059] The network communication module 552 is used to communicate with other electronic devices via one or more (wired or wireless) network interfaces 520, and exemplary network interfaces 520 include Bluetooth, wireless compatibility authentication (WiFi), universal serial bus (USB), and the like.
[0060] In some embodiments, the object recognition model training apparatus provided by the embodiments of the present application can be implemented in a software manner, and FIG. 3 shows an object recognition model training apparatus 555 stored in the memory 550, which can be software in the form of programs and plug-ins, including the following software modules: an acquisition module 5551, a sample recognition module 5552, and a training module 5553. These modules are logical, and thus can be combined or further split according to the functions implemented. The functions of each module will be described below.
[0061] In some embodiments, the object recognition apparatus provided by the embodiments of the present application can be implemented in a hardware manner. For example, the object recognition apparatus provided by the embodiments of the present application can be a hardware decoding processor programmed to execute the object recognition method provided by the embodiments of the present application. For example, the hardware decoding processor can be one or more application specific integrated circuits (ASICs), DSPs, programmable logic devices (PLDs), complex programmable logic devices (CPLDs), field-programmable gate arrays (FPGAs), or other electronic elements.
[0062] In some embodiments, the terminal or server can implement the object recognition method provided by the embodiments of the present application by running a computer program or computer executable instructions. For example, the computer program can be a native program (e.g., a dedicated object recognition program) in an operating system or a software module, for example, an object recognition module embedded in any program (e.g., an instant messaging client, a photo album program, an electronic map client, a navigation client), or a native (Native) application (APP), i.e., a program that needs to be installed in an operating system to run. In summary, the computer program can be any form of application program, module, or plug-in.
[0063] The object recognition method provided by the embodiments of the present application will be described in combination with exemplary applications and implementations of the server or terminal provided by the embodiments of the present application.
[0064] Referring to FIG. 4, which is a flowchart of the object recognition method provided by the embodiments of the present application, the steps 101 to 103 shown in FIG. 4 will be described. The object recognition method provided by the embodiments of the present application can be implemented by a server or a terminal alone or by a server and a terminal in cooperation. In the following, the implementation by a terminal alone will be described as an example.
[0065] In step 101, a sound signal is sent in a time period, which includes a first time period and a second time period. The sound signal in the first time period is a sound signal in a first frequency band, and the sound signal in the second time period is a sound signal in a second frequency band.
[0066] In some embodiments, the object to be identified can remain stationary in the first period, and can also perform a prescribed action in the second period. In another embodiment, the object to be identified can also perform a prescribed action in the first period, and can also remain stationary in the second period.
[0067] In some embodiments, the above-mentioned sending of the sound signal in a certain period can be achieved by sending a first frequency band sound signal to the object to be identified in the first period, and sending a second frequency band sound signal to the object to be identified in the second period, the first sound signal and the second sound signal corresponding to different signal parameters. The signal parameters include but are not limited to frequency and amplitude, for example, the frequency of the first sound signal can be different from the frequency of the second sound signal, and the amplitude of the first sound signal can be different from the amplitude of the second sound signal.
[0068] In some embodiments, the above-mentioned sound signal is mainly designed from the following aspects: signal waveform, frequency, and duration and interval, which are described below.
[0069] In some embodiments, since the distances from different facial regions to the device are different, the first feature of the face depth can be obtained from the echo signal of the face by using the characteristics of the sound signal sensitive to the distance. In addition, the frequency of the sound signal will change due to the motion of the observer and the sound source, resulting in a Doppler shift, so the action information of the face can be obtained by detecting the Doppler shift in the echo signal.
[0070] In some embodiments, in order to detect the first feature of the face depth, the present application utilizes a frequency-modulated continuous wave (FMCW) signal widely used in distance measurement as the above-mentioned first frequency band sound signal:
[0071] wherein e s represents the first frequency band sound signal, A represents the amplitude of the first frequency band sound signal, f l represents the starting frequency of the first frequency band sound signal, c represents the frequency to be increased for each sample, i represents the sample amount, and sp represents the sampling rate (44.1KHz).
[0072] In some embodiments, in order to detect the second feature of the face, the present application utilizes a second sound signal with a constant frequency that is more likely to exhibit a Doppler shift: e D = 2Asin(2πft) (2)
[0073] wherein e DThe sound signal for representing the second frequency band, f represents the signal frequency of the sound signal for representing the second frequency band, A represents the amplitude of the sound signal for representing the second frequency band, and t represents time.
[0074] In some embodiments, for the frequency of the sound signal, since the upper limit of the audible frequency of an adult is about 15-17KHz, the frequency range of the sound signal e s for detecting the first feature is set to 16-22KHz, and the frequency of the sound signal e D for detecting the second feature is set to 20KHz.
[0075] In some embodiments, in the process of sending the sound signal within a certain period, for the setting of the duration and interval of the above period, according to the investigation, the distance between the user's face and the device is about 25-50cm, and the corresponding time delay is 1.4-2.8ms. Since the first feature extraction needs to locate the position of the face in the signal, the duration of the detection signal e s of the first feature is set to 60 samples (about 1ms, the sampling rate sp is 44.1KHz), so as to minimize the overlap between the face echo signal and the next transmitted signal; in addition, in order to minimize the interference of the echo signal from other objects at a distance on the next transmitted signal, the interval is set to 1042 samples (about 24ms) in the embodiments of the present application, corresponding to a distance range of 408cm; for the extraction of the first feature, a total of N continuous wave signals are transmitted and collected, and the period T is 1102 samples (1042+60=1102, about 25ms), and the total duration is about t s =N*0.025 seconds. For the face signal e D for detecting the second feature, the signal duration is t d without interval, because the continuous signal can capture the more continuous second feature.
[0076] In some embodiments, the first prompt information is output within the first period, and the sound signal for prompting is a sound signal of a first frequency band.
[0077] In some embodiments, the first prompt information for prompting the to-be-identified object to keep still is output within the first period, and the sound signal of the first frequency band is sent to the to-be-identified object within the first period, so that the to-be-identified object receives and reflects the first sound signal in a still state.
[0078] As an example, the above output of the first prompt information within the first period can be playing the first prompt voice within the first period: please keep still.
[0079] In some embodiments, the second prompting information is outputted in the second time period, and the second prompting information is used to prompt the to-be-identified object to perform the specified action.
[0080] In some embodiments, the second prompting information is outputted in the second time period, and the second prompting information is used to prompt the to-be-identified object to perform the specified action.
[0081] As an example, the second prompting information outputted in the second time period can be a second prompting voice: please turn your head to the left.
[0082] In this way, the sound signal of the first frequency band is sent to the to-be-identified object in the first time period, and the position and state of the to-be-identified object can be identified by the sound signal of the first frequency band. The sound signal of the first frequency band can be received by a receiver and converted into an electrical signal, which can be used to locate and detect the to-be-identified object. The signal parameters corresponding to the sound signal of the first frequency band and the sound signal of the second frequency band are different, which can make the system have higher reliability and stability. Since the sound signal is dependent on the environment and environmental conditions, if the parameters of different signals are different, the interference and distortion of the signal can be reduced, so that the to-be-identified object can be more accurately identified. The to-be-identified object remains stationary in the first time period, which means that the sound signal of the first frequency band can more accurately locate the position of the to-be-identified object. This is very useful for application scenarios that require precise control and positioning. The to-be-identified object performs the specified action in the second time period, which means that the sound signal can be used to detect and track the action of the to-be-identified object. The sound signal of the second frequency band can be used to detect the motion state of the to-be-identified object, so that the state of the system can be detected and adjusted in real time.
[0083] In step 102, the first echo signal corresponding to the sound signal of the first frequency band and the second echo signal corresponding to the sound signal of the second frequency band are obtained.
[0084] In some embodiments, the echo signal corresponding to the sound signal refers to the echo formed when the sound signal propagates in space and is reflected back by an obstacle. The echo signal can be received by a receiver and converted into an electrical signal, which can be used to locate and identify the source and direction of the sound signal. In a sound recognition and positioning system, the echo signal can provide additional information to help the system better identify and locate the sound signal.
[0085] In some embodiments, a sound signal is sent to the object to be identified within a specific time period, and then an echo signal of the sound signal is obtained through a receiver. The echo signal can provide position and direction information of the object to be identified, which can be used to locate and identify the object to be identified. In many sound identification and positioning applications, the echo signal is an indispensable part, which can help the system to more accurately identify and locate the sound signal.
[0086] In some embodiments, the step 102 can be implemented by receiving an initial echo signal corresponding to the sound signal in the time period, and obtaining a target signal frequency of the sound signal in the time period; performing denoising processing on the initial echo signal based on the target signal frequency; dividing the denoised echo signal based on the first time period and the second time period to obtain the first echo signal and the second echo signal.
[0087] In some embodiments, the initial echo signal corresponding to the sound signal in the time period can be implemented by determining the distance between the sound signal emitting device and the object to be identified; and intercepting the initial echo signal from the received audio signal based on the time period and the distance. The sound signal emitting device can be any terminal device for object authentication of the object to be identified, for example, the sound signal emitting device can be a terminal held by the object to be identified, or the sound signal emitting device can be a terminal held by another user.
[0088] In some embodiments, the initial echo signal can be intercepted from the received audio signal based on the distance and the propagation speed of the sound signal, the propagation time can be calculated; the receiving time of the initial echo signal can be determined according to the time period and the propagation time; and the audio signal corresponding to the receiving time can be intercepted as the initial echo signal. The calculation formula of the propagation time can be t=2s / v, where t represents the propagation time, s represents the distance, and v represents the propagation speed. Each time point in the receiving time is the sum of each time point in the time period and the propagation time. For example, the time period is 10:00:00-10:03:00, and the propagation time is 1 second, then the receiving time is 10:00:01-10:03:01.
[0089] The embodiment can reasonably estimate the receiving time by combining the distance and the time period, and then quickly intercept the initial echo signal from the audio signal through the receiving time.
[0090] In some embodiments, the initial echo signal can be denoised based on the target signal frequency by deleting the signals in the initial echo signal that are not equal to the target signal frequency to obtain the denoised echo signal.
[0091] In some embodiments, the initial echo signal corresponding to the sound signal received by the signal receiver can be filtered by a high-pass filter to remove environmental noise below 16KHz, since the frequency of environmental background noise is usually below 12KHz, while the designed echo signal frequency is above 16KHz, in order to effectively eliminate the interference caused by background noise.
[0092] Thus, receiving the initial echo signal corresponding to the sound signal and obtaining the target signal frequency of the sound signal means that after receiving the echo signal, the frequency of the echo signal is analyzed to determine the frequency of the sound signal of the object to be identified. Based on the target signal frequency, the initial echo signal is denoised to obtain the denoised echo signal, which means that the initial echo signal is filtered or denoised according to the target signal frequency to eliminate or reduce noise interference and obtain a purer echo signal, thereby improving the quality and accuracy of the echo signal and more accurately identifying and locating the sound signal.
[0093] In some embodiments, the first echo signal corresponding to the sound signal of the first frequency band corresponds to the object to be identified in a stationary state, and the second echo signal corresponding to the sound signal of the second frequency band corresponds to the object to be identified in a moving state.
[0094] In some embodiments, the denoised echo signal is signal-divided based on the first time period and the second time period to obtain the first echo signal and the second echo signal, which means that the echo signal is divided into two parts according to the time period of the sound signal transmission: the echo signal of the first time period is the first echo signal, and the echo signal of the second time period is the second echo signal. The first echo signal represents the object to be identified remaining stationary during the sound signal transmission, while the second echo signal represents the moving state of the object to be identified after the sound signal transmission. This signal division can help further analyze the state and position of the object to be identified, thereby more accurately identifying and locating the sound signal.
[0095] In some embodiments, the above signal division of the denoised echo signal based on the first time period and the second time period to obtain the first echo signal and the second echo signal can be achieved by the following way: obtaining a first time period ratio of the first time period to the time period and a second time period ratio of the second time period to the time period; determining a first signal duration of the first echo signal based on the first time period ratio and a signal duration of the denoised echo signal; determining a second signal duration of the second echo signal based on the second time period ratio and the signal duration of the denoised echo signal; and signal-dividing the denoised echo signal based on the first signal duration and the second signal duration to obtain the first echo signal and the second echo signal.
[0096] In some embodiments, the determination of the first signal duration of the first echo signal based on the first time period ratio and the signal duration of the de-noised echo signal can be achieved by multiplying the first time period ratio and the signal duration of the de-noised echo signal to obtain the first signal duration of the first echo signal.
[0097] In some embodiments, the determination of the second signal duration of the second echo signal based on the second time period ratio and the signal duration of the de-noised echo signal can be achieved by multiplying the second time period ratio and the signal duration of the de-noised echo signal to obtain the second signal duration of the second echo signal.
[0098] In some embodiments, the first time period ratio and the second time period ratio are obtained by analyzing and measuring the de-noised echo signal to determine the proportion of different time periods in the de-noised echo signal. Based on the first time period ratio and the signal duration of the de-noised echo signal, the first signal duration of the first echo signal, i.e., the length of time the object to be identified remains stationary in the first time period, can be determined. Based on the second time period ratio and the signal duration of the de-noised echo signal, the second signal duration of the second echo signal, i.e., the length of time the object to be identified performs the specified action in the second time period, can be determined. Finally, according to the first signal duration and the second signal duration, the de-noised echo signal can be signal-divided to obtain the first echo signal and the second echo signal, thereby more accurately identifying and locating the sound signal.
[0099] As an example, the expression of the first time period ratio can be:
[0100] wherein T1 represents the first time period ratio, t1 represents the first time period, and T represents the time period.
[0101] As an example, the expression of the second time period ratio can be:
[0102] wherein T2 represents the second time period ratio, t2 represents the second time period, and T represents the time period.
[0103] Thus, the ratio of the first time period to the time period and the ratio of the second time period to the time period are obtained by analyzing and measuring the de-noised echo signal, and the proportions of different time periods in the de-noised echo signal are obtained. Based on the first time period ratio and the signal length of the de-noised echo signal, the first signal length of the first echo signal can be determined, that is, the length of time that the object to be identified remains stationary in the first time period. Based on the second time period ratio and the signal length of the de-noised echo signal, the second signal length of the second echo signal can be determined, that is, the length of time that the object to be identified performs the specified action in the second time period. Finally, according to the first signal length and the second signal length, the de-noised echo signal can be signal-divided to obtain the first echo signal and the second echo signal, so as to more accurately identify and locate the sound signal.
[0104] In some embodiments, the above-mentioned signal division of the de-noised echo signal based on the first signal length and the second signal length to obtain the first echo signal and the second echo signal can be realized by: obtaining a time period relationship between the first time period and the second time period, the time period relationship being used to indicate whether the first time period is before the second time period; and based on the time period relationship, the first signal length and the second signal length, signal-dividing the de-noised echo signal to obtain the first echo signal and the second echo signal.
[0105] In some embodiments, obtaining the time period relationship between the first time period and the second time period, that is, the order between the two, can be used to determine whether the first time period is before the second time period. Based on this time period relationship, the first signal length and the second signal length, the de-noised echo signal can be signal-divided to obtain the first echo signal and the second echo signal. If the first time period is before the second time period, the first echo signal will contain the echo signal in the first time period, and the second echo signal will contain the echo signal in the second time period. This signal division can help further analyze the state and action of the object to be identified, so as to more accurately identify and locate the sound signal.
[0106] In some embodiments, the signal division of the de-noised echo signal based on the time period relationship, the first signal duration and the second signal duration to obtain the first echo signal and the second echo signal can be realized in the following manner: if the time period relationship indicates that the first time period is before the second time period, the sub-echo signal between the start time and the first time in the de-noised echo signal is determined as the first echo signal, and the sub-echo signal between the first time and the end time in the de-noised echo signal is determined as the second echo signal; if the time period relationship indicates that the second time period is before the first time period, the sub-echo signal between the start time and the second time in the de-noised echo signal is determined as the second echo signal, and the sub-echo signal between the second time and the end time in the de-noised echo signal is determined as the first echo signal.
[0107] In some embodiments, the time period length from the start time to the first time is equal to the time period length of the first time period, and the time period length from the start time to the second time is equal to the time period length of the second time period.
[0108] In some embodiments, if the time period relationship indicates that the first time period is before the second time period, the sub-echo signal between the start time and the first time in the de-noised echo signal is determined as the first echo signal, because this signal represents that the to-be-identified object remains stationary in the first time period. The sub-echo signal between the first time and the end time in the de-noised echo signal is determined as the second echo signal, because this signal represents that the to-be-identified object performs a specified action in the second time period. If the time period relationship indicates that the second time period is before the first time period, the sub-echo signal between the start time and the second time in the de-noised echo signal is determined as the second echo signal, because this signal represents that the to-be-identified object has started to move at the beginning of the second time period. The sub-echo signal between the second time and the end time in the de-noised echo signal is determined as the first echo signal, because this signal represents that the to-be-identified object has ended the stationary state in the first time period, so that the signal division can help to further understand the state and action of the to-be-identified object in different time periods.
[0109] Therefore, if the first time period is before the second time period, the sub-echo signal between the start time and the first time in the de-noised echo signal is determined as the first echo signal, and the sub-echo signal between the first time and the end time in the de-noised echo signal is determined as the second echo signal. In this way, more detailed information can be provided to describe the state of the to-be-identified object in the first time period. Specifically, if the first time period is before the second time period, the second echo signal provides the motion information of the to-be-identified object in the second time period, which is helpful for real-time monitoring and adjusting the state of the system. The first echo signal provides the position information of the to-be-identified object in the first time period, which is helpful for more accurately identifying and positioning the sound signal. Therefore, the state and action of the to-be-identified object in different time periods can be comprehensively understood, and the reliability and accuracy of the system are improved.
[0110] In some embodiments, if the time period relationship indicates that the second time period is before the first time period, the sub-echo signal between the start time and the first time in the de-noised echo signal is determined as the first echo signal, and the sub-echo signal between the first time and the end time in the de-noised echo signal is determined as the second echo signal; if the time period relationship indicates that the first time period is before the first time period, the sub-echo signal between the start time and the first time in the de-noised echo signal is determined as the second echo signal, and the sub-echo signal between the first time and the end time in the de-noised echo signal is determined as the first echo signal.
[0111] In some embodiments, the length of the time period from the start time to the second time is equal to the length of the first time period, and the length of the time period from the start time to the first time is equal to the length of the second time period.
[0112] In some embodiments, the second echo signal can be a sub-echo signal between the start time and the first time (including the first time) in the de-noised echo signal, and the first echo signal can be a sub-echo signal between the first time (not including the first time) and the end time in the de-noised echo signal. The second echo signal can be a sub-echo signal between the start time and the first time (not including the first time) in the de-noised echo signal, and the first echo signal can be a sub-echo signal between the first time (including the first time) and the end time in the de-noised echo signal. Whether the first echo signal and the second echo signal include the echo signal of the first time does not constitute a limitation on the embodiments of the present application.
[0113] In some embodiments, if the second time period is before the first time period, the sub-echo signal between the start time and the first time in the de-noised echo signal is determined as the second echo signal, because this part of the signal represents that the to-be-identified object has started to move at the beginning of the second time period. The sub-echo signal between the first time and the end time in the de-noised echo signal is determined as the first echo signal, because this part of the signal represents that the to-be-identified object remains stationary in the first time period. Similarly, if the first time period is before the second time period, the sub-echo signal between the start time and the first time in the de-noised echo signal is determined as the first echo signal, because this part of the signal represents that the to-be-identified object ends the stationary state in the first time period. These information helps to further understand the state and action of the to-be-identified object in different time periods.
[0114] In this way, if the second time period is before the first time period, the sub-echo signal between the start time and the first time in the de-noised echo signal is determined as the second echo signal, and the sub-echo signal between the first time and the end time is determined as the first echo signal, which can better reflect the action state of the to-be-identified object in the second time period. The second echo signal can provide motion information of the to-be-identified object in the second time period, and the first echo signal can provide position information of the to-be-identified object in the first time period. In this way, the state and action of the to-be-identified object in different time periods can be more comprehensively understood, thereby improving the reliability and accuracy of the system. Similarly, if the first time period is before the second time period, the sub-echo signal between the start time and the first time in the de-noised echo signal is determined as the first echo signal, and the sub-echo signal between the first time and the end time is determined as the second echo signal, which can better reflect the stationary state of the to-be-identified object in the first time period and the action state of the to-be-identified object in the second time period. In this way, the sound signal can be more accurately identified and positioned, thereby improving the efficiency and precision of the system.
[0115] In other embodiments, the signal division of the echo signal corresponding to the sound signal in the time period to obtain the first echo signal and the second echo signal can be achieved by the following manner: obtaining a first preset text corresponding to the sound signal in the first frequency band, and obtaining a second preset text corresponding to the sound signal in the second frequency band; performing speech recognition on the echo signal to obtain a text recognition result; determining the echo signal with the text recognition result as the first preset text as the first echo signal, and determining the echo signal with the text recognition result as the second preset text as the second echo signal. The first preset text is used to indicate the text for the to-be-identified object to remain stationary, for example, the first preset text can be set as: please remain stationary. The second preset text is used to prompt the to-be-identified object to perform a specified action, for example, the second preset text can be set as: please turn your head to the left.
[0116] In the embodiment, the first echo signal and the second echo signal are extracted from the echo signal based on the relationship between the first preset text and the text recognition result and the relationship between the second preset text and the text recognition result, so that the first echo signal and the second echo signal do not include noise, and the accuracy of the first echo signal and the second echo signal is improved.
[0117] In step 103, an object recognition result is obtained based on the first echo signal and the second echo signal.
[0118] In some embodiments, the object recognition result is used to indicate whether the to-be-recognized object and the authentication object are the same object, that is, the object recognition result is used to indicate whether the to-be-recognized object passes the object authentication. If the to-be-recognized object and the authentication object are the same object, the to-be-recognized object passes the object authentication. If the to-be-recognized object and the authentication object are not the same object, the to-be-recognized object does not pass the object authentication.
[0119] In some embodiments, step 103 can be implemented in the following manner: the first echo signal is subjected to signal processing to obtain first echo information, and the second echo signal is subjected to signal processing to obtain second echo information; and the object recognition result is obtained based on the first echo information and the second echo information.
[0120] In some embodiments, the object recognition result is used to indicate whether the to-be-recognized object and the authentication object are the same object.
[0121] In some embodiments, the first echo signal is subjected to signal processing to obtain first echo information, which can be implemented in the following manner: the first echo signal is subjected to signal purification to obtain a first purified signal, and the first purified signal is subjected to time-domain signal processing to obtain first time-domain information of the first purified signal; the first purified signal is subjected to frequency-domain signal processing to obtain first frequency-domain information of the first purified signal; and the first time-domain information and the first frequency-domain information are subjected to information fusion to obtain the first echo information.
[0122] In some embodiments, the first echo signal is subjected to signal purification to obtain a first purified signal, which means that the first echo signal r s The signal is segmented into N segments in order to better locate the face echo. Specifically, a peak detection algorithm is used to find the peak p1 of the first wave, and then all the peaks P = {p1, …, pn} of the waves are found according to the period of the signal. N i = p1+ (i-1) *T, and the position of the first 30 samples of the peak is taken as the starting point of each segment, because there is still a segment of the signal with values before the peak of each segment. Direct transmission removal, in order to further accurately locate the face echo, it is also necessary to preliminarily confirm the position of the direct transmission and remove it. Specifically, first, the Hilbert transform function is used to calculate the envelope curve of each segment, and then the left and right troughs of the first peak of the envelope curve are calculated as the starting position and end point of the direct transmission signal, after locating the direct transmission signal, it can be removed.
[0123] In some embodiments, by face echo positioning, generally speaking, the position of the first peak after direct transmission is the position of the face echo, in order to enhance robustness, a face echo position is calculated for all segments of a signal, and the final face echo position is equal to the average of the face echo positions of all segments. After obtaining the face echo position, 60 sample signals are intercepted as the face echo signal.
[0124] In some embodiments, the above-mentioned time domain signal processing of the first purified signal to obtain the first time domain information of the first purified signal refers to obtaining time-frequency characteristics by short-time Fourier transform. The FFT characteristics are usually the frequency domain representation of the global range of the face echo signal, ignoring the local frequency spectrum information in the time domain of the face signal. In order to further improve the accuracy, the time-frequency characteristics F s2 of the face signal are obtained by using short-time Fourier transform, which supplements the local frequency spectrum information in the time domain of the face signal. The finally obtained time-frequency characteristics F s2 have a dimension of 2*30*60, wherein the first dimension also represents two values of amplitude and phase, the second dimension represents time resolution, and the third dimension is frequency resolution.
[0125] In some embodiments, the above-mentioned frequency domain signal processing of the first purified signal to obtain the first frequency domain information of the first purified signal can obtain the first frequency domain information by fast Fourier transform. After obtaining the face echo signal, it is converted into the FFT feature representation in the frequency domain by using the fast Fourier transform algorithm (FFT). Since a total of 30 segments are included, the finally obtained FFT feature representation F s1 has a dimension of 2*30*60, wherein the first dimension represents two values of amplitude and phase, the second dimension represents the face echo segment, and the third dimension is the feature length of each segment.
[0126] Thus, the first echo signal is signal purified to obtain a first purified signal, which can eliminate noise interference and improve the quality of the first echo signal. Then, the first purified signal is processed in the time domain to obtain first time domain information, such as position information of the to-be-identified object in the first time period. At the same time, the first purified signal is processed in the frequency domain to obtain first frequency domain information, such as frequency distribution and energy distribution of the signal. The first time domain information and the first frequency domain information are fused to obtain more comprehensive and accurate static echo information, which is of great significance for further analyzing the state and action of the to-be-identified object.
[0127] In some embodiments, the above-mentioned signal processing of the second echo signal to obtain the second echo information can be achieved by: performing frequency shift signal processing on the second echo signal to obtain second frequency shift information of the second echo signal; performing phase signal processing on the second echo signal to obtain second phase information of the second echo signal; and fusing the second frequency shift information and the second phase information to obtain the second echo information.
[0128] In some embodiments, since the second echo signal r d is a constant frequency signal, in order to extract the signal of the specified frequency (20KHz) and retain the frequency shift brought by the second feature, a bandwidth filter of 19.9-20.1KHz is used to filter the second feature signal. Short-time Fourier transform, this step uses the short-time Fourier algorithm (STFT) to extract the time-frequency distribution of the second signal. The time-frequency diagram can clearly describe the relationship between the signal frequency and time, and reflect the Doppler shift brought by the second feature. Normalization, in order to make the obtained Doppler feature more robust, the time-frequency diagram is normalized in the frequency domain. After normalization, the influence of amplitude difference on the feature can be reduced.
[0129] In some embodiments, the above-mentioned frequency shift signal processing of the second echo signal to obtain the second frequency shift information of the second echo signal can be achieved by Doppler shift feature extraction. In order to make the Doppler shift feature more prominent, the original frequency part is removed and only the frequency change part is retained. The final obtained Doppler shift feature representation F d1 has a dimension of 100*67, where 100 represents that the signal is divided into 100 segments, representing the time resolution, and 67 represents the length of the FFT feature of each segment, also representing the frequency resolution.
[0130] In some embodiments, the phase signal processing on the second echo signal to obtain the second phase information of the second echo signal can be implemented by phase change feature extraction. In addition to the Doppler shift feature, a second feature can also cause a phase change in the signal. Assuming that φ(t) represents the phase of the second echo signal at time point t, then Δφ(t) = φ(t) - φ(t-1), and the final phase change feature is represented as F d2 = {Δφ(t1), Δφ(t2), …, Δφ(t n )}.
[0131] In this way, the second echo signal is important information describing the action state of the to-be-identified object in the second time period. The frequency shift signal processing on the second echo signal can obtain the second frequency shift information, which can provide the motion frequency change of the to-be-identified object in the second time period. At the same time, the phase signal processing on the second echo signal can obtain the second phase information, which can reflect the motion phase change of the to-be-identified object. The information fusion of the second frequency shift information and the second phase information can obtain more comprehensive and more accurate second echo information. These information has important significance for further understanding the motion state and action of the to-be-identified object. Through the frequency shift and phase signal processing on the second echo signal, the sound signal can be more accurately identified and positioned, and the efficiency and accuracy of the system are improved.
[0132] In some embodiments, the object recognition is implemented by an object recognition model, and the object recognition model includes a first feature extraction layer, a second feature extraction layer, and a recognition layer. The object recognition result based on the first echo information and the second echo information can be obtained by the following method: calling the first feature extraction layer to perform feature extraction on the first echo information to obtain first echo features; calling the second feature extraction layer to perform feature extraction on the second echo information to obtain second echo features; and calling the recognition layer to perform object recognition on the to-be-identified object based on the first echo features and the second echo features to obtain the object recognition result.
[0133] In other embodiments, if the object recognition result indicates that the to-be-identified object and the authentication object are the same object, a third preset text is displayed. A voice signal of the to-be-identified object based on the third preset text is obtained, and a first timbre feature of the to-be-identified object is extracted from the voice signal. A second timbre feature of the authentication object is obtained from the database. If the similarity between the first timbre feature and the second timbre feature is less than a preset similarity threshold, the object recognition result is updated. If the similarity between the first timbre feature and the second timbre feature is greater than or equal to the preset similarity threshold, the object recognition result is not updated.
[0134] The type of the third preset text can be set according to the needs of the to-be-recognized object, for example, the third preset text can be any Chinese, and the third preset text can also be any English. The updating the object recognition result includes updating the object recognition result to a result indicating that the to-be-recognized object and the authentication object are not the same object.
[0135] In the embodiment, when the object recognition result indicates that the to-be-recognized object and the authentication object are the same object, whether the object recognition result is correct is determined by further comparing whether the similarity of the first timbre feature and the second timbre feature is greater than or equal to a preset similarity threshold. In this way, when the similarity of the first timbre feature and the second timbre feature is less than the preset similarity threshold, it is indicated that the object recognition result is inaccurate. Therefore, by updating the object recognition result, the accuracy of the object recognition result is improved.
[0136] In some other embodiments, if the object recognition result indicates that the to-be-recognized object and the authentication object are not the same object, the object recognition result is not updated.
[0137] As an example, referring to FIG. 6, FIG. 6 is a schematic diagram of an object recognition model according to an embodiment of the present application. The object recognition model shown in FIG. 6 includes a first feature extraction layer 71, a second feature extraction layer 72, and a recognition layer 73. The first feature extraction layer 71 is called to perform feature extraction on the first echo information Fs to obtain a first echo feature. The second feature extraction layer 72 is called to perform feature extraction on the second echo information Fd1 and Fd2 to obtain a second echo feature. The recognition layer 73 is called to perform object recognition on the to-be-recognized object based on the first echo feature Fs and the second echo features Fd1 and Fd2 to obtain an object recognition result.
[0138] In some embodiments, the first feature extraction layer is called to perform feature extraction on the first echo information to obtain a first echo feature. The first echo feature can include the frequency, amplitude, phase, and other characteristics of the first echo signal. Next, the second feature extraction layer is called to perform feature extraction on the second echo information to obtain a second echo feature. The second echo feature can include the frequency change, phase change, signal energy, and other characteristics of the second echo signal. Finally, the recognition layer is called to perform object recognition on the to-be-recognized object based on the first echo feature and the second echo feature to obtain an object recognition result. This process can be implemented through machine learning algorithms such as classification, clustering, deep learning, etc. to achieve accurate recognition results.
[0139] In this way, the first echo feature and the second echo feature respectively provide information of the to-be-identified object in static and dynamic states, including frequency, amplitude, phase, and other characteristics of the signal. After extraction and processing, these characteristics can be used to train and identify the model, so as to better distinguish different types of objects. Meanwhile, object identification using these characteristics can provide more accurate and reliable identification results, improving the efficiency and accuracy of the system.
[0140] Referring to FIG. 5, which is a flowchart of a method for training an object identification model according to an embodiment of the present application, steps 201 to 203 will be described in conjunction with FIG. 5. The method for training an object identification model according to an embodiment of the present application can be implemented by a server or a terminal alone, or by a server and a terminal in cooperation. The following will be described by taking the implementation of the server alone as an example.
[0141] In step 201, an echo signal sample of a sample object is obtained.
[0142] In some embodiments, the echo signal sample includes a first echo signal sample and a second echo signal sample.
[0143] In some embodiments, step 201 can be implemented by sending a sound signal to the sample object and receiving an echo signal sample of the sample object in response to the sound signal.
[0144] In some embodiments, the first echo signal and the second echo signal are obtained by dividing the echo signal sample.
[0145] In step 202, an initial object identification model is called, and an object identification result of the sample object is obtained based on the first echo signal sample and the second echo signal sample.
[0146] In some embodiments, the initial object identification model includes an initial first feature extraction layer, an initial second feature extraction layer, and an initial identification layer. The initial object identification model is called, and an object identification result of the sample object is obtained based on the first echo signal sample and the second echo signal sample. This can be implemented by calling the initial first feature extraction layer to extract features from the first echo signal sample to obtain first echo sample features, calling the initial second feature extraction layer to extract features from the second echo signal sample to obtain second echo sample features, and calling the identification layer to identify the sample object based on the first echo sample features and the second echo sample features to obtain the object identification result of the sample object.
[0147] In step 203, the initial object identification model is trained based on the object identification result of the sample object and the sample label carried by the echo signal sample to obtain an object identification model.
[0148] In some embodiments, the object recognition model is configured to perform object recognition on the object to be recognized based on the first echo signal and the second echo signal of the object to be recognized.
[0149] In some embodiments, the training set D d contains a dataset of n legitimate users as positive samples, and for each user, there is corresponding attack data as negative samples, therefore, where d represents normal data, d ′ represents attack data.
[0150] The cross-entropy error is used as the loss function for model training, and the calculation formula is:
[0151] where N represents the total number of samples, p i represents the probability that sample i is predicted as a positive sample, and y i represents the label of sample i (1 for positive samples and 0 for negative samples).
[0152] In this way, the echo signal samples of the sample object, including the first echo signal samples and the second echo signal samples, provide rich data resources for subsequent object recognition. Secondly, based on the first echo signal samples and the second echo signal samples, the initial object recognition model is called to perform object recognition on the sample object, and the preliminary recognition result can be obtained. These results can be used to verify and improve the model, and improve the accuracy and reliability of the model. Finally, based on the object recognition result of the sample object and the sample label carried by the echo signal sample, the initial object recognition model can be trained to obtain a more accurate and reliable recognition model. This training process can further improve the performance of the model, so that it can better perform object recognition on the object to be recognized.
[0153] In this way, by sending a sound signal to the object to be recognized within a time period, the first echo signal corresponding to the sound signal of the first frequency band and the second echo signal of the second frequency band are obtained, and the object recognition result is obtained based on the first echo signal and the second echo signal. In this way, on the one hand, the sound signal is more difficult to add sensitive information than image signals and other forms of signals, and object recognition is performed through the sound signal, thereby effectively preventing the influence of sensitive information on recognition. On the other hand, based on the first echo signal and the second echo signal participating in the recognition process, compared with single-dimensional recognition, the first echo signal and the second echo signal participating in the recognition process can enrich the recognition dimension of object recognition, thereby realizing double recognition, so that the accuracy of object recognition can be effectively improved, thereby effectively improving the accuracy of object recognition.
[0154] In the following, an exemplary application of the embodiments of the present application in an actual object recognition application scenario will be described.
[0155] Face authentication is a kind of identity verification method based on face biometric features, which confirms the identity of a user by matching the input face image with the pre-registered face template. Compared with traditional password or card verification methods, face authentication has the advantages of unforgeability, convenience and high efficiency, and therefore face authentication systems have been widely used in daily life. However, existing face authentication systems are vulnerable to various spoofing attacks, and it is particularly important to ensure the security of face authentication. Currently, the attack means against face authentication systems are increasingly diversified and sophisticated. Attackers can use various means to deceive face authentication systems, such as using fake faces such as photos / videos, 3D masks, and adversarial samples to impersonate legitimate users and bypass the authentication process. For example, traditional 2D face authentication systems are vulnerable to basic attacks such as image or video attacks. In order to enhance the security of face authentication systems, face anti-spoofing technology has attracted widespread attention.
[0156] In some embodiments, referring to FIG. 7, which is a schematic diagram of the object recognition method provided by the embodiments of the present application. The embodiments of the present application aim to use acoustic signals to realize face anti-spoofing, and obtain the first feature (depth feature) and the second feature (expression, nodding action feature) of the face as clues through acoustic perception, and research and explore an acoustic face liveness detection scheme based on dynamic and static combined features, thereby improving the security of the face authentication system. The technical scheme adopted by the embodiments of the present application can be realized through the steps shown in FIG. 7. Step one: acoustic signal perception, that is, using a loudspeaker to play a pre-designed ultrasonic signal, and using a microphone to collect echo signals containing user face depth information and second features. Step two: acoustic signal preprocessing, that is, performing preprocessing operations such as environmental noise removal, signal synchronization and segmentation on the collected acoustic signals. Step three: first feature extraction, including signal segmentation, removing direct transmission from the loudspeaker to the microphone, locating and extracting face echo signals, using a fast Fourier transform algorithm to obtain a global feature representation, using a short-time Fourier transform algorithm to obtain a local feature representation, and feature combination. Step four: second feature extraction, including using a bandwidth filter to extract signals in a specific frequency range, using a short-time Fourier transform algorithm to obtain a Doppler shift spectrogram of the signals, normalizing and generating a Doppler shift feature representation, and extracting phase change features. Step five: decision module, designing a deep learning network, training a deep learning model using the first feature and the second feature, and using the trained model to judge the legitimacy of the user.
[0157] In some embodiments, the above step one can be implemented by the following way: acoustic signal design, the embodiments of the present application mainly design acoustic signal from three aspects: signal waveform, frequency and duration and interval. Signal waveform, due to the different distance from different facial regions to the device, based on the characteristics of the acoustic signal sensitive to the distance, the first feature about the depth of the face can be obtained from the echo signal of the face. And the frequency of the acoustic signal will change because of the motion of the observer and the sound source, producing Doppler shift, so the action information of the face can be obtained by detecting the Doppler shift in the echo signal. In order to detect the first feature of the face depth, the embodiments of the present application use the frequency modulated continuous wave (FMCW) signal widely used in distance measurement Where f l represents the starting frequency, c represents the frequency to be increased by one sample, and sp is the sampling rate (44.1KHz); and in order to detect the second feature of the face, the embodiments of the present application use the constant frequency acoustic signal e D = 2Asin(2πft). Frequency, since the average frequency of the upper limit of the audible sound of adults is about 15-17KHz, the frequency range of the acoustic signal e s for detecting the first frequency band is set to 16-22KHz, and the frequency of the acoustic signal e D for detecting the second frequency band is set to 20KHz. Duration and interval, according to the investigation, the distance between the user's face and the device is about 25-50cm, and the corresponding time delay is 1.4-2.8ms. Since the extraction of the first feature needs to locate the position of the face in the signal, the duration of the first feature detection signal e s is set to 60 samples (about 1ms, the sampling rate sp is 44.1KHz), the purpose is to minimize the overlap between the face echo signal and the next transmitted signal; in addition, in order to minimize the interference of the echo signal from other objects at a distance from the next transmitted signal, the embodiments of the present application set the interval to 1042 samples (about 24ms), corresponding to a distance range of 408cm; for the extraction of the first feature, a total of N pulse continuous wave signals are transmitted and collected, the period T is 1102 samples (1042+60=1102, about 25ms), and the total duration is about t s =N*0.025 seconds. And for the face signal e D detected in the second period, the signal duration is t d , there is no interval length, because the continuous signal can capture the more continuous second feature. Acoustic signal perception, the designed acoustic signal is played using a loudspeaker, and the microphone is called to collect the echo signal.
[0158] In some embodiments, the above step two can be implemented in the following way: ambient noise removal, according to the investigation, the frequency of the ambient background noise is usually below 12KHz, while the frequency of the designed acoustic signal is above 16KHz; in order to effectively eliminate the interference caused by the background noise, a high-pass filter is used to filter out the ambient noise below 16KHz. Signal synchronization and segmentation, due to the time delay between the echo signal and the transmitted signal, in order to reduce the impact of time delay on the subsequent processing operation, the cross-correlation function is used to align the echo signal and the transmitted signal. Since the designed acoustic signal contains two signals for detecting the first feature and the second feature, according to the duration of the two signals, the echo signal r is segmented into a signal containing the first feature r s and a signal containing the second feature r d .
[0159] In some embodiments, the above step three can be implemented in the following way: signal segmentation, after step two, the first echo signal r s still contains the direct transmission signal from the loudspeaker to the microphone, the face echo signal and the environmental reflection signal. In order to better locate the face echo, the signal is segmented into N segments at this stage. Specifically, the peak value detection algorithm is used to find the peak value p1 of the first wave, and then all the peak values P = {p1, …, pi, …, p N} of all the waves are found according to the period of the signal, where p i = p1+ (i-1)*T, and the position of the first 30 samples of the peak value of the wave is taken as the starting point of each segment, because there is still a segment of signal with value before the peak value of each segment. Direct transmission removal, in order to further accurately locate the face echo, it is also necessary to preliminarily confirm the position of the direct transmission and remove it. Specifically, first, the Hilbert transform function is used to calculate the envelope curve of each segment, and then the left and right troughs of the first peak value of the envelope curve are calculated as the starting position and end point of the direct transmission signal, after locating the direct transmission signal, it can be removed. Face echo positioning, usually, the position of the first peak value after the direct transmission is the position of the face echo, in order to enhance the robustness, a face echo position is calculated for all segments of a signal, and the final face echo position is equal to the average of the face echo positions of all segments. After obtaining the face echo position, 60 sample signals are intercepted as the face echo signal. Fast Fourier transform, after obtaining the face echo signal through face echo positioning, the fast Fourier transform algorithm (FFT) is used to convert it into the FFT feature representation in the frequency domain. Since there are 30 segments in total, the final obtained FFT feature representation F s1The dimension is 2*30*60, where the first dimension represents two values of amplitude and phase, the second dimension represents the face echo segment, and the third dimension is the feature length of each segment. The short-time Fourier transform obtains the time-frequency feature, and the FFT feature is usually a frequency domain representation of the global range of the face echo signal, ignoring the information in the time domain of the face signal. In order to further improve the accuracy, the short-time Fourier transform is used to obtain the time-frequency feature F s2 of the face signal, which supplements the local spectral information in the time domain of the face signal. The finally obtained time-frequency feature F s2 has a dimension of 2*30*60, where the first dimension also represents two values of amplitude and phase, the second dimension represents time resolution, and the third dimension is frequency resolution. Feature combination, obtain the global frequency domain feature F s1 and the time-frequency feature F s2 of the face echo signal, which can be combined into a 4*30*60 final first feature F s . This feature represents a combination of global and local spectral features, contains rich face depth information, and further improves the final performance of the system.
[0160] In some embodiments, the above step four can be implemented in the following way: band-pass filtering, after step two, since the second echo signal r d detected in the second feature is a constant frequency signal, in order to extract the signal of the specified frequency (20KHz) and retain the frequency shift brought by the second feature, a bandwidth filter of 19.9-20.1KHz is used to filter the second feature signal. Short-time Fourier transform, this step uses the short-time Fourier algorithm (STFT) to extract the time-frequency distribution of the second signal, which can clearly describe the relationship between signal frequency and time, and reflect the Doppler shift brought by the second feature. Normalization, in order to make the obtained Doppler feature more robust, the time-frequency graph is normalized in the frequency domain, which can reduce the influence of amplitude difference on the feature after normalization. Doppler shift feature extraction, in order to make the Doppler shift feature more prominent, the original frequency part is removed and only the frequency change part is retained, and the finally obtained Doppler shift feature representation F d1 has a dimension of 100*67, where "100" means that the signal is divided into 100 segments, representing time resolution, and "67" represents the length of the FFT feature of each segment, also representing frequency resolution. Phase change feature extraction, in addition to the Doppler shift feature, the second feature will also cause a change in the phase of the signal, assuming that φ(t) represents the phase of the echo signal at time point t, then Δφ(t) = φ(t)-φ(t-1), and the finally obtained phase change feature is represented as F d2 ={Δφ(t1),Δφ(t2),…,Δφ(t n )}.
[0161] In some embodiments, the above step five can be implemented by designing a model that has the following differences in processing sound signals compared to processing image signals: (1) images are usually three-dimensional data, while audio signals are usually two-dimensional, so a three-dimensional feature representation symbol needs to be extracted from the audio signal before inputting into the CNN model, for example, the first extracted feature F s ; (2) image signals usually do not have timing, while audio signals are data with timing, so the timing information needs to be paid attention to, and therefore an LSTM model (one of the commonly used models in signal processing) is introduced to learn the timing information in the dynamic feature F D ; (3) image data usually needs a large parameter amount model for processing, while signal data requires a model with less parameter amount, reducing the calculation cost.
[0162] Based on the above considerations, the first extracted facial feature F s , the second extracted feature F d1 , and F d2 are used as inputs of the model. Referring to FIG. 6, the first feature F s is input into a convolutional neural network (CNN), and sequentially passes through a convolutional layer CONV1, a pooling layer POOLING1, a convolutional layer CONV2, a pooling layer POOLING2, a convolutional layer CONV3, a pooling layer POOLING3, a convolutional layer CONV4, and a fully connected layer FC1, and finally a 256-dimensional feature vector fea s is obtained. In order to learn the timing feature in the second feature, F d1 is input into a long short-term memory network (LSTM), the size of the hidden layer is set to 512, the output of the LSTM is input into a fully connected layer FC2, and a 512-dimensional feature vector fea d1 is obtained. Meanwhile, F d2 is input into a long short-term memory network (LSTM), the size of the hidden layer is set to 128, the output of the LSTM is input into a fully connected layer FC3, and a 128-dimensional feature vector fea d2 is obtained. The feature vectors fea d1 and fea d2 are connected to form a second feature vector fea d , which is input into a fully connected layer FC4. Finally, the feature vectors fea s and fea d are connected and input into a fully connected layer FC5, and finally a Softmax layer determines which category the current feature belongs to.
[0163] In some embodiments, the training set D dThe dataset containing n legitimate users is taken as positive samples, and for each user, there is corresponding attack data as negative samples, thus, where d represents normal data, d ′ represents attack data.
[0164] The cross-entropy error is used as the loss function for model training, and the calculation formula is:
[0165] where N represents the total number of samples, p i represents the probability that sample i is predicted as a positive sample, and y i represents the label of sample i (1 for a positive sample and 0 for a negative sample).
[0166] In some embodiments, the test set D t contains m users different from D d , and similar to the training set, for each user, there is corresponding attack data as negative samples, thus After the model training is completed, the dataset D t of the user who does not participate in the training is tested, and for each sample, if the model outputs 1, it means that the sample is a legitimate user, and if the model outputs 0, it means that the sample is an attacker.
[0167] In some embodiments, referring to FIG. 8, which is a schematic diagram of the principle of sound signal processing provided in an embodiment of the present application, a high-pass filter is performed on a sound signal r to obtain a filtered signal rs, the filtered signal rs is segmented, then direct transmission removal is performed, and then face echo positioning is performed, and then a fast Fourier transform is performed to obtain a result Fs. The sound signal r is segmented to obtain a segmentation result rd, then bandwidth filtering is performed, then a short-time Fourier transform is performed, then normalization is performed, and then Doppler feature extraction is performed to obtain a result Fd.
[0168] In this way, the embodiments of the present application are mainly used in the field of face anti-spoofing. Existing face anti-spoofing technologies are mostly based on vision, but these schemes have poor privacy (may need to collect additional visual data containing faces), poor robustness (easily affected by environmental light), and rely on additional hardware (depth camera). The existing face anti-spoofing technology based on acoustics only uses the depth information of the face and can only resist 2D face presentation attacks, without considering the threat of 3D face presentation attacks. In order to solve the above shortcomings, the embodiments of the present application use the first feature and the second feature of the face perceived by the acoustic signal to perform face anti-spoofing, without collecting sensitive visual face data, robust to environmental noise and light, without additional hardware, and can resist 2D and 3D face presentation attacks. It is a robust, practical, and highly secure face anti-spoofing technology.
[0169] Thus, the embodiment of the present application utilizes an acoustic signal to perceive a face to implement living body detection. Compared with a scheme based on vision, the embodiment of the present application does not need to collect visual data (for example, a face image, etc.) containing sensitive information, thereby reducing the concern of a user for personal privacy. The embodiment of the present application relies on a microphone and a loudspeaker, which are commonly used in traditional devices, and does not need an additional sensor (for example, a depth camera, etc.), thereby being high in practicability and popularization. The embodiment of the present application is high in security. By combining a first feature (a depth feature) and a second feature (an expression, a nod, etc.) of a face, the embodiment of the present application can resist current mainstream face display attacks, including 2D display attacks and 3D display attacks, while other existing face living body detection schemes based on acoustics cannot resist 3D display attacks. The embodiment of the present application uses ultrasonic waves to perceive a face, and is not affected by environmental light and environmental noise, thereby being high in robustness.
[0170] It can be understood that, in the embodiment of the present application, data related to a sound signal is involved. When the embodiment of the present application is applied to a specific product or technology, permission or consent of a user needs to be obtained, and collection, use and processing of related data need to comply with relevant laws, regulations and standards in a country or region.
[0171] The following continues to describe an example structure of the object recognition apparatus 455 provided by the embodiment of the present application, which is implemented as a software module. In some embodiments, as shown in FIG. 2, the software module stored in the object recognition apparatus 455 in the memory 450 can include: a sending module 4551, configured to send a sound signal in a certain period of time, the target period of time including a first period of time and a second period of time; wherein the sound signal in the first period of time is a sound signal of a first frequency band, and the sound signal in the second period of time is a sound signal of a second frequency band; a division module 4552, configured to obtain a first echo signal corresponding to the sound signal of the first frequency band and a second echo signal corresponding to the sound signal of the second frequency band; and an object recognition module 4553, configured to obtain an object recognition result based on the first echo signal and the second echo signal.
[0172] In some embodiments, the object recognition apparatus 455 described above further includes: an output module, further configured to output first prompt information in the first period of time, the first prompt information being used to prompt that the sound signal is a sound signal of a first frequency band; and output second prompt information in the second period of time, the second prompt information being used to prompt the to-be-recognized object to perform the specified action.
[0173] In some embodiments, the division module 4552 is further configured to receive an initial echo signal corresponding to the sound signal in the time period, and obtain a target signal frequency of the sound signal in the time period; perform de-noising processing on the initial echo signal based on the target signal frequency; and divide the de-noising processed echo signal based on the first time period and the second time period to obtain the first echo signal and the second echo signal.
[0174] In some embodiments, the division module 4552 is further configured to determine a distance between a device that emits the sound signal and the object to be identified, and cut the initial echo signal from the received audio signal based on the time period and the distance.
[0175] In some embodiments, the division module 4552 is further configured to calculate a propagation time based on the distance and a propagation speed of the sound signal, determine a receiving time of the initial echo signal according to the time period and the propagation time, and cut the audio signal corresponding to the receiving time as the initial echo signal.
[0176] In some embodiments, the division module 4552 is further configured to obtain a first time period ratio of the first time period to the time period, and a second time period ratio of the second time period to the time period, determine a first signal duration of the first echo signal based on the first time period ratio and a signal duration of the de-noising processed echo signal, determine a second signal duration of the second echo signal based on the second time period ratio and the signal duration of the de-noising processed echo signal, and perform signal division on the de-noising processed echo signal based on the first signal duration and the second signal duration to obtain the first echo signal and the second echo signal.
[0177] In some embodiments, the division module 4552 is further configured to obtain a time period relationship between the first time period and the second time period, the time period relationship being used to indicate whether the first time period is before the second time period, and perform signal division on the de-noising processed echo signal based on the time period relationship, the first signal duration and the second signal duration to obtain the first echo signal and the second echo signal.
[0178] In some embodiments, the division module 4552 is further configured to, if the time period relationship indicates that the first time period is before the second time period, determine a sub-echo signal between a start time and a first time in the denoised echo signal as the first echo signal, and determine a sub-echo signal between the first time and an end time in the denoised echo signal as the second echo signal; if the time period relationship indicates that the second time period is before the first time period, determine a sub-echo signal between the start time and a second time in the denoised echo signal as the second echo signal, and determine a sub-echo signal between the second time and the end time in the denoised echo signal as the first echo signal; wherein a time period length from the start time to the first time is equal to a time period length of the first time period, and a time period length from the start time to the second time is equal to a time period length of the second time period.
[0179] In some embodiments, the division module 4552 is further configured to obtain a first preset text corresponding to a sound signal of the first frequency band, and obtain a second preset text corresponding to a sound signal of the second frequency band; perform speech recognition on the echo signal in the time period to obtain a text recognition result; determine an echo signal with the text recognition result being the first preset text as the first echo signal, and determine an echo signal with the text recognition result being the second preset text as the second echo signal.
[0180] In some embodiments, the division module 4552 is further configured to perform signal processing on the first echo signal to obtain first echo information, and perform signal processing on the second echo signal to obtain second echo information; and obtain the object recognition result based on the first echo information and the second echo information; wherein the object recognition result is used to indicate whether the to-be-identified object and the authentication object are the same object.
[0181] In some embodiments, the division module 4552 is further configured to perform signal purification on the first echo signal to obtain a first purified signal, and perform time domain signal processing on the first purified signal to obtain first time domain information of the first purified signal; perform frequency domain signal processing on the first purified signal to obtain first frequency domain information of the first purified signal; and perform information fusion on the first time domain information and the first frequency domain information to obtain the first echo information.
[0182] In some embodiments, the division module 4552 is further configured to perform frequency shift signal processing on the second echo signal to obtain second frequency shift information of the second echo signal, perform phase signal processing on the second echo signal to obtain second phase information of the second echo signal, and fuse the second frequency shift information and the second phase information to obtain the second echo information.
[0183] In some embodiments, the object recognition is implemented by an object recognition model including a first feature extraction layer, a second feature extraction layer, and a recognition layer. The object recognition module 4553 is further configured to call the first feature extraction layer to perform feature extraction on the first echo information to obtain first echo features, call the second feature extraction layer to perform feature extraction on the second echo information to obtain second echo features, and call the recognition layer to obtain the object recognition result based on the first echo features and the second echo features.
[0184] In some embodiments, the object recognition module 4553 is further configured to display a third preset text if the object recognition result indicates that the to-be-recognized object and the authentication object are the same object, obtain a voice signal of the to-be-recognized object based on the third preset text, extract first vocal characteristics of the to-be-recognized object from the voice signal, obtain second vocal characteristics of the authentication object from a database, and update the object recognition result if a similarity between the first vocal characteristics and the second vocal characteristics is less than a preset similarity threshold.
[0185] The following continues to illustrate an exemplary structure of the object recognition apparatus 555 provided by the embodiments of the present application, which is implemented as a software module. In some embodiments, as shown in FIG. 3, the software module stored in the training apparatus 555 of the object recognition model of the memory 550 can include: an obtaining module 5551 configured to obtain echo signal samples of sample objects, the echo signal samples including first echo signal samples and second echo signal samples; a sample recognition module 5552 configured to call an initial object recognition model to obtain an object recognition result of the sample objects based on the first echo signal samples and the second echo signal samples; and a training module 5553 configured to train the initial object recognition model based on the object recognition result of the sample objects and sample labels carried by the echo signal samples to obtain an object recognition model, wherein the object recognition model is configured to obtain an object recognition result based on first echo signals and second echo signals of a to-be-recognized object.
[0186] The embodiment of the present application provides a computer program product, which comprises a computer program or computer executable instructions stored in a computer readable storage medium. The processor of the electronic device reads the computer executable instructions from the computer readable storage medium, and the processor executes the computer executable instructions, so that the electronic device executes the object identification method and the training method of the object identification model provided in the embodiment of the present application.
[0187] The embodiment of the present application provides a computer readable storage medium storing computer executable instructions, wherein the computer executable instructions are stored in the computer readable storage medium. When the computer executable instructions are executed by the processor, the processor will execute the object identification method and the training method of the object identification model provided in the embodiment of the present application, for example, the object identification method shown in FIG. 4.
[0188] In some embodiments, the computer readable storage medium can be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disc, or CD-ROM; and can also be various electronic devices comprising one or any combination of the above memories.
[0189] In some embodiments, the computer executable instructions can be in the form of programs, software, software modules, scripts or codes, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and can be deployed in any form, including being deployed as independent programs or being deployed as modules, components, subroutines or other units suitable for use in a computing environment.
[0190] As an example, the computer executable instructions can but not necessarily correspond to files in a file system, can be stored in a part of a file storing other programs or data, for example, stored in one or more scripts in a hyper text markup language (HTML, Hyper Text Markup Language) document, stored in a single file dedicated to the program in question, or stored in multiple cooperative files (for example, files storing one or more modules, subroutines or code parts).
[0191] As an example, the computer executable instructions can be deployed to execute on one electronic device, or on multiple electronic devices located in one place, or on multiple electronic devices distributed in multiple places and interconnected through a communication network.
[0192] In summary, the embodiment of the present application has the following beneficial effects:
[0193] (1) By sending a sound signal within a certain period of time, the first frequency band sound signal corresponding to the first echo signal and the second frequency band sound signal corresponding to the second echo signal are obtained, and the object recognition result is obtained based on the first echo signal and the second echo signal. In this way, on the one hand, the sound signal is more difficult to add sensitive information than the image signal and other forms of signals, and the object recognition is carried out through the sound signal, thereby effectively preventing the influence of sensitive information on the recognition, on the other hand, based on the first echo signal and the second echo signal participating in the identification process, compared with single dimension identification, the first echo signal and the second echo signal participating in the identification process can enrich the identification dimension of object recognition, thereby realizing double identification, thereby the accuracy of object recognition can be effectively improved, thereby the accuracy of object recognition can be effectively improved.
[0194] (2) The initial echo signal corresponding to the sound signal in the period is received, and the target signal frequency of the sound signal in the period is obtained, which means that after receiving the initial echo signal, the frequency of the sound signal of the object to be identified is determined by analyzing the frequency of the sound signal in the period. Based on the target signal frequency, the initial echo signal is denoised, which means that according to the target signal frequency, the initial echo signal is filtered or denoised to eliminate or reduce noise interference, so as to obtain a more pure echo signal, thereby improving the quality and accuracy of the echo signal, so as to more accurately identify and locate the sound signal.
[0195] (3) The first frequency band sound signal is sent to the object to be identified within the first period, which can be used to identify the position and state of the object to be identified. The sound signal can be received by the receiver and converted into an electrical signal, which can be used to locate and detect the object to be identified. The signal parameters corresponding to the first frequency band sound signal and the second frequency band sound signal are different, which can make the system have higher reliability and stability. Since the sound signal is dependent on the environment and environmental conditions, if the parameters of different signals are different, the interference and distortion of the signal can be reduced, thereby more accurately identifying the object to be identified. The object to be identified remains stationary within the first period, which means that the position of the object to be identified can be more accurately located by the first frequency band sound signal. This is very useful for application scenarios that require precise control and positioning. The object to be identified performs a prescribed action within the second period, which means that the second frequency band sound signal can be used to detect and track the action of the object to be identified. The second frequency band sound signal can be used to detect the motion state of the object to be identified, so that the state of the system can be detected and adjusted in real time.
[0196] (4) The ratio of the first time period to the time period and the ratio of the second time period to the time period are obtained by analyzing and measuring the echo signal to obtain the proportion of different time periods in the echo signal. Based on the first time period ratio and the signal length of the echo signal, the first signal length of the first echo signal can be determined, that is, the length of time that the object to be identified remains stationary in the first time period. Based on the second time period ratio and the signal length of the echo signal, the second signal length of the second echo signal can be determined, that is, the length of time that the object to be identified performs the specified action in the second time period. Finally, according to the first signal length and the second signal length, the echo signal can be signal divided to obtain the first echo signal and the second echo signal, so as to more accurately identify and locate the sound signal.
[0197] (5) When the first time period is before the second time period, the sub-echo signal between the start time and the first time in the echo signal is determined as the first echo signal, and the sub-echo signal between the first time and the end time is determined as the second echo signal. This processing method can provide more detailed information to describe the state of the object to be identified in the first time period. Specifically, when the first time period is before the second time period, the second echo signal provides the motion information of the object to be identified in the second time period, which helps to monitor and adjust the state of the system in real time. The first echo signal provides the position information of the object to be identified in the first time period, which helps to more accurately identify and locate the sound signal. Therefore, this processing method can more comprehensively understand the state and action of the object to be identified in different time periods, thereby improving the reliability and accuracy of the system.
[0198] (6) When the second time period is before the first time period, the sub-echo signal between the start time and the first time in the echo signal is determined as the second echo signal, and the sub-echo signal between the first time and the end time is determined as the first echo signal. This processing method can better reflect the action state of the object to be identified in the second time period. The second echo signal can provide the motion information of the object to be identified in the second time period, and the first echo signal can provide the position information of the object to be identified in the first time period. In this way, the state and action of the object to be identified in different time periods can be more comprehensively understood, thereby improving the reliability and accuracy of the system. Similarly, when the first time period is before the second time period, the sub-echo signal between the start time and the first time in the echo signal is determined as the first echo signal, and the sub-echo signal between the first time and the end time is determined as the second echo signal. This processing method can better reflect the stationary state of the object to be identified in the first time period and the action state of the object to be identified in the second time period. In this way, the sound signal can be more accurately identified and located, and the efficiency and precision of the system are improved.
[0199] (7) The first echo signal is signal purified to obtain a first purified signal. This step can eliminate noise interference and improve the quality of the first echo signal. Then, the first purified signal is processed in the time domain to obtain first time domain information, such as the position information of the object to be identified in the first time period. At the same time, the first purified signal is processed in the frequency domain to obtain first frequency domain information, such as the frequency distribution and energy distribution of the signal. The first time domain information and the first frequency domain information are fused to obtain more comprehensive and accurate first echo information, which is of great significance for further analyzing the state and action of the object to be identified.
[0200] (8) The second echo signal is important information describing the action state of the object to be identified in the second time period. The second echo signal is processed in the frequency shift domain to obtain second frequency shift information, which can provide the motion frequency change of the object to be identified in the second time period. At the same time, the second echo signal is processed in the phase domain to obtain second phase information, which can reflect the motion phase change of the object to be identified. The dynamic frequency shift information and the second phase information are fused to obtain more comprehensive and accurate second echo information. These information are of great significance for further understanding the motion state and action of the object to be identified. By processing the second echo signal in the frequency shift and phase domains, the sound signal can be more accurately identified and located, improving the efficiency and accuracy of the system.
[0201] (9) The first echo feature and the second echo feature respectively provide information of the object to be identified in static and dynamic states, including frequency, amplitude, phase and other features of the signal. These features can be extracted and processed to train and identify the model, so as to better distinguish different types of objects. At the same time, using these features for object identification can provide more accurate and reliable identification results, improving the efficiency and accuracy of the system.
[0202] (10) The embodiments of the present application are mainly used in the field of face anti-spoofing. Existing face anti-spoofing technologies are mostly based on vision, but these solutions have poor privacy (may need to collect additional visual data containing faces), poor robustness (easily affected by environmental light), and rely on additional hardware (depth camera). Existing face anti-spoofing technologies based on acoustics only use the depth information of the face and can only resist 2D face display attacks, without considering the threat of 3D face display attacks. In order to solve the above shortcomings, the embodiments of the present application use the first feature and the second feature of the face perceived by the acoustic signal to perform face anti-spoofing, without collecting sensitive visual face data, robust to environmental noise and light, without additional hardware, and can resist 2D and 3D face display attacks. It is a robust, practical and highly secure face anti-spoofing technology.
[0203] (11) The embodiment of the present application utilizes an acoustic signal to perceive a face to realize living body detection. Compared with a scheme based on vision, the embodiment of the present application does not need to collect visual data (for example, a face image, etc.) containing sensitive information, thereby reducing the worry of users about personal privacy. The sensor relied on by the embodiment of the present application is a microphone and a loudspeaker commonly existing in traditional devices, without the need of an additional sensor (for example, a depth camera, etc.), thereby being high in practicability and popularization. The embodiment of the present application is high in security, and can resist current mainstream face display attacks, including 2D display attacks and 3D display attacks, by combining the use of a first feature (a depth feature) and a second feature (expression, nodding, etc.) of a face. However, other existing face living body detection schemes based on acoustics cannot resist 3D display attacks. The embodiment of the present application adopts ultrasonic waves to perceive a face, and is not affected by environmental light and environmental noise, thereby being high in robustness.
[0204] (12) The echo signal samples of the sample object, including the first echo signal samples and the second echo signal samples, provide rich data resources for subsequent object recognition. Secondly, based on the first echo signal samples and the second echo signal samples, the initial object recognition model is called to perform object recognition on the sample object, and a preliminary recognition result can be obtained. These results can be used to verify and improve the model, and improve the accuracy and reliability of the model. Finally, based on the object recognition result of the sample object and the sample label carried by the echo signal sample, the initial object recognition model can be trained to obtain a more accurate and reliable recognition model. This training process can further improve the performance of the model, so that it can better perform object recognition on the object to be recognized.
[0205] The above merely describes the embodiments of the present application, and is not used to limit the protection scope of the present application. Any modification, equivalent replacement and improvement made within the spirit and scope of the present application shall be included in the protection scope of the present application.
Claims
A method for object recognition, the method comprising: sending a sound signal within a time period, the time period comprising a first time period and a second time period; wherein the sound signal within the first time period is a sound signal of a first frequency band, and the sound signal within the second time period is a sound signal of a second frequency band; obtaining a first echo signal corresponding to the sound signal of the first frequency band and a second echo signal corresponding to the sound signal of the second frequency band; obtaining an object recognition result based on the first echo signal and the second echo signal. The method of claim 1, wherein the obtaining the first echo signal corresponding to the sound signal of the first frequency band and the second echo signal corresponding to the sound signal of the second frequency band comprises: receiving an initial echo signal corresponding to the sound signal within the time period, and obtaining a target signal frequency of the sound signal within the time period; performing denoising processing on the initial echo signal based on the target signal frequency; dividing the denoised echo signal based on the first time period and the second time period to obtain the first echo signal and the second echo signal. The method of claim 2, wherein the receiving the initial echo signal corresponding to the sound signal within the time period comprises: determining a distance between a device emitting the sound signal within the time period and an object to be recognized; cutting the initial echo signal from the received audio signal based on the time period and the distance. The method of claim 3, wherein the cutting the initial echo signal from the received audio signal based on the time period and the distance comprises: calculating a propagation time based on the distance and a propagation speed of the sound signal within the time period; determining a receiving time of the initial echo signal according to the time period and the propagation time; cutting the audio signal corresponding to the receiving time as the initial echo signal. The method of claim 2, wherein the dividing the denoised echo signal based on the first time period and the second time period to obtain the first echo signal and the second echo signal comprises: obtaining a first time period ratio of the first time period to the time period, and a second time period ratio of the second time period to the time period; determining a first signal duration of the first echo signal based on the first time period ratio and a signal duration of the denoised echo signal; determining a second signal duration of the second echo signal based on the second time period ratio and the signal duration of the denoised echo signal; dividing the denoised echo signal based on the first signal duration and the second signal duration to obtain the first echo signal and the second echo signal. The method of claim 5, wherein the dividing the denoised echo signal based on the first signal duration and the second signal duration to obtain the first echo signal and the second echo signal comprises: obtaining a time period relationship between the first time period and the second time period, the time period relationship being used to indicate whether the first time period is before the second time period; The de-noised echo signal is divided based on the time period relationship, the first signal duration and the second signal duration, to obtain the first echo signal and the second echo signal. The method according to claim 6, wherein the de-noised echo signal is divided based on the time period relationship, the first signal duration and the second signal duration, to obtain the first echo signal and the second echo signal, comprises: if the time period relationship indicates that the first time period is before the second time period, a sub-echo signal between the start time and a first time in the de-noised echo signal is determined as the first echo signal, and a sub-echo signal between the first time and the end time in the de-noised echo signal is determined as the second echo signal; if the time period relationship indicates that the second time period is before the first time period, a sub-echo signal between the start time and a second time in the de-noised echo signal is determined as the second echo signal, and a sub-echo signal between the second time and the end time in the de-noised echo signal is determined as the first echo signal; wherein a time period length from the start time to the first time is equal to a time period length of the first time period, and a time period length from the start time to the second time is equal to a time period length of the second time period. The method according to claim 1, further comprising: obtaining a first preset text corresponding to the voice signal of the first frequency band, and obtaining a second preset text corresponding to the voice signal of the second frequency band; performing speech recognition on the echo signal corresponding to the voice signal in the time period to obtain a text recognition result; determining the echo signal with the text recognition result being the first preset text as the first echo signal, and determining the echo signal with the text recognition result being the second preset text as the second echo signal. The method according to claim 1, wherein the object recognition result is obtained based on the first echo signal and the second echo signal, comprising: performing signal processing on the first echo signal to obtain first echo information, and performing signal processing on the second echo signal to obtain second echo information; obtaining the object recognition result based on the first echo information and the second echo information; wherein the object recognition result is used to indicate whether the to-be-identified object and the authentication object are the same object. The method according to claim 9, wherein the first echo information is obtained by performing signal processing on the first echo signal, comprising: performing signal purification on the first echo signal to obtain a first purified signal, and performing time domain signal processing on the first purified signal to obtain first time domain information of the first purified signal; performing frequency domain signal processing on the first purified signal to obtain first frequency domain information of the first purified signal; performing information fusion on the first time domain information and the first frequency domain information to obtain the first echo information. The method according to claim 9, wherein the signal processing of the second echo signal to obtain second echo information comprises: frequency shift signal processing of the second echo signal to obtain second frequency shift information of the second echo signal; phase signal processing of the second echo signal to obtain second phase information of the second echo signal; and information fusion of the second frequency shift information and the second phase information to obtain the second echo information. The method according to claim 9, wherein the object recognition is implemented by an object recognition model, the object recognition model comprising a first feature extraction layer, a second feature extraction layer and a recognition layer; and the obtaining of the object recognition result based on the first echo information and the second echo information comprises: calling the first feature extraction layer to perform feature extraction on the first echo information to obtain first echo features; calling the second feature extraction layer to perform dynamic feature extraction on the second echo information to obtain second echo features; and calling the recognition layer to perform object recognition on the object to be recognized based on the first echo features and the second echo features to obtain the object recognition result. The method according to claim 1, after obtaining the object recognition result, the method further comprises: if the object recognition result indicates that the object to be recognized and the authentication object are the same object, displaying a third preset text; obtaining a voice signal of the object to be recognized based on the third preset text, and extracting first timbre features of the object to be recognized from the voice signal; obtaining second timbre features of the authentication object from a database; and if a similarity between the first timbre features and the second timbre features is less than a preset similarity threshold, updating the object recognition result. A method for training an object recognition model, the method comprising: obtaining echo signal samples of sample objects and sample labels of the echo signal samples, the echo signal samples comprising first echo signal samples and second echo signal samples; calling an initial object recognition model to obtain object recognition results of the sample objects based on the first echo signal samples and the second echo signal samples; training the initial object recognition model based on the object recognition results of the sample objects and the sample labels of the echo signal samples to obtain an object recognition model; and wherein the object recognition model is used to obtain object recognition results based on first echo signals and second echo signals of an object to be recognized. An object recognition device, the device comprising: a sending module configured to send sound signals within a certain period of time, the period of time comprising a first period of time and a second period of time; wherein the sound signals in the first period of time are sound signals of a first frequency band, and the sound signals in the second period of time are sound signals of a second frequency band; a division module configured to obtain first echo signals corresponding to the sound signals of the first frequency band and second echo signals corresponding to the sound signals of the second frequency band; an object recognition module configured to obtain object recognition results based on the first echo signals and the second echo signals. A training device of an object recognition model, the device comprising: an acquisition module, configured to acquire an echo signal sample of a sample object and a sample label of the echo signal sample, the echo signal sample comprising a first echo signal sample and a second echo signal sample; a sample recognition module, configured to call an initial object recognition model, and obtain an object recognition result of the sample object based on the first echo signal sample and the second echo signal sample; a training module, configured to train the initial object recognition model based on the object recognition result of the sample object and the sample label of the echo signal sample, and obtain an object recognition model; wherein the object recognition model is configured to obtain an object recognition result based on a first echo signal and a second echo signal of a to-be-recognized object. An electronic device, the electronic device comprising: a memory, configured to store computer executable instructions or computer programs; a processor, configured to execute the computer executable instructions or the computer programs stored in the memory, and when the computer executable instructions or the computer programs are executed by the processor, the processor is caused to perform the method according to any one of claims 1 to 13 or the method according to claim 14. A computer readable storage medium, storing computer executable instructions, when the computer executable instructions are executed by a processor in an electronic device, the electronic device is caused to perform the method according to any one of claims 1 to 13 or the method according to claim 14. A computer program product, comprising computer programs or computer executable instructions, when the computer programs or the computer executable instructions are executed by a processor in an electronic device, the electronic device is caused to perform the method according to any one of claims 1 to 13 or the method according to claim 14.
Citation Information
Patent Citations
Indoor positioning method and device and computer readable storage medium
CN110286358A
Object recognition method and device based on ultrasonic echoes and storage medium
CN114814800A
Underwater target identification method and system based on double-frequency echo signal characteristics
CN115878982A
Controlling sensitivity of presence detection using ultrasonic signals
US11982737B1
System and Method Associated with User Authentication Based on an Acoustic-Based Echo-Signature
US20200309930A1