Object Recognition Method, Device and Storage Medium Based on Ultrasonic Echo

Through the ultrasonic echo recognition method, the correspondence of ultrasonic signals is used to solve the problem that image recognition is affected by scenes, and accurate object recognition and terminal security improvement in diverse scenarios are achieved.

CN114814800BActive Publication Date: 2025-05-27TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110071159.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-19
Publication Date
2025-05-27
Estimated Expiration
2041-01-19

AI Technical Summary

Technical Problem

In the prior art, the object recognition process relies on image recognition, which is affected by scene diversity, resulting in a decrease in recognition accuracy and terminal security.

Method used

The ultrasonic echo recognition method is used to transmit ultrasonic signals through the terminal, receive the echo signals of the object to be identified, perform vector standardization processing, extract ultrasonic echo features, and call the target model for feature dimension conversion and face translation to generate face image information.

Benefits of technology

It improves the accuracy of object recognition and terminal security, makes the recognition process suitable for diverse scenarios, and enhances the security performance of terminals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114814800B_ABST
    Figure CN114814800B_ABST
Patent Text Reader

Abstract

The present application discloses an object recognition method, device and storage medium based on ultrasonic echoes, which are applied to the computer vision technology of artificial intelligence. An ultrasonic signal is transmitted to the object to be recognized through the sound wave transmitting device of the terminal; then the echo signal reflected by the object to be recognized is received through the sound wave receiving device; the ultrasonic echo features of the echo signal are extracted; and the ultrasonic echo features are subjected to feature dimension conversion to obtain target dimension features; furthermore, face translation processing is performed to obtain the face image information corresponding to the object to be recognized. Thus, the object recognition process based on ultrasonic signals is realized. Since the echo signal based on ultrasonic signals corresponds to the object to be recognized, the object recognition process is applicable to different scenarios, ensuring the accuracy of object recognition and improving the security of the terminal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to an object recognition method, device and storage medium based on ultrasonic echo. Background Art

[0002] With the development of artificial intelligence technology, object recognition has been widely used in fields such as security and finance. In the process of object recognition, the accuracy of application scenario permission management and the security performance of the terminal are improved by identifying the objects used.

[0003] In the related art, the object recognition process is based on image recognition, that is, by collecting image information of the object or the object to be recognized, image features are extracted, and the result of image recognition is used as the recognition result of the object.

[0004] However, during the image recognition process of the terminal, due to the diversity of scenes, image features may not cover all complex scenes, that is, the recognition process based on image features will be affected by scene changes, resulting in errors in object recognition, affecting the accuracy of object recognition and the security of the corresponding terminal. Summary of the invention

[0005] In view of this, the present application provides an object recognition method based on ultrasonic echo, which can effectively improve the accuracy of object recognition.

[0006] The first aspect of the present application provides an object recognition method based on ultrasonic echo, which can be applied to a system or program including an object recognition function based on ultrasonic echo in a terminal device, specifically comprising:

[0007] Transmitting an ultrasonic signal to the object to be identified through the sound wave transmitting device of the terminal;

[0008] Receiving an echo signal reflected by the object to be identified through a sound wave receiving device of the terminal, wherein the echo signal corresponds to the ultrasonic wave signal;

[0009] Performing vector normalization processing on the echo signal to extract ultrasonic echo features corresponding to the echo signal;

[0010] Calling the target model to perform feature dimension conversion on the ultrasonic echo feature to obtain a target dimension feature for characterizing the object to be identified;

[0011] The target dimensional features of the object to be identified are subjected to face translation processing to obtain face image information corresponding to the object to be identified.

[0012] Optionally, in some possible implementations of the present application, the method further includes:

[0013] Acquire training image data and training echo data collected for the object to be identified;

[0014] Inputting the training image data into an image feature extraction network in a preset model to determine training image features in response to the setting of the target dimension;

[0015] generating corresponding frequency spectrum information based on the training echo data to determine training echo characteristics;

[0016] Inputting the training echo features into the acoustic wave conversion network in the preset model to convert the training echo data based on the target dimension to obtain training acoustic wave features, wherein the training acoustic wave features have the same dimension as the training image features;

[0017] Using the training image features as training targets to adjust the training sound wave features to obtain a similarity loss;

[0018] The parameters of the sound wave conversion network in the preset model are adjusted based on the similarity loss to obtain the target model.

[0019] Optionally, in some possible implementations of the present application, adjusting parameters of the sound wave conversion network in the preset model based on the similarity loss to obtain the target model includes:

[0020] Fixing the parameters of the image feature extraction network in the preset model;

[0021] In response to the parameters of the image feature extraction network in the preset model being fixed, the parameters of the sound wave conversion network in the preset model are adjusted based on the similarity loss to obtain the target model.

[0022] Optionally, in some possible implementations of the present application, the method further includes:

[0023] Acquire pre-training data, where the pre-training data is used to indicate a correspondence between a pre-training image and a pre-training feature;

[0024] The image feature extraction network in the preset model is trained based on the pre-training data to fix the parameters of the image feature extraction network in the preset model.

[0025] Optionally, in some possible implementations of the present application, the method further includes:

[0026] Controlling the sound wave emitting device to emit a single-channel ultrasonic wave toward the object to be identified based on a preset frequency characteristic;

[0027] Acquire the sound wave data obtained by the reflection received by the microphone to determine the training echo data;

[0028] In response to receiving the training echo data, calling an image acquisition device in the terminal to acquire training image data corresponding to the object to be identified;

[0029] Aligning the training echo data with the training image data based on execution time to generate training sample pairs;

[0030] The step of adjusting parameters of the sound wave conversion network in the preset model based on the similarity loss to obtain the target model includes:

[0031] The similarity loss is determined based on the training sample pair to adjust parameters of the sound wave conversion network in the preset model to obtain a target model.

[0032] Optionally, in some possible implementations of the present application, the determining the training echo data based on the sound wave data obtained by receiving the reflection by the sound wave receiving device includes:

[0033] The sound wave data obtained by receiving the reflection based on the sound wave receiving device;

[0034] Filtering the sound wave data according to the preset frequency characteristics to determine the sound wave data reflected by the object to be identified;

[0035] The training echo data is determined based on the sound wave data reflected by the object to be identified.

[0036] Optionally, in some possible implementations of the present application, filtering the sound wave data according to the preset frequency feature to determine the sound wave data reflected by the object to be identified includes:

[0037] Dividing the sound wave data into a plurality of fixed-length sound wave sequences according to the preset frequency characteristics;

[0038] Filtering the sound wave sequence based on a preset filter to obtain a filtered sequence;

[0039] The filter sequence is normalized according to a sliding window to determine the sound wave data reflected by the object to be identified.

[0040] Optionally, in some possible implementations of the present application, aligning the training echo data with the training image data based on execution time to generate a training sample pair includes:

[0041] In response to receiving the training echo data, calling an image acquisition device to acquire a video stream corresponding to the object to be identified;

[0042] Obtaining a timestamp corresponding to each video frame in the video stream;

[0043] The training echo data and the training image data are aligned according to the timestamp to generate the training sample pairs.

[0044] Optionally, in some possible implementations of the present application, transmitting an ultrasonic signal to the object to be identified by a sound wave transmitting device of the terminal includes:

[0045] determining the privacy operation authority in response to the identification instruction;

[0046] The sound wave emitting device is called based on the private operation authority so that the sound wave emitting device emits a single-channel ultrasonic wave.

[0047] Optionally, in some possible implementations of the present application, the method further includes:

[0048] Acquire a preset image saved in a target application corresponding to the target model;

[0049] Comparing the facial image information with the preset image to obtain comparison information;

[0050] Instructing execution of the target application based on the comparison information.

[0051] Optionally, in some possible implementations of the present application, the method further includes:

[0052] Obtaining position information and depth of field information corresponding to the facial image information;

[0053] Verifying the facial image information based on the position information and the depth of field information to determine object recognition information;

[0054] The execution of the target application is instructed according to the object identification information.

[0055] Optionally, in some possible implementations of the present application, the ultrasonic information is single-channel ultrasonic information, the terminal is a mobile phone, the sound wave transmitting device is a speaker, and the sound wave receiving device is a microphone.

[0056] A second aspect of the present application provides an object recognition device based on ultrasonic echo, comprising:

[0057] A transmitting unit, used to transmit an ultrasonic signal to an object to be identified through a sound wave transmitting device of the terminal;

[0058] a receiving unit, configured to receive an echo signal reflected by the object to be identified through a sound wave receiving device of the terminal, wherein the echo signal corresponds to the ultrasonic wave signal;

[0059] an extraction unit, used for performing vector normalization processing on the echo signal to extract ultrasonic echo features corresponding to the echo signal;

[0060] A conversion unit, used for calling a target model to perform feature dimension conversion on the ultrasonic echo feature, so as to obtain a target dimension feature for characterizing the object to be identified;

[0061] The recognition unit is used to perform face translation processing on the target dimensional features of the object to be recognized to obtain the face image information corresponding to the object to be recognized.

[0062] Optionally, in some possible implementations of the present application, the recognition unit is specifically used to obtain training image data and training echo data collected for the object to be recognized;

[0063] The recognition unit is specifically used to input the training image data into an image feature extraction network in a preset model to determine the training image features in response to the setting of the target dimension;

[0064] The identification unit is specifically used to generate corresponding frequency spectrum information based on the training echo data to determine the training echo characteristics;

[0065] The recognition unit is specifically configured to input the training echo feature into the acoustic wave conversion network in the preset model, so as to convert the training echo data based on the target dimension to obtain a training acoustic wave feature, wherein the training acoustic wave feature has the same dimension as the training image feature;

[0066] The recognition unit is specifically used to adjust the training sound wave feature by taking the training image feature as a training target to obtain a similarity loss;

[0067] The recognition unit is specifically used to adjust the parameters of the sound wave conversion network in the preset model based on the similarity loss to obtain the target model.

[0068] Optionally, in some possible implementations of the present application, the recognition unit is specifically used to fix the parameters of the image feature extraction network in the preset model;

[0069] The recognition unit is specifically used to adjust the parameters of the sound wave conversion network in the preset model based on the similarity loss in response to the fixed parameters of the image feature extraction network in the preset model to obtain the target model.

[0070] Optionally, in some possible implementations of the present application, the recognition unit is specifically used to obtain pre-training data, where the pre-training data is used to indicate a correspondence between a pre-training image and a pre-training feature;

[0071] The recognition unit is specifically used to train the image feature extraction network in the preset model based on the pre-training data to fix the parameters of the image feature extraction network in the preset model.

[0072] Optionally, in some possible implementations of the present application, the identification unit is specifically used to control the sound wave emitting device to emit a single-channel ultrasonic wave to the object to be identified based on a preset frequency feature;

[0073] The recognition unit is specifically used to obtain the sound wave data obtained by the reflection received by the microphone to determine the training echo data;

[0074] The recognition unit is specifically configured to, in response to receiving the training echo data, call the image acquisition device in the terminal to acquire the training image data corresponding to the object to be identified;

[0075] The recognition unit is specifically configured to align the training echo data with the training image data based on execution time to generate a training sample pair;

[0076] The recognition unit is specifically used to determine the similarity loss based on the training sample pair to adjust the parameters of the sound wave conversion network in the preset model to obtain the target model.

[0077] Optionally, in some possible implementations of the present application, the identification unit is specifically configured to receive the sound wave data obtained by the sound wave receiving device based on the reflection;

[0078] The recognition unit is specifically used to filter the sound wave data according to the preset frequency characteristics to determine the sound wave data reflected by the object to be recognized;

[0079] The recognition unit is specifically configured to determine the training echo data based on the sound wave data reflected by the object to be recognized.

[0080] Optionally, in some possible implementations of the present application, the recognition unit is specifically used to divide the sound wave data into a plurality of fixed-length sound wave sequences according to the preset frequency characteristics;

[0081] The recognition unit is specifically used to filter the sound wave sequence based on a preset filter to obtain a filtered sequence;

[0082] The recognition unit is specifically used to perform normalization processing on the filter sequence according to the sliding window to determine the sound wave data reflected by the object to be recognized.

[0083] Optionally, in some possible implementations of the present application, the recognition unit is specifically configured to, in response to receiving the training echo data, call an image acquisition device to acquire a video stream corresponding to the object to be identified;

[0084] The identification unit is specifically used to obtain a timestamp corresponding to each video frame in the video stream;

[0085] The recognition unit is specifically configured to align the training echo data with the training image data according to the timestamp to generate the training sample pair.

[0086] Optionally, in some possible implementations of the present application, the transmitting unit is specifically configured to determine the private operation authority in response to the identification instruction;

[0087] The transmitting unit is specifically used to call the sound wave transmitting device based on the private operation authority, so that the sound wave transmitting device transmits a single-channel ultrasonic wave.

[0088] Optionally, in some possible implementations of the present application, the recognition unit is specifically used to obtain a preset image saved in a target application corresponding to the target model;

[0089] The recognition unit is specifically used to compare the facial image information with the preset image to obtain comparison information;

[0090] The identification unit is specifically configured to instruct execution of the target application based on the comparison information.

[0091] Optionally, in some possible implementations of the present application, the recognition unit is specifically used to obtain position information and depth of field information corresponding to the face image information;

[0092] The recognition unit is specifically configured to verify the face image information based on the position information and the depth of field information to determine the object recognition information;

[0093] The identification unit is specifically configured to instruct execution of the target application according to the object identification information.

[0094] The third aspect of the present application provides a computer device, including: a memory, a processor and a bus system; the memory is used to store program code; the processor is used to execute the object recognition method based on ultrasonic echo as described in the first aspect or any one of the first aspects according to the instructions in the program code.

[0095] A fourth aspect of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the object recognition method based on ultrasonic echo as described in the first aspect or any one of the first aspects.

[0096] According to one aspect of the present application, a computer program product or a computer program is provided, the computer program product or the computer program comprising computer instructions, the computer instructions being stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the object recognition method based on ultrasonic echo provided in the first aspect or various optional implementations of the first aspect.

[0097] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0098] The ultrasonic signal is transmitted to the object to be identified through the terminal's sound wave transmitting device; then the echo signal reflected by the object to be identified is received through the terminal's sound wave receiving device, and the echo signal corresponds to the ultrasonic signal; the echo signal is further vector-normalized to extract the ultrasonic echo feature corresponding to the echo signal; and the target model is called to perform feature dimension conversion on the ultrasonic echo feature to obtain the target dimension feature used to characterize the object to be identified; and then the target dimension feature of the object to be identified is subjected to face translation processing to obtain the face image information corresponding to the object to be identified. Thus, the object recognition process based on the ultrasonic signal is realized. Since the echo signal based on the ultrasonic signal corresponds to the object to be identified, the object recognition process is applicable to different scenarios, ensuring the accuracy of object recognition and improving the security of the terminal. BRIEF DESCRIPTION OF THE DRAWINGS

[0099] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0100] Figure 1 A diagram of the network architecture that operates the ultrasound echo-based object recognition system;

[0101] Figure 2 A flowchart of an ultrasonic echo-based object recognition process according to an embodiment of the present application;

[0102] Figure 3 A flowchart of an object recognition method based on ultrasonic echo provided in an embodiment of the present application;

[0103] Figure 4 A schematic diagram of a scene of an object recognition method based on ultrasonic echo provided in an embodiment of the present application;

[0104] Figure 5 A schematic diagram of another scene of an object recognition method based on ultrasonic echo provided in an embodiment of the present application;

[0105] Figure 6 A schematic diagram of another scene of an object recognition method based on ultrasonic echo provided in an embodiment of the present application;

[0106] Figure 7 A schematic diagram of another scene of an object recognition method based on ultrasonic echo provided in an embodiment of the present application;

[0107] Figure 8 A flowchart of another object recognition method based on ultrasonic echo provided in an embodiment of the present application;

[0108] Fig. 9 A flowchart of another object recognition method based on ultrasonic echo provided in an embodiment of the present application;

[0109] Fig.10 A schematic diagram of another scene of an object recognition method based on ultrasonic echo provided in an embodiment of the present application;

[0110] Fig.11 A schematic diagram of another scene of an object recognition method based on ultrasonic echo provided in an embodiment of the present application;

[0111] Fig.12 A schematic diagram of the structure of an object recognition device based on ultrasonic echo provided in an embodiment of the present application;

[0112] Fig.13 A schematic diagram of the structure of a terminal device provided in an embodiment of the present application;

[0113] Fig.14 A schematic diagram of the structure of a server provided in an embodiment of the present application. DETAILED DESCRIPTION

[0114] The embodiment of the present application provides an object recognition method, device and storage medium based on ultrasonic echo, wherein an ultrasonic signal is transmitted to an object to be recognized through a sound wave transmitting device of a terminal; then an echo signal reflected by the object to be recognized is received through a sound wave receiving device of the terminal, and the echo signal corresponds to the ultrasonic signal; the echo signal is further vector-normalized to extract the ultrasonic echo feature corresponding to the echo signal; and a target model is called to perform feature dimension conversion on the ultrasonic echo feature to obtain a target dimension feature for characterizing the object to be recognized; and then the target dimension feature of the object to be recognized is subjected to face translation processing to obtain the face image information corresponding to the object to be recognized. Thus, an object recognition process based on ultrasonic signals is realized. Since the echo signal based on ultrasonic signals corresponds to the object to be recognized, the object recognition process is applicable to different scenarios, ensuring the accuracy of object recognition and improving the security of the terminal.

[0115] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein, for example. In addition, the terms "including" and "corresponding to" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0116] First, some terms that may appear in the embodiments of the present application are explained.

[0117] Sound wave echo: Sound wave that is reflected from an object after it is emitted.

[0118] Face feature extractor: A network that extracts face features from images.

[0119] Sound wave conversion network: A network that extracts features from input sound wave samples and generates facial features.

[0120] Face image translation network: a network that translates high-dimensional face features into corresponding face images.

[0121] It should be understood that the object recognition method based on ultrasonic echo provided in the present application can be applied to a system or program in a terminal device that includes an object recognition function based on ultrasonic echo, such as a password-protected application. Specifically, the object recognition system based on ultrasonic echo can be run on, for example, Figure 1 In the network architecture shown in Figure 1 As shown in FIG. 1 , it is a network architecture diagram of the operation of the object recognition system based on ultrasonic echo. As can be seen from the figure, the object recognition system based on ultrasonic echo can provide an object recognition process based on ultrasonic echo with multiple information sources, that is, the ultrasonic echo information is collected by the terminal, and the image of the ultrasonic echo information is restored on the server side, so as to perform the corresponding object recognition process; it can be understood that Figure 1 A variety of terminal devices are shown in FIG. 1 . The terminal devices may be mobile terminals such as mobile phones, computer devices, etc. In actual scenarios, more or fewer types of terminal devices may participate in the process of object recognition based on ultrasonic echoes. The specific number and type depend on the actual scenario and are not limited here. In addition, Figure 1 One server is shown in the figure, but in an actual scenario, multiple servers may be involved, and the specific number of servers depends on the actual scenario.

[0122] In this embodiment, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected through wired or wireless communication, and the terminal and the server can be connected to form a blockchain network, which is not limited in this application.

[0123] It can be understood that the above-mentioned ultrasonic echo-based object recognition system can run on a personal mobile terminal, for example, as an application such as a password-protected application, or it can be run on a server, or it can be run on a third-party device to provide ultrasonic echo-based object recognition, so as to obtain the ultrasonic echo-based object recognition processing result of the information source; the specific ultrasonic echo-based object recognition system can be run in the above-mentioned device in the form of a program, or it can be run as a system component in the above-mentioned device, or it can be used as a cloud service program. The specific operation mode depends on the actual scenario and is not limited here.

[0124] Artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines so that machines have the functions of perception, reasoning and decision-making.

[0125] Artificial intelligence technology is a comprehensive discipline that covers a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, mechatronics and other technologies. Artificial intelligence software technology mainly includes computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0126] The solution provided in the embodiment of the present application relates to computer vision technology of artificial intelligence. Computer vision is a science that studies how to make machines "see". To put it more concretely, it refers to machine vision such as using cameras and computers to replace human eyes to identify, locate and measure targets, and further perform graphic processing to make computer processing into images that are more suitable for human eye observation or transmission to instrument detection. As a scientific discipline, computer vision studies related theories and technologies, and attempts to establish an artificial intelligence system that can obtain information from images or multidimensional data. Computer vision technology generally includes image processing, image restoration, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, three-dimensional object reconstruction, 3D (3-Dimension) technology, virtual reality, augmented reality, synchronous positioning and map construction, and other technologies, as well as common biometric recognition technologies such as face recognition and object recognition.

[0127] With the development of artificial intelligence technology, object recognition has been widely used in fields such as security and finance. In the process of object recognition, the accuracy of application scenario permission management and the security performance of the terminal are improved by identifying the objects used.

[0128] In the related art, the object recognition process is based on image recognition, that is, by collecting image information of the object or the object to be recognized, image features are extracted, and the result of image recognition is used as the recognition result of the object.

[0129] However, during the image recognition process of the terminal, due to the diversity of scenes, image features may not cover all complex scenes, that is, the recognition process based on image features will be affected by scene changes, resulting in errors in object recognition, affecting the accuracy of object recognition and the security of the corresponding terminal.

[0130] In order to solve the above problems, this application proposes an object recognition method based on ultrasonic echo, which is applied to Figure 2 In the process framework of object recognition based on ultrasonic echo, as shown in Figure 2As shown, it is a process architecture diagram of object recognition based on ultrasonic echo provided in an embodiment of the present application. The user performs relevant private operations through the interface layer. In order to ensure that the private operations are performed by the user himself, the object recognition process can be triggered in the background, that is, the terminal transmits an ultrasonic signal to the object to be identified by controlling the sound wave transmitting device (such as a speaker), and receives the ultrasonic signal reflected by the object to be identified, that is, the echo signal, through the sound wave receiving device (such as a microphone); and then inputs the collected echo signal into the target model for identification. Since the target model is obtained by comparing and training the echo information associated with the user with the image information, it has the ability to restore the image information from the echo information, thereby ensuring the echo signal recognition process, and then outputting the image information corresponding to the echo signal.

[0131] In a possible scenario, the terminal is a mobile phone, that is, a single-channel ultrasonic wave can be emitted to the object to be identified through a speaker that controls the position of the front camera, and then the single-channel ultrasonic wave reflected by the object to be identified is received through a microphone, and the user's face video or face image is recorded through the front camera. Specifically, the video can be a 2D face image sequence of the user. Further, the terminal will send the above-mentioned sound wave and image information to the server for feature extraction, that is, firstly standardize the detected single ultrasonic echo vector, and use the high-dimensional features obtained by the 2D face image through the face feature extractor as the target feature, so that the ultrasonic vector passes through the sound wave conversion network to train the sound wave conversion network to approach the corresponding face features, and restore the 2D image with the face features obtained from the sound wave. Finally, the sound wave conversion network can learn the mapping relationship between the reflected sound wave and the face image features, and realize the construction of the corresponding face image from the reflected sound wave; thereby, the echo information collected by the terminal is recognized to obtain a recognition image.

[0132] It is understandable that the embodiment of the present application can extract the physical information contained in the ultrasound echo by learning the mapping relationship between the ultrasound echo and the high-dimensional features of the 2D face image. The extracted image features may serve as the basis for subsequent ultrasound-based learning of other tasks (such as identity verification, etc.).

[0133] It can be understood that the method provided in the present application can be written as a program as a processing logic in a hardware system, or as an object recognition device based on ultrasonic echo, and the above processing logic is implemented in an integrated or external manner. As an implementation method, the object recognition device based on ultrasonic echo transmits an ultrasonic signal to the object to be recognized through the sound wave transmitting device of the terminal; then the echo signal reflected by the object to be recognized is received through the sound wave receiving device of the terminal, and the echo signal corresponds to the ultrasonic signal; the echo signal is further vector-normalized to extract the ultrasonic echo feature corresponding to the echo signal; and the target model is called to perform feature dimension conversion on the ultrasonic echo feature to obtain the target dimension feature for characterizing the object to be recognized; and then the target dimension feature of the object to be recognized is subjected to face translation processing to obtain the face image information corresponding to the object to be recognized. Thereby, the object recognition process based on ultrasonic signal is realized. Since the echo signal based on ultrasonic signal has correspondence with the object to be recognized, the object recognition process is applicable to different scenarios, ensuring the accuracy of object recognition and improving the security of the terminal.

[0134] The solution provided in the embodiments of the present application relates to computer vision technology of artificial intelligence, which is specifically described by the following embodiments:

[0135] In combination with the above process architecture, the object recognition method based on ultrasonic echo in this application will be introduced below. Figure 3 , Figure 3 The flowchart of the object recognition method based on ultrasonic echo provided in the embodiment of the present application, the management method can be executed by the terminal device, can be executed by the server, or can be executed by the terminal device and the server together, and the embodiment of the present application at least includes the following steps:

[0136] 301. Transmit an ultrasonic wave signal to an object to be identified through a sound wave transmitting device of a terminal.

[0137] In this embodiment, the action of the terminal emitting ultrasonic waves requires hardware support. For example, when the terminal is a mobile phone, the hardware devices used are speakers and microphones, and the auxiliary device is the front camera currently configured in all smart phones. Taking this embodiment as an example, the speaker on the top of the mobile phone is used as a sound wave transmitting device, and the microphone is used as a sound wave receiving device, and the front camera is used as an image acquisition device, that is, for image and video acquisition.

[0138] In one possible scenario, Figure 4 FIG. 1 is a schematic diagram of a scene of an object recognition method based on ultrasonic echo provided in an embodiment of the present application; wherein Figure 4 (1)(2)(3) are collection terminals with different configurations. Figure 4(1) shows a speaker A1, a microphone A2 and a front camera A3; Figure 4 (2) shows a speaker B1, a microphone B2 and a front camera B3; Figure 4 (3) shows a speaker C1, a microphone C2 and a front camera C3. Specifically, the speaker is responsible for transmitting ultrasonic waves; the microphone is responsible for receiving ultrasonic echoes; and the front camera is responsible for capturing and recording images and videos of the target object.

[0139] Understandably, Figure 4 The locations of sensors on three types of mobile phone structures (Android phones, iPhone series and special-shaped screen phones) are listed in the figure. The locations of the three sensors on smartphones with other structures depend on the actual scenario and are not limited here.

[0140] Optionally, the ultrasonic signal emitted by the speaker in this embodiment can be a multi-channel ultrasonic signal. When considering the applicability of the mobile phone terminal, the object recognition process can be achieved only by a monophonic ultrasonic signal, thereby significantly reducing costs and improving the applicability of the technology.

[0141] Specifically, since monophonic ultrasonic waves will produce reflection and diffuse reflection when encountering obstacles or planes during propagation, the generated echo contains not only distance information, but also shape and material information of the reflecting surface or obstacle itself.

[0142] 302. Receive an echo signal reflected by the object to be identified through a sound wave receiving device of the terminal, where the echo signal corresponds to the ultrasonic wave signal.

[0143] In this embodiment, the echo signal is a sound wave signal reflected by the ultrasonic signal from the object to be identified. Specifically, the sound wave receiving device is the microphone described in step 301. The relevant description can be referred to and will not be repeated here.

[0144] Optionally, the sound wave receiving device in this embodiment can also be used in the process of acquiring training echo data, that is, the process of training the target model through training echo data and training image data for recognition. Figure 5 , which is a schematic diagram of another scene of an object recognition method based on ultrasonic echo provided by an embodiment of the present application; the mobile phone may be placed facing the face of a person so that the front of the face appears at the center of the front camera, and ultrasonic waves (higher frequency sound waves that are difficult for the human ear to accept) are emitted by the top speaker of the mobile phone, such as Figure 5 Left: Sound waves propagate in the air and produce echoes when they encounter obstacles. When they propagate to a person's face, the sound waves encounter obstacles and form reflections, such as Figure 5In the right picture, the sound waves are reflected on different planes of the face, and the reflected waves propagate to the position of the phone and are received by the microphone on the top. While the sound waves are being transmitted and received, the front camera records the video and collects the training image information at the corresponding moment.

[0145] Optionally, this embodiment proposes to use the microphone and speaker on the top of the mobile phone in conjunction with the front camera, but the solution is not limited to this. Any speaker and any microphone of the same mobile phone device can be used to emit or collect sound waves. The specific hardware location depends on the actual scenario and is not limited here.

[0146] In addition, this embodiment is based on monophonic sound wave transmission and recovery, considering cost control and scalability. Since multi-channel arrays can often provide richer information, on the mobile phone side, it can be manifested as using multiple speakers and microphones to transmit sound waves and receive echoes at the same time; it can also be manifested as the addition of external devices. Changes in the number and location of devices should be regarded as an expansion of this embodiment.

[0147] Through the description of the above embodiments, training echo data and training image data can be obtained through the terminal. The specific data scale can be for multiple objects (for example, face recognition of 100 people), or for multiple objects in a certain group (for example, face recognition of male gender), or it can be an enhanced recognition process for specific users (for example, obtaining the user's face data when the mobile phone is initialized to adjust the parameters of the target model in the mobile phone, so that the target model can recognize the user more accurately). The specific scale of training data and application method depend on the actual scenario and are not limited here.

[0148] 303. Perform vector normalization processing on the echo signal to extract ultrasonic echo features corresponding to the echo signal.

[0149] In this embodiment, the process of vector normalization processing can be to generate spectral information corresponding to the echo signal, such as generating a spectrum diagram, and processing it to obtain frequency domain and time domain information, thereby extracting features in the frequency domain and time domain information to obtain ultrasonic echo features corresponding to the echo signal.

[0150] It can be understood that since the distance between the user and the terminal is limited, the corresponding transmission distance of the ultrasound is also limited. Therefore, in the process of vector normalization of the echo signal, the time domain information within a certain time range can be intercepted, thereby avoiding the recognition interference caused by the ultrasonic echoes of other objects, thereby improving the accuracy of feature extraction.

[0151] 304. Call the target model to perform feature dimension conversion on the ultrasonic echo feature to obtain a target dimension feature for characterizing the object to be identified.

[0152] In this embodiment, the target model is a model that converts echo information into image information, wherein the target model that has not been adjusted in parameters can be called a preset model. The following is an explanation of the process of obtaining the target model by training the preset model; the preset model includes an image feature extraction network and an acoustic wave conversion network, wherein the image feature extraction network is used to extract training image features corresponding to training image data, and the acoustic wave conversion network is used to extract training acoustic wave features corresponding to training echo data; and the parameter adjustment process is the training process of the preset model, which is performed based on the training process of the acoustic wave conversion network, that is, the present application mainly adjusts the parameters of the acoustic wave conversion network so that the echo data matches the image data; in addition, the training process of the acoustic wave conversion network is performed according to the similarity loss obtained by comparing the training image features and the training acoustic wave features, that is, the difference between the training acoustic wave features and the training image features is minimized.

[0153] Specifically, for the training process of the preset model, that is, the matching process of the training acoustic wave features and the training image features, the matching process is the matching of multiple feature vector dimensions. In order to ensure the accuracy of the matching, the training acoustic wave features and the training image features can be represented by high dimensions (for example, 2048 dimensions). Specifically, the training image data can first be input into the image feature extraction network in the preset model to determine the training image features in response to the setting of the target dimension; then the corresponding spectrum information is generated based on the training echo data (for example, a spectrum graph is generated, and the spectrum is processed to obtain frequency domain and time domain information); next, the spectrum information is input into the acoustic wave conversion network in the preset model to process the training echo data into the training acoustic wave features with the same dimension as the training image feature based on the target dimension; and the training acoustic wave features are adjusted with the training image features as the training target to obtain the similarity loss; and then the parameters of the acoustic wave conversion network in the preset model are adjusted based on the similarity loss to obtain the target model, thereby ensuring the accuracy of the matching process of the training acoustic wave features and the training image features, that is, improving the accuracy of the similarity loss, and ensuring the ability of the target model to convert the acoustic wave features into image features.

[0154] Optionally, since this application is mainly aimed at the ability of converting the acoustic wave features of the target model into image features, the parameters of the image feature extraction network in the preset model can be fixed; then in response to the fixed parameters of the image feature extraction network in the preset model, the parameters of the acoustic wave conversion network in the preset model are adjusted based on the similarity loss to obtain the target model. This avoids the training process based on the similarity loss for the image feature extraction network.

[0155] Optionally, the process of converting sound wave features into image features involves an image translation network, that is, the preset model also includes an image translation network, so the parameters of the image translation network are also fixed during the training process based on similarity loss.

[0156] Specifically, the training process of the image feature extraction network and the image translation network can be directly used after pre-training with an additional data set, that is, first obtain pre-training data, and the pre-training data is used to indicate the correspondence between the pre-training image and the pre-training features; then train the image feature extraction network in the preset model based on the pre-training data to fix the parameters of the image feature extraction network in the preset model.

[0157] It is understandable that in order to ensure the privacy of subsequent identification data, the data samples used for training are not the same as the data used for subsequent testing, and it is ensured that the training and test sets do not overlap to avoid information leakage.

[0158] In a possible scenario, the training image data of the present application is face data, and the corresponding target model (or preset model) can be found in Figure 6 The scene architecture shown is as follows: Figure 6 As shown, it is a scene schematic diagram of another object recognition method based on ultrasonic echo provided by an embodiment of the present application; the figure shows a face feature extractor (image feature extraction network), an acoustic wave conversion network and a face image translator (image translation network); specifically, for the model training process, the pre-trained face feature extractor can be used to perform feature processing on the face image to obtain high-dimensional face features as training targets. At the same time, the ultrasonic echo signal is processed (for example, a spectrum diagram is generated, etc.) to obtain frequency domain and time domain information, which is input into the acoustic wave conversion network to convert it into high-dimensional acoustic wave features, and the high-dimensional acoustic wave features are compared with the above-mentioned processed high-dimensional face features to minimize the difference between the two. Then the extracted high-dimensional acoustic wave features are input into the face image translator to obtain a restored face.

[0159] Specifically, for the training process of the model, the facial feature extractor processes the input 2D image into high-dimensional features (such as a 2048-dimensional vector), and the facial image translator uses the features to reconstruct the 2D facial image, and the sound wave conversion network processes the sound wave into a vector of the same dimension as the facial features; and the similarity loss calculates the difference between the two variables (facial image features and sound wave features), and updates the sound wave conversion network through back propagation.

[0160] It is understandable that the parameters of the fixed part, including the sound wave conversion network, are used directly after pre-training with additional data sets, and remain fixed during the model training process. The data samples used for training are different from those used for testing, and the training and test sets are guaranteed not to overlap to avoid information leakage.

[0161] In one possible scenario, after the target model is trained, an object recognition process based on single-channel ultrasound can be performed, such as face recognition; the specific recognition object can be an object related to humans (or animals), such as faces, whole bodies, etc.; it can also be an object related to objects, such as flowers, trees, etc. The specific object depends on the actual scenario.

[0162] It is understandable that the object recognition process needs to correspond to the training data used in the training process, that is, if the object recognition is for a face, the training data is training echo data and training image data based on the face.

[0163] In another possible scenario, the recognition process of the echo signal (ultrasonic echo signal) is as follows: Figure 7 As shown, it is a scene schematic diagram of another object recognition method based on ultrasonic echo provided in an embodiment of the present application; that is, after processing the newly collected sound wave data (extraction of echo features, such as frequency, intensity threshold, etc.), the ultrasonic echo features are input into the sound wave conversion network to obtain high-dimensional sound wave features (collected sound wave features), and the high-dimensional sound wave features are input into the face image translator, and then the image is directly generated from the ultrasonic echo signal (restored face image).

[0164] 305. Perform face translation processing on the target dimension features of the object to be identified to obtain the face image information corresponding to the object to be identified.

[0165] In this embodiment, since the above-mentioned model, which has been trained extensively, already has the ability to restore the corresponding image from the echo, the fixed model in the above process can be directly used to restore the corresponding image using the collected ultrasonic echo. Specifically, after the training is stable, the model can achieve a relatively clear restoration of a clean face. And the restored image is very similar to the collected image, with little difference in physical position. It can be proved that the single-channel ultrasonic echo has the ability to restore the corresponding image.

[0166] It is understandable that, in addition to verifying that the model has the ability to extract physical information from echoes and generate facial images, the restored images or intermediate layer features can also directly serve other task scenarios. For example, the facial image restored by the sound wave can be directly connected to the classic face recognition network for recognition, or it can be compared with the face captured by the camera to see if they are the same; at the same time, if the user uses a photo or screen face to try to deceive the object recognition, since the ultrasonic reflection wave contains rich physical position and depth of field information, the restored facial image can also show a large difference from the photographed image.

[0167] Optionally, in the final generated image, this embodiment generates an RGB three-channel color image. In principle, ultrasonic echoes do not contain color information, and the image color learned by the model can be considered as the prior of the training data. In some scenarios, using a grayscale image as the target for restoration does not change the process, so the color selection of the target image can also be regarded as the recognition parameter of this embodiment.

[0168] In one possible scenario, the target model in this application can be applied to identity authentication, e-commerce, data security, financial risk control, smart hardware, etc., for example, as an identity authentication model for e-commerce; as a transaction object identification model in financial risk control. The specific use of the model depends on the actual scenario and is not limited here.

[0169] In combination with the above embodiments, it can be known that an ultrasonic signal is transmitted to the object to be identified through the sound wave transmitting device of the terminal; then the echo signal reflected by the object to be identified is received through the sound wave receiving device of the terminal, and the echo signal corresponds to the ultrasonic signal; the echo signal is further vector-normalized to extract the ultrasonic echo feature corresponding to the echo signal; and the target model is called to perform feature dimension conversion on the ultrasonic echo feature to obtain the target dimension feature for characterizing the object to be identified; and then the target dimension feature of the object to be identified is subjected to face translation processing to obtain the face image information corresponding to the object to be identified. Thus, the object recognition process based on the ultrasonic signal is realized. Since the echo signal based on the ultrasonic signal corresponds to the object to be identified, the object recognition process is applicable to different scenarios, ensuring the accuracy of object recognition and improving the security of the terminal.

[0170] The above embodiment introduces the training and use process of the target model, wherein the training of the target model uses training echo data and training image data. In some scenarios, the training echo data and training image data can be used as samples to train the input prediction model. The following introduces the scenario. Figure 8 FIG. 1 is a schematic diagram of another scene of an object recognition method based on ultrasonic echo provided in an embodiment of the present application; the scene includes the following steps:

[0171] Step 1: In response to the start of the target application, an object recognition process is executed.

[0172] In this embodiment, the start of the ultrasonic echo-based object recognition method can be triggered by the software application dimension, that is, the object recognition process is triggered in response to the startup of the target application, for example, the object recognition process is triggered in response to the download of a password-protected application, or it can be triggered by the configuration process of the password-protected application.

[0173] In addition, the start of the ultrasonic echo-based object recognition method can also be triggered in the hardware device dimension, such as triggering the object recognition process during the initialization of a mobile phone, or triggering the object recognition process after detecting that the mobile phone has changed the card.

[0174] Step 2: Call the training image data and training echo data based on the object to be identified, and process the samples.

[0175] In this embodiment, the data collection and processing of training image data (image signals) and training echo data (acoustic wave signals) are for generating sample pairs for training a prediction model, which is an untrained target model. Specifically, first, based on a preset frequency feature, the acoustic wave transmitting device in the terminal is controlled to transmit a single-channel ultrasonic wave to the object to be identified. Then, the acoustic wave receiving device in the terminal receives the reflected acoustic wave data to determine the training echo data. In response to the reception of the training echo data, the image acquisition device in the terminal is called to acquire the training image data corresponding to the object to be identified. The training echo data and the training image data are aligned based on the execution time to generate a training sample pair. Then, the preset model is trained based on the training sample pair to obtain the target model.

[0176] It can be understood that the purpose of emitting single-channel ultrasonic waves based on preset frequency characteristics is to facilitate the filtering of sound wave signals. This is because the microphone (sound wave receiving device) will collect the echo reflected from the face, and can also directly receive the ultrasonic waves emitted by the speaker (sound wave transmitting device). Attenuation will occur in the process of reflecting the echo, but this will not happen when the ultrasonic waves are directly received from the speaker. Therefore, the sound wave signal can be filtered according to the preset frequency characteristics to ensure the accuracy of the signal.

[0177] Specifically, the screening process first receives the reflected sound wave data based on the sound wave receiving device in the terminal; then filters the sound wave data according to the preset frequency characteristics to determine the sound wave data reflected by the object to be identified; and determines the training echo data based on the sound wave data reflected by the object to be identified.

[0178] In addition, in the specific sound wave filtering process, it can be based on the filter, that is, firstly, the sound wave data is divided into a plurality of fixed-length sound wave sequences according to the preset frequency characteristics; then, the sound wave sequence is filtered based on the preset filter to obtain a filtered sequence; and then, the filtered sequence is normalized according to the sliding window to determine the sound wave data reflected by the object to be identified. Since the normalization process is performed between the sequences, the continuity of the sound wave data is guaranteed.

[0179] It can be understood that since the image can be acquired through video, the processing of the image data can respond to the reception of the training echo data, call the image acquisition device in the terminal to acquire the video stream corresponding to the object to be identified; then obtain the timestamp corresponding to each video frame in the video stream; and then align the training echo data with the training image data according to the timestamp to generate a training sample pair, thereby ensuring the correspondence between the acoustic wave data and the image data.

[0180] Step 3: Build the target model based on the training image data and training echo data, and train and learn it.

[0181] In this embodiment, the model is constructed with reference to Figure 6 The scene architecture shown is not described in detail here; as for the training and learning process, the sample pair <image, reflected sound wave> is input into the preset model, wherein the image in the sample pair (training image data) is input into the image feature extraction network, and the reflected sound wave (training echo data) is input into the sound wave conversion network, thereby realizing the input of multiple sample pairs to train the prediction model and obtain the target model.

[0182] By inputting the training echo data and the training image data into the preset model in the form of sample pairs for training, the correspondence between the training data is ensured, the accuracy of the similarity difference is improved, and the effectiveness of the target model training is ensured.

[0183] Step 4: In response to receiving the echo reflected by the identification object, a corresponding face image is identified.

[0184] In this embodiment, the reception of the echo can be obtained after the ultrasonic emission is automatically triggered when the private operation is initiated. The specific private operation can correspond to scenarios such as identity authentication, e-commerce, data security, financial risk control, and smart hardware authentication.

[0185] Step 5: Perform the target task based on the recognized face image.

[0186] In this embodiment, the target task can be an image recognition task, that is, the facial image restored by the sound wave can be directly connected to the classic face recognition network for recognition, or it can be compared with the face captured by the camera to see if they are the same; at the same time, if the user uses a photo or screen face to try to deceive object recognition, since the ultrasonic reflection wave contains rich physical position and depth of field information, the restored facial image can also show a large difference from the photographed image.

[0187] In one possible scenario, during the initialization of applications with high security requirements (such as payment applications, privacy applications, etc.) in the mobile phone, the training process of the target model of steps 1 to 3 above can be performed, and then each time the application is opened, the starting user object can be automatically identified, and the identified image can be compared with the preset user image. If the similarity of the image reaches a preset value (for example, the similarity is 95%), it is determined that the object using the application is the preset user, thereby ensuring the security of the application data.

[0188] It can be seen from the above embodiments that the embodiment of the present application uses the reflected wave returned from the face, which contains more accurate and rich physical information, sufficient to support the recovery of facial image details. The embodiment of the present application can perform imaging only through a single-channel acoustic wave echo, so that the single-channel echo obtained on an ordinary mobile phone can also be used for imaging, which significantly reduces costs and improves the applicability of the technology. In addition, the ultrasonic echo from the face is used for imaging, which contains richer and more accurate physical information, and is assisted by the existing mature facial feature extractor. The characteristics of the neural network can be used to flexibly model complex scenes for high-precision imaging.

[0189] In one possible scenario, the generation process of <image, reflected sound wave> can be adopted Fig. 9 The process shown, Fig. 9 A flowchart of another object recognition method based on ultrasonic echo provided in an embodiment of the present application, the embodiment of the present application at least includes the following steps:

[0190] 901. Obtain a reflected sound wave reflected by an object to be identified.

[0191] It should be noted that the image sequence of the target object and the reflected sound wave sequence reflected by the target object here all come from the target object. The image sequence includes multiple images, and each image corresponds to a timestamp. In one possible implementation, the terminal obtains the image sequence of any target object in the following manner: the terminal collects at least one image of any target object at a fixed time interval, and sorts the at least one image in the order of the collection time to obtain the image sequence of any target object. The embodiment of the present application does not limit the fixed time interval, and can be determined according to the image collection frame rate of the image collection device of the terminal. For example, assuming that the image collection frame rate of the image collection device of the terminal is 50fps (frames per second), the fixed time interval is 20ms (milliseconds).

[0192] In a possible implementation, the terminal obtains the reflected sound wave reflected by any target object in the following manner: the terminal periodically transmits sound waves to the any target object, and uses the received reflected sound waves as the reflected sound waves reflected by the any target object. Since the terminal periodically transmits sound waves, the received reflected sound waves include the reflected sound waves of the sound waves emitted by the any target object in each period. In a possible implementation, the terminal periodically transmits sound waves to the any target object means that the terminal transmits a sound wave to the any target object once every period of time. The embodiment of the present application does not limit the interval between the emission of two sound waves, and can be set according to the need to perform the reference distance between the object to be detected and the object identification terminal and the sound wave propagation speed. For example, the interval between the emission of two sound waves can be set to 5ms. The reference distance between the object to be detected and the object identification terminal that needs to be detected can be obtained based on experience.

[0193] 902. Split the reflected sound wave into a plurality of reflected sub-sound waves.

[0194] Since the reflected sound waves reflected by any target object include the reflected sound waves of the sound waves emitted by any target object in each cycle, the reflected sound waves can be segmented according to the significant features to segment the reflected sound waves into at least one reflected sub-sound wave, and each reflected sub-sound wave is regarded as the reflected sound wave of the sound waves emitted by any target object in one cycle. In a possible implementation, for the interval time between the emission of two sound waves, the reference distance between the object to be detected and the object recognition terminal and the sound wave propagation speed are set as needed, and the significant feature can be the cycle time. It should be noted that the emitted sound wave itself has an emission time, and the cycle time is the sum of an emission time and an interval time. For example, assuming that a transmission time is 1ms and an interval time is 5ms, the cycle time is 6ms. After the reflected sound wave is segmented according to the cycle time of 6ms, each reflected sub-sound wave obtained is a high-frequency sound wave fragment of 6ms. After the reflected sound wave is segmented into at least one reflected sub-sound wave, each reflected sub-sound wave corresponds to time information. The time information corresponding to each reflected sub-sound wave may refer to the starting timestamp of each reflected sub-sound wave, or the ending timestamp of each reflected sub-sound wave, or the timestamp of a certain position (for example, the middle position) of each reflected sub-sound wave, or the timestamp range of each reflected sub-sound wave, etc. The embodiments of the present application are not limited to this.

[0195] 903. Filter and normalize the reflected sub-sound wave.

[0196] In this embodiment, preprocessing the reflected sound waves can improve the reliability of the reflected sound waves in the object recognition process. The operation of preprocessing the reflected sound waves can be set according to experience, and the embodiment of the present application is not limited to this. Exemplarily, the operation of preprocessing the reflected sound waves includes at least one of filtering, normalization, and wavelet transformation. Optionally, the process of filtering the reflected sound waves can be performed using a filter. The embodiment of the present application does not limit the type of the filter. Exemplarily, the filter is a time domain filter, a frequency domain filter, or a Kalman filter. Optionally, the process of normalizing the reflected sound waves can be implemented based on a sliding window.

[0197] 904. Obtain an image sequence of the object to be identified.

[0198] In this embodiment, the image sequence may be composed of adjacent video frames within a preset time period, and the preset time period corresponding to the video frame is a time period for echo acquisition.

[0199] 905. Determine the timestamp of the image object in the image sequence.

[0200] In this embodiment, the timestamp of the image object in the image sequence is to ensure the correspondence between the image and the sound wave. Specifically, the video data can be regarded as an image stream with a short interval, so it is intercepted according to the timestamp to obtain each frame of the image and its timestamp, and then the face is intercepted (face detection, center cutting, etc.) on the image.

[0201] 906. Perform an augmentation operation on the image.

[0202] In one possible implementation, preprocessing the image may refer to performing an augmentation operation on the image to increase the reliability of the image in the object recognition process. The embodiment of the present application does not limit the augmentation operation, and illustratively, the augmentation operation includes one or more of rotation, color change, blur, adding random noise, center cutting, and resolution reduction.

[0203] It is understandable that the acoustic wave processing method or image augmentation method used in the image processing process is only a part of the examples, and other feature processing methods (such as using continuous wavelet transform to obtain more features) may also be used, which is not limited here.

[0204] 907. Align the processed reflected sub-sound waves with the processed image to determine a corresponding relationship between the image and the reflected sound waves.

[0205] The collection frame rate of the reflected sub-sound wave is different from the collection frame rate of the images in the image sequence. For example, assuming that each reflected sub-sound wave is a high-frequency sound wave segment of 6ms, the collection frame rate of the reflected sub-sound wave is 167 reflected sub-sound waves per second, while the collection frame rate of the image is usually 30-60 images per second. The collection frame rate of the reflected sub-sound wave is significantly different from the collection frame rate of the images in the image sequence. It is necessary to align at least one acquired reflected sub-sound wave with at least one image in the image sequence to obtain a reflected sub-sound wave aligned with each image.

[0206] In a possible implementation, the process of aligning at least one reflected sub-sound wave with at least one image in an image sequence to obtain the reflected sub-sound wave aligned with the at least one image is as follows: determining the reflected sub-sound wave aligned with the at least one image according to the timestamp of at least one image in the image sequence and the time information of the at least one reflected sub-sound wave. In a possible implementation, the process of determining the reflected sub-sound wave aligned with the at least one image according to the timestamp of at least one image in the image sequence and the time information of the at least one reflected sub-sound wave is as follows: for any image, taking the reflected sub-sound wave whose time information matches the timestamp of the any image as the reflected sub-sound wave aligned with the any image.

[0207] It should be noted that the number of reflected sub-sound waves aligned with any image may be one or more, which is not limited in the embodiments of the present application. There may or may not be intersections in the reflected sub-sound waves aligned with two adjacent images, which is not limited in the embodiments of the present application. In the case where there is no intersection in the reflected sub-sound waves aligned with two adjacent images, there may be reflected sub-sound waves that are not aligned with any image, and these reflected sub-sound waves are discarded.

[0208] The condition for determining whether the time information matches the timestamp of the image can be set based on experience. For example, the condition for determining whether the time information matches the timestamp of the image can be: determining whether the absolute value of the difference between the timestamp indicated by the time information and the timestamp of the image is not greater than a reference threshold; when the absolute value of the difference between the timestamp indicated by the time information and the timestamp of the image is not greater than the reference threshold, it indicates that the time information matches the timestamp of the image; when the absolute value of the difference between the timestamp indicated by the time information and the timestamp of the image is greater than the reference threshold, it indicates that the time information does not match the timestamp of the image. It should be noted that when the time information is a timestamp, the timestamp indicated by the time information is the timestamp; when the time information is a timestamp range, the timestamp indicated by the time information may refer to a timestamp at a reference position (e.g., a middle position) in the timestamp range. The reference threshold can be set based on experience or flexibly adjusted according to the application scenario, and the embodiments of the present application do not limit this.

[0209] In one possible implementation, before aligning at least one reflected sub-sound wave with at least one image in an image sequence, at least one reflected sub-sound wave and at least one image in the image sequence may be preprocessed respectively, and then the preprocessed at least one reflected sub-sound wave may be aligned with the preprocessed at least one image.

[0210] In one possible implementation, preprocessing an image may refer to performing an augmentation operation on the image. The embodiment of the present application does not limit the augmentation operation. Exemplarily, the augmentation operation includes one or more of rotation, color change, blur, adding random noise, center cutting, and resolution reduction. The operation of preprocessing the reflected sub-sound wave can be set according to experience, and the embodiment of the present application does not limit this. Exemplarily, the operation of preprocessing the reflected sub-sound wave includes at least one of filtering, normalization, and wavelet transformation. Optionally, the process of filtering the reflected sub-sound wave can be performed using a filter. The embodiment of the present application does not limit the type of filter. Exemplarily, the filter is a time domain filter, a frequency domain filter, or a Kalman filter. Optionally, the process of normalizing the reflected sub-sound wave can be implemented based on a sliding window.

[0211] 908. Use the <image, reflected sound wave> pair for the same image as a training sample of the object to be identified.

[0212] After obtaining the reflected sub-sound waves respectively aligned with at least one image, for any image, the reflected sound waves corresponding to the any image are constructed based on the reflected sub-sound waves respectively aligned with the any image. In one possible implementation, the reflected sound waves corresponding to any image are constructed based on the reflected sub-sound waves respectively aligned with the any image as follows: the reflected sub-sound waves respectively aligned with the any image are spliced ​​in the order of the timestamps indicated by the time information of the reflected sub-sound waves to obtain the reflected sound waves corresponding to the any image. According to this method, the reflected sound waves respectively corresponding to at least one image can be obtained based on the reflected sub-sound waves respectively aligned with at least one image.

[0213] After obtaining the reflected sound wave corresponding to any image, any training sample corresponding to any target object is formed based on the any image and the reflected sound wave corresponding to the any image. It should be noted that after forming any training sample corresponding to any target object based on the any image and the reflected sound wave corresponding to the any image, the any training sample includes a <image, reflected sound wave> pair for the same image from the same target object.

[0214] After obtaining the reflected sound waves corresponding to at least one image, each image and the reflected sound waves corresponding to the image can constitute a training sample corresponding to any target object, thereby obtaining at least one training sample corresponding to any target object constituted by at least one image.

[0215] It is understandable that in the above data collection process, only single-channel ultrasound is used to generate 2D facial images. Specifically, when both the sound-emitting and receiving devices are stationary, single-channel ultrasound cannot perform 2D imaging. Moreover, ultrasound basically does not contain color information of objects. The embodiment of the present application uses the ability of the face feature extractor to extract key point features of the face, and the fitting ability of the deep neural network to perform 2D imaging based on single-channel ultrasound, and can also perform color imaging. This greatly reduces the cost of using a multi-channel ultrasound device. The method of the embodiment of the present application can be used to extract features on an ordinary mobile phone that only has a single-channel sound wave device.

[0216] The embodiment of the present application can perform imaging only through a single-channel acoustic wave echo, so that a single-channel echo obtained on an ordinary mobile phone can also be used for imaging, which significantly reduces costs and improves the applicability of the technology.

[0217] By collecting the reflected sound waves (echoes) and image sequences and performing corresponding preprocessing processes, the correspondence and accuracy of the <image, reflected sound wave> pairs are guaranteed, and the effectiveness of the preset model training is improved.

[0218] Next, we will introduce the object recognition process by taking the application scenario of Turing Shield owner recognition as an example. Turing Shield is an electronic encrypted digital economy application product for payment derived from the underlying program of Ethereum smart contracts. First, we will briefly introduce Turing Shield's owner recognition based on ultrasonic echo. When a user enters a password or verification code on a mobile terminal, it is assumed that the user's face is facing the camera on the top of the phone most of the time. At this time, the receiver at the camera position on the top of the phone emits sound waves, and the microphone at the camera position on the top of the phone receives the sound wave echo. After the sound wave echo data is reported, it is used to train the user's identity verification model, which is used to verify whether the password entered is the owner of the phone. The mobile terminal also authenticates the user. The identity verification includes two levels of verification. One level is to verify whether the user's facial features match the facial features of the owner of Turing Shield, and the other level is to verify whether the user's password or verification code input habits match the input habits of the owner of Turing Shield.

[0219] The password input interface during Turing Shield owner identification is as follows Fig.10 As shown, a password guard test is performed on the password input interface, and the number of tests and successful identifications are counted. Fig.10The number 171718 is a given password, and the user needs to enter the given password in the digital input position. When the user enters the password through the numeric keypad, the mobile terminal will collect the user's input habits, such as the tilt angle of the mobile terminal during the input process, the length of time the user's finger stays during the input process, etc., and then match the collected input habits with the pre-stored input habits of the owner of the Turing Shield to verify whether the user's password or verification code input habits match the input habits of the owner of the Turing Shield. This process can achieve unconscious object recognition of the user at a low user interaction cost while the user enters the password.

[0220] In another possible scenario, in addition to the above-mentioned user verification process, that is, preventing intruders from occupying the background operation permissions, this application can also be used for user authentication during the password input process, that is, to ensure that the executor of the password input is the target user, as shown in the following example. Fig.11 As shown, Fig.11 A flowchart of another object recognition method based on ultrasonic echo provided in an embodiment of the present application includes the following steps:

[0221] 1101. In response to a password input operation, control the speaker to send a single-channel ultrasonic wave to the face.

[0222] In this embodiment, the response to the password input operation may be during the password input process, when the relevant password page appears, or after the password input is completed, for example Fig.11 When the user clicks the password confirmation button D1, the server will be triggered to perform user authentication and password verification.

[0223] 1102. Receive reflected echo information.

[0224] In this embodiment, receiving the reflected echo information is a process in which the sound wave transmitting device in the calling terminal emits a single-channel ultrasonic wave in response to the execution of the private operation to obtain the echo signal through the sound wave receiving device in the terminal, and then the echo signal is input into the sound wave conversion network in the target model to obtain the collected sound wave characteristics.

[0225] 1103. Input the echo information into the target model to obtain a recognition image.

[0226] In this embodiment, for the recognition process of echo information, see Figure 3 The description of step 304 of the illustrated embodiment is not repeated here.

[0227] 1104. Compare the recognized image with the stored user image to determine comparison information.

[0228] In this embodiment, the comparison process can be performed based on a saved user image, that is, obtaining a preset image saved in the terminal for the target application; then comparing the recognized image with the preset image to obtain comparison information; and then instructing the execution of the target application based on the comparison information.

[0229] In addition, in order to prevent intruders from using photos or pictures to replace faces for recognition, the position information and depth of field information corresponding to the recognition image can also be obtained; the recognition image is verified based on the position information and depth of field information to determine the object recognition information; and then the execution of the target application is instructed according to the object recognition information. This is because the echo information contains the corresponding position information and depth of field information during the training process, thus avoiding the use of photos or pictures to replace faces for recognition.

[0230] 1105. The comparison information indicates that the recognized image is a target user.

[0231] In this embodiment, if the recognition image is consistent with the user image, it means that the object performing the password operation is the target user. Fig.11 The operation is successful as shown in D2.

[0232] 1106. The comparison information indicates that the recognized image is not the target user.

[0233] In this embodiment, if the recognition image is inconsistent with the user image, it means that the object performing the password operation is not the target user and further identity verification is required. Fig.11 Operation abnormality D3 is displayed in the middle. By identifying the operation object during the password input process, the security of private operations such as password input is guaranteed.

[0234] It is understandable that when both the sound wave transmitting device and the sound wave receiving device are stationary, single-channel ultrasound cannot perform 2D imaging. Moreover, ultrasound basically does not contain the color information of the object. This embodiment uses the ability of the face feature extractor to extract key features of the face and the fitting ability of the deep neural network to perform 2D imaging based on single-channel ultrasound, and can also perform color imaging. The cost of using a multi-channel ultrasound device is reduced. The method of this solution can be used to extract features on ordinary mobile phones that only have a single-channel sound wave device.

[0235] In order to better implement the above solution of the embodiment of the present application, the following also provides related devices for implementing the above solution. Fig.12 , Fig.12 This is a schematic diagram of the structure of an object recognition device based on ultrasonic echo provided in an embodiment of the present application. The recognition device 1200 includes:

[0236] The transmitting unit 1201 is used to transmit an ultrasonic signal to the object to be identified through the sound wave transmitting device of the terminal;

[0237] A receiving unit 1202, configured to receive an echo signal reflected by the object to be identified through a sound wave receiving device of the terminal, wherein the echo signal corresponds to the ultrasonic wave signal;

[0238] The extraction unit 1203 is used to perform vector normalization processing on the echo signal to extract the ultrasonic echo feature corresponding to the echo signal;

[0239] The conversion unit 1204 is used to call the target model to perform feature dimension conversion on the ultrasonic echo feature to obtain a target dimension feature for characterizing the object to be identified;

[0240] The recognition unit 1205 is used to perform face translation processing on the target dimensional features of the object to be recognized to obtain the face image information corresponding to the object to be recognized.

[0241] Optionally, in some possible implementations of the present application, the identification unit 1205 is specifically used to obtain training image data and training echo data collected for the object to be identified;

[0242] The recognition unit 1205 is specifically used to input the training image data into an image feature extraction network in a preset model to determine the training image features in response to the setting of the target dimension;

[0243] The identification unit 1205 is specifically configured to generate corresponding frequency spectrum information based on the training echo data to determine the training echo features;

[0244] The recognition unit 1205 is specifically configured to input the training echo feature into the acoustic wave conversion network in the preset model, so as to convert the training echo data based on the target dimension to obtain a training acoustic wave feature, wherein the training acoustic wave feature has the same dimension as the training image feature;

[0245] The recognition unit 1205 is specifically configured to adjust the training sound wave feature by taking the training image feature as a training target to obtain a similarity loss;

[0246] The recognition unit 1205 is specifically configured to adjust parameters of the sound wave conversion network in the preset model based on the similarity loss to obtain the target model.

[0247] Optionally, in some possible implementations of the present application, the recognition unit 1205 is specifically used to fix the parameters of the image feature extraction network in the preset model;

[0248] The recognition unit 1205 is specifically configured to adjust the parameters of the sound wave conversion network in the preset model based on the similarity loss in response to the fixed parameters of the image feature extraction network in the preset model to obtain the target model.

[0249] Optionally, in some possible implementations of the present application, the recognition unit 1205 is specifically used to obtain pre-training data, where the pre-training data is used to indicate a correspondence between a pre-training image and a pre-training feature;

[0250] The recognition unit 1205 is specifically used to train the image feature extraction network in the preset model based on the pre-training data to fix the parameters of the image feature extraction network in the preset model.

[0251] Optionally, in some possible implementations of the present application, the identification unit 1205 is specifically configured to control the sound wave emission device to emit a single-channel ultrasonic wave to the object to be identified based on a preset frequency feature;

[0252] The recognition unit 1205 is specifically used to obtain the sound wave data obtained by the reflection received by the microphone to determine the training echo data;

[0253] The recognition unit 1205 is specifically configured to, in response to receiving the training echo data, call an image acquisition device in the terminal to acquire training image data corresponding to the object to be recognized;

[0254] The identification unit 1205 is specifically configured to align the training echo data with the training image data based on execution time to generate a training sample pair;

[0255] The recognition unit 1205 is specifically configured to determine the similarity loss based on the training sample pair, so as to adjust parameters of the sound wave conversion network in the preset model, so as to obtain a target model.

[0256] Optionally, in some possible implementations of the present application, the identification unit 1205 is specifically configured to receive the sound wave data obtained by the sound wave receiving device based on the reflection;

[0257] The identification unit 1205 is specifically configured to filter the sound wave data according to the preset frequency characteristics to determine the sound wave data reflected by the object to be identified;

[0258] The identification unit 1205 is specifically configured to determine the training echo data based on the sound wave data reflected by the object to be identified.

[0259] Optionally, in some possible implementations of the present application, the identification unit 1205 is specifically configured to divide the sound wave data into a plurality of fixed-length sound wave sequences according to the preset frequency characteristics;

[0260] The identification unit 1205 is specifically configured to filter the sound wave sequence based on a preset filter to obtain a filtered sequence;

[0261] The identification unit 1205 is specifically configured to perform normalization processing on the filter sequence according to a sliding window to determine the sound wave data reflected by the object to be identified.

[0262] Optionally, in some possible implementations of the present application, the identification unit 1205 is specifically configured to, in response to receiving the training echo data, call an image acquisition device to acquire a video stream corresponding to the object to be identified;

[0263] The identification unit 1205 is specifically used to obtain a timestamp corresponding to each video frame in the video stream;

[0264] The identification unit 1205 is specifically configured to align the training echo data with the training image data according to the timestamp to generate the training sample pair.

[0265] Optionally, in some possible implementations of the present application, the transmitting unit 1201 is specifically configured to determine the private operation authority in response to the identification instruction;

[0266] The transmitting unit 1201 is specifically configured to call the sound wave transmitting device based on the private operation authority so that the sound wave transmitting device transmits a single-channel ultrasonic wave.

[0267] Optionally, in some possible implementations of the present application, the recognition unit 1205 is specifically used to obtain a preset image saved in a target application corresponding to the target model;

[0268] The recognition unit 1205 is specifically used to compare the facial image information with the preset image to obtain comparison information;

[0269] The identification unit 1205 is specifically configured to instruct execution of the target application based on the comparison information.

[0270] Optionally, in some possible implementations of the present application, the recognition unit 1205 is specifically used to obtain position information and depth of field information corresponding to the face image information;

[0271] The recognition unit 1205 is specifically configured to verify the recognition image based on the position information and the depth information to determine the object recognition information;

[0272] The identification unit 1205 is specifically configured to instruct execution of the target application according to the object identification information.

[0273] The ultrasonic signal is transmitted to the object to be identified through the terminal's sound wave transmitting device; then the echo signal reflected by the object to be identified is received through the terminal's sound wave receiving device, and the echo signal corresponds to the ultrasonic signal; the echo signal is further vector-normalized to extract the ultrasonic echo feature corresponding to the echo signal; and the target model is called to perform feature dimension conversion on the ultrasonic echo feature to obtain the target dimension feature used to characterize the object to be identified; and then the target dimension feature of the object to be identified is subjected to face translation processing to obtain the face image information corresponding to the object to be identified. Thus, the object recognition process based on the ultrasonic signal is realized. Since the echo signal based on the ultrasonic signal corresponds to the object to be identified, the object recognition process is applicable to different scenarios, ensuring the accuracy of object recognition and improving the security of the terminal.

[0274] The present application also provides a terminal device, such as Fig.13 The figure is a schematic diagram of the structure of another terminal device provided by the embodiment of the present application. For the convenience of explanation, only the part related to the embodiment of the present application is shown. For the specific technical details not disclosed, please refer to the method part of the embodiment of the present application. The terminal can be any terminal device including a mobile phone, a tablet computer, a personal digital assistant (PDA), a point of sales (POS), a car computer, etc., taking the mobile phone as an example:

[0275] Fig.13 FIG. 1 is a block diagram showing a partial structure of a mobile phone related to a terminal provided in an embodiment of the present application. Fig.13 The mobile phone includes: a radio frequency (RF) circuit 1310, a memory 1320, an input unit 1330, a display unit 1340, a sensor 1350, an audio circuit 1360, a wireless fidelity (WiFi) module 1370, a processor 1380, and a power supply 1390. Those skilled in the art will understand that Fig.13 The mobile phone structure shown in the figure does not constitute a limitation on the mobile phone, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0276] Combine the following Fig.13 A detailed introduction to the various components of the mobile phone:

[0277] The RF circuit 1310 can be used for receiving and sending signals during the process of sending and receiving information or making calls. In particular, after receiving the downlink information of the base station, it is sent to the processor 1380 for processing; in addition, the designed uplink data is sent to the base station. Generally, the RF circuit 1310 includes but is not limited to an antenna, at least one amplifier, a transceiver, a coupler, a low noise amplifier (LNA), a duplexer, etc. In addition, the RF circuit 1310 can also communicate with the network and other devices through wireless communication. The above wireless communication can use any communication standard or protocol, including but not limited to the global system of mobile communication (GSM), general packet radio service (GPRS), code division multiple access (CDMA), wideband code division multiple access (WCDMA), long term evolution (LTE), email, short messaging service (SMS), etc.

[0278] The memory 1320 can be used to store software programs and modules. The processor 1380 executes various functional applications and data processing of the mobile phone by running the software programs and modules stored in the memory 1320. The memory 1320 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, an application required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), etc. In addition, the memory 1320 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0279] The input unit 1330 can be used to receive input digital or character information, and to generate key signal input related to the user settings and function control of the mobile phone. Specifically, the input unit 1330 may include a touch panel 1331 and other input devices 1332. The touch panel 1331, also known as a touch screen, can collect the user's touch operation on or near it (such as the user's operation on the touch panel 1331 or near the touch panel 1331 using any suitable object or accessory such as a finger, stylus, etc., and the air touch operation within a certain range on the touch panel 1331), and drive the corresponding connection device according to a pre-set program. Optionally, the touch panel 1331 may include two parts: a touch detection device and a touch controller. Among them, the touch detection device detects the user's touch orientation, detects the signal brought by the touch operation, and transmits the signal to the touch controller; the touch controller receives the touch information from the touch detection device, converts it into contact point coordinates, and then sends it to the processor 1380, and can receive and execute commands sent by the processor 1380. In addition, the touch panel 1331 can be implemented in various types such as resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1331, the input unit 1330 can also include other input devices 1332. Specifically, the other input devices 1332 can include but are not limited to one or more of a physical keyboard, a function key (such as a volume control key, a switch key, etc.), a trackball, a mouse, a joystick, etc.

[0280] The display unit 1340 can be used to display information input by the user or information provided to the user and various menus of the mobile phone. The display unit 1340 may include a display panel 1341. Optionally, the display panel 1341 may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), etc. Further, the touch panel 1331 may cover the display panel 1341. When the touch panel 1331 detects a touch operation on or near it, it is transmitted to the processor 1380 to determine the type of touch event. Subsequently, the processor 1380 provides a corresponding visual output on the display panel 1341 according to the type of touch event. Although in Fig.13 In the embodiment, the touch panel 1331 and the display panel 1341 are used as two independent components to realize the input and output functions of the mobile phone, but in some embodiments, the touch panel 1331 and the display panel 1341 can be integrated to realize the input and output functions of the mobile phone.

[0281] The mobile phone may also include at least one sensor 1350, such as a light sensor, a motion sensor, and other sensors. Specifically, the light sensor may include an ambient light sensor and a proximity sensor, wherein the ambient light sensor may adjust the brightness of the display panel 1341 according to the brightness of the ambient light, and the proximity sensor may turn off the display panel 1341 and / or the backlight when the mobile phone is moved to the ear. As a type of motion sensor, the accelerometer sensor can detect the magnitude of acceleration in all directions (generally three axes), and can detect the magnitude and direction of gravity when stationary. It can be used for applications that identify the posture of the mobile phone (such as horizontal and vertical screen switching, related games, magnetometer posture calibration), vibration recognition related functions (such as pedometer, tapping), etc.; as for other sensors that can be configured in the mobile phone, such as gyroscopes, barometers, hygrometers, thermometers, infrared sensors, etc., they will not be repeated here.

[0282] The audio circuit 1360, the speaker 1361, and the microphone 1362 can provide an audio interface between the user and the mobile phone. The audio circuit 1360 can transmit the received audio data to the speaker 1361 after converting the received audio data into an electrical signal, which is converted into a sound signal for output; on the other hand, the microphone 1362 converts the collected sound signal into an electrical signal, which is received by the audio circuit 1360 and converted into audio data, and then the audio data is output to the processor 1380 for processing, and then sent to another mobile phone through the RF circuit 1310, or the audio data is output to the memory 1320 for further processing.

[0283] WiFi is a short-range wireless transmission technology. The mobile phone can help users send and receive emails, browse web pages and access streaming media through the WiFi module 1370. It provides users with wireless broadband Internet access. Fig.13 A WiFi module 1370 is shown, but it is understandable that it is not an essential component of the mobile phone and can be omitted as needed without changing the essence of the invention.

[0284] The processor 1380 is the control center of the mobile phone. It uses various interfaces and lines to connect various parts of the entire mobile phone. It executes various functions of the mobile phone and processes data by running or executing software programs and / or modules stored in the memory 1320 and calling data stored in the memory 1320. Optionally, the processor 1380 may include one or more processing units; optionally, the processor 1380 may integrate an application processor and a modem processor, wherein the application processor mainly processes the operating system, user interface and application programs, etc., and the modem processor mainly processes wireless communications. It is understandable that the above-mentioned modem processor may not be integrated into the processor 1380.

[0285] Specifically, the processor 1380 is specifically configured to transmit an ultrasonic wave signal to the object to be identified through a sound wave transmitting device of the terminal;

[0286] The processor 1380 is specifically configured to receive an echo signal reflected by the object to be identified through a sound wave receiving device of the terminal, wherein the echo signal corresponds to the ultrasonic wave signal;

[0287] The processor 1380 is specifically used to perform vector normalization processing on the echo signal to extract the ultrasonic echo feature corresponding to the echo signal;

[0288] The processor 1380 is specifically used to call the target model to perform feature dimension conversion on the ultrasonic echo feature to obtain a target dimension feature for characterizing the object to be identified;

[0289] The processor 1380 is specifically used to perform facial translation processing on the target dimensional features of the object to be identified to obtain facial image information corresponding to the object to be identified.

[0290] Optionally, in a possible scenario, the processor 1380 is specifically used to obtain training image data and training echo data collected for the object to be identified;

[0291] The processor 1380 is specifically used to input the training image data into an image feature extraction network in a preset model to determine the training image features in response to the setting of the target dimension;

[0292] The processor 1380 is specifically configured to generate corresponding frequency spectrum information based on the training echo data to determine training echo features;

[0293] The processor 1380 is specifically configured to input the training echo feature into the acoustic wave conversion network in the preset model, so as to convert the training echo data based on the target dimension to obtain a training acoustic wave feature, wherein the training acoustic wave feature has the same dimension as the training image feature;

[0294] The processor 1380 is specifically configured to adjust the training sound wave feature by taking the training image feature as a training target to obtain a similarity loss;

[0295] The processor 1380 is specifically used to adjust the parameters of the sound wave conversion network in the preset model based on the similarity loss to obtain the target model.

[0296] Optionally, in a possible scenario, the processor 1380 is specifically used to fix the parameters of the image feature extraction network in the preset model;

[0297] The processor 1380 is specifically configured to adjust the parameters of the sound wave conversion network in the preset model based on the similarity loss in response to the fixed parameters of the image feature extraction network in the preset model to obtain the target model.

[0298] Optionally, in a possible scenario, the processor 1380 is specifically used to obtain pre-training data, where the pre-training data is used to indicate a corresponding relationship between a pre-training image and a pre-training feature;

[0299] The processor 1380 is specifically used to train the image feature extraction network in the preset model based on the pre-training data to fix the parameters of the image feature extraction network in the preset model.

[0300] Optionally, in a possible scenario, the processor 1380 is specifically configured to control the sound wave emitting device to emit a single-channel ultrasonic wave to the object to be identified based on a preset frequency feature;

[0301] The processor 1380 is specifically used to obtain the sound wave data obtained by the reflection received by the microphone to determine the training echo data;

[0302] The processor 1380 is specifically configured to, in response to receiving the training echo data, call the image acquisition device in the terminal to acquire the training image data corresponding to the object to be identified;

[0303] The processor 1380 is specifically configured to align the training echo data with the training image data based on execution time to generate a training sample pair;

[0304] The processor 1380 is specifically used to adjust the parameters of the sound wave conversion network in the preset model based on the similarity loss to obtain the target model, including:

[0305] The processor 1380 is specifically used to determine the similarity loss based on the training sample pair to adjust the parameters of the sound wave conversion network in the preset model to obtain the target model.

[0306] Optionally, in a possible scenario, the processor 1380 is specifically configured to receive the sound wave data obtained by reflection based on the sound wave receiving device;

[0307] The processor 1380 is specifically configured to filter the sound wave data according to the preset frequency characteristics to determine the sound wave data reflected by the object to be identified;

[0308] The processor 1380 is specifically configured to determine the training echo data based on the sound wave data reflected by the object to be identified.

[0309] Optionally, in a possible scenario, the processor 1380 is specifically configured to divide the sound wave data into a plurality of fixed-length sound wave sequences according to the preset frequency characteristics;

[0310] The processor 1380 is specifically configured to filter the sound wave sequence based on a preset filter to obtain a filtered sequence;

[0311] The processor 1380 is specifically configured to perform normalization processing on the filter sequence according to a sliding window to determine the sound wave data reflected by the object to be identified.

[0312] Optionally, in a possible scenario, the processor 1380 is specifically configured to, in response to receiving the training echo data, call an image acquisition device to acquire a video stream corresponding to the object to be identified;

[0313] The processor 1380 is specifically used to obtain a timestamp corresponding to each video frame in the video stream;

[0314] The processor 1380 is specifically configured to align the training echo data with the training image data according to the timestamp to generate the training sample pair.

[0315] Optionally, in a possible scenario, the processor 1380 is specifically configured to determine the private operation authority in response to the identification instruction;

[0316] The processor 1380 is specifically configured to call the sound wave emitting device based on the private operation authority so that the sound wave emitting device emits a single-channel ultrasonic wave.

[0317] Optionally, in a possible scenario, the processor 1380 is specifically configured to obtain a preset image stored in a target application corresponding to the target model;

[0318] The processor 1380 is specifically used to compare the facial image information with the preset image to obtain comparison information;

[0319] The processor 1380 is specifically configured to instruct execution of the target application based on the comparison information.

[0320] Optionally, in a possible scenario, the processor 1380 is specifically used to obtain position information and depth of field information corresponding to the facial image information;

[0321] The processor 1380 is specifically configured to verify the recognition image based on the position information and the depth of field information to determine the object recognition information;

[0322] The target application is instructed to execute according to the object identification information. The mobile phone also includes a power supply 1390 (such as a battery) for supplying power to various components. Optionally, the power supply can be logically connected to the processor 1380 through a power management system, so that the power management system can manage charging, discharging, power consumption and other functions.

[0323] Although not shown, the mobile phone may also include a camera (image acquisition device), a microphone (sound wave receiving device), a speaker (sound wave transmitting device), a Bluetooth module, etc., which will not be described in detail here.

[0324] In the embodiment of the present application, the processor 1380 included in the terminal also has the function of executing each step of the above-mentioned page processing method.

[0325] The present application also provides a server. Fig.14 , Fig.14 14 is a schematic diagram of the structure of a server provided in an embodiment of the present application. The server 1400 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPU) 1422 (for example, one or more processors) and memory 1432, and one or more storage media 1430 (for example, one or more mass storage devices) storing application programs 1442 or data 1444. Among them, the memory 1432 and the storage medium 1430 may be short-term storage or permanent storage. The program stored in the storage medium 1430 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server. Furthermore, the central processing unit 1422 may be configured to communicate with the storage medium 1430 to execute a series of instruction operations in the storage medium 1430 on the server 1400.

[0326] The server 1400 may also include one or more power supplies 1426, one or more wired or wireless network interfaces 1450, one or more input and output interfaces 1458, and / or one or more operating systems 1441, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0327] The server 1400 is used to train the target model and perform the recognition process after training. For the training process, a partially dimensional processed <sound wave, face image> pair is input. First, a pre-trained face feature extractor is used to perform feature processing on the face image to obtain high-dimensional face features as the training target. At the same time, the ultrasonic echo signal is processed (for example, a spectrogram is generated, etc.) to obtain frequency domain and time domain information, which is input into the sound wave conversion network to convert it into high-dimensional sound wave features. The high-dimensional sound wave features are compared with the above-processed high-dimensional face features to minimize the difference between the two. Then the extracted high-dimensional sound wave features are input into the face image translator to obtain the restored face.

[0328] It can be understood that the face feature extractor processes the input image into high-dimensional features, the face image translator uses the features to reconstruct the face image, and the sound wave conversion network processes the sound waves into vectors of the same dimension as the face features; while the similarity loss calculates the difference between the two variables and updates the sound wave conversion network through back propagation.

[0329] The trainable parts of the target model mainly include the sound wave conversion network; while other parts (such as the face feature extractor and face image translator) are directly used after pre-training with additional data sets and remain fixed during the model training process. In addition, the data samples used for training are different from those used for testing, and the training and test sets are guaranteed not to overlap to avoid information leakage.

[0330] The steps performed by the management device in the above embodiment can be based on the Fig.14 The server structure shown.

[0331] In an embodiment of the present application, a computer-readable storage medium is further provided. The computer-readable storage medium stores an object recognition instruction based on ultrasonic echo. When the computer-readable storage medium is run on a computer, the computer executes the above-mentioned Figures 3 to 11 The illustrated embodiment describes the steps performed by the ultrasonic echo-based object recognition device in the method.

[0332] In an embodiment of the present application, a computer program product including instructions for object recognition based on ultrasonic echoes is also provided, which, when executed on a computer, enables the computer to execute the above-mentioned Figures 3 to 11 The illustrated embodiment describes the steps performed by the ultrasonic echo-based object recognition device in the method.

[0333] The present application also provides an object recognition system based on ultrasonic echoes. The object recognition system based on ultrasonic echoes may include Fig.12 The object recognition device based on ultrasonic echo in the described embodiment, or Fig.13 The terminal device in the described embodiment, or Fig.14The server being described.

[0334] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0335] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0336] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0337] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0338] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions to enable a computer device (which can be a personal computer, an object recognition device based on ultrasonic echo, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.

[0339] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An object recognition method based on ultrasonic echo, characterized in that, it includes: Transmitting an ultrasonic signal to the object to be recognized through the acoustic wave transmitting device of the terminal; Receiving the echo signal reflected by the object to be recognized through the acoustic wave receiving device of the terminal, and the echo signal corresponds to the ultrasonic signal; Performing vector normalization processing on the echo signal to extract the ultrasonic echo features corresponding to the echo signal; Invoking a target model to perform feature dimension conversion on the ultrasonic echo features to obtain target dimension features for characterizing the object to be recognized. Among them, the ultrasonic echo features are processed to obtain frequency domain and time domain information, and the frequency domain and time domain information are input into the target model to convert it into high-dimensional acoustic wave features, and the high-dimensional acoustic wave features are the target dimension features; Performing face translation processing on the target dimension features of the object to be recognized to obtain the face image information corresponding to the object to be recognized.

2. The method according to claim 1, characterized in that, the method further includes: Obtaining training image data and training echo data collected for the object to be recognized; Inputting the training image data into the image feature extraction network in the preset model to determine training image features in response to the setting of the target dimension; Generating corresponding spectrum information based on the training echo data to determine training echo features; Inputting the training echo features into the acoustic wave conversion network in the preset model to convert the training echo data based on the target dimension to obtain training acoustic wave features, and the training acoustic wave features have the same dimension as the training image features; Using the training image features as training targets to adjust the training acoustic wave features to obtain a similarity loss; Adjusting the parameters of the acoustic wave conversion network in the preset model based on the similarity loss to obtain the target model.

3. The method according to claim 2, characterized in that, the adjusting the parameters of the acoustic wave conversion network in the preset model based on the similarity loss to obtain the target model includes: Fixing the parameters of the image feature extraction network in the preset model; In response to the fixing of the parameters of the image feature extraction network in the preset model, adjusting the parameters of the acoustic wave conversion network in the preset model based on the similarity loss to obtain the target model.

4. The method according to claim 3, characterized in that, the method further includes: Obtaining pre-training data, which is used to indicate the correspondence between pre-training images and pre-training features; Training the image feature extraction network in the preset model based on the pre-training data to fix the parameters of the image feature extraction network in the preset model.

5. The method according to claim 2, characterized in that, the method further includes: Controlling the acoustic wave transmitting device to emit single-channel ultrasonic waves to the object to be recognized based on preset frequency features; Obtaining the acoustic wave data reflected and received by the microphone to determine the training echo data; In response to receiving the training echo data, call the image acquisition device in the terminal to acquire the training image data corresponding to the object to be recognized; Align the training echo data and the training image data based on the execution time to generate a training sample pair; The parameter adjustment of the acoustic wave conversion network in the preset model based on the similarity loss to obtain the target model includes: Determine the similarity loss based on the training sample pair, and perform parameter adjustment on the acoustic wave conversion network in the preset model to obtain the target model.

6. The method according to claim 5, wherein, The determining the training echo data based on the acoustic wave data reflected by the acoustic wave receiving device includes: Based on the acoustic wave data reflected by the acoustic wave receiving device; Filter the acoustic wave data according to the preset frequency characteristics to determine the acoustic wave data reflected by the object to be recognized; Determine the training echo data based on the acoustic wave data reflected by the object to be recognized.

7. The method according to claim 6, wherein, The filtering the acoustic wave data according to the preset frequency characteristics to determine the acoustic wave data reflected by the object to be recognized includes: Segment the acoustic wave data into multiple fixed-length acoustic wave sequences according to the preset frequency characteristics; Filter the acoustic wave sequences based on a preset filter to obtain a filtered sequence; Perform normalization processing on the filtered sequence according to a sliding window to determine the acoustic wave data reflected by the object to be recognized.

8. The method according to claim 5, wherein, The aligning the training echo data and the training image data based on the execution time to generate a training sample pair includes: In response to receiving the training echo data, call the image acquisition device to acquire the video stream corresponding to the object to be recognized; Obtain the time stamp corresponding to each video frame in the video stream; Align the training echo data and the training image data according to the time stamp to generate the training sample pair.

9. The method according to claim 1, wherein, The transmitting an ultrasonic signal to the object to be recognized through the acoustic wave transmitting device of the terminal includes: Determine the private operation permission in response to the recognition instruction; Call the acoustic wave transmitting device based on the private operation permission, so that the acoustic wave transmitting device transmits a single-channel ultrasonic wave.

10. The method according to any one of claims 1-9, wherein, The method further includes: Obtain a preset image saved in the target application corresponding to the target model; Compare the face image information with the preset image to obtain comparison information; Indicate the execution of the target application based on the comparison information.

11. The method according to claim 10, wherein, The method further includes: Obtain the position information and depth of field information corresponding to the face image information; Verify the face image information based on the position information and the depth of field information to determine the object recognition information; Indicate the execution of the target application according to the object recognition information.

12. The method according to claim 1, wherein, the ultrasonic information is single-channel ultrasonic, the terminal is a mobile phone, the sound wave transmitting device is a speaker, and the sound wave receiving device is a microphone.

13. An object recognition device based on ultrasonic echo, wherein, comprising: a transmitting unit configured to transmit an ultrasonic signal to an object to be recognized through a sound wave transmitting device of a terminal; a receiving unit configured to receive an echo signal reflected by the object to be recognized through the sound wave receiving device of the terminal, the echo signal corresponding to the ultrasonic signal; an extraction unit configured to perform vector normalization processing on the echo signal to extract ultrasonic echo features corresponding to the echo signal; a conversion unit configured to call a target model to perform feature dimension conversion on the ultrasonic echo features to obtain target dimension features for characterizing the object to be recognized, wherein the ultrasonic echo features are processed to obtain frequency domain and time domain information, and the frequency domain and time domain information are input into the target model to be converted into high-dimensional sound wave features, and the high-dimensional sound wave features are the target dimension features; an identification unit configured to perform face translation processing on the target dimension features of the object to be recognized to obtain face image information corresponding to the object to be recognized.

14. The device according to claim 13, wherein, the identification unit is specifically configured to obtain training image data and training echo data collected for the object to be recognized; input the training image data into an image feature extraction network in a preset model to determine training image features in response to the setting of a target dimension; generate corresponding spectrum information based on the training echo data to determine training echo features; input the training echo features into a sound wave conversion network in the preset model to perform conversion on the training echo data based on the target dimension to obtain training sound wave features, the training sound wave features having the same dimension as the training image features; use the training image features as a training target to adjust the training sound wave features to obtain a similarity loss; and adjust parameters of the sound wave conversion network in the preset model based on the similarity loss to obtain the target model.

15. The device according to claim 14, wherein, the identification unit is specifically configured to: fix parameters of the image feature extraction network in the preset model; in response to the fixing of the parameters of the image feature extraction network in the preset model, adjust parameters of the sound wave conversion network in the preset model based on the similarity loss to obtain the target model.

16. The device according to claim 15, wherein, the identification unit is specifically configured to: obtain pre-training data for indicating the correspondence between pre-training images and pre-training features; train the image feature extraction network in the preset model based on the pre-training data to fix parameters of the image feature extraction network in the preset model.

17. The device according to claim 14, It is characterized in that The recognition unit is specifically configured to: Control the acoustic wave emitting device to emit single-channel ultrasonic waves to the object to be recognized based on a preset frequency feature; Obtain the acoustic wave data obtained by reflection received by the microphone to determine the training echo data; In response to the reception of the training echo data, call the image acquisition device in the terminal to acquire the training image data corresponding to the object to be recognized; Align the training echo data and the training image data based on the execution time to generate a training sample pair; The recognition unit adjusts the parameters of the acoustic wave conversion network in the preset model based on the similarity loss to obtain the target model, specifically for: Determine the similarity loss based on the training sample pair to adjust the parameters of the acoustic wave conversion network in the preset model to obtain the target model.

18. The device according to claim 17, It is characterized in that The recognition unit is specifically configured to: Based on the acoustic wave data obtained by reflection received by the acoustic wave receiving device; Filter the acoustic wave data according to the preset frequency feature to determine the acoustic wave data reflected by the object to be recognized; Determine the training echo data based on the acoustic wave data reflected by the object to be recognized.

19. The device according to claim 18, It is characterized in that The recognition unit is specifically configured to: Segment the acoustic wave data into multiple fixed-length acoustic wave sequences according to the preset frequency feature; Filter the acoustic wave sequences based on a preset filter to obtain a filtered sequence; Perform normalization processing on the filtered sequence according to a sliding window to determine the acoustic wave data reflected by the object to be recognized.

20. The device according to claim 17, It is characterized in that The recognition unit is specifically configured to: In response to the reception of the training echo data, call the image acquisition device to acquire the video stream corresponding to the object to be recognized; Obtain the time stamps corresponding to each video frame in the video stream; Align the training echo data and the training image data according to the time stamps to generate the training sample pair.

21. The device according to claim 13, It is characterized in that The transmitting unit is specifically configured to: Determine the private operation permission in response to the recognition instruction; Call the acoustic wave emitting device based on the private operation permission so that the acoustic wave emitting device emits single-channel ultrasonic waves.

22. The device according to any one of claims 13-21, It is characterized in that The recognition unit is specifically configured to: Obtain the preset image saved in the target application corresponding to the target model; Compare the face image information with the preset image to obtain comparison information; Indicate the execution of the target application based on the comparison information.

23. The device according to claim 22, It is characterized in that The recognition unit is specifically configured to: Obtain the position information and depth of field information corresponding to the face image information; Verify the face image information based on the position information and the depth of field information to determine the object recognition information; Indicate the execution of the target application according to the object recognition information.

24. A computer device, characterized in that, the computer device includes a processor and a memory: the memory is used for storing program code; the processor is used for executing the object recognition method based on ultrasonic echo according to any one of claims 1 to 12 according to the instructions in the program code.

25. A computer-readable storage medium, in which instructions are stored, and when the instructions run on a computer, the computer is caused to execute the object recognition method based on ultrasonic echo according to any one of claims 1 to 12 above.

26. A computer program product, characterized in that, the computer program product includes computer instructions, and the processor of the computer device executes the computer instructions, so that the computer device executes the object recognition method based on ultrasonic echo according to any one of claims 1 to 12 above.

Citation Information

Patent Citations

  • Imaging method and device based on ultrasonic echo signals, storage medium and electronic device

    CN111444830A