Semantic recognition method and related apparatus

By acquiring user vocal cord vibration frequency data by radar and performing time-frequency transformation and semantic recognition, the problem of recognizing user semantics in noisy or dimly lit environments by electronic devices is solved, enabling accurate recognition of user semantics in these environments and improving call quality.

CN119274556BActive Publication Date: 2026-04-07HONOR DEVICE CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-06
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Electronic devices struggle to accurately interpret user semantics in noisy, dimly lit, or obstructed environments, leading to a decline in call quality.

Method used

It uses radar to acquire the vibration frequency data of the user's vocal cords, extracts useful signals through time-frequency transformation and preprocessing techniques, and combines convolutional neural networks for semantic recognition. It supports local or cloud processing to improve recognition accuracy and privacy.

Benefits of technology

Accurately identify user semantics in noisy, dimly lit, or obstructed environments to improve call quality and enhance user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119274556B_ABST
    Figure CN119274556B_ABST
Patent Text Reader

Abstract

The semantic recognition method and related apparatus provided in this application relate to the field of terminal technology. The method includes: acquiring data such as the vibration frequency of a user's vocal cords based on the radar of an electronic device, and parsing the radar-acquired data to obtain the semantics the user wants to express. Using radar to detect vocal cord vibration offers good privacy and transmission capabilities. This allows the electronic device to not rely entirely on processing voice signals or lip movements, enabling more accurate semantic recognition even in noisy, dimly lit, or obstructed environments, improving call quality and thus enhancing the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of terminal technology, and in particular to semantic recognition methods and related devices. Background Technology

[0002] In some scenarios, electronic devices can acquire a user's voice and perform corresponding functions, such as in phone calls and live streaming.

[0003] However, due to environmental factors, the user's voice received by the electronic device may be mixed with other noises; or the user's voice may be too soft, making it difficult for the electronic device to clearly capture the user's voice, thus reducing the call quality. Summary of the Invention

[0004] The semantic recognition method and related apparatus provided in this application can acquire data such as the vibration frequency of a user's vocal cords based on the radar of an electronic device, and analyze the data acquired by the radar to obtain the semantics that the user wants to express, thereby improving call quality and enhancing user experience.

[0005] In a first aspect, the semantic recognition method provided in the embodiments of this application is applied to an electronic device, including a radar, and the method includes:

[0006] During the collection of user audio and video data, the vibration frequency of the user's vocal cords is obtained based on radar; a first sound recognition result is obtained by semantically recognizing the vibration frequency of the vocal cords; a target sound recognition result is obtained based on the first sound recognition result; the target sound recognition result is either the first sound recognition result, or the target sound recognition result is obtained by processing the first sound recognition result and the second audio recognition result, or the target sound recognition result is obtained by processing the first sound recognition result, the second audio recognition result, and the third image recognition result, wherein the second audio recognition result is obtained by semantically recognizing the audio in the audio and video data, and the third image recognition result is obtained by lip-reading recognition of the user in the audio and video data; the target sound recognition result is displayed. In this way, using radar to detect vocal cord vibration has good privacy and transmission capability, and electronic devices do not need to rely entirely on the processing of speech signals or lip movements, enabling relatively accurate recognition of user semantics even in noisy, dimly lit, or obstructed environments.

[0007] In one possible implementation, before obtaining the first sound recognition result by semantically recognizing the vibration frequency of the vocal cords, the method further includes: preprocessing the vibration frequency of the vocal cords in the range dimension to obtain a time-frequency map corresponding to the vibration frequency of the vocal cords; and performing semantic recognition on the time-frequency map corresponding to the vibration frequency of the vocal cords. In this way, the radar can search for the vibration frequency of the vocal cords in different range dimensions and accumulate the searched information, thereby reducing the dimensionality of the data to obtain a more complete set of vocal cord vibration frequencies, facilitating further extraction of useful signals from the received signal.

[0008] In one possible implementation, the radar is used to acquire electromagnetic wave signals, which include the vibration frequency of the vocal cords. Preprocessing of the vocal cord vibration frequency in the range dimension includes: determining a first location with the highest energy in the electromagnetic wave signal within the radar search range; searching for a second location centered on the first location towards the boundary of the radar search range; retaining the signal corresponding to the second location if the Rayleigh entropy of the second location is less than a preset value; or, not retaining the signal corresponding to the second location if the Rayleigh entropy of the second location is greater than or equal to the preset value; and outputting a time-frequency map based on the retained signal, which includes the time-frequency map corresponding to the vocal cord vibration frequency. In this way, the radar can select the location with the strongest energy in the received signal as the target center point and search around the target center point to obtain more complete and useful information, facilitating further extraction of the vocal cord vibration frequency signal from the received signal.

[0009] In one possible implementation, before determining the first location of maximum energy in the electromagnetic wave signal, the process further includes: performing time-frequency transformation on different distance dimensions of the electromagnetic wave signal to obtain a time-frequency distribution map (STFT) for each distance dimension; and determining a preset value based on the Rayleigh entropy of the FTFR, where the preset value includes the average of the sum of the Rayleigh entropies of some or all FTFRs. Thus, after time-frequency transformation, the target signal exhibits better local focusing in time and frequency, making the distinction between the target signal and clutter more obvious, thereby obtaining a clearer and more concentrated energy distribution of the target signal.

[0010] In one possible implementation, before preprocessing the vocal cord vibration frequency in the distance dimension, the electromagnetic wave signal is further processed by performing one or more of the following: moving target detection, constant false alarm rate (CFAR) detection, multipath suppression, and time-frequency focusing. Moving target detection filters signals with frequencies of zero or less than a preset frequency value from the electromagnetic wave signal; CFAR detection filters signals with low frequency energy from the electromagnetic wave signal; multipath suppression filters echo signals from the electromagnetic wave signal, including signals reflected back from radar transmission signals when encountering obstacles; and time-frequency focusing enhances the time-frequency focusing capability of the vocal cord vibration frequency in the electromagnetic wave signal. Thus, moving target detection can filter signals with frequencies of zero or near the frequency period from the electromagnetic wave signal, thereby highlighting the target signal. CFAR detection can filter signals with low frequency energy from the electromagnetic wave signal and suppress clutter in the received signal, thereby more accurately identifying the target signal. Multipath suppression can filter echo signals from the electromagnetic wave signal, reducing interference from echo signals to the target signal. Time-frequency focusing processing can utilize the time-frequency focusing difference between the target signal and clutter to enhance the time-frequency focusing capability of the vocal cord vibration frequency in the electromagnetic wave signal, thereby improving the time-frequency focusing capability of the target signal and further enhancing the time-frequency extraction effect of the target signal.

[0011] In one possible implementation, before obtaining the first sound recognition result of semantic recognition of the vocal cord vibration frequency, the method further includes: displaying a first interface, the first interface including one or more of the following options: a first option, a second option, a third option, and a fourth option; wherein, the first option is used to indicate the real-time requirement for semantic recognition of the vocal cord vibration frequency, the first option includes a first level and a second level, the first level having a higher real-time requirement than the second level; the second option is used to indicate the accuracy requirement for semantic recognition of the vocal cord vibration frequency, the second option includes a third level and a fourth level, the third level having a higher accuracy requirement than the fourth level; the third option is used to indicate the accuracy requirement for semantic recognition of the vocal cord vibration frequency. For the privacy requirements of semantic recognition based on vibration frequency, the third option includes levels five and six, with level five having higher privacy requirements than level six. The fourth option indicates the bandwidth consumption requirements for semantic recognition of vocal cord vibration frequency, including levels seven and eight, with level seven having higher bandwidth requirements than level eight. When level one, four, five, or seven is selected, levels two, three, six, and / or eight are unselectable. Similarly, when level two, three, six, and / or eight is selected, levels one, four, five, and / or seven are unselectable. This can lead to conflicts between options on the first interface, for example, requiring both high real-time performance and strict accuracy. To reduce these conflicts, after a user selects an option, the electronic device can disable conflicting options, such as by graying them out, preventing the user from simultaneously selecting conflicting options and improving the rationality of the electronic device's execution logic.

[0012] In one possible implementation, upon receiving a selection operation for a first, fourth, fifth, or seventh level, a first sound recognition result is obtained by semantically recognizing the vibration frequency of the vocal cords. This includes: using a convolutional neural network to perform semantic recognition of the vocal cord vibration frequency to obtain the first sound recognition result. In this way, the electronic device does not need to upload data to a cloud server, saving the bandwidth consumed by data uploads. Since local processing does not require additional time spent uploading to a cloud server, data processing on the electronic device itself allows for faster completion of call content recognition.

[0013] In one possible implementation, upon receiving a selection operation for a second, third, sixth, and / or eighth level, a first sound recognition result is obtained by semantically recognizing the vibration frequency of the vocal cords. This includes: uploading the vibration frequency of the vocal cords to a cloud server, and obtaining the first sound recognition result from the cloud server, wherein the first sound recognition result is obtained by the cloud server using a convolutional neural network to perform semantic recognition on the vibration frequency of the vocal cords. Thus, due to the strong computing power and / or storage capacity of the cloud server, its data processing capabilities are stronger, enabling more accurate analysis and recognition of the user's speech content, thereby improving the user experience.

[0014] In one possible implementation, before obtaining the first sound recognition result for semantic recognition of vocal cord vibration frequency, the method further includes: performing principal component analysis (PCA) on the vocal cord vibration frequency to extract characteristic data related to the recognition semantics. This characteristic data includes one or more of the following: the highest frequency of vocal cord vibration, the lowest frequency of vocal cord vibration, the ratio of the highest to the lowest frequency, and the duration of vocal cord vibration frequency. Obtaining the first sound recognition result for semantic recognition of vocal cord vibration frequency includes obtaining the first sound recognition result for semantic recognition of the characteristic data. In this way, PCA can filter out more important features from the data, thereby reducing the amount of data while preserving as much information as possible from the original data. This allows for the extraction of user-related characteristic data and the filtering of user-irrelevant features, thus enabling faster extraction of useful information from the received signal.

[0015] Secondly, embodiments of this application provide a semantic recognition apparatus, which may be an electronic device, a chip within an electronic device, or a chip system. The apparatus may include a processing unit and a display unit. The processing unit is used to implement any processing-related method executed by the electronic device in the first aspect or any possible implementation of the first aspect. The display unit is used to implement any display-related method executed by the electronic device in the first aspect or any possible implementation of the first aspect. When the apparatus is an electronic device, the processing unit may be a processor. The apparatus may further include a storage unit, which may be a memory. The storage unit is used to store instructions, and the processing unit executes the instructions stored in the storage unit to cause the electronic device to implement the methods described in the first aspect or any possible implementation of the first aspect. When the apparatus is a chip within an electronic device, the processing unit may be a processor. The processing unit executes the instructions stored in the storage unit to cause the electronic device to implement the methods described in the first aspect or any possible implementation of the first aspect. The storage unit may be a storage unit within the chip (e.g., a register, cache, etc.), or a storage unit located outside the chip within the electronic device (e.g., read-only memory, random access memory, etc.).

[0016] For example, the processing unit is used to acquire the vibration frequency of the user's vocal cords; it is also used to obtain a first sound recognition result by semantically recognizing the vibration frequency of the vocal cords; specifically, it is also used to obtain a target sound recognition result of the audio and video data based on the first sound recognition result. The display unit is used to display the target sound recognition result.

[0017] In one possible implementation, the processing unit is used to preprocess the vibration frequency of the vocal cords in the distance dimension to obtain a time-frequency diagram corresponding to the vibration frequency of the vocal cords; it is also used to perform semantic recognition on the time-frequency diagram corresponding to the vibration frequency of the vocal cords.

[0018] In one possible implementation, the processing unit is configured to determine a first position with the maximum energy in the electromagnetic wave signal within the radar search range; and to search for a second position centered on the first position in the direction of the boundary of the radar search range; specifically, if the Rayleigh entropy of the second position is less than a preset value, then retain the signal corresponding to the second position; or, if the Rayleigh entropy of the second position is greater than or equal to the preset value, then do not retain the signal corresponding to the second position; and to output a time-frequency diagram based on the retained signal.

[0019] In one possible implementation, the processing unit is used to perform time-frequency transformation on different distance dimensions of the electromagnetic wave signal to obtain a time-frequency distribution map STFT for each distance dimension; it is also used to determine a preset value based on Rayleigh entropy of FTFR.

[0020] In one possible implementation, the processing unit is used to perform one or more of the following processes on the electromagnetic wave signal: moving target detection, constant false alarm rate detection, multipath suppression processing, and time-frequency focusing processing.

[0021] In one possible implementation, a display unit is used to display the first interface.

[0022] In one possible implementation, the processing unit is used to perform semantic recognition on the vibration frequency of the vocal cords using a convolutional neural network to obtain a first sound recognition result.

[0023] In one possible implementation, the processing unit is used to upload the vibration frequency of the vocal cords to a cloud server and obtain a first sound recognition result from the cloud server for semantic recognition of the vibration frequency of the vocal cords.

[0024] In one possible implementation, the processing unit is used to perform principal component analysis (PCA) on the vibration frequency of the vocal cords to extract characteristic data related to the semantic recognition; it is also used to obtain a first sound recognition result based on the semantic recognition of the characteristic data.

[0025] Thirdly, embodiments of this application provide an electronic device including one or more processors and a memory, the memory being coupled to one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, and one or more processors calling the computer instructions to cause the electronic device to perform the methods described in the first aspect or any possible implementation of the first aspect.

[0026] Fourthly, this application provides a chip or chip system applied to an electronic device. The chip or chip system includes one or more processors and a communication interface. The communication interface and at least one processor are interconnected via a circuit. The one or more processors are used to invoke computer instructions to cause the electronic device to execute the methods described in the first aspect or any possible implementation thereof. The communication interface in the chip can be an input / output interface, pins, or circuits, etc.

[0027] In one possible implementation, the chip or chip system described above in this application further includes at least one memory storing instructions. The memory can be an internal storage unit of the chip, such as a register or cache, or it can be a storage unit of the chip itself (e.g., read-only memory, random access memory, etc.).

[0028] Fifthly, embodiments of this application provide a computer-readable storage medium including computer instructions that, when executed on an electronic device, cause the electronic device to perform the methods described in the first aspect or any possible implementation thereof.

[0029] In a sixth aspect, embodiments of this application provide a computer program product including computer program code, which, when run on an electronic device, causes the electronic device to perform the methods described in the first aspect or any possible implementation thereof.

[0030] It should be understood that the second to sixth aspects of this application correspond to the technical solutions of the first aspect of this application, and the beneficial effects achieved by each aspect and the corresponding feasible implementation are similar, and will not be repeated here. Attached Figure Description

[0031] Figure 1 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0032] Figure 2 A schematic diagram of the software structure of an electronic device provided in an embodiment of this application;

[0033] Figure 3 A schematic diagram illustrating radar acquisition of the vibration frequency of the vocal cords, provided as an embodiment of this application;

[0034] Figure 4 A flowchart illustrating a semantic recognition method provided in an embodiment of this application;

[0035] Figure 5 A schematic diagram of a radar function interface provided in an embodiment of this application;

[0036] Figure 6 A flowchart for determining the correlation between a target signal and a suspected multipath signal is provided in an embodiment of this application;

[0037] Figure 7 A schematic diagram illustrating a time-frequency extraction range provided in an embodiment of this application;

[0038] Figure 8 A flowchart of time-frequency extraction provided in an embodiment of this application;

[0039] Figure 9 A schematic diagram illustrating a semantic recognition method provided in an embodiment of this application;

[0040] Figure 10 This is a schematic diagram of the structure of a chip provided in an embodiment of this application. Detailed Implementation

[0041] To facilitate a clear description of the technical solutions in the embodiments of this application, some terms and technologies involved in the embodiments of this application will be briefly introduced below:

[0042] 1. Terminology

[0043] In the embodiments of this application, terms such as "first" and "second" are used to distinguish identical or similar items with substantially the same function and purpose. For example, "first chip" and "second chip" are used only to distinguish different chips and do not limit their order of execution. Those skilled in the art will understand that terms such as "first" and "second" do not limit the quantity or execution order, and that "first" and "second" do not necessarily imply that they are different.

[0044] It should be noted that, in the embodiments of this application, the terms "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design scheme described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0045] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can represent: a, b, c, ab, a--c, bc, or abc, where a, b, and c can be single or multiple.

[0046] 2. Electronic equipment

[0047] The electronic devices in this application embodiment can also be any form of terminal device. For example, electronic devices may include: mobile phones, tablet computers, handheld computers, laptops, mobile internet devices (MIDs), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminals in industrial control, wireless terminals in self-driving, wireless terminals in remote medical surgery, wireless terminals in smart grids, wireless terminals in transportation safety, wireless terminals in smart cities, wireless terminals in smart homes, cellular phones, cordless phones, session initiation protocol (SIP) phones, wireless local loop (WLL) stations, personal digital assistants (PDAs), handheld devices with wireless communication capabilities, computing devices or other processing devices connected to a wireless modem, in-vehicle devices, wearable devices, electronic devices in 5G networks, or future evolved public land mobile communication networks (PLANs). The embodiments of this application do not limit the scope of electronic devices in a mobile network (PLMN).

[0048] By way of example and not limitation, in this embodiment, the electronic device can also be a wearable device. Wearable devices, also known as wearable smart devices, are a general term for devices that utilize wearable technology to intelligently design and develop everyday wearables, such as glasses, gloves, watches, clothing, and shoes. Wearable devices are portable devices that are worn directly on the body or integrated into the user's clothing or accessories. Wearable devices are not merely hardware devices, but also achieve powerful functions through software support, data interaction, and cloud interaction. Broadly speaking, wearable smart devices include those that are feature-rich, large in size, and can achieve complete or partial functions without relying on a smartphone, such as smartwatches or smart glasses, as well as those that focus on a specific type of application function and require the use of other devices such as smartphones, such as various smart bracelets and smart jewelry for vital sign monitoring.

[0049] Furthermore, in this application embodiment, the electronic device can also be an electronic device in the Internet of Things (IoT) system. IoT is an important part of the future development of information technology. Its main technical feature is to connect objects to the network through communication technology, thereby realizing an intelligent network of human-machine interconnection and object-to-object interconnection.

[0050] The electronic equipment in the embodiments of this application may also be referred to as: user equipment (UE), mobile station (MS), mobile terminal (MT), access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent, or user device, etc.

[0051] In this embodiment, the electronic device or various network devices include a hardware layer, an operating system layer running on top of the hardware layer, and an application layer running on top of the operating system layer. The hardware layer includes hardware such as a central processing unit (CPU), a memory management unit (MMU), and memory (also called main memory). The operating system can be any one or more computer operating systems that implement business processing through processes, such as Linux, Unix, Android, iOS, or Windows. The application layer includes applications such as browsers, address books, word processing software, and instant messaging software.

[0052] For example, Figure 1 A schematic diagram of the electronic device is shown.

[0053] The electronic device may include a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, a headphone jack 170D, a sensor module 180, buttons 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0054] It is understood that the structures illustrated in the embodiments of this application do not constitute a specific limitation on the electronic device. In other embodiments of this application, the electronic device may include more or fewer components than illustrated, or combine some components, or split some components, or have different component arrangements. The illustrated components may include hardware, software, or a combination of software and hardware.

[0055] Processor 110 may include one or more processing units, such as application processors (APs), modem processors, graphics processing units (GPUs), image signal processors (ISPs), controllers, video codecs, digital signal processors (DSPs), baseband processors, and / or neural network processing units (NPUs). These different processing units may be independent devices or integrated into one or more processors. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.

[0056] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can directly retrieve it from the aforementioned memory. This avoids repeated access, reduces the waiting time of the processor 110, and thus improves the efficiency of the system. For example, in the embodiments of this application, the processor 110 can be used for processing radar data, determining real-time requirements, and recognizing call semantics, etc.

[0057] It is understood that the interface connection relationships between the modules illustrated in the embodiments of this application are merely illustrative and do not constitute a limitation on the structure of the electronic device. In other embodiments of this application, the electronic device may also employ different interface connection methods or combinations of multiple interface connection methods as described in the above embodiments.

[0058] Internal memory 121 can be used to store computer executable program code, including instructions. Internal memory 121 may include a program storage area and a data storage area. The program storage area may store the operating system, applications required for at least one function, etc. The data storage area may store data created during the use of the electronic device, etc. Furthermore, internal memory 121 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, universal flash storage (UFS), etc. Processor 110 executes various functional applications and data processing of the electronic device by running instructions stored in internal memory 121 and / or instructions stored in memory disposed in the processor. For example, in this embodiment, internal memory 121 may be used to store radar data and related code for processing the radar data, etc.

[0059] Camera 193 is used to capture still images or videos. In some embodiments, the electronic device may include one or N cameras 193, where N is a positive integer greater than 1. For example, in this embodiment, the camera can acquire changes in the user's lip movements to identify the semantics corresponding to the lip movements.

[0060] Figure 2This is a software structure block diagram of an electronic device according to an embodiment of this application. The layered architecture divides the software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers, from top to bottom: the application layer, the application framework layer, the Android runtime and system libraries, the hardware adaptation layer (HAL), and the kernel layer.

[0061] The application layer, also known as the application layer, can include a series of application packages. For example... Figure 2 As shown, the application package can include applications such as phone, music, and camera. Applications can include system applications and third-party applications.

[0062] The application framework layer, also known as the framework layer, provides application programming interfaces (APIs) and programming frameworks for applications in the application layer. The framework layer can include some predefined functions.

[0063] like Figure 2 As shown, the Framework layer can include the Activity Manager, Window Manager, Resource Manager, Notification Manager, Content Provider, and View System, etc.

[0064] The Android runtime consists of core libraries and a virtual machine. The Android runtime is responsible for the control and management of the Android system.

[0065] The core library consists of two parts: one part is the functionalities that need to be called by the Java language, and the other part is the Android core library.

[0066] The application layer and framework layer run in a virtual machine. The virtual machine executes the Java files of the application layer and framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection. For example, in the embodiments of this application, the virtual machine can be used to perform functions such as radar data acquisition and radar data analysis.

[0067] The system library, also known as the native layer, can include multiple functional modules. Examples include media libraries, function libraries, and graphics processing libraries.

[0068] The Hardware Abstraction Layer (HAL) is a layer of abstraction located between the kernel layer and the Android runtime. The HAL can be a wrapper around hardware drivers, providing a unified interface for calls from upper-layer applications.

[0069] The kernel layer is the layer between hardware and software. The kernel layer can include display drivers, camera drivers, audio drivers, etc.

[0070] It should be noted that the embodiments of this application are only illustrated using the Android system. In other operating systems (such as Windows system, iOS system, etc.), as long as the functions implemented by each functional module are similar to those in the embodiments of this application, the solution of this application can also be implemented.

[0071] In scenarios such as phone calls or live streaming, electronic devices may be affected by environmental factors, causing the received user's voice to be mixed with other noise, thus reducing call quality. Alternatively, when making calls in a quiet environment, users may speak softly to avoid disturbing others, and the electronic device may not be able to clearly pick up the user's voice, also reducing call quality.

[0072] In some implementations, electronic devices can use camera-based lip-reading recognition technology to identify user semantics in order to improve call quality. However, in low-light conditions or when users are wearing masks or other obstructions, cameras cannot recognize user semantics using lip-reading recognition technology, thus affecting the accuracy of lip-reading recognition.

[0073] In view of this, the semantic recognition method provided in this application can acquire data such as the vibration frequency of a user's vocal cords based on the radar of an electronic device, and analyze the data acquired by the radar to obtain the semantics that the user wants to express. Using radar to detect vocal cord vibration has good privacy and transmission capabilities. Thus, the electronic device does not need to rely entirely on processing voice signals or lip movements, enabling more accurate recognition of the user's semantics even in noisy, dimly lit, or obstructed environments, improving call quality and thus enhancing the user experience.

[0074] Taking a phone call as an example, Figure 3 A schematic diagram illustrating radar acquiring the vibration frequency of the vocal cords is shown. Electronic devices may include radar, which includes a transmitting antenna and a receiving antenna. During a call, the transmitting antenna can be used to transmit electromagnetic wave signals, and the receiving antenna can be used to receive the returned electromagnetic wave signals.

[0075] In some scenarios, the transmitted electromagnetic wave signal can be called the transmitted signal, and the returned electromagnetic wave signal can be called the echo data or the received signal. For ease of description, the following explanation will use transmitted and received signals as examples.

[0076] It is understandable that the received signal can include the reflected transmitted signal, as well as the vibration frequencies of multiple parts of the user's body, such as the vibration frequency of the user's vocal cords, breathing, and chest cavity during a conversation. The received signal can also include the vibration frequencies of the surrounding environment, such as the vibration frequencies reflected by static objects like walls. In some scenarios, vibration frequency can also be referred to as vibration velocity; for ease of description, vibration frequency will be used as an example in the following explanations.

[0077] Electronic devices can acquire data such as the vibration frequency of a user's vocal cords based on radar, and then analyze this data to obtain the semantics that the user wants to express.

[0078] For example, Figure 4 A flowchart of a semantic recognition method according to an embodiment of this application is shown.

[0079] S401. Turn on radar function and select mode.

[0080] The semantic recognition method of this application embodiment can be applied not only to call scenarios but also to live streaming scenarios, silent call scenarios, etc. A silent call scenario can be understood as a scenario where the electronic device cannot obtain sound or the obtained sound is too low to recognize content through sound. Therefore, the interface of the electronic device can include different mode options such as call mode, live streaming mode, and silent call mode. On the one hand, users can select the appropriate mode according to their needs. Based on the different modes selected by the user, the electronic device can execute different semantic recognition processes. On the other hand, the electronic device can also recognize the current usage scenario; for example, the electronic device can recognize that the current scenario is a live streaming scenario or a call scenario, and depending on the usage scenario, the electronic device can execute different semantic recognition processes.

[0081] For ease of description, the following explanation will use a phone call scenario as an example.

[0082] When using radar to identify user semantics, electronic devices can provide users with options related to radar functionality. For example... Figure 5 As shown, the interface 501 of the electronic device is the radar function interface, which can include relevant options for radar functions. For example, it can include real-time options, accuracy options, privacy options, and data-saving mode options. The electronic device or the user can select the appropriate options according to the needs of the actual scenario.

[0083] The real-time options can include fast, medium, and slow. For example, a fast real-time option indicates a high demand for real-time performance; the electronic device can identify user semantics from radar data as quickly as possible, but this may consume more data traffic or result in relatively lower accuracy. A slow real-time option indicates a lower demand for real-time performance; the electronic device can accurately identify user semantics from radar data, but the identification speed may be slower. A medium real-time option means the electronic device's recognition performance falls between the fast and slow options, and will not be elaborated further.

[0084] Accuracy options can include strict, accurate, and auxiliary options. For example, when the accuracy option is strict, it indicates a high requirement for accuracy in recognizing call content based on radar data, such as in silent call scenarios. In this case, the electronic device can upload the data to a cloud server for relatively accurate semantic recognition based on the radar data. When the accuracy option is auxiliary, it indicates a lower requirement for accuracy in recognizing call content based on radar data. The electronic device can primarily rely on the voice signals received by the microphone for call content recognition, while semantic recognition based on radar data can be used as an auxiliary method. In this case, the accuracy of semantic recognition based on radar data by the electronic device can be relatively low. When the real-time option is accurate, the recognition effect of the electronic device on radar data can fall between the strict and auxiliary options, which will not be elaborated further.

[0085] Privacy options may include local processing and cloud processing. Cloud processing can be understood as uploading data to a cloud server for processing. Since cloud processing requires uploading radar data from the electronic device to a cloud server, there is a possibility of data leakage during data transmission. Therefore, when privacy requirements are high, radar data processing can be performed locally on the electronic device. When accurate radar data identification is required and privacy requirements are not too high, radar data processing can be performed in the cloud.

[0086] The data-saving mode options can include options to turn it on and off. Since cloud processing requires uploading radar data from the electronic device to a cloud server, which consumes data, radar data processing can be performed locally on the electronic device when data saving is needed. When accurate radar data identification is required and data saving is not a primary concern, radar data processing can be performed in the cloud.

[0087] It is understandable that the selection of various options in interface 501 may conflict, for example, requiring both high real-time performance and strict accuracy. In a possible implementation, to reduce conflicts between options, after a user selects an option, the electronic device can set other conflicting options to an unselectable state, such as graying them out, preventing the user from simultaneously selecting conflicting options, thereby improving the rationality of the electronic device's execution logic. Of course, the electronic device can also use other methods to reduce conflicts between options, and this application embodiment does not limit this approach.

[0088] It is understood that interface 501 is an exemplary interface, and different electronic devices may have different radar function interfaces. Interface 501 may include more or less content. The specific content included in interface 501 is not limited in this application embodiment.

[0089] S402, Radar Data Analysis.

[0090] Electronic devices can analyze and process radar data. In possible implementations, the radar's transmitted signal can take the form of a triangular wave pulse or a sawtooth wave pulse, etc., and this application's embodiments are not limited to these forms. Taking a sawtooth wave pulse as an example, the radar's transmitted signal S... T (t) can satisfy the following formula:

[0091]

[0092] Among them, A T denoted as the amplitude value of the transmitted signal, j as the imaginary unit, f0 as the starting frequency of the transmitted signal, t as the transmission time, and S as the frequency modulation slope of the transmitted signal.

[0093] Then the phase of the corresponding transmitted signal The following formula can be satisfied:

[0094]

[0095] Radar received signal S R (t) can satisfy the following formula:

[0096]

[0097] Among them, A R τ represents the amplitude of the received signal, and τ represents the time delay of the received signal relative to the transmitted signal.

[0098] It is understandable that, since the received signal is later than the transmitted signal, there can be a certain time delay between the transmitted and received signals. This time delay τ can satisfy the following formula:

[0099]

[0100] Where r is the radial distance between the target and the radar, and in this embodiment, the target can be understood as the user. v is the vibration frequency of the vocal cords, which can also be understood as the radial velocity of the target relative to the radar, and c is the propagation speed of electromagnetic waves, which can be taken as the speed of light c = 3 × 10 8 m / s.

[0101] Then the phase of the corresponding received signal The following formula can be satisfied:

[0102]

[0103] It is understandable that, due to signal loss during transmission, the amplitude of the received signal may not be equal to the amplitude of the transmitted signal, and there may be a phase difference between the transmitted and received signals. In some scenarios, this phase difference can also be called the difference-frequency phase. For ease of description, the difference-frequency phase will be used as an example in the following explanation.

[0104] For example, the difference frequency phase The following formula can be satisfied:

[0105]

[0106] Difference frequency signal S IF (t) can satisfy the following formula:

[0107]

[0108] Among them, A IF This represents the amplitude value of the difference frequency signal.

[0109] Taking the vibration frequency v of the vocal cords in the range of 100 Hz to 1000 Hz as an example, since the vibration frequency v of the human vocal cords differs significantly from the speed of light c, after squaring the time delay τ, the difference between τ and τc becomes significant. 2 The orders of magnitude difference between them are even greater, therefore, the above equation S can be simplified. IF The last term in (t) is the difference frequency signal S. IF (t) can satisfy the following formula:

[0110] s IF (t)=A IF exp(j2π(f o τ+Stτ)).

[0111] For the difference frequency signal S IF Performing a Fourier transform on (t) yields the difference frequency signal in the spectrum, f. IF The following formula can be satisfied:

[0112]

[0113] Where λ is the wavelength of the electromagnetic wave.

[0114] Since the vibration frequency v of the vocal tract differs significantly from the speed of light c by orders of magnitude, the first term in the above equation can be simplified to obtain the difference frequency signal f in the following frequency spectrum. IF :

[0115]

[0116] Furthermore, a fast fourier transform (FFT) in the range dimension can be performed on the radar's received signal to obtain the spatial distribution of the received signal. Then, an FFT in the velocity dimension can be performed to obtain the velocity information of the received signal, which may include the vibration frequency of the vocal cords.

[0117] It is understandable that during a call, a user's vocal cords can vibrate at different frequencies when pronouncing different words. Therefore, electronic devices can interpret the semantics that the user wants to express based on the vibration frequency of the vocal cords.

[0118] For example, the vibration frequency v of the vocal cords i The following formula can be satisfied:

[0119]

[0120] Among them, T c It is the duration of each sawtooth pulse. It refers to the difference frequency phase of the vocal cords, which can also be understood as the difference frequency phase mentioned above.

[0121] It is understandable that since the received signal may include the vibration frequencies of multiple parts of the user's body, as well as the vibration frequencies of the surrounding environment, and in a call scenario, more attention is paid to the vibration frequency of the vocal cords, therefore, in this embodiment, the vibration frequency of the vocal cords can be referred to as the target signal, and other vibration frequencies can be referred to as the vibration frequencies of clutter.

[0122] In this embodiment, the electronic device can process the received signal to enhance the target signal as much as possible, such as by performing time-frequency focusing enhancement, and to remove as much noise vibration frequency as possible from the received signal, such as by performing moving target display, constant false alarm rate detection, and multipath suppression. In this way, the electronic device can more accurately analyze the call semantics based on the target signal in the received signal, thereby improving call quality.

[0123] S403, Moving target display.

[0124] When a radar transmitting antenna emits a signal into the surrounding space, the received signal may include the vibration frequencies of clutter reflected by static objects in the surrounding space. For example, in an indoor setting, due to the limited space, the received signal may contain the vibration frequencies of clutter from indoor furniture, walls, floors, ceilings, and other static objects. These clutter frequencies mix with the target signal, potentially masking it and interfering with the identification of the target signal by electronic devices. Therefore, embodiments of this application can perform clutter suppression processing on the received signal.

[0125] It is understandable that static objects hardly vibrate, therefore their vibration frequency can be zero or close to zero. Based on this, embodiments of this application can employ a moving target cancellation method, utilizing the difference in Doppler frequencies between the target signal and the static object to filter out the zero-frequency component of the static object. In other words, signals with frequencies of zero or close to zero can be filtered out, thereby highlighting the target signal.

[0126] In a possible implementation, electronic devices can use recursive cancellers to describe the dynamic behavior and response of signals. For example, the output result of the previous moment can be fed back into the input result of the next moment. By changing the feedback coefficient k, clutter can be suppressed, thereby achieving signal filtering and denoising.

[0127] For example, the output of the previous time step and the input of the next time step can satisfy the following difference equation:

[0128] y(n+1)=x(n+1)-x(n)+ky(n).

[0129] Where n is the previous time step, n+1 is the next time step, x is the input result, y is the output result, and k is the feedback coefficient. It is understood that k can be set by the electronic device based on empirical values; for example, k can take a value between 0 and 1, or other possible values. The specific value of k is not limited in this embodiment.

[0130] The above difference equation can be transformed into an expression in the Z-domain. That is, by expressing the variables in the difference equation in terms of powers of Z, we obtain the following expression in the Z-domain:

[0131] zY(z)=zX(z)-X(z)+kY(z).

[0132] The transfer function H(z) can satisfy the following formula:

[0133]

[0134] The corresponding frequency response can satisfy the following formula:

[0135]

[0136] Where T is the duration and ω is the frequency in the received signal.

[0137] It is understood that the electronic device can multiply the received signal by the frequency response. When the frequency ω is 0, sinω is also 0, and therefore the received signal is also 0. When the frequency ω is close to 0, sinω is also close to 0, and therefore the received signal is also close to 0. ω can also take values ​​near the frequency period, such as π or 2π, and this embodiment does not limit the values. In this way, the electronic device can filter signals with frequencies of zero or close to zero from the received signal.

[0138] S404, Constant False Alarm Rate (CFAR) algorithm.

[0139] After performing step S403 above and filtering the vibration frequency of the static object included in the received signal, the electronic device can further suppress the clutter in the received signal through a constant false alarm rate (CFAR) detection algorithm, thereby more accurately identifying the target signal.

[0140] In a possible implementation, the CFAR algorithm can select a detection unit in the received signal based on experience or randomly. The range near the detection unit can be called the protection unit range, and the area outside the protection unit range can include one or more reference unit ranges. It is understood that the size of the proximity range can be set by the electronic device based on experience; for example, the proximity range can include a range of 1 meter (m) around the detection unit. The specific size of the proximity range is not limited in this embodiment. The size of the reference unit range can also be set by the electronic device based on experience; for example, each reference unit range can include a range of 7 centimeters (cm). The specific size of the reference unit range is not limited in this embodiment.

[0141] It is understandable that the detection unit can be compared with units within the reference unit range, but not with units within the protection unit range.

[0142] In a call scenario, since the energy of the target signal is relatively large compared to the frequency energy of the surrounding environment, when the detection unit compares the frequency energy with that of the units within the reference unit range, if the frequency energy of the detection unit is larger, it is more likely that the detection unit includes the target signal; if the frequency energy of the units within the reference unit range is larger, it is more likely that the units within the reference unit range include the target signal.

[0143] Once the target signal range is determined, the portion outside the target signal range can be suppressed, thereby suppressing clutter.

[0144] For example, the CFAR algorithm can select sampling reference points from each reference cell range and sort multiple sampling reference points according to their frequency energy. The CFAR algorithm can also multiply the sampling reference point by a scaling factor to obtain a clutter detection threshold for the reference cell range containing that sampling reference point. This clutter detection threshold is randomly varied and can be adaptively adjusted based on the average power of its neighborhood.

[0145] The scaling factor can be determined by the electronic device based on experience or scenario. The scaling factor can include a range from 0 to 1, for example, a scaling factor of 0.5. Of course, the scaling factor can also include other values. The specific value of the scaling factor is not limited in this application embodiment.

[0146] Understandably, a larger scaling factor results in more signal filtering, potentially obscuring some target signals and reducing the accuracy of semantic recognition. Conversely, a smaller scaling factor filters out less signal, potentially retaining more clutter and resulting in poor clutter suppression. Therefore, electronic devices can empirically set an appropriate scaling factor to filter out clutter while retaining the target signal, thus achieving both good clutter suppression and high semantic recognition accuracy.

[0147] When a detection unit is compared with a certain reference unit range, if the frequency energy of the detection unit is greater than or equal to the clutter detection threshold of the reference unit range, the signal of the detection unit can be retained; if the frequency energy of the detection unit is less than the clutter detection threshold of the reference unit range, the signal of the detection unit can be suppressed.

[0148] Optionally, after sorting multiple sampling reference points according to the magnitude of their frequency energy, the sampling reference points corresponding to the larger and / or smaller frequency energies in the sort can be removed first. In this way, when calculating the clutter detection threshold, the influence of extreme data on the threshold value can be sorted, thereby setting the clutter detection threshold within a more reasonable range and achieving better detection performance in clutter environments.

[0149] S405, multipath suppression.

[0150] During transmission, radar signals may undergo reflection, refraction, and scattering due to the influence of the surrounding environment, resulting in one or more echo signals. This phenomenon is known as multipath effect. In some scenarios, these echo signals may also be referred to as multipath signals. For ease of description, multipath signals will be used as an example in the following explanations.

[0151] It is understandable that since the target signal contains relatively rich micro-motion feature information, and the multipath signal has similar micro-motion feature information to the target signal, the target signal and the multipath signal have a strong correlation. As a result, the multipath signal will affect the recognition of the micro-motion feature of the target signal. Therefore, it is necessary to eliminate these multipath signals and further reduce the interference of multipath signals on the target signal.

[0152] Micro-motion feature information can be understood as the feature information of minute movements or vibrations contained in the target signal. For example, in fields such as radar, communication, or sonar, the minute movements of the target signal can be described by micro-motion feature information. This micro-motion feature information may include the target's vibration frequency, phase changes, or other motion-related features, such as minute swaying of the limbs or torso, chest rise and fall, throat vibration, etc., which are not limited in the embodiments of this application.

[0153] In a possible implementation, electronic devices can use a correlation coefficient-based multipath suppression method to calculate the correlation coefficient between the multipath signal and the target signal, and determine the similarity between the multipath signal and the target signal.

[0154] like Figure 6 As shown, in the range-velocity (RD) graph of each frame, the electronic device can acquire the target signal. and suspected multipath signals The correlation between the two signals is then calculated. The signal closer to the radar can be considered the target signal, while the signal farther from the radar can be considered a suspected multipath signal. Both the target signal and the suspected multipath signal consist of vectors with range and velocity dimensions. In some scenarios, the RD diagram can also be called the RD spectrum; for ease of description, the RD diagram will be used as an example in the following explanations.

[0155] Electronic devices can normalize the vectors of the target signal and the suspected multipath signal by taking their modulus, and then calculate the correlation between the target signal and the suspected multipath signal to obtain the Pearson correlation coefficient.

[0156] For example, the Pearson correlation coefficient between the target signal and the suspected multipath signal can satisfy the following formula:

[0157]

[0158] in, Let σ be the covariance between the target signal and the suspected multipath signal. X Let σ be the standard deviation of the target signal. Y The standard deviation is the suspected multipath signal.

[0159] The covariance can satisfy the following formula:

[0160]

[0161] Where E represents the expected value, μ x μ is the expected value of the target signal. Y This represents the expected value for a suspected multipath signal.

[0162] Based on the above formula, the correlation coefficient between the target signal and the suspected multipath signal can be obtained, and this correlation coefficient satisfies the following formula:

[0163]

[0164] The electronic device can compare the correlation coefficient of the target signal and the suspected multipath signal with a correlation coefficient threshold, and determine whether the suspected multipath signal is indeed a multipath signal. The correlation coefficient threshold can be set by the electronic device based on experience; for example, the correlation coefficient threshold can be set to 0.7. The specific value of the correlation coefficient threshold is not limited in this embodiment.

[0165] If the correlation coefficient between the target signal and the suspected multipath signal is less than the correlation coefficient threshold, it indicates that the suspected multipath signal is not a multipath signal, and the suspected multipath signal will not be eliminated. If the correlation coefficient between the target signal and the suspected multipath signal is greater than or equal to the correlation coefficient threshold, it indicates that the suspected multipath signal is a multipath signal, and the electronic equipment can eliminate the multipath signal, thereby reducing the influence of the multipath signal on the target signal.

[0166] S406, Time-Frequency Focus Enhancement.

[0167] Time-frequency concentration (TFC) can be understood as the ability of an analysis window to simultaneously focus on a specific frequency component of a signal at different times and frequencies during time-frequency analysis, in order to describe the trend of that component's change over time and frequency.

[0168] Understandably, compared to clutter, the energy of a target signal is more concentrated around a certain frequency, giving it more pronounced time-frequency local characteristics and better time-frequency focusing ability. Clutter, on the other hand, exhibits more complex or dispersed time-frequency characteristics due to multipath effects, terrain conditions, and weather conditions, resulting in poor focusing ability. Therefore, when performing time-frequency analysis on received signals, the difference in time-frequency focusing ability between the target signal and clutter can be utilized to adopt corresponding processing and optimization strategies, thereby improving the time-frequency focusing ability of the target signal and further enhancing the time-frequency extraction effect.

[0169] In a possible implementation, Rayleigh entropy (rényi) can be used to determine time-frequency focusing properties during time-frequency analysis. For example, a smaller Rayleigh entropy indicates that the signal has better local energy focusing properties in the time and frequency domains; a larger Rayleigh entropy indicates that the signal has poorer local energy focusing properties in the time and frequency domains.

[0170] For example, Rayleigh entropy R α (x) can satisfy the following formula:

[0171]

[0172] Where x is a random variable, α is the exponent of Rayleigh entropy, and the range of α can include [0,1] and (1,+∞), P i Let be the probability distribution of variable x, and n be the number of possible values ​​of variable x.

[0173] The embodiments of this application can use Rayleigh entropy to describe the time-frequency representation (TFR).

[0174]

[0175] The Time-Frequency Representation (TFR) can be composed of multiple consecutive frames and is used to describe the frequency variation over time. For each frame's RD diagram, the corresponding time-frequency information can be obtained, denoted as the frame-based time-frequency representation (FTFR). Therefore, by calculating the Rayleigh entropy of the FTFR, the energy distribution of the corresponding distance dimension of the current frame can be obtained. Furthermore, the STFT processing result can be adjusted using a synchronous compression function through Fourier-based synchrosqueezing (FSST), and the instantaneous frequency information ω0(t,ω) can satisfy the following formula:

[0176]

[0177] Where j is the imaginary unit, ω0(t,ω) is the instantaneous frequency in the time-frequency plane, and F(t,ω) is the time-frequency distribution coefficient. The time-frequency information TFR energy generated by the STFT is distributed in [ω0-Δ, ω0+Δ], where Δ can be understood as a certain range around the frequency ω0. When F(t,ω) is not equal to 0, ω0(t,ω) equals ω0, thus concentrating the TFR energy. This can also be understood as concentrating the energy distributed in [ω0-Δ, ω0+Δ] onto ω0, enhancing the ability to use the target signal.

[0178] Understandably, the above time-frequency focusing processing enables the target signal to have better local focusing in time and frequency, making the distinction between the target signal and clutter more obvious, thereby obtaining a clearer and more concentrated energy distribution of the target signal.

[0179] S407, Adaptive Time-Frequency Extraction.

[0180] In the process of identifying the vibration frequency of the vocal cords in the received signal, when the indoor space is large, the radar can acquire useful information within a certain range around the human body. For example, the radar can select the area with the strongest energy in the received signal as the target center point and search around the target center point to obtain more complete useful information. In other words, adaptive time-frequency extraction can further extract the target signal from the received signal.

[0181] It's understandable that a person cannot be considered a point target relative to radar; different body parts have different radar cross sections (RCS). RCS can be understood as the cross section of a target's radar waves at a specific frequency, polarization, and angle of incidence. Therefore, a received signal containing a target signal may cross multiple range gates, and this range gate varies with the radial distance between the target and the radar. This range gate can be understood as the radar's resolution or the minimum distance the radar can resolve; for example, a radar resolution of 7 cm can also be described as a range gate of 7 cm.

[0182] like Figure 7 As shown, the angle θ between the farthest and closest points of the human body from the radar can satisfy the following formula:

[0183]

[0184] Among them, H rader X represents the radar's altitude. min This represents the distance from the nearest point on the human body to the radar.

[0185] The distance X from the furthest point of the human body to the radar max The following formula can be satisfied:

[0186]

[0187] The displacement difference ΔX between the farthest and the nearest points can satisfy the following formula:

[0188] ΔX=H rader ·sinθ.

[0189] As can be seen from the above formula, the time-frequency extraction range is related to the radar's altitude H. rader This is related to the angle θ between the farthest and closest points of the human body relative to the radar. The displacement difference ΔX between the farthest and closest points can also be understood as the time-frequency extraction range.

[0190] For example, assuming the calculated time-frequency extraction range is 1m, when the radar determines that the distance to the center point of the detected target is 5m, the radar can also extract time-frequency information between 4m and 6m, thus obtaining a more complete target signal.

[0191] The adaptive time-frequency extraction method of this application embodiment can improve the time-frequency focusing of the signal by using synchronous compression transformation. The time-frequency focusing at the range gate containing the target signal is better than the time-frequency focusing at the range gate containing only clutter.

[0192] In a possible implementation, electronic devices can use the Rayleigh entropy, a time-frequency focusing index, to reflect the difference in time-frequency focusing between the target signal and clutter, filter out range gates containing the target signal, and extract time-frequency information.

[0193] Figure 8 The flowchart for time-frequency extraction is shown.

[0194] S801, Obtain the RD diagram of the current data.

[0195] The electronic equipment can convert the information acquired by the radar into an RD map and execute step S802.

[0196] S802, synchronous compression transformation for time-frequency transformation of different distance dimensions.

[0197] It should be noted that step S802 can be understood as the execution of step S406 described above. The electronic device can perform time-frequency focusing enhancement, calculate the range Doppler spectrum for each frame of signal, and perform inverse Fourier transform to obtain the signal variation information over time in the range dimension. The electronic device can also use short-time Fourier transform to transform from the time domain to the frequency domain, thereby improving the focus of the target signal and increasing the time-frequency convergence of the target signal and clutter.

[0198] S803. Record the Rayleigh entropy for each dimension and sort them.

[0199] Electronic devices can calculate the Rayleigh entropy at different locations in space and sort the Rayleigh entropies. Rayleigh entropy can be used to identify the concentration intensity of signals; it is understood that the Rayleigh entropy of the target signal is less than the Rayleigh entropy of clutter. The Rayleigh entropy is described in step S406 above and will not be repeated here.

[0200] S804. Take the mean of a portion of the Rayleigh entropy data and use it as a spatial noise background threshold.

[0201] Electronic devices can also take the mean of a portion of Rayleigh entropy data, for example, by taking the mean of the second half of the sorted Rayleigh array as a spatial noise background threshold, which can also be called the Rayleigh entropy threshold.

[0202] If the Rayleigh entropy at a certain location in space is less than the Rayleigh entropy threshold, it means that the location contains the target signal and needs to be retained. If the Rayleigh entropy at a certain location in space is greater than or equal to the Rayleigh entropy threshold, it means that the location does not contain the target signal and can be filtered out.

[0203] S805: Record the maximum energy distance dimension position Xmax in the received signal, and perform time-frequency transformation and synchronous compression transformation.

[0204] It is understandable that in the RD diagram corresponding to the received signal, the location of the maximum energy, Xmax, can be regarded as the location of the center point of the target signal. The electronic device can search around based on the location of this center point to obtain a more complete target signal.

[0205] The electronic device can also perform time-frequency transformation and synchronous compression transformation on the RD diagram at the center point to obtain the time-frequency diagram corresponding to that position, which can include relevant information of the target signal.

[0206] S806, Position X is updated to Xmax+1, search backwards.

[0207] After processing the position Xmax, the electronic device can further search the surrounding area. In this way, by incrementally incrementing by 1 to search the surrounding area, the entire search range can be traversed, thereby obtaining a more complete target signal.

[0208] Optionally, position X can also be updated to Xmax-1, i.e., searching forward. The specific search direction is not limited in this embodiment.

[0209] S807. Is position X less than the rear boundary of the movement range?

[0210] To avoid blind searching, the electronic device can determine whether the position X is less than the rear boundary of the movement range. This rear boundary of the movement range can be set by the electronic device based on experience; for example, in a call scenario, the rear boundary of the movement range can be set to 10m, or other values. The specific range of the rear boundary of the movement range is not limited in this embodiment.

[0211] If position X is less than the back boundary of the movement range, it means that no place outside the search range has been found yet, and the search can continue. Then step S808 can be executed.

[0212] If position X is greater than or equal to the rear boundary of the movement range, it means that the search has reached a place outside the search range and it is necessary to change the direction of the search. Then step S812 can be executed.

[0213] S808 performs time-frequency transformation and synchronous compression transformation on position X.

[0214] The electronic device can also perform time-frequency transformation and synchronous compression transformation on the RD diagram of position X to obtain the time-frequency diagram corresponding to position X, which can include relevant information of the target signal.

[0215] S809, Calculate Rayleigh entropy.

[0216] It is understandable that if, in step S803 above, after calculating the Rayleigh entropy at different locations in space, the Rayleigh entropy at each location is still stored, then steps S808 and S809 can be omitted. Instead, the Rayleigh entropy at the current location X can be obtained from the stored Rayleigh entropy at each location. This can save the computing power of electronic devices and improve the efficiency of time-frequency extraction.

[0217] S810. Determine whether Rayleigh entropy is less than the Rayleigh entropy threshold.

[0218] If the Rayleigh entropy at position X is less than the Rayleigh entropy threshold, it indicates that the position contains the target signal, and step S811 can be executed.

[0219] When the Rayleigh entropy at position X is greater than or equal to the Rayleigh entropy threshold, it indicates that the position does not contain the target signal, and step S812 can be executed.

[0220] S811. Accumulate the current distance dimension time-frequency information, update position X to X+1, and record the number of accumulations.

[0221] Electronic devices can accumulate and superimpose time and frequency information corresponding to various locations of the target signal, and record the number of accumulations, thereby continuously enriching the time and frequency information related to the target signal, which facilitates subsequent identification and analysis of the target signal.

[0222] S812, position X updated to Xmax-1, search forward.

[0223] After completing a search in one direction around position Xmax, the electronic device can switch directions and continue searching around position Xmax. This allows the entire search range to be covered, resulting in a more complete target signal.

[0224] It is understandable that if position X is updated to Xmax-1 in step S806, then position X in the current step can also be updated to Xmax+1, i.e., searching backward. The specific search direction is not limited in this embodiment.

[0225] S813, Is position X greater than the front boundary of the range of motion?

[0226] To avoid blind searching, the electronic device can determine whether the position X is greater than the front boundary of the movement range. This front boundary of the movement range can be set by the electronic device based on experience; for example, in a call scenario, the front boundary of the movement range can be set to 10m, or it can be set to other values. The front boundary of the movement range can be set to be equal to or different from the rear boundary of the movement range. The specific setting of the front boundary of the movement range is not limited in this embodiment.

[0227] If position X is greater than the front boundary of the movement range, it means that no place outside the search range has been found yet, and the search can continue. Then step S814 can be executed.

[0228] If position X is less than or equal to the front boundary of the movement range, it means that the search has been conducted outside the search range and a different direction needs to be searched. In this case, step S818 can be executed.

[0229] S814. Perform time-frequency transformation and synchronous compression transformation on position X.

[0230] The electronic device can also perform time-frequency transformation and synchronous compression transformation on the RD diagram of position X to obtain the time-frequency diagram corresponding to position X, which can include relevant information of the target signal.

[0231] S815, Calculate Rayleigh entropy.

[0232] Similarly to step S809, if in step S803, after calculating the Rayleigh entropy at different locations in space, the Rayleigh entropy at each location is still stored, then steps S814 and S815 can be omitted. Instead, the Rayleigh entropy at the current location X can be obtained from the stored Rayleigh entropy at each location. This can save the computing power of electronic devices and improve the efficiency of time-frequency extraction.

[0233] S816. Determine whether Rayleigh entropy is less than the Rayleigh entropy threshold.

[0234] When the Rayleigh entropy at position X is less than the Rayleigh entropy threshold, it indicates that position X includes the target signal, and step S817 can be executed.

[0235] When the Rayleigh entropy at position X is greater than or equal to the Rayleigh entropy threshold, it means that position X does not include the target signal. It is understandable that, since the signals around the human body are continuous, when position X does not include the target signal, it means that the boundary of the target signal has been reached, and there is no need to continue searching forward or backward; therefore, step S818 can be executed.

[0236] S817. Accumulate the current distance dimension time-frequency information, update position X to X-1, and record the number of accumulated information.

[0237] The same principle applies to step S811, so it will not be repeated here.

[0238] S818. Normalize the data based on the accumulated number of items.

[0239] Because the number of time-frequency information accumulations varies at different times, it is impossible to uniformly and accurately reflect the target signal.

[0240] Understandably, normalization can average the accumulated time-frequency information, thus reflecting the time-frequency information more accurately after normalization.

[0241] S819, Output the current frame time-frequency graph.

[0242] After searching each location in the RD diagram, the electronic device can obtain full-dimensional information containing the target signal.

[0243] S408, Binarization of the time Doppler matrix.

[0244] Electronic devices can extract the vocal cord vibration frequency range from the time-frequency graph to obtain the time-frequency graph of vocal cord vibration, and then perform binarization processing on the time-frequency graph. This facilitates the subsequent recognition of vocal cord vibration by electronic devices in different scenarios.

[0245] After binarization, electronic devices can identify vocal cord vibrations according to different scenarios.

[0246] If the current scenario is a live broadcast, the electronic device can execute step S409 to determine the user's need for real-time performance and execute different processes based on the real-time performance needs.

[0247] If the current scenario is a call or a silent call, the electronic device can execute step S413 to determine the size of the data to be analyzed and perform different data processing based on the data size.

[0248] In addition, in the case of a silent call, since the electronic device cannot collect voice data, the need to recognize the user's speech is higher. In this case, the electronic device can also execute step S417 to obtain the data collected by the camera and perform lip-reading recognition, thereby improving the accuracy of recognition.

[0249] It is understood that the execution order of steps S413 and S417 is not important. Step S417 can be executed before step S413, for example, during the execution of steps S401 to S408. Step S417 can also be executed after step S413, or it can be executed simultaneously with step S413. The specific timing of the execution of step S417 is not limited in this embodiment.

[0250] S409. Is the real-time requirement high?

[0251] Electronic devices can determine the data processing flow based on real-time requirements.

[0252] For example, when real-time requirements are high, such as in the above... Figure 5 In interface 501, when the user selects the real-time option as fast, the electronic device can execute step S410 to reduce the amount of data through principal component analysis, thereby speeding up the data processing and improving the real-time performance of recognizing the user's speech.

[0253] When real-time requirements are low, such as in the above... Figure 5 In interface 501, when the user selects the real-time option as slow, the electronic device can execute step S415 to upload the data to the cloud server for processing.

[0254] Optionally, before performing step S415, step S414 can be performed first to perform principal component analysis on the data. This application embodiment does not limit this.

[0255] S410. Principal component analysis (PCA).

[0256] PCA is a commonly used data dimensionality reduction technique used to discover patterns and structures in data. For example, PCA can project the original data into a new coordinate system through a linear transformation, maximizing the variance of the projected data. This allows for the selection of more important features, reducing the amount of data while preserving as much information as possible from the original data. In other words, PCA can extract user-relevant features and filter out features irrelevant to the user.

[0257] Among them, user-related characteristic data can be understood as relevant features used to identify semantics, such as the highest frequency of vocal cord vibration, the lowest frequency of vocal cord vibration, the ratio of the highest frequency to the lowest frequency, and / or the duration of vocal cord vibration frequency, etc.

[0258] After completing the principal component analysis, the electronic device can perform step S411 to locally process the user-related characteristic data.

[0259] S411, Local processing.

[0260] Principal component analysis can extract useful information from received signals more quickly, reducing the amount of data to be processed. Furthermore, local processing does not require additional time to upload to the cloud server. Therefore, data processing on the electronic device itself can quickly complete the identification of call content.

[0261] Understandably, due to limitations in local electronic components and memory size, electronic devices have less capacity to analyze and process data than cloud servers, resulting in lower accuracy in recognizing call content compared to cloud servers. Therefore, when real-time requirements are high, call content recognition can be performed locally on the electronic device to ensure recognition speed, although this may sacrifice some accuracy.

[0262] In the possible implementations, in the above Figure 5In interface 501, when the user selects the accuracy option as auxiliary, it means that the user has low requirements for the accuracy of the call content recognition, and the electronic device can perform data processing locally.

[0263] In the possible implementations, in the above Figure 5 In interface 501, when the user selects the privacy option as local processing, the electronic device can perform data processing locally.

[0264] In the possible implementations, in the above Figure 5 In interface 501, when the user selects the data saving mode option as enabled, the electronic device can process data locally. This eliminates the need for the electronic device to upload data to a cloud server, saving the bandwidth used for data uploads.

[0265] S412, a lightweight convolutional neural network.

[0266] Electronic devices can use lightweight convolutional neural networks to process the dimensionality-reduced vocal cord vibration data, thereby determining the content of the user's speech.

[0267] Understandably, lightweight convolutional neural networks (CNNs) can be understood as CNN models designed for operation in electronic devices or embedded systems. These models can have fewer parameters and lower computational cost to run in resource-constrained environments. They can employ specific techniques, such as depthwise separable convolutions, network pruning, and quantization, to reduce model size and computational requirements while maintaining high accuracy.

[0268] S413. Is the data volume greater than or equal to the threshold?

[0269] Understandably, step S408 can yield a relatively large matrix containing a large amount of data.

[0270] When the data volume is greater than or equal to the threshold, it indicates that the matrix contains a large amount of data. The electronic device can then execute step S414 to perform principal component analysis on the data. This reduces the amount of data uploaded to the cloud server, saves bandwidth, and alleviates the pressure on the cloud server to process the data.

[0271] When the data volume is less than the threshold, it indicates that the matrix contains a relatively small amount of data. In this case, the electronic device can skip principal component analysis and instead execute step S415 to upload the data to the cloud server for processing. This improves code execution efficiency and saves computing power.

[0272] The threshold can be set by the electronic device based on experience or based on real-time requirements. The specific value of the threshold is not limited in this embodiment.

[0273] For example, when real-time requirements are high, such as in the above... Figure 5 In interface 501, when the user selects the real-time option as fast, the electronic device can set the threshold to a relatively small value. This increases the probability that the amount of data exceeds the threshold, so step S414 can be executed to reduce the amount of data through principal component analysis, thereby speeding up the data processing and improving the real-time performance of recognizing the user's speech.

[0274] When real-time requirements are low, such as in the above... Figure 5 In interface 501, when the user selects the real-time option as slow, the electronic device can set the threshold to a relatively large value, thereby reducing the probability that the amount of data exceeds the threshold. Then, step S415 can be executed to upload the data to the cloud server for processing.

[0275] S414, Principal Component Analysis.

[0276] Principal component analysis can be performed with reference to the relevant description in step S410 above, and will not be repeated here.

[0277] After completing the principal component analysis, step S415 can be executed to upload the user-related characteristic data to the cloud server for processing.

[0278] Optionally, after completing the principal component analysis, the data may not be uploaded to the cloud server, but may be processed locally on the electronic device. This application embodiment does not limit this.

[0279] S415, Upload data to the cloud server for processing.

[0280] Because cloud servers have strong computing power and / or storage capacity, they are better able to process data and analyze and recognize user speech more accurately.

[0281] Since uploading to the cloud server takes time, the data processing speed of uploading to the cloud server is relatively slower than the data processing speed of the electronic device locally. When the real-time requirement is low, in order to ensure the recognition accuracy, the call content can be uploaded to the cloud server for recognition, but some time may be sacrificed.

[0282] In the possible implementations, in the above Figure 5 In interface 501, when the user selects the accuracy option as "accurate", it means that the user has relatively high requirements for the accuracy of the call content recognition. In this case, the electronic device can upload the data to the cloud server for processing.

[0283] In the possible implementations, in the above Figure 5In interface 501, when the user selects the privacy option as cloud processing, the electronic device can upload data to the cloud server for processing.

[0284] In the possible implementations, in the above Figure 5 In interface 501, when the user selects the data saving mode as off, it indicates that the user is not concerned about data usage. Optionally, the electronic device can upload data to a cloud server for processing. This allows for more accurate analysis and recognition of the user's speech, improving the user experience.

[0285] S416 uses a convolutional neural network to process radar data.

[0286] Cloud servers can include convolutional neural network models derived from big data analysis of radar data. Understandably, because cloud servers have more computing resources and storage space, they can support larger and more complex network models with less concern for resource consumption. This allows cloud-based convolutional neural network models to focus more on accuracy and complexity, without being limited by device resources. Therefore, cloud-based convolutional neural network models can perform more accurate analysis and processing of radar data, thereby recognizing the user's speech.

[0287] S417, Acquire camera data.

[0288] In silent call scenarios, electronic devices can acquire camera data and perform lip-reading recognition, thereby improving the accuracy of user semantic recognition.

[0289] In the possible implementations, in the above Figure 5 In interface 501, when the user selects the accuracy option as strict, it means that the user has high requirements for the accuracy of call content recognition. In this case, the electronic device can combine radar data and data collected by the camera to conduct a comprehensive analysis of the call content in the cloud server, thereby obtaining a more accurate recognition result.

[0290] S418, a convolutional neural network, processes camera data.

[0291] The cloud server may include a convolutional neural network model obtained by analyzing camera data through big data analysis. Based on this model, the camera data can be processed to recognize semantics based on the user's lip movements.

[0292] S419. Judgment of recognition results.

[0293] In the case of a silent call, the cloud server can fuse the recognition results of radar data and camera data, and execute step S420 to output the recognition result of the call content.

[0294] S420, output call text subtitles or other processing.

[0295] The methods of this application will be described in detail below through specific embodiments. The following embodiments can be combined with each other or implemented independently, and the same or similar concepts or processes may not be described again in some embodiments.

[0296] Figure 9 A semantic recognition method according to an embodiment of this application is illustrated. The method includes:

[0297] S901. During the process of collecting users' audio and video data, the vibration frequency of the user's vocal cords is obtained based on radar.

[0298] In this embodiment of the application, the process of collecting users' audio and video data may include scenarios using a microphone and / or a camera, such as voice call scenarios, video call scenarios, silent call scenarios, live streaming scenarios, etc.

[0299] S902, Obtain the first sound recognition result by semantically recognizing the vibration frequency of the vocal cords.

[0300] In this embodiment of the application, the process for obtaining the first voice recognition result can be as described above. Figure 4 The relevant descriptions of steps S412 or S415 in the corresponding embodiments will not be repeated here.

[0301] S903. Obtain the target sound recognition result of the audio and video data based on the first sound recognition result; the target sound recognition result is the first sound recognition result, or the target sound recognition result is obtained by processing the first sound recognition result and the second audio recognition result, or the target sound recognition result is obtained by processing the first sound recognition result, the second audio recognition result and the third image recognition result, wherein the second audio recognition result is obtained by semantic recognition of the audio in the audio and video data, and the third image recognition result is obtained by lip-reading recognition of the user in the audio and video data.

[0302] In this embodiment of the application, the target sound recognition result may include text obtained based on the first sound recognition result, such as subtitles.

[0303] It is understood that electronic devices may obtain target sound recognition results based solely on the first sound recognition result, or they may simultaneously consider semantic recognition of audio and / or lip-reading recognition of the user in the image. This application does not limit the scope of the embodiments.

[0304] For example, in voice call scenarios, since no camera is used to capture images, electronic devices can perform semantic recognition based solely on the vibration frequency of the vocal cords, or they can perform semantic recognition by considering both the vibration frequency of the vocal cords and audio data.

[0305] In silent call scenarios, since the microphone cannot receive the user's voice, electronic devices can perform semantic recognition based solely on the vibration frequency of the vocal cords, or they can simultaneously consider the vibration frequency of the vocal cords and the user's lip movements in the image for semantic recognition.

[0306] In scenarios such as video calls and live streaming, since the vibration frequency of the vocal cords, audio information and / or image information can be obtained, electronic devices can obtain the target sound recognition result based solely on the first sound recognition result, or simultaneously consider the semantic recognition of the audio and / or the lip-reading recognition of the user in the image.

[0307] S904. Display the target sound recognition result.

[0308] The semantic recognition method provided in this application can acquire data such as the vibration frequency of a user's vocal cords based on the radar of an electronic device, and analyze the radar-acquired data to obtain the semantics the user wants to express. Using radar to detect vocal cord vibration offers good privacy and transmission capabilities. This allows the electronic device to not rely entirely on processing speech signals or lip movements, enabling more accurate semantic recognition even in noisy, dimly lit, or obstructed environments, improving call quality and enhancing the user experience.

[0309] Optional, in Figure 9 Based on the corresponding embodiment, before obtaining the first sound recognition result of semantic recognition of the vibration frequency of the vocal cords, the method further includes: preprocessing the vibration frequency of the vocal cords in the distance dimension to obtain the time-frequency diagram corresponding to the vibration frequency of the vocal cords; and performing semantic recognition on the time-frequency diagram corresponding to the vibration frequency of the vocal cords.

[0310] In this embodiment, the preprocessing of the vocal cord vibration frequency in the distance dimension can be referred to the above. Figure 4 The description of step S407 in the corresponding embodiment will not be repeated here.

[0311] Radar can search for the vibration frequency of the vocal cords at different distance dimensions and accumulate the searched information, thereby reducing the dimensionality of the data to obtain a more complete vibration frequency of the vocal cords, which facilitates the further extraction of useful signals from the received signals.

[0312] Optional, in Figure 9Based on the corresponding embodiment, the radar is used to acquire electromagnetic wave signals, which include the vibration frequency of the vocal cords. Preprocessing of the vocal cord vibration frequency in the range dimension includes: determining a first position with maximum energy in the electromagnetic wave signal within the radar search range; searching for a second position centered on the first position and moving towards the boundary of the radar search range; retaining the signal corresponding to the second position if the Rayleigh entropy of the second position is less than a preset value; or, not retaining the signal corresponding to the second position if the Rayleigh entropy of the second position is greater than or equal to the preset value; and outputting a time-frequency graph based on the retained signal, which includes the time-frequency graph corresponding to the vibration frequency of the vocal cords.

[0313] In this embodiment of the application, the first position can be understood as described above. Figure 4 In the corresponding embodiment, the location Xmax of the maximum energy in the RD diagram corresponding to the received signal. The first location can be understood as the center point of the vocal cord vibration frequency.

[0314] The boundary direction can include the boundary directions of each direction of the radar search range, and this application embodiment does not limit it. For example, the boundary direction can include the above-mentioned... Figure 4 In the corresponding embodiments, the direction of backward or forward search.

[0315] The second position may include position Xmax+1 in the backward search, or position Xmax-1 in the forward search, etc., and this application does not limit the implementation.

[0316] The preset value may include the value set by the electronic device based on experience, or it may include the value obtained based on Rayleigh entropy. For example, the preset value may include the average value of the sum of Rayleigh entropy at each or some locations. This application does not limit the implementation.

[0317] If the Rayleigh entropy of the second position is less than the preset value, it means that the position includes the vibration frequency signal of the vocal cords, and the signal corresponding to the second position can be retained; if the Rayleigh entropy of the second position is greater than or equal to the preset value, it means that the position does not include the vibration frequency signal of the vocal cords, and the signal corresponding to the second position is not retained.

[0318] Understandably, the preprocessing of the vocal cord vibration frequency in the distance dimension can be referred to the above. Figure 8 The relevant descriptions of the corresponding embodiments will not be repeated here.

[0319] Radar can select the point with the strongest energy in the received signal as the target center point and search around the target center point to obtain more complete and useful information, which facilitates the further extraction of the vocal cord vibration frequency signal in the received signal.

[0320] Optional, in Figure 9Based on the corresponding embodiment, before determining the first position of the maximum energy in the electromagnetic wave signal, the method further includes: performing time-frequency transformation on different distance dimensions of the electromagnetic wave signal to obtain a time-frequency distribution map STFT for each distance dimension; and determining a preset value based on the Rayleigh entropy of FTFR, wherein the preset value includes the average value of the sum of the Rayleigh entropies of some or all FTFRs.

[0321] In this embodiment, time-frequency transformation of electromagnetic wave signals across different distance dimensions can be performed as described above. Figure 8 The relevant description of step S802 in the corresponding embodiment will not be repeated here.

[0322] Understandably, time-frequency transformation processing enables the target signal to have better local focus in time and frequency, making the distinction between the target signal and clutter more obvious, thus obtaining a clearer and more concentrated energy distribution of the target signal.

[0323] Optional, in Figure 9 Based on the corresponding embodiment, before preprocessing the vibration frequency of the vocal cords in the distance dimension, the method further includes performing one or more of the following processing on the electromagnetic wave signal: moving target detection, constant false alarm rate detection, multipath suppression processing, and time-frequency focusing processing; wherein, moving target detection is used to filter signals in the electromagnetic wave signal with a frequency of zero or a frequency less than a preset frequency value, constant false alarm rate detection is used to filter signals in the electromagnetic wave signal with low frequency energy, multipath suppression processing is used to filter echo signals in the electromagnetic wave signal, the echo signals include signals reflected back when the radar's transmitted signal encounters an obstacle, and time-frequency focusing processing is used to enhance the time-frequency focusing capability of the vibration frequency of the vocal cords in the electromagnetic wave signal.

[0324] In this embodiment of the application, moving target detection can refer to the above. Figure 4 The relevant description of step S403 in the corresponding embodiment will not be repeated here. Constant false alarm rate (CFAR) detection can be referred to the above. Figure 4 The description of step S404 in the corresponding embodiment will not be repeated here. Multipath suppression processing can be referred to the above. Figure 4 The description of step S405 in the corresponding embodiment will not be repeated here. The time-frequency focusing processing can be referred to the above. Figure 4 The relevant description of step S406 in the corresponding embodiment will not be repeated here.

[0325] It is understood that the preset frequency value can be a value close to zero, or a value near the frequency period such as π or 2π. This application embodiment does not limit this.

[0326] Moving target detection can filter out signals with frequencies near zero or the frequency period from electromagnetic wave signals, thus highlighting the target signal. Constant false alarm rate (CFAR) detection can filter out signals with low frequency energy from electromagnetic wave signals and suppress clutter in the received signal, thereby more accurately identifying the target signal. Multipath suppression can filter out echo signals from electromagnetic wave signals, reducing the interference of echo signals on the target signal. Time-frequency focusing processing can utilize the time-frequency focusing difference between the target signal and clutter to enhance the time-frequency focusing ability of the vocal cord vibration frequency in the electromagnetic wave signal, thereby improving the time-frequency focusing ability of the target signal and further enhancing the time-frequency extraction effect of the target signal.

[0327] Optional, in Figure 9 Based on the corresponding embodiment, before obtaining the first sound recognition result of semantic recognition of the vocal cord vibration frequency, the method further includes: displaying a first interface, the first interface including one or more of the following options: a first option, a second option, a third option, and a fourth option; wherein, the first option is used to indicate the real-time requirement for semantic recognition of the vocal cord vibration frequency, the first option includes a first level and a second level, the first level having a higher real-time requirement than the second level; the second option is used to indicate the accuracy requirement for semantic recognition of the vocal cord vibration frequency, the second option includes a third level and a fourth level, the third level having a higher accuracy requirement than the fourth level; the third option is used to indicate the accuracy requirement for semantic recognition of the vocal cord vibration frequency. The privacy requirements for semantic recognition of the vibration frequency of the vocal cords are specified in the third option, which includes levels five and six, with level five having higher privacy requirements than level six. The fourth option indicates the bandwidth requirements for semantic recognition of the vibration frequency of the vocal cords, and includes levels seven and eight, with level seven having higher bandwidth requirements than level eight. When level one, four, five, or seven is selected, levels two, three, six, and / or eight are unselectable.

[0328] In this embodiment of the application, the first interface may include the above-mentioned... Figure 5 The interface 501. For example, the first option may include a real-time option, the second option may include an accuracy option, the third option may include a privacy option, and the fourth option may include a data-saving mode option. The first interface may also include other options, which are not limited in this embodiment.

[0329] The first option can include multiple levels, such as fast, medium, and slow in interface 501. When the first level is fast, the second level can be medium or slow, and when the first level is medium, the second level can be slow.

[0330] The second option can include multiple levels, such as strict, accurate, and auxiliary in interface 501. When the third level is strict, the fourth level can be accurate or auxiliary, and when the third level is accurate, the fourth level can be auxiliary.

[0331] The third option can include multiple levels. For example, in interface 501, the fifth level is local processing and the sixth level is cloud processing.

[0332] The fourth option can include multiple levels. For example, in interface 501, the seventh level is on and the eighth level is off.

[0333] It's understandable that the selection of various options on the first interface may conflict, for example, requiring both high real-time performance and strict accuracy. To reduce conflicts between options, after a user selects an option, the electronic device can set other conflicting options to an unselectable state, such as graying them out, so that the user cannot select conflicting options simultaneously, thereby improving the rationality of the electronic device's execution logic.

[0334] Optional, in Figure 9 Based on the corresponding embodiments, upon receiving a selection operation for the first level, the fourth level, the fifth level, or the seventh level, a first sound recognition result is obtained by semantically recognizing the vibration frequency of the vocal cords, including: using a convolutional neural network to perform semantic recognition on the vibration frequency of the vocal cords to obtain the first sound recognition result.

[0335] In this embodiment of the application, when the user selects the accuracy option as auxiliary, it indicates that the user has low requirements for the accuracy of call content recognition. Alternatively, when the user selects the privacy option as local processing, or when the user selects the data saving mode option as enabled, the electronic device can perform data processing locally.

[0336] In this way, electronic devices do not need to upload data to cloud servers, saving the bandwidth used for data uploads. Since local processing does not require additional time for uploading to cloud servers, data processing on the electronic device itself allows for faster recognition of call content.

[0337] Optional, in Figure 9 Based on the corresponding embodiments, upon receiving a selection operation for the second level, third level, sixth level, and / or eighth level, a first sound recognition result for semantic recognition of the vibration frequency of the vocal cords is obtained, including: uploading the vibration frequency of the vocal cords to a cloud server, and obtaining the first sound recognition result for semantic recognition of the vibration frequency of the vocal cords from the cloud server. The first sound recognition result is obtained by the cloud server using a convolutional neural network to perform semantic recognition of the vibration frequency of the vocal cords.

[0338] In this embodiment of the application, when the user selects the accuracy option as "accurate", it indicates that the user has relatively high requirements for the accuracy of call content recognition. Or when the user selects the privacy option as "cloud processing", or when the user selects the data saving mode option as "off", it indicates that the user is not very concerned about data traffic usage. In this case, the electronic device can upload the data to the cloud server for processing.

[0339] Thus, because cloud servers have strong computing power and / or storage capacity, they are better able to process data and analyze and recognize user speech more accurately, thereby improving user experience.

[0340] Optional, in Figure 9 Based on the corresponding embodiment, before obtaining the first sound recognition result of semantic recognition of the vocal cord vibration frequency, the method further includes: performing principal component analysis (PCA) on the vocal cord vibration frequency to extract characteristic data related to the recognition semantics. The characteristic data includes one or more of the following: the highest frequency of vocal cord vibration, the lowest frequency of vocal cord vibration, the ratio of the highest frequency to the lowest frequency, and the duration of the vocal cord vibration frequency. Obtaining the first sound recognition result of semantic recognition of the vocal cord vibration frequency includes: obtaining the first sound recognition result of semantic recognition of the characteristic data.

[0341] In this embodiment, the specific PCA can be referred to the above. Figure 4 The relevant descriptions in step S410 of the corresponding embodiment will not be repeated here.

[0342] PCA can filter out more important features from the data, thereby reducing the amount of data while preserving as much information as possible from the original data. This allows for the extraction of user-relevant features and the filtering of irrelevant features, leading to a faster extraction of useful information from the received signal.

[0343] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0344] The foregoing primarily describes the solutions provided by the embodiments of this application from a methodological perspective. To achieve the aforementioned functions, it includes corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should readily recognize that, based on the method steps of the examples described in conjunction with the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0345] This application embodiment can divide the apparatus for implementing the method into functional modules based on the above method examples. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.

[0346] like Figure 10 The diagram shows a schematic of a chip provided in an embodiment of this application. The chip 1000 includes one or more processors 1001, a communication line 1002, a communication interface 1003, and a memory 1004.

[0347] In some implementations, memory 1004 stores elements such as executable modules or data structures, or subsets thereof, or extended sets thereof.

[0348] The methods described in the embodiments of this application can be applied to or implemented by the processor 1001. The processor 1001 may be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 1001 or by instructions in the form of software. The processor 1001 may be a general-purpose processor (e.g., a microprocessor or conventional processor), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gates, transistor logic devices, or discrete hardware components. The processor 1001 can implement or execute the various processing-related methods, steps, and logic block diagrams disclosed in the embodiments of this application.

[0349] The steps of the method disclosed in the embodiments of this application can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software modules in the decoding processor. The software modules can be located in mature storage media in the art, such as random access memory, read-only memory, programmable read-only memory, or electrically erasable programmable read-only memory (EEPROM). This storage medium is located in memory 1004, and processor 1001 reads information from memory 1004 and, in conjunction with its hardware, completes the steps of the above method.

[0350] The processor 1001, memory 1004 and communication interface 1003 can communicate with each other via communication line 1002.

[0351] In the above embodiments, the instructions stored in the memory for execution by the processor can be implemented in the form of a computer program product. This computer program product can be pre-written into the memory, or it can be downloaded and installed into the memory as software.

[0352] This application also provides a computer program product comprising one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the flow or function according to the embodiments of this application is generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions may be transmitted from a website site, computer, server, or data center to another website site, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a server or data center that integrates one or more available media. For example, available media may include magnetic media (e.g., floppy disk, hard disk, or magnetic tape), optical media (e.g., digital versatile disc (DVD)), or semiconductor media (e.g., solid-state disk (SSD)).

[0353] This application also provides a computer-readable storage medium. The methods described in the above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any combination thereof. The computer-readable medium may include computer storage media and communication media, and may also include any medium capable of transferring a computer program from one place to another. The storage medium can be any target medium accessible by a computer.

[0354] As one possible design, computer-readable media may include compact disc read-only memory (CD-ROM), RAM, ROM, EEPROM, or other optical disc storage; computer-readable media may also include disk storage or other disk storage devices. Furthermore, any connecting cable may also be appropriately referred to as computer-readable media. For example, if software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave, then coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technologies such as infrared, radio, and microwave are included in the definition of media. As used herein, disks and optical discs include optical discs (CD), laser discs, optical discs, digital versatile discs (DVD), floppy disks, and Blu-ray discs, where disks typically reproduce data magnetically, while optical discs optically reproduce data using lasers.

[0355] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processing unit of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processing unit of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

Claims

1. A semantic recognition method, characterized in that, Applied to an electronic device, the electronic device including radar, the method includes: During the process of collecting user audio and video data, the vibration frequency of the user's vocal cords is obtained based on the radar; the radar is used to acquire electromagnetic wave signals, which include the vibration frequency of the vocal cords. A first sound recognition result is obtained by semantically recognizing the vibration frequency of the vocal cords; The target sound recognition result of the audio and video data is obtained based on the first sound recognition result; the target sound recognition result is the first sound recognition result, or the target sound recognition result is obtained by processing the first sound recognition result and the second audio recognition result, or the target sound recognition result is obtained by processing the first sound recognition result, the second audio recognition result and the third image recognition result, wherein the second audio recognition result is obtained by semantic recognition of the audio in the audio and video data, and the third image recognition result is obtained by lip-reading recognition of the user in the audio and video data; Display the target sound recognition result; Before obtaining the first sound recognition result of semantic recognition of the vibration frequency of the vocal cords, the method further includes: Within the radar search range, determine the first location of the maximum energy in the electromagnetic wave signal; Using the first position as the center, a search is performed towards a second position in the direction of the boundary of the radar search range; If the Rayleigh entropy at the second position is less than a preset value, the signal corresponding to the second position is retained; or, if the Rayleigh entropy at the second position is greater than or equal to the preset value, the signal corresponding to the second position is not retained. Based on the retained signal output time-frequency diagram, the time-frequency diagram includes the time-frequency diagram corresponding to the vibration frequency of the vocal cords; Semantic recognition is performed on the time-frequency graph corresponding to the vibration frequency of the vocal cords; Before obtaining the first sound recognition result of semantic recognition of the vibration frequency of the vocal cords, the method further includes: Display a first interface, which includes one or more of the following options: first option, second option, third option, and fourth option; The first option is used to indicate the real-time requirement for semantic recognition of the vibration frequency of the vocal cords. The first option includes a first level and a second level, with the real-time requirement of the first level being higher than that of the second level. The second option is used to indicate the accuracy requirement for semantic recognition of the vibration frequency of the vocal cords. The second option includes a third level and a fourth level, wherein the accuracy requirement of the third level is higher than that of the fourth level. The third option is used to indicate the privacy requirements for semantic recognition of the vibration frequency of the vocal cords. The third option includes a fifth level and a sixth level, with the fifth level having higher privacy requirements than the sixth level. The fourth option is used to indicate the flow consumption requirement for semantic recognition of the vibration frequency of the vocal cords. The fourth option includes a seventh level and an eighth level, wherein the flow consumption requirement of the seventh level is higher than that of the eighth level. Upon receiving a selection operation for the first level, the fourth level, the fifth level, or the seventh level, obtaining a first sound recognition result that semantically identifies the vibration frequency of the vocal cords includes: A convolutional neural network is used to perform semantic recognition on the vibration frequency of the vocal cords to obtain the first sound recognition result; Upon receiving a selection operation for the second level, the third level, the sixth level, and / or the eighth level, obtaining a first sound recognition result that semantically identifies the vibration frequency of the vocal cords includes: The vibration frequency of the vocal cords is uploaded to a cloud server, and the first sound recognition result of semantic recognition of the vibration frequency of the vocal cords is obtained from the cloud server. The first sound recognition result is obtained by the cloud server using a convolutional neural network to perform semantic recognition of the vibration frequency of the vocal cords.

2. The method according to claim 1, characterized in that, Before determining the first location of the maximum energy in the electromagnetic wave signal, the method further includes: The electromagnetic wave signal is subjected to time-frequency transformation in different distance dimensions to obtain the time-frequency distribution map STFT for each distance dimension; The preset value is determined based on the Rayleigh entropy of FTFR, and the preset value includes the average of the sum of the Rayleigh entropies of some or all FTFRs.

3. The method according to claim 1 or 2, characterized in that, Before determining the first location of the maximum energy in the electromagnetic wave signal within the radar search range, the method further includes: The electromagnetic wave signal is subjected to one or more of the following processing: moving target detection, constant false alarm rate detection, multipath suppression processing, and time-frequency focusing processing; The moving target detection is used to filter signals with a frequency of zero or less than a preset frequency value in the electromagnetic wave signal; the constant false alarm rate (CFAR) detection is used to filter signals with low frequency energy in the electromagnetic wave signal; the multipath suppression processing is used to filter echo signals in the electromagnetic wave signal, the echo signals including the signals reflected back when the radar's transmitted signal encounters an obstacle; and the time-frequency focusing processing is used to enhance the time-frequency focusing capability of the vocal cord vibration frequency in the electromagnetic wave signal.

4. The method according to claim 1 or 2, characterized in that, When the first level, the fourth level, the fifth level, or the seventh level is selected, the second level, the third level, the sixth level, and / or the eighth level are unselectable. When the second level, the third level, the sixth level, and / or the eighth level are selected, the first level, the fourth level, the fifth level, and / or the seventh level are unselectable.

5. The method according to claim 4, characterized in that, Upon receiving a selection operation for the first level, the fourth level, the fifth level, or the seventh level, obtaining a first sound recognition result that semantically identifies the vibration frequency of the vocal cords includes: A convolutional neural network is used to perform semantic recognition on the vibration frequency of the vocal cords to obtain the first sound recognition result.

6. The method according to any one of claims 1-2 and 5, characterized in that, Before obtaining the first sound recognition result of semantic recognition of the vibration frequency of the vocal cords, the method further includes: Principal component analysis (PCA) is performed on the vibration frequency of the vocal cords to extract characteristic data related to semantic recognition. The characteristic data includes one or more of the following: the highest frequency of vocal cord vibration, the lowest frequency of vocal cord vibration, the ratio of the highest frequency to the lowest frequency, and the duration of vocal cord vibration frequency. The first sound recognition result obtained by semantically recognizing the vibration frequency of the vocal cords includes: The first sound recognition result is obtained by semantically recognizing the characteristic data.

7. An electronic device, characterized in that, The electronic device includes: one or more processors and memory; The memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the electronic device to perform the method as described in any one of claims 1-6.

8. A chip system, characterized in that, The chip system is applied to an electronic device, the chip system including one or more processors, the one or more processors being used to invoke computer instructions to cause the electronic device to perform the method as described in any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1-6.

10. A computer program product, characterized in that, The computer program product includes computer program code that, when run on an electronic device, causes the electronic device to perform the method as described in any one of claims 1-6.

Citation Information

Patent Citations

  • Multi-modal voice recognition method and system, and computer readable storage medium

    CN113744731A

  • Non-line-of-sight path speech recognition method and system based on millimeter wave radar

    CN115775556A

  • Multimodal speech recognition system and method based on millimeter wave radar

    CN116416996A