Speech translation method and related equipment

By shooting user videos and translating voice data in real time on a foldable electronic device, combined with a microphone array and lip movement recognition, the problem of users being unable to communicate through eye contact and facial expressions during the translation process is solved, thereby improving the communication experience.

CN120671687APending Publication Date: 2025-09-19HONOR DEVICE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410289633.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-11
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

When a user uses a foldable electronic device for voice translation, he or she needs to check the translation information of the other user's speech displayed on the display screen, which makes it inconvenient to communicate with the other user through eye contact and facial expressions, affecting the communication experience.

Method used

A camera is used to capture user video data and display it on the screen. At the same time, a microphone array is used to determine the location of the sound source. The video and voice data of the other user are translated and displayed in real time. The speaking user is determined by combining lip movements and voiceprint recognition to achieve AR translation.

Benefits of technology

Real-time voice translation is achieved during the conversation, allowing users to communicate through eye contact and facial expressions, improving the user's communication experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671687A_ABST
    Figure CN120671687A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a speech translation method and related equipment. The method comprises the following steps: acquiring speech data through at least one microphone; shooting video data of a user in front of the first display screen through the first camera, shooting video data of a user in front of the second display screen through the second camera, displaying the video data shot by the first camera on the second display screen, and displaying the video data shot by the second camera on the first display screen; determining a target display screen corresponding to the voice data according to the sound source position of the voice data; converting the voice data into text data, and obtaining translation information corresponding to the voice data based on the text data; and displaying the translation information on the target display screen. According to the embodiment of the invention, the translation information of the video and voice data of the opposite user can be displayed on the display screen viewed by each user in real time, so that the dialogue user can perform eye expression and expression communication conveniently.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of intelligent terminal technology, and in particular to a speech translation method and related equipment. Background Art

[0002] Foldable electronic devices are increasingly being used, for example, in translation scenarios. These devices, which include both foldable and non-foldable displays, can translate speech data from users speaking different languages ​​in real time. The translated information is then displayed on each user's display, facilitating conversations between users who speak different languages.

[0003] However, since users need to view the translation information of the other user's speech displayed on the display screen during the conversation, it is inconvenient for users to communicate with each other through eye contact and facial expressions, thereby affecting the user's communication experience. Summary of the Invention

[0004] In view of the above, it is necessary to provide a speech translation method and related equipment to solve the problem that the translation information of the conversation voice data is displayed on the display screen, which makes it inconvenient for users to communicate with each other through eye contact and facial expressions during the conversation.

[0005] In a first aspect, the present application provides a speech translation method, which is applied to an electronic device, wherein the electronic device includes a first camera, a second camera, a first display screen, a second display screen and at least one microphone, and the method includes: collecting speech data through the at least one microphone; shooting video data of a user in front of the first display screen through the first camera, shooting video data of a user in front of the second display screen through the second camera, and displaying the video data shot by the first camera on the second display screen, and displaying the video data shot by the second camera on the first display screen; determining a target display screen corresponding to the speech data based on a sound source position of the speech data; converting the speech data into text data, and obtaining translation information corresponding to the speech data based on the text data; and displaying the translation information on the target display screen.

[0006] Based on the above technical solution, the embodiment of the present application can shoot the video of the conversing users during the conversation, and at the same time perform real-time translation of the voice data of users speaking different languages, and display the video of the other user and the translation information of the voice data in real time on the display screen viewed by each user, thereby realizing AR translation, facilitating the conversing users to communicate through eye contact and facial expressions, and effectively improving the user's communication experience.

[0007] In one possible implementation, the electronic device includes multiple microphones, and the multiple microphones form a microphone array. The method also includes: determining the sound source position based on the voice data, including: calculating the distance between the sound source of the voice data and the microphone array; if the distance between the sound source and the microphone array is less than or equal to a preset distance, using a near-field sound source localization model to locate the sound source and determine the sound source position of the voice data; or if the distance between the sound source and the microphone array is greater than the preset distance, using a far-field sound source localization model to locate the sound source and determine the sound source position of the voice data.

[0008] Based on the above technical solution, the embodiment of the present application can use a near-field sound source localization model to locate the sound source when the sound source corresponding to the voice data is close to the electronic device, and use a far-field sound source localization model to locate the sound source when the sound source corresponding to the voice data is far from the electronic device, thereby accurately determining the position of the sound source.

[0009] In one possible implementation, the calculation of the distance between the sound source of the voice data and the microphone array includes: obtaining the timestamp when each microphone collects the voice data, and determining the delay difference between the multiple microphones in collecting the voice data based on the timestamp when each microphone collects the voice data; calculating the energy of the voice signal corresponding to the voice data collected by each microphone, and determining the energy difference between the voice data collected by the multiple microphones; calculating the distance between the sound source and each microphone based on the delay difference between the voice data collected by the multiple microphones in the microphone array and the energy difference between the voice signals collected by the multiple microphones; and determining the distance between the sound source of the voice data and the microphone array based on the distance between the sound source and each microphone.

[0010] Based on the above technical solution, the embodiment of the present application can accurately calculate the distance between the sound source and the microphone array.

[0011] In a possible implementation, the preset distance is 2D 2 / λ, D=n*d, d is the distance between adjacent elements in the microphone array, n is the number of array spacings, and λ is the wavelength of the highest frequency speech of the sound source.

[0012] Based on the above technical solution, the embodiment of the present application can determine whether the distance between the sound source and the microphone array is far or close based on the setting of the preset distance, and select the corresponding sound source localization model to determine the sound source position according to the judgment result.

[0013] In one possible implementation, the sound source is located using a near-field sound source localization model to determine the sound source position of the voice data, including: extracting a phase spectrum of the voice data; inputting the phase spectrum of the voice data into the near-field sound source localization model, and extracting the phase features of the voice data through the convolutional layer of the near-field sound source localization model; inputting the phase features of the voice data into a feedforward layer of the near-field sound source localization model, and extracting the intermediate features of the voice data based on the phase features of the voice data through the feedforward layer; time-averaging the intermediate features of the voice data to obtain a feature vector of the voice data; inputting the feature vector of the voice data into the affine layer of the near-field sound source localization model, and outputting the arrival direction of the voice signal corresponding to the voice data through the affine layer.

[0014] Based on the above technical solution, the embodiment of the present application can accurately locate the sound source according to the near-field sound source localization model.

[0015] In one possible implementation, the sound source is located using a far-field sound source localization model to determine the sound source position of the voice data, including: extracting audio features of the voice data; performing cross-correlation calculation on the voice data using a generalized cross-correlation-phase transformation algorithm to obtain a cross-correlation function of the voice data; performing phase transformation weighting on the cross-correlation function of the voice data; estimating the time delay between the multiple microphones collecting the voice data based on the peak value of the cross-correlation function, calculating the sound source angle based on the time delay and the distance between the multiple microphones, and using the sound source angle as the sound source position of the voice data.

[0016] Based on the above technical solution, the embodiment of the present application can accurately locate the sound source according to the far-field sound source localization model.

[0017] In one possible implementation, determining the sound source position based on the voice data includes: identifying facial areas in the first video data captured by the first camera and the second video data captured by the second camera; performing lip movement detection on the facial areas in the first video data and the second video data, respectively; if the first video data contains lip movements, determining that the sound source position is toward the first display screen; or if the second video data contains the lip movements, determining that the sound source position is toward the second display screen.

[0018] Based on the above technical solution, the embodiment of the present application can use lip movement detection to determine the speaking user corresponding to the voice data, thereby accurately determining the position of the speaking user relative to the electronic device.

[0019] In one possible implementation, the identifying of facial areas in the first video data captured by the first camera and the second video data captured by the second camera includes: inputting each video frame in the first video data and the second video data into a face recognition model, and identifying the facial areas of each video frame in the first video data and the second video data through the face recognition model.

[0020] Based on the above technical solution, the embodiment of the present application can accurately identify the face area in the video data.

[0021] In one possible implementation, the lip movement detection is performed on the face areas in the first video data and the second video data respectively, including: outputting multiple feature point coordinates of the lip area in the face area of ​​each video frame in the first video data and the second video data through the face recognition model; comparing the multiple feature point coordinates of the lip area of ​​multiple video frames in the first video data and the second video data respectively; if the multiple feature point coordinates of the lip area of ​​any two video frames are different, determining that the corresponding video data contains lip movements.

[0022] Based on the above technical solution, the embodiment of the present application can accurately obtain lip movement detection results.

[0023] In one possible implementation, the method further includes: determining whether there are multiple users in front of the first display screen based on the video data captured by the first camera, and determining whether there are multiple users in front of the second display screen based on the video data captured by the second camera; if the number of users in front of the first display screen and / or the number of users in front of the second display screen is multiple, determining the speaking user of the voice data based on lip movement detection, and determining the target display screen corresponding to the voice data based on the speaking user.

[0024] Based on the above technical solution, the embodiment of the present application can use lip movement detection to determine the speaking user of the voice data when there are multiple users on either side or both sides of the electronic device, thereby accurately determining the target display screen corresponding to the voice data.

[0025] In one possible implementation, the method further includes: determining whether there are multiple users in front of the first display screen based on the video data captured by the first camera, and determining whether there are multiple users in front of the second display screen based on the video data captured by the second camera; if the number of users in front of the first display screen and / or the number of users in front of the second display screen is multiple, determining the speaking user of the voice data based on lip movement detection and voiceprint recognition, and determining the target display screen corresponding to the voice data based on the speaking user.

[0026] Based on the above technical solution, the embodiment of the present application can combine lip movement detection and voiceprint recognition to determine the speaking user of the voice data when there are multiple users on either side or both sides of the electronic device, thereby accurately determining the target display screen corresponding to the voice data.

[0027] In one possible implementation, the determining the speaking user of the voice data based on lip movement detection and voiceprint recognition, and determining the target display screen corresponding to the voice data according to the speaking user, includes: pre-acquiring target voiceprint features of the voice data of multiple users in the video data displayed on the first display screen or the second display screen respectively; determining the user who makes lip movements in the video data based on lip movement detection, and extracting the actual voiceprint features of the voice data through a voiceprint recognition model, comparing the actual voiceprint features with the target voiceprint features of each user, and determining the user corresponding to the actual voiceprint features; if the user who makes lip movements in the video data and the user corresponding to the actual voiceprint features are the same user, then determining that the user who makes lip movements in the video data is the speaking user of the voice data.

[0028] Based on the above technical solution, the embodiment of the present application can accurately locate the speaking user of the voice data when there are multiple users on either side or both sides of the electronic device.

[0029] In a possible implementation, converting the voice data into text data includes: performing voice recognition on the voice data using a voice recognition model, and converting the voice data into the text data.

[0030] Based on the above technical solution, the embodiment of the present application first converts the voice data into text data to improve the translation efficiency of the voice data.

[0031] In one possible implementation, obtaining translation information corresponding to the voice data based on the text data includes: using a machine translation model to translate the text data into a target language to obtain translation information corresponding to the voice data, wherein the target language is determined based on a target display screen of the translation information.

[0032] Based on the above technical solution, the embodiment of the present application can translate the text data corresponding to the voice data into a language that the user can understand, thereby improving communication efficiency.

[0033] In a possible implementation, displaying the translation information on the target display screen includes: displaying the translation information in a window on the target display screen, where the window is displayed near a portrait area in the video data displayed on the target display screen.

[0034] Based on the above technical solution, the embodiment of the present application can optimize the display effect of the translation information.

[0035] In a possible implementation, displaying the translation information on the target display screen includes: if the video data displayed on the target display screen includes multiple portraits, displaying the translation information on the target display screen and marking the speaking user corresponding to the translation information.

[0036] Based on the above technical solution, when there are multiple users on one side of the electronic device, the embodiment of the present application can indicate the speaking user corresponding to the voice data through a window, making it easier for the user to determine the speaking user of the voice data.

[0037] In a second aspect, the present application provides an electronic device, comprising a memory and a processor: wherein the memory is used to store program instructions; the processor is used to read and execute the program instructions stored in the memory, and when the program instructions are executed by the processor, the electronic device performs the above-mentioned speech translation method.

[0038] In a third aspect, the present application provides a chip coupled to a memory in an electronic device, wherein the chip is used to control a processor of the electronic device to execute the above-mentioned speech translation method.

[0039] In a fourth aspect, the present application provides a computer storage medium storing program instructions. When the program instructions are executed on an electronic device, the processor of the electronic device executes the above-mentioned speech translation method.

[0040] In addition, the technical effects brought about by the second to fourth aspects can be found in the descriptions of the methods of each design in the above method section, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a schematic diagram of a voice translation interface in the related art.

[0042] Figure 2 It is a three-dimensional diagram of an electronic device provided in one embodiment of the present application.

[0043] Figure 3 This is a software architecture diagram of an electronic device provided in one embodiment of the present application.

[0044] Figure 4 This is another three-dimensional diagram of an electronic device provided by an embodiment of the present application.

[0045] Figure 5 This is an application scenario diagram of the speech translation method provided in one embodiment of the present application.

[0046] Figure 6This is a flowchart of a speech translation method provided in one embodiment of the present application.

[0047] Figure 7 This is a schematic diagram of the interface of a speech translation application provided in one embodiment of the present application.

[0048] Figure 8 This is a flowchart for determining the sound source location of voice data provided by an embodiment of the present application.

[0049] Figure 9 It is a schematic diagram of the sound source angle provided in one embodiment of the present application.

[0050] Figure 10 This is a flowchart for determining the sound source location of voice data provided by another embodiment of the present application.

[0051] Figure 11 This is a flowchart of a speech translation method provided by another embodiment of the present application.

[0052] Figure 12 This is another interface diagram of the voice translation application provided in one embodiment of the present application.

[0053] Figure 13 This is a hardware architecture diagram of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0054] The terms "first" and "second" involved in the embodiments of the present application are for descriptive purposes only and are not to be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more features. In the description of the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way.

[0055] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. The terms used in the specification of this application are for the purpose of describing specific embodiments only and are not intended to limit this application. It should be understood that, unless otherwise specified in this application, " / " means or. For example, A / B can mean A or B. "And / or" in this application is merely a way to describe the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. "At least one" means one or more. "Multiple" means two or more than two. For example, at least one of a, b or c can mean: a, b, c, a and b, a and c, b and c, a, b and c. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.

[0056] Currently, foldable electronic devices are increasingly being used, for example, in translation scenarios. Foldable electronic devices, which include foldable and non-foldable displays, can translate voice data from users speaking different languages ​​in real time and display the translated voice data of the other user on the display screens viewed by each user, facilitating conversations between users who speak different languages. However, since users need to view the translated voice data of the other user displayed on the display screen during a conversation, it is difficult for users to communicate through eye contact and facial expressions, which affects the user's communication experience.

[0057] See Figure 1 As shown in FIG, a schematic diagram of a voice translation interface in related art. In related art, a user can communicate with other users who speak different languages ​​through a voice translation application pre-installed in an electronic device. During the conversation, the electronic device can record the voice of the other user, convert the voice into text through the voice translation application, and translate the text into the language corresponding to the electronic device user. Figure 1 As shown, in voice translation apps, translations are typically presented as single text, and there's a delay between when the user speaks and when the translation appears on the electronic device, hindering communication efficiency. During conversations, users' attention is primarily focused on viewing the translated content displayed on the electronic device, hindering eye contact and facial expressions, which in turn impacts the user's communication experience.

[0058] See Figure 2, which is a perspective view of an electronic device provided in an embodiment of the present application. The electronic device 100 provided in an embodiment of the present application is a foldable electronic device, comprising a first display screen 10 and a second display screen 20. The first display screen 10 is the outer screen of the electronic device 100, and the second display screen 20 is the inner screen of the electronic device 100. The second display screen 20 is a foldable display screen.

[0059] In the Figure 2 When the foldable electronic device shown is used in a translation scenario, the foldable electronic device can be unfolded and placed between users who need to communicate, so that the first display screen faces the first user and the second display screen faces the second user. During the conversation, the electronic device can display the translation content corresponding to the first user's voice data on the second display screen for the second user to view, and display the translation content corresponding to the second user's voice data on the first display screen for the first user to view. However, since the user needs to pay attention to the translation content of the other user's speech displayed on the display screen during the conversation, it is impossible to keep an eye on the other user, which makes it inconvenient for the users to communicate with each other through eye contact and facial expressions, thereby affecting the user's communication experience.

[0060] In order to avoid affecting the communication experience of users in dialogue scenarios that require translation, the embodiment of the present application provides a voice translation method, which can shoot the video of the dialogue users during the dialogue, and simultaneously translate the voice data of users speaking different languages ​​in real time, and display the video of the other user and the translation information of the voice data on the display screen viewed by each user in real time, realizing AR (Augmented Reality) translation, facilitating the dialogue users to communicate through eye contact and facial expressions, and effectively improving the user's communication experience. The voice translation method is applied to electronic devices, and the following is combined with Figure 3 A diagram illustrating the software architecture of an electronic device.

[0061] See Figure 3 , which is a software architecture diagram of an electronic device provided in an embodiment of the present application. A layered architecture divides software into several layers, each with clear roles and division of labor. Layers communicate with each other through software interfaces. For example, the Android system is divided into four layers: application layer 101, framework layer 102, Android runtime and system library 103, hardware abstraction layer 104, kernel layer 105, and hardware layer 106.

[0062] The application layer 101 may include a series of application packages, such as camera, gallery, calendar, call, map, navigation, WLAN, Bluetooth, music, video, short message, device control service, etc.

[0063] The framework layer 102 provides an application programming interface (API) and programming framework for the applications in the application layer. The application framework layer includes some predefined functions. For example, the application framework layer may include a window manager, content provider, view system, telephony manager, resource manager, notification manager, etc.

[0064] The window manager manages window programs. It can obtain the display size, determine whether a status bar exists, lock the screen, and take screenshots. The content provider stores and retrieves data and makes it accessible to applications. This data can include video, images, audio, incoming and outgoing calls, browsing history and bookmarks, and the phone book. The view system includes visual controls, such as those for displaying text and images. The view system can be used to build applications. The display interface can consist of one or more views. For example, a display interface containing a text notification icon can include a view for displaying text and a view for displaying images. The call manager provides communication functions for electronic devices, such as managing call status (including connected and ended calls). The resource manager provides applications with various resources, such as localized strings, icons, images, layout files, and video files. The notification manager enables applications to display notifications in the status bar. These messages can be displayed briefly and then disappear automatically, without requiring user interaction. For example, the notification manager is used to notify download completions and message reminders. The notification manager can also be a notification that appears in the system's top status bar in the form of an icon or scrolling text bar, such as a notification from an application running in the background, or a notification that appears on the screen in the form of a dialog window. For example, a text message may be displayed in the status bar, a notification sound may be emitted, an electronic device may vibrate, an indicator light may flash, etc.

[0065] The Android Runtime consists of a core library and a virtual machine. The Android runtime is responsible for scheduling and managing the Android system. The core library consists of two parts: one for the Java language's callable functions and the other for the Android core library.

[0066] The application layer 101 and the framework layer 102 run in a virtual machine. The virtual machine executes the Java files in the application layer and the framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0067] The system library 103 may include multiple functional modules, such as a surface manager, a media library, a 3D graphics processing library (such as OpenGL ES), a 2D graphics engine (such as SGL), and the like.

[0068] The surface manager manages the display subsystem and provides fusion of 2D and 3D layers for multiple applications. The media library supports playback and recording of a variety of common audio and video formats, as well as static image files. The media library supports a variety of audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG. The 3D graphics processing library implements 3D graphics drawing, image rendering, compositing, and layer processing. The 2D graphics engine is the drawing engine for 2D drawing.

[0069] The hardware abstraction layer 104 runs in the user space, encapsulates the kernel layer driver, and provides a calling interface to the upper layer.

[0070] The kernel layer 105 is a layer between hardware and software and includes at least a display driver, a camera driver, an audio driver, and a sensor driver.

[0071] The kernel layer 105 is the core of the electronic device's operating system. It is the first layer of software expansion based on the hardware, providing the most basic operating system functions and the foundation of the operating system. It is responsible for managing the system's processes, memory, device drivers, files, and network systems, and determines the system's performance and stability. For example, the kernel can determine the time an application can operate on a certain piece of hardware.

[0072] The kernel layer 105 includes programs closely related to the hardware, such as interrupt handlers and device drivers. It also includes basic, common, and high-frequency modules, such as the clock management module and the process scheduling module, as well as key data structures. The kernel layer can be set in the processor or stored in internal memory.

[0073] The hardware layer 106 includes the hardware of the electronic device, such as a display screen, buttons, a microphone array, etc.

[0074] like Figure 2 As shown, the electronic device 100 further includes a first camera 1931 and a plurality of microphones 1701. The first camera 1931 is disposed at the upper edge of the first display screen 10, that is, the first camera 1931 is disposed on the opposite side of the second display screen 20. Figure 4, which is another perspective view of an electronic device provided in an embodiment of the present application, the electronic device 100 further includes a second camera 1932, which is disposed at an edge of the second display screen 20, that is, the second camera 1932 is disposed on the opposite side of the first display screen 10. In one embodiment of the present application, the multiple microphones 1701 include three microphones, which are respectively disposed near the first camera 1931, near the second camera 1932, and at the bottom of the electronic device 100. The three microphones constitute a microphone array. In other embodiments of the present application, the number of the multiple microphones 1701 may also be 2, 4, 5, or other values, and the present application embodiment does not limit this.

[0075] See Figure 5 FIG2 is a schematic diagram of an application scenario of a speech translation method provided by an embodiment of the present application. An electronic device 100 is placed between users who need to communicate, wherein the first display screen 10 and the second display screen 20 of the electronic device 100 face the users who need to communicate, and the first display screen 10 and the second display screen 20 respectively display the translation information of the video data and speech data of the opposite user, thereby facilitating communication between multiple users of different languages. The opposite user can be observed through video, facilitating eye contact and facial expression communication between the users, and effectively improving the user communication experience.

[0076] See Figure 6 FIG. 1 is a flow chart of a speech translation method provided in an embodiment of the present application. The method is applied in an electronic device and includes:

[0077] S101, collecting voice data through a microphone.

[0078] In one embodiment of the present application, a speech translation application is installed in an electronic device. When a user desires to perform a speech translation operation through the electronic device, the user can trigger an icon of the speech translation application. In response to the user triggering the speech translation application icon, the electronic device launches the speech translation application and displays a display interface of the speech translation application. After the speech translation application is launched, the electronic device controls multiple microphones of the electronic device to collect real-time voice data of the user during the conversation.

[0079] S102, capture video data of a user in front of a first display screen through a first camera, capture video data of a user in front of a second display screen through a second camera, and display the video data captured by the first camera on the second display screen, and display the video data captured by the second camera on the first display screen.

[0080] In one embodiment of the present application, in a scenario where real-time translation of voice data from a conversation using an electronic device is performed, the electronic device is placed between the conversation users, with a first display screen and a second display screen facing the conversation users respectively. After the voice translation application is started, the first camera and the second camera are respectively controlled to shoot video of the scene in front in real time, so that the first camera shoots video data of the user in front of the first display screen, and the second camera shoots video data of the user in front of the second display screen, and the captured video data are displayed on the corresponding display screens. Specifically, the video data shot by the first camera is displayed on the second display screen, and the video data shot by the second camera is displayed on the first display screen, so that users on both sides of the electronic device can see the video data of the user on the opposite side on the display screens.

[0081] S103: Determine a target display screen corresponding to the voice data according to the sound source position of the voice data.

[0082] In one embodiment of the present application, the sound source position of the voice data is determined based on the voice data, including: calculating the distance between the sound source of the voice data and a microphone array composed of multiple microphones, and selecting a near-field sound source localization model or a far-field sound source localization model to locate the sound source according to the distance between the sound source and the microphone array to determine the sound source position of the voice data. If the distance between the sound source and the microphone array is less than or equal to the preset distance, the sound source is located according to the near-field sound source localization model to determine the sound source position of the voice data; if the distance between the sound source and the microphone array is greater than the preset distance, the sound source is located according to the far-field sound source localization model to determine the sound source position of the voice data. The preset distance is 2D 2 / λ, D = n*d, where d is the distance between adjacent elements in the microphone array, n is the number of array pitches, and λ is the wavelength of the highest-frequency speech from the source, or the minimum wavelength of the speech output by the source. If the distances between all microphones in the array are the same, d can be the distance between any two microphones; if the distances between all microphones in the array are different, d can be the average distance between the two microphones.

[0083] In another embodiment of the present application, the user's lip movements can also be detected based on the video data captured by the first camera and the second camera. When lip movements are detected in the video data, the sound source position is determined to be the side of the camera that captured the video data containing the lip movements.

[0084] In one embodiment of the present application, after determining the sound source position of the voice data, the display screen to which the sound source position is directed is determined based on the sound source position of the voice data, and the display screen opposite to the display screen to which the sound source position is directed is used as the target display screen corresponding to the voice data.

[0085] In one embodiment of the present application, the sound source position determined by the near-field sound source localization model or the far-field sound source localization model is the angle (hereinafter referred to as the azimuth) of the sound source relative to the microphone array, that is, the electronic device, and the azimuth of the sound source relative to the microphone array is the angle between the line between the sound source and the microphone array (for example, the center position of the microphone array) and the shorter edge of the first display screen. For example, if the azimuth of the sound source relative to the electronic device is within (0°, 180°), the sound source is determined to be in front of the first display screen, that is, the display screen facing the sound source position is the first display screen, and the display screen opposite the first display screen, that is, the second display screen, is used as the target display screen for the translation information. If the azimuth of the sound source relative to the electronic device is within (180°, 360°), the sound source is determined to be in front of the second display screen, that is, the display screen facing the sound source position is the second display screen, and the display screen opposite the second display screen, that is, the first display screen, is used as the target display screen for the translation information.

[0086] In another embodiment of the present application, the sound source position determined by lip movement detection is the display screen toward which the sound source is directed. If the display screen toward which the sound source is directed is the first display screen, the display screen opposite the first display screen, i.e., the second display screen, is used as the target display screen for the translation information. If the display screen toward which the sound source is directed is the second display screen, the display screen opposite the second display screen, i.e., the first display screen, is used as the target display screen for the translation information.

[0087] S104: Convert the voice data into text data, and obtain translation information corresponding to the voice data based on the text data.

[0088] In one embodiment of the present application, speech recognition is performed on speech data, the speech data is converted into text data, the target language to be translated into the speech data is obtained, the text data is translated into the target language, and translation information corresponding to the speech data is obtained.

[0089] In one embodiment of the present application, a speech recognition model can be used to perform speech recognition on speech data and convert the speech data into text data. The speech recognition model uses speech features and the mapping relationship between speech features and text as training data for training and generation. The speech recognition model can be a hidden Markov model, a mixed Gaussian model or a recurrent neural network model. The voice data collected by the microphone is preprocessed, for example, the preprocessing includes removing silent segments, denoising, enhancing, framing, windowing, etc.; the preprocessed voice data is subjected to feature extraction, for example, the Mel Frequency Cepstrum Coefficient (MFCC) of the voice data is extracted as the feature of the voice data, wherein the Mel Frequency Cepstrum Coefficient of the voice data is extracted including: performing a fast Fourier transform on the voice signal corresponding to the voice data to obtain the frequency domain feature of the voice signal, then calculating the square of the amplitude of each frame of the voice signal to obtain the power spectrum density of each frame of the audio signal, filtering the power spectrum density of each frame of the audio signal through a filter bank to obtain the filtered energy, taking the logarithm of the filtered energy value, performing an inverse Fourier transform on the logarithmic energy spectrum to obtain the cepstrum coefficient, and selecting the cepstrum coefficients of the first preset number (for example, 12 or 13) as the final Mel cepstrum coefficient. In other embodiments of the present application, the feature of the extracted voice data may also be a linear prediction cepstrum parameter (LPCC). The features of the extracted speech data are input into the speech recognition model, which then outputs the text units (such as phonemes, words, or sentences) to which the multiple features of the speech data belong, as well as the probability of belonging to each text unit. The text units are then decoded and post-processed based on the language model and other post-processing techniques to obtain the text data corresponding to the speech data. For example, post-processing includes grammar verification, word order adjustment, word sense disambiguation, etc.

[0090] In one embodiment of the present application, a machine translation model can be used to translate text data into a target language to obtain translation information corresponding to the speech data. The machine translation model uses text features and the mapping relationship between text features and translation information in multiple languages ​​as training data for training and generation. The machine translation model can be a statistical machine translation model, a neural network machine translation model, a zero-shot translation model, or a reinforcement learning translation model. The neural network machine translation model can be a recurrent neural network model or a transformer model.

[0091] Taking the machine translation model as the converter model as an example, the translation process of text data is explained: the text data corresponding to the speech data is input into the converter model, the text features and position features of each word in the text data are extracted through the converter model to obtain the feature vector of each word, the feature vector of each word is input into the encoder block of the converter model to obtain the encoding matrix of each word, the encoding matrix of each word is input into the decoder block of the converter model, and each word is translated according to the encoding matrix of each word through the decoder block, that is, each word is translated into the target language.

[0092] In one embodiment of the present application, the electronic device can set the language corresponding to the translation information displayed on each display screen according to user operations, and the language corresponding to the translation information displayed on each display screen is the same as the language of the voice data of the user to which the display screen is facing. Among them, the language of the translation information displayed on the first display screen is the same as the language of the voice data of the first user to which the first display screen is facing, and the language of the translation information displayed on the second display screen is the same as the language of the voice data of the second user to which the second display screen is facing. The target language of the voice data to be translated is obtained according to the target display screen of the translation information. If the target display screen of the translation information is the first display screen, the target language of the voice data to be translated is the language of the translation information displayed on the first display screen; if the target display screen of the translation information is the second display screen, the target language of the voice data to be translated is the language of the translation information displayed on the second display screen.

[0093] For example, if the language of the voice data of the first user facing the first display screen is Chinese, and the target display screen for the translation information is the first display screen, then the target language of the currently collected voice data to be translated is Chinese. If the language of the voice data of the second user facing the second display screen is English, and the target display screen for the translation information is the second display screen, then the target language of the currently collected voice data to be translated is English.

[0094] S105: Display the translated information on the target display screen.

[0095] In one embodiment of the present application, after the voice data is translated into translation information corresponding to the target language, the translation information is displayed on the target display screen, so that the user viewing the target display screen can view the translation information of the voice data of the opposite user. For example, if the target display screen is a second display screen, the translation information of the voice data is displayed on the second display screen, so that the second user facing the second display screen can see the translation information of the first user's speech content.

[0096] See Figure 7, which is a schematic diagram of the interface of a speech translation application provided in one embodiment of the present application. In one embodiment of the present application, the translation information corresponding to each acquired speech data is displayed in the form of a window on the target display screen, and the display position of the window is near the portrait area or face area in the video data displayed on the target display screen.

[0097] In one embodiment of the present application, the time information corresponding to each translation information may be marked above each translation information so that the user can intuitively see the time point corresponding to the translation information.

[0098] In one embodiment of the present application, when the translation information in the first display screen or the second display screen is too much and fills the display screen, the translation information with an earlier translation time can be hidden on the corresponding display screen, and a scroll bar can be displayed at a designated position so that the user can scroll to view the hidden translation information.

[0099] In one embodiment of the present application, in addition to displaying the translation information of the other user on the display screen, the text information corresponding to the voice data of the current user can also be displayed on the display screen. Taking the first display screen as an example, not only the translation information of the second user can be displayed on the first display screen, but also the text data corresponding to the voice data of the first user can be displayed on the first display screen, so that the first user can check whether the text data converted from the voice data is correct.

[0100] In one embodiment of the present application, text data corresponding to the first user's voice data can be scrolled and displayed in a separate window on the first display screen. For example, it can be displayed floating on the window displaying the translation information without blocking the window displaying the translation information, or it can be displayed side by side with the window displaying the translation information.

[0101] In this embodiment of the present application, after receiving voice data, the electronic device can determine which user the voice data comes from, thereby determining the target display screen between the first and second displays. The electronic device then translates the voice data and displays the translated information on the target display screen. Furthermore, the electronic device can also display video data of the opposite user on the display screen, allowing the user to see the opposite user's facial expressions and eyes while viewing the translated information on the display screen, effectively enhancing the user's communication experience.

[0102] See Figure 8 As shown, it is a flowchart of determining the sound source location of speech data provided by an embodiment of the present application.

[0103] S201, calculating the distance between the sound source of the speech data and the microphone array.

[0104] In one embodiment of the present application, after each microphone in the microphone array collects voice data, it also records the timestamp of the voice data collected, obtains the timestamp of the voice data collected by each microphone, and determines the delay difference of the voice data collected by multiple microphones based on the timestamp of the voice data collected by each microphone. For example, the microphone array includes microphone A, microphone B and microphone C, and the timestamp of the voice data collected by microphone A is t A , the time stamp of the voice data collected by microphone B is t B , the time stamp of the voice data collected by microphone C is t C , the delay difference Δt between microphone A and microphone B when collecting voice data AB =t A -t B , the delay difference Δt between microphone A and microphone C when collecting voice data AC =t A -t C Delay difference Δt AB and Δt AC The difference between the distance PA between the sound source P and microphone A and the distance PB between the sound source position P and microphone B, and the difference between the distance PA between the sound source P and microphone A and the distance PC between the sound source P and microphone C may be reflected.

[0105] In one embodiment of the present application, after each microphone in the microphone array collects voice data, the voice data collected by each microphone is obtained, and the energy of the voice signal corresponding to the voice data collected by each microphone is calculated. For example, it can be the short-time average energy. The energy of the voice signal can reflect the amplitude of the voice signal. For example, the energy of the voice signal collected by microphone A is E A , the energy of the speech signal collected by microphone B is E B , the energy of the speech signal collected by microphone C is E C , the energy difference ΔE between the speech signals collected by microphone A and microphone B AB =E A -E B , the energy difference ΔE between the speech signals collected by microphone A and microphone C AC =E A -E C Energy difference ΔE AB and ΔE ACIt can also reflect the difference between the distance PA between the sound source P and microphone A and the distance PB between the sound source position P and microphone B, as well as the difference between the distance PA between the sound source P and microphone A and the distance PC between the sound source P and microphone C. According to the principle of spherical wave transmission of near-field sound waves, the energy attenuates by 6dB every time the distance doubles. In this way, the difference between the distances can be determined based on the energy difference. In other words, according to the energy difference ΔE AB and ΔE AC The relationship between distance PA and distance PB can be calculated as follows: PB = aPA, and the relationship between distance PA and distance PC is: PC = bPA. AB When ΔE is 6dB, a=2, that is, the PB distance is twice the PA distance; when ΔE AC When it is 12dB, b=4, that is, the PC distance is 4 times the PA distance.

[0106] In one embodiment of the present application, the distance between the sound source and each microphone in the microphone array is calculated based on the delay difference in the voice data collected by multiple microphones and the energy difference in the voice signals collected by multiple microphones. The calculation formula for the distance between the sound source P and microphones A, B, and C is:

[0107] PB-PA=(a-1)PA=Δt AB *v (1);

[0108] PC-PA=(b-1)PA=Δt AC *v (2);

[0109] In the calculation formulas (1) and (2), v is the propagation speed of sound waves in air.

[0110] In one embodiment of the present application, the distance between the sound source of the voice data and the microphone array is determined based on the distance between the sound source and each microphone, including: obtaining the distance between the sound source of the voice data and the microphone array by calculating the average of the distance PA between the sound source P and microphone A, the distance PB between the sound source P and microphone B, and the distance PC between the sound source P and microphone C.

[0111] S202: Determine whether the distance between the sound source and the microphone array is less than or equal to a preset distance. If the distance between the sound source and the microphone array is less than or equal to the preset distance, the process proceeds to S203; if the distance between the sound source and the microphone array is greater than the preset distance, the process proceeds to S204.

[0112] In one embodiment of the present application, the preset distance is 2D 2 / λ, D = n*d, d is the distance between adjacent elements in the microphone array, n is the number of array spacings, and λ is the wavelength of the highest frequency speech of the sound source, that is, the minimum wavelength of the speech output by the sound source.

[0113] S203: Localize the sound source using a near-field sound source localization model to determine the sound source position of the speech data.

[0114] In one embodiment of the present application, a near-field sound source localization model is used to calculate the azimuth and / or elevation angle of the sound source relative to the microphone, that is, the electronic device, to obtain the sound source position of the voice data.

[0115] In one embodiment of the present application, a near-field sound source localization model is generated by training speech features and the mapping relationship between speech features and the arrival direction of speech signals as training data. The near-field sound source localization model is a convolutional neural network model. First, the phase spectrum of the speech data is extracted, including: preprocessing the speech signal corresponding to the speech data, such as preprocessing including denoising, filtering, etc.; performing fast Fourier transform on the preprocessed speech signal to extract the frequency domain features of the speech signal; extracting the phase information of the speech signal from the frequency domain features of the speech signal, such as calculating the phase of each frequency point of the speech signal through an inverse tangent function; performing a preset processing operation on the extracted phase information to obtain the phase spectrum of the speech signal, such as the preset processing operation including smoothing, filtering or interpolation, etc.

[0116] In one embodiment of the present application, the phase spectrum of the speech data is input into the near-field sound source localization model, and the phase features of the speech data are extracted through the convolutional layer of the near-field sound source localization model; the phase features of the speech data are input into the feedforward layer of the near-field sound source localization model, and the intermediate features of the speech data are extracted based on the phase features of the speech data through the feedforward layer; the intermediate features of the speech data are time-averaged to obtain a feature vector of the speech data; the feature vector of the speech data is input into the affine layer (AffineLayer) of the near-field sound source localization model, and the direction of arrival (DOA) of the speech signal corresponding to the speech data is output through the affine layer, where the direction of arrival of the speech signal is the azimuth of the sound source corresponding to the speech data relative to the electronic device.

[0117] S204: Use a far-field sound source localization model to locate the sound source and determine the sound source position of the voice data.

[0118] In one embodiment of the present application, the far-field sound source localization model can be a sound source angle detection model, and the sound source angle detection model is a generalized cross correlation-phase transform (GCC-PHAT) algorithm. After collecting the voice data, the audio features of the voice data are extracted, and the voice data is cross-correlated by the generalized cross correlation-phase transform algorithm to obtain a cross-correlation function of the voice data. Among them, the audio signals x1 and x2 received by two microphones (for example, two microphones in a microphone array of an electronic device) are respectively:

[0119] x1(t)=α1s(t-τ1)+n1(t) (3);

[0120] x2(t)=α2s(t-τ2)+n2(t) (4);

[0121] In calculation formulas (3) and (4), the audio signals x1 and x2 are the audio signals obtained by fast Fourier transform of the audio data collected by the microphones, s(t) is the sound source signal, n1(t) and n2(t) are the environmental noises, and τ1 and τ2 are the propagation times of the sound source signal from the sound source to the two microphones.

[0122] The cross-correlation function between the audio signals x1 and x2 received by the two microphones is subjected to phase transformation weighting. The calculation formula (5) of the phase transformation weighting is:

[0123]

[0124] The calculation formula (6) of the cross-correlation function between the phase-shifted weighted audio signals x1 and x2 is:

[0125]

[0126] In the calculation formula (6), x1 and x2 are the audio signals collected by two microphones, α1 and α2 are the phases of the audio signals collected by two microphones, τ is the current moment, and τ 12 The time delay between the two microphones collecting the audio signal is obtained by estimating the peak value of the cross-correlation function. 12 , according to the time delay τ between the audio signals collected by the two microphones 12 , the distance d between the two microphones is calculated to obtain the sound source angle, and the sound source angle is used as the sound source position of the speech data.

[0127] See Figure 9The figure shows a schematic diagram of the sound source angle provided by an embodiment of the present application. The distance between the two microphones of the electronic device is d, the sound source angle is θ, Δr is the distance difference between the two microphones and the sound source, Δr = τ 12 ×c, c is the propagation speed of the audio signal. Among them, the sound source angle is θ=arcsin(τ 12 ×c / d).

[0128] Through the above-mentioned embodiments of the present application, when the sound source corresponding to the voice data is close to the electronic device, the near-field sound source localization model can be used to locate the sound source; when the sound source corresponding to the voice data is far away from the electronic device, the far-field sound source localization model can be used to locate the sound source, thereby accurately determining the position of the sound source.

[0129] See Figure 10 FIG. 2 is a flowchart of determining the sound source location of speech data provided in another embodiment of the present application.

[0130] S301, identifying a face area in first video data captured by a first camera and second video data captured by a second camera.

[0131] In one embodiment of the present application, each video frame in the first video data and the second video data is input into a face recognition model, and the face region in each video frame in the first video data and the second video data is identified by the face recognition model. The face recognition model is trained and generated using facial features as training data. The face recognition model can be a convolutional neural network model, a cascade classifier based on Haar features, a Dlib face detection algorithm, or the like.

[0132] In step S302, lip movement detection is performed on the facial regions in the first and second video data, respectively, to determine whether the first video data or the second video data contains lip movement. If the first video data contains lip movement, the process proceeds to step S303; if the second video data contains lip movement, the process proceeds to step S304.

[0133] In one embodiment of the present application, the face recognition model is further used to output the coordinates of multiple feature points of the lip area in the face area of ​​each video frame in the first video data and the second video data, and the multiple feature point coordinates of the lip area of ​​multiple video frames in the first video data and the second video data are compared respectively. If the multiple feature point coordinates of the lip area of ​​any two video frames are different, it is determined that the corresponding video data contains lip movements. That is to say, if the multiple feature point coordinates of the lip area of ​​any two video frames in the first video data are different, it is determined that the first video data contains lip movements. If the multiple feature point coordinates of the lip area of ​​any two video frames in the second video data are different, it is determined that the second video data contains lip movements. In one embodiment of the present application, the multiple feature point coordinates of the lip area may include the coordinates of four endpoint positions at the left and right ends and the upper and lower ends.

[0134] S303: Determine whether the sound source is facing the first display screen.

[0135] In one embodiment of the present application, if the first video data captured by the first camera includes lip movements, it is determined that the sound source position is in front of the first camera, that is, the sound source position is facing the first display screen.

[0136] S304: Determine whether the sound source is facing the second display screen.

[0137] In one embodiment of the present application, if the second video data captured by the second camera includes lip movements, it is determined that the sound source position is in front of the second camera, that is, the sound source position is facing the second display screen.

[0138] The above-mentioned embodiments of the present application can accurately determine the sound source position, that is, the position of the speaking user of the voice data relative to the electronic device, by performing lip movement detection on the video data captured by the camera.

[0139] See Figure 11 FIG. 1 is a flow chart of a speech translation method provided by another embodiment of the present application. The method is applied to an electronic device and includes:

[0140] S401, collecting voice data through a microphone.

[0141] S402, capture video data of a user in front of a first display screen through a first camera, capture video data of a user in front of a second display screen through a second camera, and display the video data captured by the first camera on the second display screen, and display the video data captured by the second camera on the first display screen.

[0142] At step S403, based on the video data captured by the first camera, it is determined whether there are multiple users in front of the first display screen. Based on the video data captured by the second camera, it is determined whether there are multiple users in front of the second display screen. If the number of users in front of the first display screen and the number of users in front of the second display screen are both the same, the process proceeds to step S404. If the number of users in front of the first display screen and / or the number of users in front of the second display screen are both the same, the process proceeds to step S405.

[0143] In one embodiment of the present application, face detection is performed on the video data captured by the first camera and the second camera respectively through a face recognition model, and the number of faces in the video data captured by the first camera and the second camera is determined according to the face detection results. If there is only one face in the video data captured by the first camera and the second camera, it is determined that the number of users in front of the first display screen and the number of users in front of the second display screen are one; if there are multiple faces in the video data captured by the first camera and / or the second camera, it is determined that the number of users in front of the first display screen and / or the number of users in front of the second display screen are multiple.

[0144] S404: Determine the sound source position of the voice data using a far-field sound source localization model or a near-field sound source localization model, and determine the target display screen corresponding to the voice data according to the sound source position.

[0145] In one embodiment of the present application, the distance between the sound source of the voice data and the microphone array composed of multiple microphones is calculated, and the near-field sound source localization model or the far-field sound source localization model is selected to locate the sound source according to the distance between the sound source and the microphone array to determine the sound source position of the voice data. If the distance between the sound source and the microphone is less than or equal to the preset distance, the sound source is located according to the near-field sound source localization model to determine the sound source position of the voice data; if the distance between the sound source and the microphone is greater than the preset distance, the sound source is located according to the far-field sound source localization model to determine the sound source position of the voice data. The preset distance is 2D 2 / λ, D = n*d, where d is the distance between adjacent elements in the microphone array, n is the number of array pitches, and λ is the wavelength of the highest-frequency speech from the source, or the minimum wavelength of the speech output by the source. If the distances between all microphones in the array are the same, d can be the distance between any two microphones; if the distances between all microphones in the array are different, d can be the average distance between the two microphones.

[0146] The specific method of using the far-field sound source localization model or the near-field sound source localization model to determine the sound source position of the voice data can refer to the above-mentioned method of the present application. Figure 8 The above description is made in the embodiment and will not be repeated here.

[0147] S405: Convert the voice data into text data, and obtain translation information corresponding to the voice data based on the text data.

[0148] S406: Display the translated information on the target display screen.

[0149] In one embodiment of the present application, if the number of users in front of the first display screen and the number of users in front of the second display screen are one, based on the following Figure 6 The manner shown displays the translation information corresponding to the voice data.

[0150] S407: Determine a speaker of the voice data based on lip movement detection, and determine a target display screen corresponding to the voice data according to the speaker.

[0151] In one embodiment of the present application, the user's lip movements can also be detected based on the video data captured by the first camera and the second camera. When lip movements are detected in the video data, the speaking user of the voice data is determined to be the user who produced the lip movements.

[0152] In one embodiment of the present application, after determining the speaking user of the voice data, a display screen displaying video data containing the speaking user is determined based on the speaking user of the voice data, and the display screen displaying video data containing the speaking user is used as the target display screen corresponding to the voice data. This is equivalent to determining the sound source position of the voice data by using lip movement detection, and determining the target display screen corresponding to the voice data based on the sound source position.

[0153] S408: Convert the voice data into text data, and obtain translation information corresponding to the voice data based on the text data.

[0154] The specific implementations of S401 - S402 , S405 - S406 , and S408 are the same as those of S101 - S305 and are not described in detail here.

[0155] S409: Display the translated information on the target display screen and indicate the speaker corresponding to the translated information.

[0156] See Figure 12 As shown, another interface diagram of the speech translation application provided by one embodiment of the present application is shown. When the video data displayed on the target display screen includes multiple users, in addition to displaying the translation information window near the portrait area or face area in the video data displayed on the target display screen, a marking component is also displayed on the translation information window corresponding to the currently collected speech data to indicate the speaking user corresponding to the translation information, so that other users can easily identify the speaking user corresponding to the speech data. The marking component can be a triangular component displayed at the bottom of the translation information window near the end of the speaking user.

[0157] Through the above-mentioned embodiments of the present application, in a scenario of multi-person conversation, that is, when there are multiple speakers on one side of an electronic device, lip movement detection and voiceprint recognition are used to accurately locate the speaker corresponding to the voice data, and the corresponding speaker is indicated through translation information, so that the user can determine who is currently speaking, thereby effectively improving the user's communication experience.

[0158] In another embodiment of the present application, S407 may also be replaced by: determining the speaking user of the voice data based on lip movement detection and voiceprint recognition, and determining the target display screen corresponding to the voice data according to the speaking user.

[0159] In another embodiment of the present application, target voiceprint features of the voice data of multiple users in the video data displayed on the first display screen or the second display screen are pre-acquired. Before performing voice translation, a microphone is controlled to collect the voice data of multiple users in the video data displayed on the first display screen or the second display screen. The voice data of each user is input into a voiceprint recognition model, the target voiceprint features of each user's voice data are extracted, and an association is established between the user's facial features and the target voiceprint features. The voiceprint recognition model is trained and generated using the voiceprint features of different users as training data. The voiceprint recognition model can be a Gaussian mixture model, a support vector machine model, a convolutional neural network model, etc. The voiceprint features can be Mel-frequency cepstral coefficients, resonance features, fundamental frequency, and intonation features, etc.

[0160] In another embodiment of the present application, the user who generates lip movements in the video data is determined based on lip movement detection, and the actual voiceprint features of the real-time collected voice data are extracted through a voiceprint recognition model, and the actual voiceprint features are compared with the target voiceprint features of each user to determine the user corresponding to the actual voiceprint features. If the user who generates lip movements in the video data and the user corresponding to the actual voiceprint features are the same user, then the user who generates lip movements in the video data is determined to be the speaking user of the voice data; if the user who generates lip movements in the video data and the user corresponding to the actual voiceprint features are not the same user, then the user who generates lip movements in the video data is determined to be not the speaking user of the voice data.

[0161] In another embodiment of the present application, the actual voiceprint features of the speech data are compared with the target voiceprint features of each user. If the actual voiceprint features of the speech data match the target voiceprint features of a user, the facial features corresponding to the target voiceprint features matching the actual voiceprint features of the speech data are determined based on the correlation between the user's facial features and the target voiceprint features. The facial features of the user who makes lip movements in the video data are compared with the facial features corresponding to the target voiceprint features matching the actual voiceprint features. If the facial features of the user who makes lip movements in the video data match the facial features corresponding to the target voiceprint features matching the actual voiceprint features, it is determined that the user who makes lip movements in the video data and the user corresponding to the actual voiceprint features are the same user; if the facial features of the user who makes lip movements in the video data do not match the facial features corresponding to the target voiceprint features matching the actual voiceprint features, it is determined that the user who makes lip movements in the video data and the user corresponding to the actual voiceprint features are not the same user.

[0162] In another embodiment of the present application, if the similarity between the actual voiceprint features of the speech data and the target voiceprint features of a user is greater than or equal to a first preset percentage, it is determined that the actual voiceprint features of the speech data match the target voiceprint features of the user; if the similarity between the actual voiceprint features of the speech data and the target voiceprint features of all users is less than the first preset percentage, it is determined that the actual voiceprint features of the speech data do not match the target voiceprint features of all users. For example, the first preset percentage is 93%, 95%, 97%, or other percentage values. The similarity between the actual voiceprint features and the target voiceprint features can be represented by Euclidean distance or cosine similarity.

[0163] In another embodiment of the present application, if the similarity between the facial features of the user who makes the lip movement in the video data and the facial features corresponding to the target voiceprint features that match the actual voiceprint features is greater than or equal to a second preset percentage, it is determined that the facial features of the user who makes the lip movement in the video data and the facial features corresponding to the target voiceprint features that match the actual voiceprint features match; if the similarity between the facial features of the user who makes the lip movement in the video data and the facial features corresponding to the target voiceprint features that match the actual voiceprint features is less than a second preset percentage, it is determined that the facial features of the user who makes the lip movement in the video data and the facial features corresponding to the target voiceprint features that match the actual voiceprint features do not match. For example, the second preset percentage is 93%, 95%, 97% or other percentage values. The similarity between the facial features of the user who makes the lip movement in the video data and the facial features corresponding to the target voiceprint features that match the actual voiceprint features can be represented by Euclidean distance or cosine similarity.

[0164] The present application also provides an electronic device 100. Figure 13As shown, the electronic device 100 can be a mobile phone, a tablet computer, a desktop computer, a laptop computer, a handheld computer, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a cellular phone, a personal digital assistant (PDA), an augmented reality (AR) device, a virtual reality (VR) device, an artificial intelligence (AI) device, a wearable device, an in-vehicle device, a smart home device and / or a smart city device. The embodiment of the present application does not impose any special restrictions on the specific type of the electronic device 100.

[0165] The electronic device 100 may include a processor 110, an external memory interface 120, an internal memory 121, a Universal Serial Bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a speaker 170A, a receiver 170B, a microphone 170C, an earphone interface 170D, a sensor module 180, a button 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a Subscriber Identification Module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, an air pressure sensor 180C, a magnetic sensor 180D, an acceleration sensor 180E, a distance sensor 180F, a proximity light sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0166] It should be understood that the structure illustrated in the embodiments of the present invention does not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown, or may combine or separate certain components, or arrange the components differently. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.

[0167] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors.

[0168] The controller can generate operation control signals according to the instruction operation code and timing signal to complete the control of instruction fetching and execution.

[0169] Processor 110 may also include a memory for storing instructions and data. In one embodiment of the present application, the memory in processor 110 is a cache memory. The memory can store instructions or data that have just been used or are being recycled by processor 110. If processor 110 needs to use an instruction or data again, it can directly call it from the memory. This avoids repeated accesses, reduces processor 110 latency, and thus improves system efficiency.

[0170] In one embodiment of the present application, the processor 110 may include one or more interfaces. The interfaces may include an Inter-integrated Circuit (I2C) interface, an Inter-integrated Circuit Sound (I2S) interface, a Pulse Code Modulation (PCM) interface, a Universal Asynchronous Receiver / Transmitter (UART) interface, a Mobile Industry Processor Interface (MIPI), a General-Purpose Input / Output (GPIO) interface, a Subscriber Identity Module (SIM) interface, and / or a Universal Serial Bus (USB) interface.

[0171] The I2C interface is a bidirectional synchronous serial bus that includes a serial data line (SDA) and a serial clock line (SCL). In one embodiment of the present application, the processor 110 may include multiple I2C buses. The processor 110 may be coupled to the touch sensor 180K, charger, flash, camera 193, etc. through different I2C bus interfaces. For example, the processor 110 may be coupled to the touch sensor 180K through the I2C interface, so that the processor 110 and the touch sensor 180K communicate through the I2C bus interface, thereby realizing the touch function of the electronic device 100.

[0172] The I2S interface can be used for audio communication. In one embodiment of the present application, the processor 110 can include multiple I2S buses. The processor 110 can be coupled to the audio module 170 via the I2S bus to enable communication between the processor 110 and the audio module 170. In one embodiment of the present application, the audio module 170 can transmit audio signals to the wireless communication module 160 via the I2S interface, enabling the function of answering calls through a Bluetooth headset.

[0173] The PCM interface can also be used for audio communication, sampling, quantizing and encoding analog signals. In one embodiment of the present application, the audio module 170 and the wireless communication module 160 can be coupled via a PCM bus interface. In one embodiment of the present application, the audio module 170 can also transmit audio signals to the wireless communication module 160 via the PCM interface, thereby realizing the function of answering calls via a Bluetooth headset. Both the I2S interface and the PCM interface can be used for audio communication.

[0174] The UART interface is a universal serial data bus used for asynchronous communication. The bus can be a bidirectional communication bus. It converts the data to be transmitted between serial communication and parallel communication. In one embodiment of the present application, the UART interface is generally used to connect the processor 110 and the wireless communication module 160. For example, the processor 110 communicates with the Bluetooth module in the wireless communication module 160 via the UART interface to implement Bluetooth functionality. In one embodiment of the present application, the audio module 170 can transmit audio signals to the wireless communication module 160 via the UART interface, enabling the function of playing music through a Bluetooth headset.

[0175] The MIPI interface can be used to connect the processor 110 to peripheral devices such as the display screen 194 and the camera 193. MIPI interfaces include the Camera Serial Interface (CSI) and the Display Serial Interface (DSI). In one embodiment of the present application, the processor 110 and the camera 193 communicate via the CSI interface to implement the camera function of the electronic device 100. The processor 110 and the display screen 194 communicate via the DSI interface to implement the display function of the electronic device 100.

[0176] The GPIO interface can be configured through software. The GPIO interface can be configured as a control signal or a data signal. In one embodiment of the present application, the GPIO interface can be used to connect the processor 110 to the camera 193, the display 194, the wireless communication module 160, the audio module 170, the sensor module 180, etc. The GPIO interface can also be configured as an I2C interface, an I2S interface, a UART interface, a MIPI interface, etc.

[0177] The USB interface 130 is an interface that complies with USB standards and may be a Mini USB interface, a Micro USB interface, a USB Type-C interface, or the like. The USB interface 130 can be used to connect a charger to charge the electronic device 100, or to transfer data between the electronic device 100 and peripheral devices. It can also be used to connect headphones to play audio. The interface can also be used to connect other electronic devices 100, such as augmented reality devices.

[0178] It is understood that the interface connection relationship between the modules illustrated in the embodiment of the present invention is merely an illustrative illustration and does not constitute a structural limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may also adopt different interface connection methods from the above embodiments, or a combination of multiple interface connection methods.

[0179] The charging management module 140 is configured to receive charging input from a charger. The charger can be either a wireless charger or a wired charger. In some wired charging embodiments, the charging management module 140 can receive charging input from the wired charger via the USB interface 130. In some wireless charging embodiments, the charging management module 140 can receive wireless charging input via the wireless charging coil of the electronic device 100. While charging the battery 142, the charging management module 140 can also provide power to the electronic device 100 via the power management module 141.

[0180] The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives input from the battery 142 and / or the charging management module 140, and provides power to the processor 110, the internal memory 121, the display 194, the camera 193, and the wireless communication module 160. The power management module 141 can also be used to monitor parameters such as battery capacity, battery cycle count, and battery health status (leakage, impedance). In some other embodiments, the power management module 141 can also be set in the processor 110. In other embodiments, the power management module 141 and the charging management module 140 can also be set in the same device.

[0181] The wireless communication function of the electronic device 100 can be implemented through the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modem processor and the baseband processor.

[0182] Antenna 1 and Antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be reused to improve antenna utilization. For example, antenna 1 can be reused as a diversity antenna for a wireless local area network. In other embodiments, the antennas can be used in conjunction with a tuning switch.

[0183] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G applied to the electronic device 100. The mobile communication module 150 may include at least one filter, a switch, a power amplifier, a low noise amplifier (LNA), etc. The mobile communication module 150 can receive electromagnetic waves from the antenna 1, and filter, amplify, and process the received electromagnetic waves, and transmit them to the modulation and demodulation processor for demodulation. The mobile communication module 150 can also amplify the signal modulated by the modulation and demodulation processor, and convert it into electromagnetic waves for radiation through the antenna 1. In one embodiment of the present application, at least some of the functional modules of the mobile communication module 150 can be set in the processor 110. In one embodiment of the present application, at least some of the functional modules of the mobile communication module 150 can be set in the same device as at least some of the modules of the processor 110.

[0184] The modem processor may include a modulator and a demodulator. The modulator is used to modulate the low-frequency baseband signal to be transmitted into a medium-high frequency signal. The demodulator is used to demodulate the received electromagnetic wave signal into a low-frequency baseband signal. The demodulator then transmits the demodulated low-frequency baseband signal to the baseband processor for processing. After being processed by the baseband processor, the low-frequency baseband signal is passed to the application processor. The application processor outputs a sound signal through an audio device (not limited to the speaker 170A, the receiver 170B, etc.) or displays an image or video through the display screen 194. In one embodiment of the present application, the modem processor may be an independent device. In other embodiments, the modem processor may be independent of the processor 110 and be set in the same device as the mobile communication module 150 or other functional modules.

[0185] The wireless communication module 160 can provide wireless communication solutions including Wireless Local Area Networks (WLAN) (such as Wireless Fidelity (Wi-Fi) network), Bluetooth (BT), Global Navigation Satellite System (GNSS), Frequency Modulation (FM), Near Field Communication (NFC), Infrared (IR), etc., which are applied to the electronic device 100. The wireless communication module 160 can be one or more devices that integrate at least one communication processing module. The wireless communication module 160 receives electromagnetic waves via the antenna 2, frequency modulates and filters the electromagnetic wave signals, and sends the processed signals to the processor 110. The wireless communication module 160 can also receive the signal to be sent from the processor 110, frequency modulate it, amplify it, and convert it into electromagnetic waves for radiation through the antenna 2.

[0186] In one embodiment of the present application, the antenna 1 of the electronic device 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the electronic device 100 can communicate with the network and other devices through wireless communication technology. The wireless communication technology may include Global System For Mobile Communications (GSM), General Packet Radio Service (GPRS), Code Division Multiple Access (CDMA), Wideband Code Division Multiple Access (WCDMA), Time-Division Code Division Multiple Access (TD-SCDMA), Long Term Evolution (LTE), BT, GNSS, WLAN, NFC, FM, and / or IR technology. The GNSS may include a Global Positioning System (GPS), a Global Navigation Satellite System (GLONASS), a Beidou Navigation Satellite System (BDS), a Quasi-Zenith Satellite System (QZSS) and / or a Satellite Based Augmentation System (SBAS).

[0187] Electronic device 100 implements display functionality through a GPU, display screen 194, and an application processor. The GPU is a microprocessor for speech translation, connecting display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0188] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. The display panel can be a liquid crystal display (LCD), an organic light-emitting diode (OLED), an active-matrix organic light-emitting diode or an active-matrix organic light-emitting diode (AMOLED), a flexible light-emitting diode (FLED), a mini-LED, a micro-LED, a micro-OLED, a quantum dot light-emitting diode (QLED), etc. In one embodiment of the present application, the electronic device 100 may include one or N display screens 194, where N is a positive integer greater than one.

[0189] The electronic device 100 can implement a shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, and an application processor.

[0190] The ISP is used to process data fed back by camera 193. For example, when taking a photo, the shutter is opened, and light is transmitted through the lens to the camera's photosensitive element. The light signal is converted into an electrical signal, which is then transmitted to the ISP for processing and converted into an image visible to the naked eye. The ISP can also perform algorithmic optimization on image noise, brightness, and skin color. The ISP can also optimize parameters such as exposure and color temperature of the captured scene. In one embodiment of the present application, the ISP can be located in camera 193.

[0191] The camera 193 is used to capture still images or videos. The object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal oxide semiconductor (CMOS) phototransistor. The photosensitive element converts the light signal into an electrical signal, and then passes the electrical signal to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV or other format. In one embodiment of the present application, the electronic device 100 may include 1 or N cameras 193, where N is a positive integer greater than 1.

[0192] The digital signal processor is used to process digital signals. In addition to processing digital image signals, it can also process other digital signals. For example, when the electronic device 100 selects a frequency point, the digital signal processor is used to perform Fourier transform on the frequency point energy.

[0193] Video codecs are used to compress or decompress digital video. Electronic device 100 may support one or more video codecs. This allows electronic device 100 to play or record videos in various encoding formats, such as Moving Picture Experts Group (MPEG) 1, MPEG2, MPEG3, and MPEG4.

[0194] The NPU is a neural network (NN) computing processor that rapidly processes input information and continuously self-learns by drawing on the structure of biological neural networks, such as the transmission patterns between neurons in the human brain. The NPU enables intelligent cognitive applications in electronic device 100, such as image recognition, face recognition, speech recognition, and text comprehension.

[0195] The internal memory 121 may include one or more random access memories (RAM) and one or more non-volatile memories (NVM).

[0196] Random access memory may include static random-access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM, for example, the fifth generation DDR SDRAM is generally referred to as DDR5 SDRAM), etc.

[0197] Non-volatile memory may include disk storage devices and flash memory.

[0198] Flash memory can be divided into NOR FLASH, NAND FLASH, 3D NAND FLASH, etc. according to the operating principle; it can be divided into single-level cell (SLC), multi-level cell (MLC), triple-level cell (TLC), quad-level cell (QLC), etc. according to the storage specification; it can be divided into universal flash storage (UFS), embedded multi media card (eMMC), etc. according to the storage specification.

[0199] The random access memory can be directly read and written by the processor 110, and can be used to store executable programs (such as machine instructions) of the operating system or other running programs, and can also be used to store user and application data.

[0200] The non-volatile memory may also store executable programs and user and application data, etc., and may be loaded into the random access memory in advance for direct reading and writing by the processor 110 .

[0201] The external memory interface 120 can be used to connect to an external non-volatile memory to expand the storage capacity of the electronic device 100. The external non-volatile memory communicates with the processor 110 via the external memory interface 120 to implement data storage. For example, files such as music and videos can be stored in the external non-volatile memory.

[0202] The internal memory 121 or the external memory interface 120 is used to store one or more computer programs. The one or more computer programs are configured to be executed by the processor 110. The one or more computer programs include multiple instructions. When the multiple instructions are executed by the processor 110, the screen display detection method executed on the electronic device 100 in the above embodiment can be implemented to realize the screen display detection function of the electronic device 100.

[0203] The electronic device 100 can implement audio functions such as music playback and recording through the audio module 170, the speaker 170A, the receiver 170B, the microphone 170C, the headphone jack 170D, and the application processor.

[0204] The audio module 170 is used to convert digital audio information into analog audio signal output, and is also used to convert analog audio input into digital audio signals. The audio module 170 can also be used to encode and decode audio signals. In one embodiment of the present application, the audio module 170 can be provided in the processor 110, or some functional modules of the audio module 170 can be provided in the processor 110.

[0205] The speaker 170A, also called a "speaker", is used to convert audio electrical signals into sound signals. The electronic device 100 can listen to music or listen to hands-free calls through the speaker 170A.

[0206] The receiver 170B, also called a "handset", is used to convert audio electrical signals into sound signals. When the electronic device 100 receives a call or a voice message, the user can place the receiver 170B close to the ear to hear the voice.

[0207] Microphone 170C, also known as "microphone" or "microphone", is used to convert sound signals into electrical signals. When making a call or sending a voice message, the user can speak by putting their mouth close to the microphone 170C to input the sound signal into the microphone 170C. The electronic device 100 can be provided with at least one microphone 170C. In other embodiments, the electronic device 100 can be provided with two microphones 170C, which can not only collect sound signals but also realize noise reduction function. In other embodiments, the electronic device 100 can also be provided with three, four or more microphones 170C to collect sound signals, reduce noise, identify the source of sound, realize directional recording function, etc.

[0208] The headphone jack 170D is used to connect a wired headphone and can be a USB interface 130 or a 3.5mm Open Mobile Terminal Platform (OMTP) standard interface or a Cellular Telecommunications Industry Association of the USA (CTIA) standard interface.

[0209] The buttons 190 include a power button, a volume button, and the like. The buttons 190 may be mechanical buttons or touch buttons. The electronic device 100 may receive key inputs and generate key signal inputs related to user settings and function control of the electronic device 100.

[0210] Motor 191 can generate vibration prompts. Motor 191 can be used for incoming call vibration prompts, and can also be used for touch vibration feedback. For example, touch operations acting on different applications (such as taking pictures, audio playback, etc.) can correspond to different vibration feedback effects. For touch operations acting on different areas of the display screen 194, motor 191 can also correspond to different vibration feedback effects. Different application scenarios (for example: time reminders, receiving messages, alarm clocks, games, etc.) can also correspond to different vibration feedback effects. The touch vibration feedback effect can also support customization.

[0211] The indicator 192 may be an indicator light, which may be used to indicate the charging status, power level changes, messages, missed calls, notifications, etc.

[0212] The SIM card interface 195 is used to connect a SIM card. The SIM card can be connected to and separated from the electronic device 100 by inserting it into or removing it from the SIM card interface 195. The electronic device 100 can support 1 or N SIM card interfaces, where N is a positive integer greater than 1. The SIM card interface 195 can support Nano SIM cards, Micro SIM cards, SIM cards, etc. Multiple cards can be inserted into the same SIM card interface 195 at the same time. The types of the multiple cards can be the same or different. The SIM card interface 195 can also be compatible with different types of SIM cards. The SIM card interface 195 can also be compatible with external memory cards. The electronic device 100 interacts with the network through the SIM card to implement functions such as calls and data communications. In one embodiment of the present application, the electronic device 100 uses an eSIM, i.e., an embedded SIM card. The eSIM card can be embedded in the electronic device 100 and cannot be separated from the electronic device 100. The embodiment of the present application further provides a computer storage medium, in which computer instructions are stored. When the computer instructions are executed on the electronic device 100, the electronic device 100 executes the above-mentioned related method steps to implement the speech translation method in the above-mentioned embodiment.

[0213] The embodiment of the present application further provides a computer program product. When the computer program product is run on a computer, the computer is caused to execute the above-mentioned related steps to implement the speech translation method in the above-mentioned embodiment.

[0214] In addition, an embodiment of the present application also provides a device, which can specifically be a chip, component or module. The device may include a connected processor and memory; wherein the memory is used to store computer-executable instructions. When the device is running, the processor can execute the computer-executable instructions stored in the memory to enable the chip to perform the speech translation method in the above-mentioned method embodiments.

[0215] Among them, the electronic device, computer storage medium, computer program product or chip provided in this embodiment is used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be repeated here.

[0216] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.

[0217] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of modules or units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0218] Units described as separate components may or may not be physically separate, and components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0219] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0220] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a device (which can be a single-chip microcomputer, chip, etc.) or a processor to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0221] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit the present application. Although the present application has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present application may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present application.

Claims

1. A speech translation method, applied to electronic equipment, characterized in that: The electronic device includes a first camera, a second camera, a first display screen, a second display screen, and at least one microphone, and the method includes: collecting voice data through the at least one microphone; capturing video data of a user in front of the first display screen through the first camera, capturing video data of a user in front of the second display screen through the second camera, and displaying the video data captured by the first camera on the second display screen, and displaying the video data captured by the second camera on the first display screen; Determining a target display screen corresponding to the voice data according to a sound source position of the voice data; Converting the speech data into text data, and obtaining translation information corresponding to the speech data based on the text data; The translated information is displayed on the target display screen.

2. The speech translation method according to claim 1, wherein: The electronic device includes a plurality of microphones, and the plurality of microphones form a microphone array. The method further includes: determining the sound source position according to the voice data, including: Calculating the distance between the sound source of the voice data and the microphone array; If the distance between the sound source and the microphone array is less than or equal to a preset distance, localize the sound source using a near-field sound source localization model to determine the sound source position of the voice data; or If the distance between the sound source and the microphone array is greater than the preset distance, the sound source is located using a far-field sound source localization model to determine the sound source position of the voice data.

3. The speech translation method according to claim 2, wherein: The calculating the distance between the sound source of the voice data and the microphone array includes: Obtaining a timestamp of when each microphone collects the voice data, and determining a time delay difference between the plurality of microphones in collecting the voice data according to the timestamp of when each microphone collects the voice data; Calculating the energy of the speech signal corresponding to the speech data collected by each microphone, and determining the energy difference of the speech data collected by the multiple microphones; Calculating the distance between the sound source and each microphone according to the delay difference of the voice data collected by the multiple microphones in the microphone array and the energy difference of the voice signals collected by the multiple microphones; The distance between the sound source of the voice data and the microphone array is determined according to the distance between the sound source and each microphone.

4. The speech translation method according to claim 2, wherein: The preset distance is 2D 2 / λ, D=n*d, d is the distance between adjacent array elements in the microphone array, n is the number of array spacings, and λ is the wavelength of the highest frequency speech of the sound source.

5. The speech translation method according to claim 2, wherein: The adopting a near-field sound source localization model to locate the sound source and determine the sound source position of the voice data includes: extracting a phase spectrum of the speech data; Inputting the phase spectrum of the speech data into the near-field sound source localization model, and extracting the phase features of the speech data through the convolution layer of the near-field sound source localization model; Inputting the phase feature of the speech data into a feedforward layer of the near-field sound source localization model, and extracting intermediate features of the speech data based on the phase feature of the speech data through the feedforward layer; Performing time averaging on the intermediate features of the speech data to obtain a feature vector of the speech data; The feature vector of the speech data is input into the affine layer of the near-field sound source localization model, and the arrival direction of the speech signal corresponding to the speech data is output through the affine layer.

6. The speech translation method according to claim 2, wherein: The adopting a far-field sound source localization model to locate the sound source and determine the sound source position of the voice data includes: Extracting audio features of the speech data; Performing cross-correlation calculation on the voice data using a generalized cross-correlation-phase transform algorithm to obtain a cross-correlation function of the voice data; performing phase transformation weighting on the cross-correlation function of the speech data; The time delay between the multiple microphones collecting the voice data is estimated based on the peak value of the cross-correlation function, the sound source angle is calculated based on the time delay and the distance between the multiple microphones, and the sound source angle is used as the sound source position of the voice data.

7. The speech translation method according to claim 2, wherein: The determining the sound source position according to the voice data includes: Identifying a face area in first video data captured by the first camera and second video data captured by the second camera; performing lip movement detection on the face areas in the first video data and the second video data respectively; If the first video data includes lip movements, determining that the sound source is facing the first display screen; or If the second video data includes the lip movement, it is determined that the sound source position is facing the second display screen.

8. The speech translation method according to claim 7, wherein: The identifying of a face area in the first video data captured by the first camera and the second video data captured by the second camera includes: Each video frame in the first video data and the second video data is input into a face recognition model respectively, and a face region of each video frame in the first video data and the second video data is identified by the face recognition model.

9. The speech translation method according to claim 7, wherein: The performing lip movement detection on the face areas in the first video data and the second video data respectively includes: Outputting, by the face recognition model, coordinates of a plurality of feature points in the lip region of the face region of each video frame in the first video data and the second video data; respectively comparing coordinates of a plurality of feature points of lip regions in a plurality of video frames in the first video data and the second video data; If the coordinates of the multiple feature points in the lip regions of any two video frames are different, it is determined that the corresponding video data contains lip movements.

10. The speech translation method according to claim 1, wherein: The method further comprises: determining whether there are multiple users in front of the first display screen based on the video data captured by the first camera, and determining whether there are multiple users in front of the second display screen based on the video data captured by the second camera; If there are multiple users in front of the first display screen and / or multiple users in front of the second display screen, the speaking user of the voice data is determined based on lip movement detection, and the target display screen corresponding to the voice data is determined based on the speaking user.

11. The speech translation method according to claim 1, wherein: The method further comprises: determining whether there are multiple users in front of the first display screen based on the video data captured by the first camera, and determining whether there are multiple users in front of the second display screen based on the video data captured by the second camera; If there are multiple users in front of the first display screen and / or multiple users in front of the second display screen, the speaking user of the voice data is determined based on lip movement detection and voiceprint recognition, and the target display screen corresponding to the voice data is determined based on the speaking user.

12. The speech translation method according to claim 11, wherein: The determining the speaking user of the voice data based on lip movement detection and voiceprint recognition, and determining the target display screen corresponding to the voice data according to the speaking user, includes: respectively pre-acquire target voiceprint features of voice data of a plurality of users in the video data displayed on the first display screen or the second display screen; Determining the user who made the lip movement in the video data based on lip movement detection, extracting the actual voiceprint features of the voice data through a voiceprint recognition model, comparing the actual voiceprint features with the target voiceprint features of each user, and determining the user corresponding to the actual voiceprint features; If the user who makes the lip movement in the video data and the user corresponding to the actual voiceprint feature are the same user, then the user who makes the lip movement in the video data is determined to be the speaking user of the voice data.

13. The speech translation method according to claim 1, wherein: The converting the voice data into text data comprises: A speech recognition model is used to perform speech recognition on the speech data, and the speech data is converted into the text data.

14. The speech translation method according to claim 1, wherein: The obtaining translation information corresponding to the voice data based on the text data includes: The text data is translated into a target language using a machine translation model to obtain translation information corresponding to the voice data, and the target language is determined according to a target display screen of the translation information.

15. The speech translation method according to claim 1, wherein: The displaying of the translation information on the target display screen includes: The translation information is displayed in the form of a window on the target display screen, and the display position of the window is near the portrait area in the video data displayed on the target display screen.

16. The speech translation method according to claim 15, wherein: The displaying of the translation information on the target display screen includes: If the video data displayed on the target display screen includes multiple users, the translation information is displayed on the target display screen, and the speaking user corresponding to the translation information is marked.

17. An electronic device, characterized in that: The electronic device comprises a memory and a processor: Wherein, the memory is used to store program instructions; The processor is configured to read and execute the program instructions stored in the memory, and when the program instructions are executed by the processor, the electronic device executes the speech translation method according to any one of claims 1 to 16.

18. A chip coupled to a memory in an electronic device, characterized in that: The chip is used to control the electronic device to execute the speech translation method according to any one of claims 1 to 16.

19. A computer storage medium, characterized in that The computer storage medium stores program instructions, and when the program instructions are executed on an electronic device, the processor of the electronic device executes the speech translation method according to any one of claims 1 to 16.