Data processing method, and device

By adding speech rate adjustment controls to the interface of electronic devices and working in conjunction with cloud servers, the problem of lacking real-time speech rate adjustment in simultaneous interpretation was solved, improving the fluency of translation and user experience.

WO2026086349A1PCT designated stage Publication Date: 2026-04-30HONOR DEVICE CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
HONOR DEVICE CO LTD
Filing Date
2025-08-06
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Existing electronic devices lack the ability to adjust speech rate in real time during simultaneous interpretation, resulting in poor translation quality.

Method used

A data processing method and device are provided, which allows users to adjust the reading speed in real time by adding a speech rate adjustment control to the interface of an electronic device, and works in conjunction with a cloud server to generate and play reading audio corresponding to the adjusted speech rate.

Benefits of technology

It enables real-time speech rate adjustment during simultaneous interpreting, improving the fluency of translation and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025112972_30042026_PF_FP_ABST
    Figure CN2025112972_30042026_PF_FP_ABST
Patent Text Reader

Abstract

The embodiments of the present application relate to the field of electronic devices. Provided are a data processing method, and a device. The method can provide a real-time speech speed adjustment function in a simultaneous interpretation function. The method may comprise: displaying a first interface, which is an interface of a first application, wherein the first interface comprises a first control, which is used for adjusting the speech speed of a read-aloud audio during playback; receiving a first operation, which is used for adjusting the read-aloud speech speed to a first speech speed by means of the first control; and playing a first read-aloud audio, wherein read-aloud content of the first read-aloud audio comprises first pre-read-aloud text, which is included in translated content of an original audio, and the read-aloud speed of the first read-aloud audio is the first speech speed.
Need to check novelty before this filing date? Find Prior Art

Description

A data processing method and apparatus

[0001] This application claims priority to Chinese Patent Application No. 202411488570.X, filed on October 23, 2024, entitled "A Data Processing Method and Apparatus", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of electronic equipment technology, and in particular to a data processing method and apparatus. Background Technology

[0003] Currently, electronic devices can provide simultaneous interpretation capabilities through applications installed on them.

[0004] For example, an electronic device can acquire ambient sound from the current environment and report this ambient sound (raw audio) to a cloud server. Correspondingly, the cloud server can perform processing on the raw audio, such as recognition, translation, and segmentation, and send the processed text data to the electronic device. Thus, the electronic device can display the original text before translation and the translated text on its screen.

[0005] Furthermore, the electronic device can send the pre-read text to a cloud server, which can then generate corresponding audio and send it to the device. The device can then play this audio to read the translated text aloud, achieving a simultaneous interpretation effect. Summary of the Invention

[0006] This application provides a data processing method and apparatus that can provide real-time speech rate adjustment functionality in simultaneous interpretation.

[0007] To achieve the above technical objectives, this application adopts the following technical solution:

[0008] Firstly, a data processing method is provided, applied to an electronic device. The electronic device has a first application installed, which provides simultaneous interpretation functionality. When the simultaneous interpretation function is activated, the electronic device plays the translated content corresponding to the acquired original audio. The method includes: displaying a first interface, which is the interface of the first application. The first interface includes a first control for adjusting the playback speed of the audio. Receiving a first operation, which adjusts the playback speed to a first speed via the first control. Playing a first audio recording, the audio recording including a first pre-read text, which is included in the translated content of the original audio. The playback speed of the first audio recording is the first speed.

[0009] Based on this solution, electronic devices can provide users with controls to adjust the speech rate while offering simultaneous interpretation functionality. Thus, after the user adjusts the speech rate using these controls (such as the first control), the electronic device can play a first audio recording corresponding to the adjusted speech rate. This achieves real-time speech rate adjustment in simultaneous interpretation.

[0010] Optionally, the first interface further includes a second control for enabling the simultaneous interpretation function. Before receiving the first operation, the method further includes receiving a second operation on the second control, the second operation being used to activate the simultaneous interpretation function.

[0011] Optionally, the first control includes: a speech rate adjustment bar and an adjustment button. The position of the adjustment button on the speech rate adjustment bar corresponds to the speech rate. The first operation corresponds to adjusting the position of the adjustment button on the speech rate adjustment bar to the position corresponding to the first speech rate.

[0012] For example, when the adjustment button is located at the far left of the speech rate adjustment bar, it corresponds to the slowest speech rate. As the user drags the adjustment button to the right, the closer the button is to the far right of the speech rate adjustment bar, the faster the speech rate.

[0013] Optionally, before playing the first audio recording, the method further includes: acquiring the original audio and sending the original audio to a cloud server; acquiring a first original text and a first translated text. The first original text corresponds to at least a portion of the original text corresponding to the original audio, and the first translated text corresponds to at least a portion of the translated text corresponding to the original audio. The first original audio uses a first language, the first original text uses the first language, and the first translated text uses a second language, wherein the first language and the second language are different.

[0014] Optionally, the raw audio may include ambient sound. In other implementations, the raw audio may also be provided by other applications installed on the electronic device, such as a calling application.

[0015] Optionally, the first language is English and the second language is Chinese.

[0016] Optionally, after obtaining the first original text and the first translated text, the method further includes: displaying the first original text and the first translated text on the first interface.

[0017] This enables the display of content in simultaneous interpretation functions.

[0018] Optionally, after obtaining the first original text and the first translated text, the method further includes: determining the first pre-reading text based on the feedback data corresponding to the first translated text. The feedback data corresponding to the first translated text is the data sent to the electronic device by the cloud server when sending the first translated text.

[0019] Optionally, after determining the first pre-read text, the method further includes: obtaining a speech rate parameter, which indicates the first speech rate; sending the first pre-read text and the speech rate parameter to a cloud server; and obtaining the first audio recording from the cloud server. Based on this, the electronic device can report the user-configured speech rate parameter along with the first pre-read text to the cloud server, thereby ensuring that the speech rate of the first audio recording generated by the cloud server corresponds to the user-configured first speech rate.

[0020] Optionally, the returned data corresponding to the first translated text includes at least first returned data and second returned data. The first returned data and the second returned data have at least partial overlap. This avoids the loss of effective data during data transmission.

[0021] Optionally, the first returned data is the current returned data, and the electronic device has received the second returned data before receiving the first returned data. The receiving first operation includes receiving the first operation during the playback of the second audio recording. The content of the second audio recording is included in the second returned data. Determining the first pre-reading text based on the returned data corresponding to the first translated text includes determining the first pre-reading text based on the first returned data.

[0022] Optionally, after acquiring the first returned data, the method further includes: if the first returned data includes at least one punctuation character, dividing the first returned data into two or more sub-data corresponding to statements based on the punctuation character. Each sub-data corresponds to one statement. This achieves the effect of real-time sentence segmentation.

[0023] Optionally, after acquiring the first returned data, the method further includes: if the first returned data does not include punctuation characters and the first returned data includes an end marker, then identifying the first returned data as data of the first pre-read text. The end marker is used to indicate the end of a paragraph.

[0024] Optionally, the method further includes: recording the first position of the first statement in the first returned data as the read-aloud position, wherein the first statement is the last statement read in the second read-aloud audio.

[0025] Optionally, the method further includes: determining a second statement in the first returned data, wherein the second statement is the statement indicated by the first position in the first returned data. Based on the consistency between the first and second statements, the already read position is continued to be recorded as the first position. Alternatively, based on the inconsistency between the first and second statements, the second position of a third statement in the first returned data is recorded as the already read position. The third statement is the statement in the first returned data with the highest similarity to the first local area. In this way, through comparison, the already read positions in the first returned data can be effectively corrected, thereby improving the accuracy of subsequently determining unread statements.

[0026] Optionally, the method further includes: if the first returned data includes an end marker, determining the data of the unread text in the first returned data as the data of the first pre-read text. The unread text is the data corresponding to the sentence following the read position in the first returned data. The end marker is used to indicate the end of the paragraph.

[0027] Optionally, if the first returned data does not include an end marker but includes one or more punctuation characters, the method further includes: dividing the unread text into pre-read text and pre-reserved text according to a first rule, wherein the unread text corresponds to the unread data; and determining the data of the pre-read text as the data of the first pre-read text.

[0028] Optionally, the first rule includes: if the number of sentences in the unread text is greater than a first threshold, configuring the last N sentences in the unread text as the pre-reserved text, and determining the sentences in the unread text that are different from the pre-reserved text as the pre-read text. N is a preset positive integer.

[0029] Optionally, before determining the data of the pre-read text as the first pre-read text, the method further includes: determining that the data length of the pre-read text is less than a second threshold, and the number of sentences in the unread text is greater than a fourth threshold. Alternatively, determining that the data length of the pre-reserved text is less than a third threshold, and the number of sentences in the unread text is greater than a fourth threshold. Alternatively, determining that the data length of the pre-read text is greater than a second threshold, and the number of sentences in the pre-reserved text is greater than a third threshold. Thus, the electronic device can determine the final first pre-read text from multiple dimensions, such as data length and sentence length, for the pre-read text, the unread text, and the pre-reserved text, thereby improving the rationality and accuracy of the determined first pre-read text.

[0030] In a second aspect, an electronic device is provided, comprising: a memory and one or more processors. The memory and the processors are coupled. The memory is used to store computer program code, which includes computer instructions that, when executed by the processor, cause the electronic device to perform the methods provided in the first aspect and any possible design thereof.

[0031] Thirdly, a chip system is provided for use in an electronic device. The chip system includes one or more interface circuits and one or more processors. The interface circuits and the processors are interconnected via lines. The interface circuits are used to receive signals from the electronic device's memory and send the signals to the processor, the signals including computer instructions stored in the memory. When the processor executes the computer instructions, the electronic device performs the methods provided in the first aspect and any possible design thereof.

[0032] Fourthly, a computer-readable storage medium is provided, including computer instructions that, when executed on an electronic device, cause the electronic device to perform the methods provided in the first aspect and any possible design thereof.

[0033] Fifthly, this application also provides a computer program product that, when run on a computer, causes the computer to execute the technical solutions provided in the first aspect and any possible implementation thereof.

[0034] In a sixth aspect, a communication system is provided, comprising the electronic device provided in the second aspect, and a cloud server. The cloud server is configured to generate a first audio recording based on a first pre-read text and speech rate parameters sent by the electronic device. The cloud server is also configured to send the first audio recording to the electronic device.

[0035] It is understood that the solutions provided in the second to sixth aspects of this application can be respectively associated with the first aspect and any of its possible designs, and therefore the beneficial effects achieved are similar, which will not be repeated here. Attached Figure Description

[0036] Figure 1 is a schematic diagram of an interface interaction provided in an embodiment of this application;

[0037] Figure 2 is a schematic diagram of an interface interaction provided in an embodiment of this application;

[0038] Figure 3 is a schematic diagram illustrating the correspondence between speech rate adjustment and reading content provided in an embodiment of this application;

[0039] Figure 4 is a schematic diagram of a communication scenario provided in an embodiment of this application;

[0040] Figure 5 is a schematic diagram of the composition of an electronic device provided in an embodiment of this application;

[0041] Figure 6 is a schematic diagram of the inter-device interaction process of a data processing method provided in an embodiment of this application;

[0042] Figure 7 is a schematic diagram of a data return provided in an embodiment of this application;

[0043] Figure 8 is a timing diagram of speech rate adjustment and data transmission provided in an embodiment of this application;

[0044] Figure 9 is a flowchart illustrating a data processing method provided in an embodiment of this application;

[0045] Figure 10 is a logical diagram of statement segmentation provided in an embodiment of this application;

[0046] Figure 11 is a logical comparison diagram of a data return provided in an embodiment of this application;

[0047] Figure 12 is a schematic diagram of the composition of an electronic device provided in an embodiment of this application;

[0048] Figure 13 is a schematic diagram of the composition of a chip system provided in an embodiment of this application. Detailed Implementation

[0049] Hereinafter, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this embodiment, unless otherwise stated, "a plurality of" means two or more.

[0050] Currently, electronic devices offer increasingly diverse services to users. Among these is simultaneous interpretation service. This service can be applied in various scenarios. For example, users often require real-time translation in international conferences, education, and legal and medical fields. Through this simultaneous interpretation service, electronic devices can provide users with timely and accurate translation services.

[0051] For example, simultaneous interpretation software can be installed in some electronic devices. By running the software, the device can translate the acquired audio in real time and display the translated text on a screen. Some devices can also play the audio corresponding to the translated text, thus enabling the reading aloud of the translated text.

[0052] Therefore, electronic devices can provide simultaneous interpretation services to users through this simultaneous interpretation software.

[0053] As an example, referring to Figure 1, a schematic diagram of an interface interaction provided in an embodiment of this application is shown. A mobile phone is used as an example of an electronic device.

[0054] A mobile phone can have one or more applications installed. These applications may include weather apps, smart assistant apps, etc. Correspondingly, icons for these applications can be displayed to the user on the electronic device's home screen (e.g., screen 01). For example, the electronic device can display icon 1 for the smart assistant on screen 01.

[0055] In this application, the smart assistant application can be a pre-installed application in an electronic device that integrates artificial intelligence (AI) capabilities. This smart assistant application can integrate multiple functions. For example, these multiple functions may include AI dialogue (or conversation), one-click video creation, simultaneous interpretation, etc.

[0056] In different implementations, one or more functions in the smart assistant (such as interpreting) can be implemented through code in the smart assistant application package; or, the smart assistant can provide corresponding functions by calling the interfaces of other applications (such as interpreting applications) in conjunction with the corresponding applications.

[0057] In this example, a simultaneous interpretation application (the first application) can be installed on the electronic device. The smart assistant application can call the application interface of the simultaneous interpretation application. Thus, the user can activate the simultaneous interpretation function through the interface of the smart assistant application.

[0058] As shown in Figure 1, users can click icon 1 on interface 01 to run the smart assistant.

[0059] Correspondingly, electronic devices can activate the smart assistant and switch to display the corresponding smart assistant interface 02 (first interface).

[0060] This interface 02 can include function tags that can provide various functions for the smart assistant. For example, these function tags can include: one-click video creation, simultaneous interpretation, AI dialogue, etc.

[0061] Take the simultaneous interpretation function tag being selected as an example.

[0062] The electronic device's display screen can include display area 1 and display area 2. Display area 1 can be used to display the text corresponding to the captured audio. Display area 2 can be used to display the translated text corresponding to the captured audio.

[0063] This interface 02 may also include a button 3 (second control) for turning the simultaneous interpretation function on / off. Users can use this button 3 to turn the simultaneous interpretation function on or off.

[0064] For example, a user can click button 3 to instruct the electronic device to activate the simultaneous interpretation function. Correspondingly, the electronic device can display the words "Activated" on button 3, thus indicating to the user that the simultaneous interpretation function is enabled. In addition, the electronic device can also acquire ambient audio signals via a microphone. By recognizing and processing this audio signal, the electronic device can obtain the corresponding text and display it in display area 1. Furthermore, by translating the text corresponding to this audio signal, the electronic device can obtain the corresponding translated text and display it in area 2.

[0065] The identification, measurement, translation, and other operations of audio signals can be performed on a cloud server that is connected to the electronic device.

[0066] Electronic devices can also read translated texts aloud through simultaneous interpretation applications.

[0067] As shown in Figure 1, in this application, the interface corresponding to the simultaneous interpretation function (such as interface 02) may also include a speech rate adjustment control (first control). This speech rate adjustment control can be used by the user to adjust the reading speed.

[0068] In this example, a speech rate adjustment control includes a speech rate adjustment bar. This speech rate adjustment bar may include adjustment buttons.

[0069] The adjustment button divides the speech rate adjustment bar into two parts. When the adjustment button is at the far left of the speech rate adjustment bar, it corresponds to the slowest speech rate. When the adjustment button is at the far right of the speech rate adjustment bar, it corresponds to the fastest speech rate.

[0070] In some embodiments, when the adjustment button is located at a position different from the end of the speech rate adjustment bar, the reading speed can be determined based on the length of the speech rate adjustment bar to the left of the adjustment button and the length of the speech rate adjustment bar to the right of the adjustment button.

[0071] Users can adjust the reading speed by dragging the adjustment button on the speed adjustment bar.

[0072] For example, when simultaneous interpretation is enabled but the electronic device has not received audio data, the user can drag the adjustment button to set the initial reading speed. This way, once the electronic device receives the audio data, it can translate, recognize, and process the audio data, then read it aloud according to the initial reading speed.

[0073] For example, during the reading process on an electronic device, the user can drag the adjustment button (the first operation) to adjust the reading speed. In this way, the electronic device can read the subsequent content according to the adjusted reading speed.

[0074] As an example, referring to Figure 2, a schematic diagram of interface interaction is provided. Taking an electronic device using simultaneous interpretation to convert English speech into Chinese for display and reading aloud as an example, the electronic device can display interface 03 according to the following processing.

[0075] As shown in Figure 2, users can read English aloud after enabling the simultaneous interpretation function (such as by turning on the simultaneous interpretation function via button 3 in Figure 1).

[0076] For example, a user can read the following content aloud in English:

[0077] “From the same cause, the idea of ​​a floating hull of an enormous wreck was given up.

[0078] There remained,then,only two possible solutions of the question.”

[0079] Correspondingly, the electronic device can receive this audio data through a microphone. In the following description, the acquired environmental audio data will be referred to as raw audio.

[0080] The electronic device can transmit the original audio to a cloud server. The cloud server can then recognize the original audio and obtain the corresponding original text (in English). The electronic device can then retrieve the original text and display it in display area 1.

[0081] For example, the electronic device can display the following content in display area 1:

[0082] “From the same cause, the idea of ​​a floating hull of an enormous wreck was given up.

[0083] There remained,then,only two possible solutions of the question.”

[0084] The cloud server can also translate the original text to obtain the corresponding translated text (in Chinese). The electronic device can then retrieve the translated text and display it in display area 2.

[0085] For example, the electronic device can display the following content in display area 2:

[0086] "For the same reason, the idea of ​​a floating hull of a huge shipwreck was abandoned."

[0087] Therefore, only two possible solutions remain for this problem.

[0088] The cloud server can interact with electronic devices to determine the content to be read aloud, and then generate the corresponding audio (in Chinese). The electronic device can then access this audio and play it through a speaker, thus reading the text aloud.

[0089] As described above regarding the speech rate adjustment control, in this application, users can also adjust the speech rate of the electronic device before or during the reading process using the speech rate adjustment control.

[0090] Take, for example, a user adjusting the reading speed of an electronic device during a reading session.

[0091] Referring to Figure 3, this is a schematic diagram illustrating the correspondence between speech speed adjustment and the reading content. Taking an initial reading speed of speed 1 as an example, this speed 1 corresponds to position 1 of the adjustment button on the speech speed adjustment bar.

[0092] In this way, the electronic device can read aloud at a speech rate of 1, "For the same reason, the idea of ​​a floating hull for a huge shipwreck was abandoned."

[0093] Users can drag the adjustment button to position 2 on the speech rate adjustment bar while the electronic device reads aloud, "For the same reason, the idea of ​​a floating hull of a huge shipwreck was abandoned." This position corresponds to speech rate 2.

[0094] In this example, when the adjustment button is in position 1, the length of the speech rate adjustment bar to the left of the adjustment button is L1. When the adjustment button is in position 2, the length of the speech rate adjustment bar to the left of the adjustment button is L2. Length L2 is greater than length L1. Correspondingly, speech rate 2 is faster than speech rate 1.

[0095] Correspondingly, the electronic device can read aloud "For the same reason, the idea of ​​a floating hull for a huge shipwreck was abandoned" at speed 1, and then read the following content at speed 2. For example, the electronic device can read aloud "Therefore, only two possible solutions to this problem remain" at speed 2.

[0096] Therefore, by using the speech rate adjustment control, users can control the reading speed of electronic devices in real time during simultaneous interpretation.

[0097] To achieve the above functions, in some embodiments of this application, as shown in Figure 4, which is a schematic diagram of a communication scenario, the electronic device can work in conjunction with a cloud server to realize simultaneous interpretation functions, as well as the adjustment of the reading speed within the simultaneous interpretation function.

[0098] It is understood that in other embodiments, the electronic device may also independently perform all the functions of the electronic device and the cloud server in the various embodiments of this application, which will not be described separately.

[0099] The electronic devices in this application embodiment may include at least one of the following: mobile phone, foldable electronic device, tablet computer, desktop computer, laptop computer, handheld computer, laptop, ultra-mobile personal computer (UMPC), netbook, cellular phone, personal digital assistant (PDA), augmented reality (AR) device, virtual reality (VR) device, artificial intelligence (AI) device, wearable device, in-vehicle device, smart home device, or smart city device. This application embodiment does not impose any special limitation on the specific type of the electronic device.

[0100] Referring to Figure 5, it is a schematic diagram of the composition of an electronic device provided in an embodiment of this application.

[0101] In this application, the software system of the electronic device can adopt a layered architecture, event-driven architecture, microkernel architecture, microservice architecture, or cloud architecture. The embodiments of this application use a layered architecture. Taking the system as an example, the software structure of the electronic device is illustrated.

[0102] A layered architecture divides software into several layers, each with a clear role and function. Layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers, from top to bottom: the application layer, the application framework layer, the Android runtime (ART) and native C / C++ libraries, the Hardware Abstraction Layer (HAL), and the kernel layer.

[0103] The application layer can include a series of application packages. The application layer can also be called the application layer or the APP layer.

[0104] As shown in Figure 5, the application package may include simultaneous interpretation, smart assistant, calling and other applications.

[0105] The application framework layer is also called the framework layer or framework layer.

[0106] The application framework layer provides application programming interfaces (APIs) and a programming framework for applications in the application layer. The application framework layer includes some predefined functions.

[0107] As shown in Figure 5, the application framework layer may include a window manager, content providers, a view system, a resource manager, a notification manager, an activity manager, an input manager, etc. The application framework layer can also be called the framework layer, the framework layer, or simply the framework.

[0108] The window manager provides Window Manager Service (WMS), which can be used for window management, window animation management, surface management, and as a relay station for the input system.

[0109] Content providers store and retrieve data, making that data accessible to applications. This data can include videos, images, audio, phone calls made and received, browsing history and bookmarks, phone books, etc.

[0110] A view system includes visual controls, such as controls for displaying text and controls for displaying images. View systems can be used to build applications. A display interface can consist of one or more views. For example, a display interface including a text notification icon could include views for displaying text and views for displaying images.

[0111] The file explorer provides applications with various resources, such as localized strings, icons, images, layout files, video files, and more.

[0112] The notification manager allows applications to display notifications in the status bar. These notifications can be used to deliver informational messages and can disappear automatically after a short pause, requiring no user interaction. For example, the notification manager can be used to notify users of completed downloads or message alerts. The notification manager can also display notifications as icons or scrolling text in the top status bar, such as notifications from background applications, or as dialog boxes on the screen. Examples include displaying text messages in the status bar, emitting sounds, vibrating electronic devices, and flashing indicator lights.

[0113] The Activity Manager Service (AMS) can be used to start, switch, and schedule system components (such as activities, services, content providers, and broadcast receivers), as well as manage and schedule application processes.

[0114] The Input Manager Service (IMS) provides input management services, which can be used to manage system inputs such as touchscreen input, keypad input, and sensor input. IMS retrieves events from input device nodes and, through interaction with the WMS (Windows Management System), distributes these events to appropriate windows.

[0115] The Android runtime consists of the core libraries and the Android runtime itself. The Android runtime is responsible for converting source code into machine code. The Android runtime primarily employs ahead-of-time (AOT) compilation and just-in-time (JIT) compilation technologies.

[0116] The core library primarily provides basic Java class library functionalities, such as libraries for fundamental data structures, mathematics, I / O, tools, databases, and networking. It also provides APIs for users to develop Android applications.

[0117] Native C / C++ libraries can include multiple functional modules. Examples include: surface manager, media framework, libc, OpenGL ES, SQLite, Webkit, etc.

[0118] The Surface Manager manages the display subsystem and provides 2D and 3D layer blending for multiple applications. The Media Framework supports playback and recording of various common audio and video formats, as well as still image files. The Media Library supports multiple audio and video encoding formats, such as MPEG4, H.264, MP3, AAC, AMR, JPG, and PNG. OpenGLES provides drawing and manipulation of 2D and 3D graphics in applications. SQLite provides a lightweight relational database for electronic device applications.

[0119] The Hardware Abstraction Layer (HAL), also known as the abstraction layer or HAL layer, runs in user space. It encapsulates kernel-level drivers and provides interfaces to higher layers. For example, the HAL may include a display module, audio module, camera module, Bluetooth module, etc.

[0120] The kernel layer is the layer between hardware and software. The kernel layer includes at least the display driver, camera driver, audio driver, and Bluetooth driver.

[0121] In this application, based on the composition of the electronic device shown in Figure 5, the smart assistant application can launch the simultaneous interpretation application under the user's operation (such as the user selecting the "simultaneous interpretation" function tab in interface 02 of Figure 1).

[0122] Simultaneous interpretation applications can use relevant modules in the framework layer to call audio drivers in the abstraction layer and kernel layer, thereby collecting raw audio through the microphone of the electronic device. In some embodiments, simultaneous interpretation applications can preprocess the raw audio, filtering out other audio besides the content to be translated, and retaining the audio data of the content to be translated.

[0123] Simultaneous interpretation applications can also control electronic devices to transmit raw audio to a cloud server. The cloud server can process the raw audio, including recognition and translation, and send the raw and translated texts to the simultaneous interpretation application on the electronic device. Correspondingly, the simultaneous interpretation application can use the raw and translated texts to call the display module, display driver, and other components of the electronic device to control the electronic device to display the corresponding information (such as the interface 03 shown in Figure 2).

[0124] Simultaneous interpreting applications can also determine the pre-read text based on preset rules and real-time sentence segmentation of the translated text. The pre-read text can include the content that the electronic device will read next.

[0125] Simultaneous interpretation applications can transmit the pre-read text and its corresponding speech rate parameters to a cloud server. The cloud server can then process the pre-read text based on the speech rate parameters to obtain the audio. The simultaneous interpretation application can then acquire this audio through an electronic device, access the device's audio module, and play it via the device's speaker or Bluetooth playback device.

[0126] In this way, the electronic device can read the subsequent text aloud according to the user's adjusted speaking speed, achieving the reading effect shown in Figure 3.

[0127] It should be noted that the composition shown in Figure 5 is merely an example and does not constitute a specific limitation on the electronic device. In other embodiments, the electronic device may be configured with more or fewer components. One or more components shown in Figure 5 may be integrated into a module, or the function of any component may be set separately. This application does not impose specific limitations on the composition of the electronic device.

[0128] Referring to Figure 6, a schematic diagram of the device-to-device interaction process of a data processing method is shown. The solution provided in Figure 6 can be applied to the communication scenario shown in Figure 4.

[0129] As shown in Figure 6, the scheme may include:

[0130] S601, Electronic devices acquire raw audio.

[0131] For example, in conjunction with the descriptions in Figures 1 to 3, in this example, the electronic device can enable simultaneous interpretation functionality at the user's command before acquiring the original audio.

[0132] In some embodiments, users can select the simultaneous interpretation function tab through interface 02 in Figure 1 and click button 3 to enable the simultaneous interpretation function.

[0133] In other embodiments, users can instruct electronic devices to run simultaneous interpretation applications and enable simultaneous interpretation functions in other ways.

[0134] Correspondingly, electronic devices can collect sounds from the environment and generate corresponding audio files.

[0135] In some embodiments, the electronic device may use the generated audio file as the original audio.

[0136] In other embodiments, the electronic device can perform preprocessing such as noise reduction on the generated audio file to obtain the original audio.

[0137] Referring to the example in Figure 2, the original audio can be in language A (first language). For example, language A can be English.

[0138] S602, Electronic devices send raw audio to a cloud server.

[0139] For example, an electronic device can send the raw audio of language A (such as English) to a cloud server. For instance, this cloud server could be an application server for a simultaneous interpretation application.

[0140] It should be noted that, in some other embodiments, the electronic device may also send information indicating translation into Language B (second language) to a cloud server so that the cloud server can perform translation processing based on the information.

[0141] S603, the cloud server performs cloud processing to obtain the original text and the translated text.

[0142] For example, the cloud processing may include: recognition processing, translation processing, etc.

[0143] In some embodiments, the recognition process may include: performing speech recognition based on the original audio to obtain text data corresponding to the original audio.

[0144] In some embodiments, the identification process may further include segmentation.

[0145] As can be understood, as shown in the example in Figure 2, the original audio will include text, punctuation, and other content, thus forming one or more sentences. The original audio may also include one or more paragraphs. Each paragraph may include at least one sentence.

[0146] In this example, the cloud server can segment the original audio / the text data corresponding to the original audio through segmentation processing.

[0147] Take, for example, the segmentation of the text data corresponding to the original audio by a cloud server.

[0148] The cloud server can segment the text data corresponding to the original audio based on audio features such as pauses, thereby obtaining the original text.

[0149] The original text may include one or more paragraphs. For example, if the original audio is in language A, the original text could also be text in language A.

[0150] In this example, the cloud server can also translate the original text. For example, the cloud server can translate language A into language B according to the instructions of the electronic device. For instance, language A can be English, and language B can be Chinese. In other embodiments, language A can be other languages, and language B can be different from Chinese; this application does not limit this.

[0151] It is understandable that if the original text contains multiple paragraphs, the corresponding translated text can also contain multiple paragraphs. Similarly, if a paragraph in the original text contains multiple sentences, that paragraph in the translated text can also contain multiple sentences.

[0152] In this application, the cloud server can add a final marker at the end of each paragraph of the original text and / or translated text. This final marker thus indicates the end of a paragraph.

[0153] S604: The cloud server sends the original text and the translated text to the electronic device.

[0154] S605. The electronic device displays the original text and the translated text.

[0155] For example, the electronic device can display content in display area 1 of Figure 2 based on the original text. The electronic device can display content in display area 2 of Figure 2 based on the translated text.

[0156] Thus, through the above S601-S605, the relevant display in the simultaneous interpretation function is realized.

[0157] In this application, the cloud server can send the acquired original text and / or translated text to the electronic device at certain intervals or in real time during the execution of S603.

[0158] In other words, even if it is the same original text / translated text, the cloud server may send (return) the corresponding content to the electronic device in multiple ways.

[0159] This can improve the timeliness of electronic devices acquiring raw / translated text.

[0160] For example, referring to Figure 7, there is a schematic diagram of a data return transmission.

[0161] In the example shown in Figure 7, a complete paragraph of the original text can include ABCDEF, etc. Any one of the letters ABCDEF can be used to represent a character or punctuation mark in language A. Correspondingly, the letter F in the original text can be configured with a Final identifier to indicate that ABCDEF corresponds to a complete paragraph.

[0162] After translation, the translated text corresponding to the complete paragraph shown in Figure 7 can include content such as 123456. Any one of the numbers 123456 can be used to represent a character or punctuation mark in language B. Correspondingly, the 6 in the translated text can be configured with a "Final" identifier to indicate that 123456 corresponds to a complete paragraph.

[0163] In the example shown in Figure 7, the cloud server can transmit the translated text corresponding to 123 in the first transmission back to the electronic device.

[0164] The cloud server can transmit the translated text of 123456 in the second transmission back to the electronic device.

[0165] The example shown in Figure 7 illustrates how the entire content can be uploaded via two uploads. In other embodiments, the cloud server can also upload the entire content via one, three, or more uploads.

[0166] In this way, the cloud server can transmit a complete paragraph of translated text to the electronic device through multiple transmissions. In these multiple transmissions, adjacent transmissions may include one or more identical (repeated) characters. This avoids character loss caused by multiple transmissions.

[0167] The transmission mechanism of the original text is similar to that of the translated text as shown in Figure 7, and will not be described in detail here.

[0168] Referring again to Figure 6, in this application, the electronic device can perform sentence segmentation on the translated text through the following steps to accurately obtain the text to be read aloud. The electronic device can also interact with a cloud server to obtain the audio corresponding to the text to be read aloud and the speech rate parameters.

[0169] S606. The electronic device performs real-time sentence segmentation on the translated text to determine the text to be read aloud.

[0170] Referring to the illustration in Figure 7, the electronic device can receive multiple data feeds from a cloud server. For example, the electronic device can receive feed 1, feed 2, and so on. Each feed 1 can be at least a portion of the translated text.

[0171] Electronic devices can perform real-time segmentation processing on the received data after each transmission from the cloud server.

[0172] For example, this real-time sentence segmentation processing may include sentence splitting, intermediate state position correction, and other processing. Specific implementation details will be provided later.

[0173] Through this real-time sentence segmentation processing, electronic devices can accurately determine the position of the sentences that have already been read aloud in the currently transmitted data.

[0174] The electronic device can also accurately determine the translation text (such as the pre-read text or the first pre-read text) corresponding to the next sentence to be read, based on preset rules and the position of the already read sentence in the current feedback data.

[0175] Referring to the example in Figure 3, in the execution logic shown in Figure 6, the electronic device can update the latest speech rate parameters before the next reading. For example, the speech rate corresponding to this parameter can correspond to the position of the adjustment button on the speech rate adjustment bar in the current speech rate adjustment control.

[0176] S607: Electronic devices send the pre-read text and speech rate parameters to the cloud server.

[0177] For example, an electronic device can send the updated speech rate parameters, along with the pre-read text (the text corresponding to language B), to a cloud server. The content of this pre-read text can be included in the translated text.

[0178] The S608 cloud server generates audio for reading aloud based on the pre-read text and speech rate parameters.

[0179] For example, the cloud server can convert the pre-read text corresponding to language B into audio data based on the speech rate parameter. This audio data is the reading audio.

[0180] It is understandable that the content of the audio recording is consistent with the pre-read text, both corresponding to language B. The speaking speed of the audio recording can be the speed indicated by the speaking speed parameter.

[0181] S609, the cloud server sends audio readings to electronic devices.

[0182] S610: Electronic devices read aloud based on audio recordings.

[0183] Therefore, the electronic device can acquire the audio reading through S606-S610 described above, and then play the audio reading through the electronic device's speaker or a voice playback device connected to the electronic device. This enables the playback of translated audio during simultaneous interpretation.

[0184] Based on the scheme shown in Figure 6, the electronic device can acquire the audio to be read aloud next time, and play the audio (the first audio) when the next audio needs to be read aloud.

[0185] Referring to Figure 8, a timing diagram of speech rate adjustment and data transmission is provided. The process shown in Figure 8 corresponds to the processing timing of S606-S610 shown in Figure 6.

[0186] As shown in Figure 8, time 1 can be the time when the pre-reading text 1 begins to be read aloud. Time 4 can be the time when the pre-reading text 1 is finished or the pre-reading text 2 begins to be read aloud. Time 4 is later than time 1. Pre-reading text 2 is the pre-reading text that follows pre-reading text 1.

[0187] Both the pre-read text 1 and the pre-read text 2 can contain one or more sentences.

[0188] The electronic device can read aloud the audio 1 it has already acquired at time 1. Before time 1, the adjustment button in the electronic device's speech rate adjustment control is located at position 1 corresponding to speech rate 1. In this way, the speech rate of the audio 1 is also speech rate 1.

[0189] Before moment 4 arrives, while the electronic device is reading the pre-read text 1, the user adjusts the position of the adjustment button from position 1 corresponding to speech speed 1 to position 2 corresponding to speech speed 2.

[0190] After time 1 and before time 4, the electronic device can determine the pre-read text 2 and the speech rate parameter 2 (speech rate 2) at time 2. The execution of this process can be referred to as S606 in Figure 6.

[0191] The electronic device can send the pre-read text 2 and speech rate parameter 2 (speech rate 2) to the cloud server at time 3, which is between time 4 and time 2.

[0192] Next, the electronic device can retrieve the pre-read text 2 and the corresponding audio 2 for the speech rate parameter 2 from the cloud server after time 3 and before time 4.

[0193] Therefore, when time 4 arrives, the electronic device can read aloud the audio 2, thereby realizing the reading of the pre-read text 2.

[0194] It should be noted that, in the example shown in Figure 6, the recognition processing, translation processing, and the processing of obtaining speech data from the text are all performed on a cloud server. In other embodiments, one or more of the above processes can also be performed internally by the electronic device. Further details will not be elaborated further.

[0195] Furthermore, in the above embodiments, the example is an electronic device acquiring ambient sound to generate raw audio, and then performing simultaneous interpretation based on that raw audio. This enables the display and reading of simultaneous interpretation in the scenario shown in Figure 2.

[0196] In other embodiments, the electronic device may also acquire the original audio through other applications.

[0197] For example, an electronic device can obtain the audio of a currently connected call from a calling application. The electronic device can implement this according to the scheme shown in Figure 6, converting the audio data of the connected call (language A) into audio data of language B for playback and / or displaying it as text data.

[0198] This application does not restrict the source of the original audio.

[0199] The specific execution process of S606 in Figure 6 will be explained below with reference to the accompanying drawings.

[0200] For example, referring to FIG9, a flowchart of a data processing method provided in an embodiment of this application is shown. Through the scheme shown in FIG9, the electronic device can accurately determine the text to be read aloud.

[0201] In some embodiments, the simultaneous interpretation application of an electronic device can trigger the execution of the scheme shown in Figure 9 based on a set of returned data received by the electronic device. For example, after receiving the first returned data as shown in Figure 7, the simultaneous interpretation application can execute the scheme shown in Figure 9 to determine the pre-read text corresponding to the returned data. Similarly, after receiving the second returned data as shown in Figure 7, the simultaneous interpretation application can execute the scheme shown in Figure 9 to determine the pre-read text corresponding to the returned data.

[0202] As shown in Figure 9, the scheme may include:

[0203] S901. Determine whether the returned data has punctuation marks and includes the final identifier.

[0204] For example, the electronic device can analyze the received data to determine whether the received data contains punctuation and whether the data includes the final identifier.

[0205] If the received data includes one or more punctuation marks, or if it does not include the final flag, then execute the following S902.

[0206] Understandably, the presence of one or more punctuation marks in the current returned data indicates that it includes multiple statements. Conversely, the absence of the `final` marker in the current returned data indicates that it does not represent the end of a complete paragraph. This allows the electronic device to continue processing and accurately pinpoint the location of the read-aloud statements.

[0207] Correspondingly, if the received data does not include punctuation but includes the `final` flag, then the received data contains a complete sentence. In this case, the electronic device can read it aloud. If all received data is determined as the pre-read text, execution jumps to S913.

[0208] S902. Determine if the returned data contains punctuation marks.

[0209] For example, if the received data includes at least one punctuation mark, the following step S903 is executed. Conversely, if the received data does not include punctuation marks, the electronic device can terminate the current process and wait for the next data transmission from the cloud server.

[0210] S903. The returned data is segmented according to punctuation marks.

[0211] For example, an electronic device can segment the returned data to obtain multiple sub-data when the returned data includes multiple punctuation marks.

[0212] In some embodiments, the electronic device can segment the returned data based on the position of punctuation characters. Punctuation and preceding text characters are segmented into the same sub-data. In this way, each sub-data can correspond to a statement.

[0213] Referring to Figure 10, a logical diagram of statement segmentation is provided.

[0214] For example, as shown in Figure 10, the returned data includes the characters corresponding to 12345. Each number in 12345 represents a character or punctuation mark in language B.

[0215] In this example, 1 and 2 represent text characters, 4 and 5 represent text characters, and 3 represents punctuation characters.

[0216] By segmenting, the electronic device can divide the returned data into sub-data 1 and sub-data 2. Sub-data 1 can include 123, and sub-data 2 can include 45.

[0217] In this way, the statement segmentation of the currently returned data can be completed through S902-S903.

[0218] S904. Retrieve the statement in the current data transmission based on the previously read position.

[0219] In this application, the electronic device can mark the last sentence that has been read aloud by the position where it was read aloud.

[0220] For example, as illustrated in Figure 7, the cloud server can transmit the entire translated text back to the electronic device through a multi-transmission mechanism.

[0221] Referring to Figure 11, a logical comparison diagram of returned data is shown. As shown in Case 1 of Figure 11, taking the statements corresponding to the first returned data as S1, S2, and S3, and the second returned data as S1, S2, S3, S4, and S5 as an example.

[0222] Each statement can correspond to a segmented sub-data. Each statement in S1, S2, S3, S4, and S5 can correspond to one or more text characters and at most one punctuation character.

[0223] Similar to the illustration in Figure 7, two adjacent data transmissions overlap by at least some extent. In terms of statements, this means two adjacent data transmissions overlap by at least one statement.

[0224] For example, in the case 1 example, the overlapping statements for the first and second data return messages may include S1 and S2.

[0225] Understandably, after the first data transmission is received, the electronic device can read aloud based on that data. For example, the electronic device can read aloud the three statements corresponding to S1, S2, and S3.

[0226] In this way, the electronic device can set the read-already position to point to the statement corresponding to S3. In this example, based on the first data transmission, the read-already position can be position 1, and position 1 can point to the third statement in the transmitted data.

[0227] In Case 1 of Figure 11, the electronic device can obtain the last sentence that has been read aloud (comparison sentence 1) based on the data transmitted back the second time and in conjunction with position 1. In this example, comparison sentence 1 can be S3.

[0228] S905. Determine whether the position has changed based on whether the statements are consistent.

[0229] The position can be the location of the last sentence that has been read aloud in the current returned data.

[0230] The two statements used to determine whether they are consistent can include: comparison statement 1 and comparison statement 2.

[0231] As explained in S904, the comparison statement 1 can be: the statement (S3) pointed to by position 1 in the current returned data, based on the multiple sub-data obtained after segmentation.

[0232] Furthermore, comparison statement 2 can be: the last statement that was read aloud in the last returned data. Referring to the example of case 1 in Figure 11, the last statement that was read aloud in the last returned data can be the statement corresponding to S3.

[0233] In this way, the electronic device can determine whether the data in the current returned data pointed to by position 1 is still the last sentence that has been read aloud, based on whether the sub-data corresponding to comparison statement 1 and comparison statement 2 are consistent.

[0234] For example, as shown in Case 1 of Figure 11, if comparison statement 1 and comparison statement 2 are the same (both are statements corresponding to S3), it indicates that for the currently returned data, position 1 can still accurately point to the last read statement. Thus, the electronic device can select a statement to be read from the statements following position 1 based on this position 1. For example, the electronic device can jump to execute S908.

[0235] Correspondingly, if the comparison statement 1 and the comparison statement 2 are inconsistent, it indicates that for the current returned data, the statement pointed to by position 1 is not the last statement that has been read aloud.

[0236] For example, as shown in Case 2 of Figure 11, let's take the current returned data (such as the second returned data) as S2, S3, S4, S5, and S6. In this case, comparison statement 2 is still the statement corresponding to S3. However, the sub-data in the current returned data pointed to by position 1 can be the statement corresponding to S4. In this case, the sub-data in the current returned data pointed to by position 1 is different from the last statement that has been read aloud.

[0237] That is, the location has changed. Therefore, the electronic device can perform position correction and update based on the currently transmitted data through the following steps S906-S907.

[0238] S906. Among the multiple sub-data of the currently returned data, find the sentence with the highest similarity to the last sentence that has been read aloud.

[0239] For example, in some embodiments, the electronic device can perform a similarity comparison between each statement in the currently transmitted data and the last statement that has been read aloud. This allows the electronic device to obtain the similarity score between each statement in the currently transmitted data and the last statement that has been read aloud. Based on the similarity scores of each statement, the electronic device can determine the statement with the highest similarity score to the last statement that has been read aloud.

[0240] For example, take scenario 2 in Figure 11. The currently returned data, after being segmented, can include statements such as S2, S3, S4, S5, and S6.

[0241] The electronic device can calculate the similarity between S2 and comparison statement 2 (such as S3) 2, the similarity between S3 and comparison statement 2 (such as S3) 3, the similarity between S4 and comparison statement 2 (such as S3) 4, the similarity between S5 and comparison statement 2 (such as S3) 5, and the similarity between S6 and comparison statement 2 (such as S3) 6 respectively.

[0242] The electronic device can determine the sentence with the highest similarity to the last sentence that has already been read, based on the above five similarity scores. For example, if similarity score 3 is higher than the other similarities, the electronic device can take S3 in the currently transmitted data (the second transmitted data in case 2 in Figure 11) as the sentence with the highest similarity to the last sentence that has already been read.

[0243] The above example illustrates how an electronic device matches all currently transmitted data with the last read-aloud data to determine their respective similarity scores.

[0244] In other embodiments, even if the statement pointed to by position 1 in the second data transmission is not the last statement that has been read aloud, its position will not be far from position 1.

[0245] Thus, in this example, the electronic device can have a preset matching range. For example, the matching range could be + / -1, + / -2, + / -3, + / -4, + / -5, or + / -6, etc. Taking a matching range of + / -1 as an example, the electronic device can select one statement before and after position 1 from the currently returned data and match it with the comparison statement 2.

[0246] For example, referring to situation 2 in Figure 11, position 1 can point to the third statement. In this way, the electronic device can select the second statement before the third statement (S3) and the fourth statement after the third statement (S5) from all the statements sent back in the second time, and match them with the comparison statement 2 respectively to determine the similarity 3 corresponding to S3 and the similarity 5 corresponding to S5.

[0247] Therefore, the electronic device can select a higher similarity (such as similarity 3) between similarity 3 and similarity 5, and determine the corresponding sentence (S3) as the sentence with the highest similarity to the last sentence that has been read aloud in the current returned data.

[0248] S907. Adjust the position of the text that has been read aloud.

[0249] According to the description in S906, the electronic device can find the sentence with the highest similarity to the last sentence that has been read aloud in the current returned data (such as the second returned data).

[0250] For example, as shown in Case 2 of Figure 11, the electronic device can find S3 in the second transmitted data within the current transmitted data, which is the sentence with the highest similarity to the last sentence that has already been read aloud.

[0251] In this way, the electronic device can update the read-aloud position based on the position of the sentence with the highest similarity in the currently transmitted data.

[0252] Referring to scenario 2 in Figure 11, position S3 in the second data transmission corresponds to position 2. For example, position 2 could point to the second sentence in the transmitted data. The electronic device can update the read-already position to position 2. Based on this updated read-already position, the electronic device can accurately determine the read-already sentences in the current transmitted data, and then select the pre-read text corresponding to the pre-read sentences from the remaining unread sentences (such as sentences after position 2).

[0253] S904-S907 can be used to perform intermediate state error correction on the currently transmitted data. This intermediate state error correction enables the electronic device to more accurately pinpoint the position of the read-out sentence in the latest received transmitted data.

[0254] After completing the above sentence segmentation and intermediate state error correction, the electronic device can determine the text to be read aloud through the following steps.

[0255] S908. Determine whether the currently returned data includes the final identifier.

[0256] For example, if the current returned data includes a final identifier, jump to execute S909a.

[0257] If the current returned data does not include the final flag, jump to execute S909b.

[0258] S909a, The data from the position of "never to be read" to "final" is used as the pre-read text.

[0259] For example, if the current returned data includes a `final` flag, it indicates that the returned data includes one or more statements that constitute a complete text segment. In this way, the electronic device can determine all statements following the read-already position in the current returned data as the text to be read aloud, based on the updated read-already position.

[0260] For example, referring to case 2 in Figure 11, the updated read-already position is position 2, pointing to the second statement. Thus, in this example of S909, if the current returned data includes the final flag, and the final flag is set after S6, the electronic device can use S4, S5, and S6 as the pre-read-already basis.

[0261] Next, the electronic device can execute S913.

[0262] S909b. Is the number of unread sentences greater than or equal to threshold 1? For example, threshold 1 can be a positive integer, such as 3, 4, etc.

[0263] For example, an electronic device can determine the unread sentences based on the already read position and the segmentation results of the currently returned data. For instance, the unread sentences could be the sentences following the already read position in the segmentation results of the currently returned data.

[0264] Referring to scenario 2 in Figure 11, the current data segmentation results include: S2, S3, S4, S5, and S6. The position that has been read aloud is position 2. Therefore, the unread sentences include S4, S5, and S6.

[0265] In this example, the electronic device can execute S910 if the number of unread statements is greater than or equal to threshold 1. If the number of unread statements is less than threshold 1, the electronic device can determine that the unread statements are short, wait for the next data transmission to read them together, and end the current process.

[0266] S910. Divide unread sentences into pre-read sentences and pre-reserved sentences.

[0267] For example, the pre-read statement can be the statement to be read aloud this time. The pre-reserved statement can be included in the statement to be read aloud next time.

[0268] In some embodiments, the electronic device may designate the last N sentences in the unread sentences as pre-reserved sentences. N is a preset threshold of 2. This preset threshold of 2 can be a positive integer, such as 1 or 2. Correspondingly, the remaining sentences in the unread sentences that are different from the pre-reserved sentences are the pre-read sentences.

[0269] Referring to the example of scenario 2 in Figure 11, the unread statements include S4, S5, and S6. Taking a threshold of 2 corresponding to N as an example, the electronic device can reserve the last two statements (such as S5 and S6) as pre-reserved statements. The remaining statement (such as S4) is used as the pre-read statement.

[0270] Therefore, the electronic device can determine the text of the pre-read statement (i.e., the sub-data of the pre-read statement) as the pre-read text. And the text of the pre-reserved statement (i.e., the sub-data of the pre-reserved statement) is used as the pre-reserved text.

[0271] In some embodiments, the electronic device may also execute the following S911-S912 to determine the pre-reserved text and the pre-read text, and further optimize the reading logic.

[0272] S911. Determine whether the length of the text to be read aloud is less than threshold 3, or whether the length of the text to be reserved is less than threshold 4. Threshold 3 and threshold 4 can correspond to byte lengths. For example, threshold 3 can be 35 and threshold 4 can be 45. In other embodiments, threshold 3 and / or threshold 4 may also be different from the examples above.

[0273] If the length of the text to be read aloud is less than threshold 3, or the length of the text to be reserved is less than threshold 4, the electronic device performs the following judgment in S912.

[0274] Correspondingly, if the length of the pre-read text is greater than threshold 3 and the length of the pre-reserved text is greater than threshold 4, the electronic device executes the following S913.

[0275] S912. Does the number of sentences in the unread text exceed a threshold of 5? For example, the threshold of 5 can be 5.

[0276] The total number of sentences in the unread text can be the sum of the number of sentences in the pre-read text and the number of sentences in the pre-reserved text.

[0277] If the total number of unread text sentences exceeds the threshold of 5, execute S913. Conversely, if the total number of unread text sentences is less than the threshold of 5, it indicates that the unread text is too short, and the process will wait for the next data transmission to be processed together, ending the process.

[0278] S913. Obtain the speech rate parameter and put the pre-read text and speech rate parameter into the reading queue.

[0279] In this example, referring to the illustrations in Figures 6-8, the electronic device can obtain the latest speech rate parameters once the text to be read is determined. For example, this speech rate parameter can correspond to the current position of the adjustment button on the speech rate adjustment bar.

[0280] Therefore, after determining the text to be read aloud, the electronic device can add the text and the corresponding speech rate parameters to the reading queue.

[0281] Correspondingly, electronic devices (such as simultaneous interpretation applications on electronic devices) can retrieve the earliest pre-reading text and corresponding speech rate parameters (such as pre-reading data) from the reading queue. The electronic device can also transmit this pre-reading data to the cloud server through the step shown in S607 of Figure 6.

[0282] Thus, as illustrated in Figure 6, the cloud server can generate the corresponding audio reading via S608. Then, as explained in S609-S610, the electronic device can read aloud based on the audio reading.

[0283] It is understood that the electronic device provided in this application embodiment includes hardware structures and / or software modules corresponding to perform each function in order to achieve the above-mentioned functions. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments disclosed herein, the embodiments of this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments of this application.

[0284] This application embodiment can divide the above-described electronic device into functional modules based on the method example described above. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated modules can be implemented in hardware or as software functional modules. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.

[0285] For example, FIG12 shows a schematic diagram of the composition of an electronic device 1200. As shown in FIG12, the electronic device 1200 may include a processor 1201 and a memory 1202. The memory 1202 is used to store computer execution instructions. For example, in some embodiments, when the processor 1201 executes the instructions stored in the memory 1202, the electronic device 1200 may perform any of the methods shown in the above embodiments.

[0286] It should be noted that all relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.

[0287] Figure 13 shows a schematic diagram of a chip system 1300. The chip system 1300 may include a processor 1301 and a communication interface 1302, used to support related devices in implementing the functions involved in the above embodiments. In one possible design, the chip system also includes a memory for storing necessary program instructions and data for the electronic device. The chip system may be composed of chips or may include chips and other discrete devices. It should be noted that in some implementations of this application, the communication interface 1302 may also be referred to as an interface circuit.

[0288] It should be noted that all relevant content of each step involved in the above method embodiments can be referenced from the functional description of the corresponding functional module, and will not be repeated here.

[0289] This application also provides a computer-readable storage medium storing a computer program thereon. When executed by a computer, the computer program implements the method flow related to the electronic device in any of the above method embodiments. Specifically, the computer can be the aforementioned electronic device.

[0290] This application also provides a computer program or a computer program product including a computer program, which, when executed on a computer, causes the computer to implement the method flow related to the electronic device in any of the above method embodiments. Specifically, the computer can be the aforementioned electronic device.

[0291] The functions, actions, operations, or steps in the above embodiments can be implemented, in whole or in part, by software, hardware, firmware, or any combination thereof. When implemented using software programs, they can be implemented, in whole or in part, in the form of a computer program product. This computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or include one or more data storage devices such as servers and data centers that can be integrated with the medium. The available media can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).

[0292] Although this application has been described in conjunction with specific features and embodiments, it is obvious that various modifications and combinations can be made thereto without departing from the scope of this application. Accordingly, this specification and drawings are merely exemplary illustrations of the application as defined by the appended claims, and are considered to cover any and all modifications, variations, combinations, or equivalents within the scope of this application. Clearly, those skilled in the art can make various alterations and modifications to this application without departing from the scope of this application. Thus, if such modifications and modifications of this application fall within the scope of the claims of this application and their equivalents, this application is also intended to include such modifications and modifications.

Claims

1. A data processing method, characterized in that, The method is applied to an electronic device, which is equipped with a first application. The first application is used to provide simultaneous interpretation function. When the simultaneous interpretation function is activated, the electronic device plays the translation content corresponding to the original audio collected. The method includes: The first interface is displayed, which is the interface of the first application; the first interface includes a first control, which is used to adjust the speech rate when playing the audio. Receive a first operation, the first operation being used to adjust the reading speed to a first speed via the first control; Play the first audio recording, the content of which includes the first pre-reading text, which is included in the translation of the original audio; the reading speed of the first audio recording is the first speech rate.

2. The method according to claim 1, characterized in that, The first interface also includes a second control, which is used to enable the simultaneous interpretation function; Before receiving the first operation, the method further includes: Receive a second operation on the second control, the second operation being used to enable the simultaneous interpretation function.

3. The method according to claim 1 or 2, characterized in that, The first control includes: a speech rate adjustment bar and an adjustment button; The position of the adjustment button on the speech rate adjustment bar corresponds to the reading speed. The first operation corresponds to adjusting the position of the adjustment button on the speech rate adjustment bar to the position corresponding to the first speech rate.

4. The method according to any one of claims 1-3, characterized in that, Before playing the first audio recording, the method further includes: Obtain the original audio and send the original audio to the cloud server; Obtain the first original text and the first translated text; Wherein, the first original text corresponds to at least a portion of the original text corresponding to the original audio, and the first translated text corresponds to at least a portion of the translated text corresponding to the original audio; The first original audio uses a first language, the first original text uses the first language, and the first translated text uses a second language, wherein the first language and the second language are different.

5. The method according to claim 4, characterized in that, The original audio includes ambient sound.

6. The method according to claim 4 or 5, characterized in that, The first language is English, and the second language is Chinese.

7. The method according to any one of claims 4-6, characterized in that, After obtaining the first original text and the first translated text, the method further includes: The first original text and the first translated text are displayed on the first interface.

8. The method according to any one of claims 4-7, characterized in that, After obtaining the first original text and the first translated text, the method further includes: The first pre-read text is determined based on the feedback data corresponding to the first translated text; the feedback data corresponding to the first translated text is the data sent to the electronic device by the cloud server when sending the first translated text.

9. The method according to claim 8, characterized in that, After determining the first pre-read text, the method further includes: Obtain speech rate parameters, wherein the speech rate parameters indicate the first speech rate; Send the first pre-read text and the speech rate parameters to the cloud server. The first audio recording is obtained from the cloud server.

10. The method according to claim 8 or 9, characterized in that, The return data corresponding to the first translated text includes at least the first return data and the second return data. The first backhaul data and the second backhaul data include at least partial overlap.

11. The method according to claim 10, characterized in that, The first back-transmission data is the current back-transmission data, and the electronic device has received the second back-transmission data before receiving the first back-transmission data; The receiving of the first operation includes: During the playback of the second audio recording, the first operation is received; the content of the second audio recording is included in the second feedback data. The step of determining the first pre-read text based on the returned data corresponding to the first translated text includes: Based on the first returned data, the first pre-read text is determined.

12. The method according to claim 11, characterized in that, After obtaining the first returned data, the method further includes: If the first returned data includes at least one punctuation character, the first returned data is divided into two or more sub-data corresponding to statements based on the punctuation character; each sub-data corresponds to one statement.

13. The method according to claim 12, characterized in that, After obtaining the first returned data, the method further includes: If the first returned data does not include punctuation characters and the first returned data includes an end marker, the first returned data is determined to be the data of the first pre-read text; the end marker is used to indicate the end of the paragraph.

14. The method according to any one of claims 11-13, characterized in that, The method further includes: The first position of the first statement in the first returned data is recorded as the read-aloud position. The first statement is the last statement read in the second read-aloud audio.

15. The method according to claim 14, characterized in that, The method further includes: Determine the second statement in the first returned data, wherein the second statement is the statement indicated by the first position in the first returned data; If the first statement and the second statement are consistent, continue recording the already read position as the first position; or... Based on the inconsistency between the first statement and the second statement, the second position of the third statement in the first returned data is recorded as the read-already position; the third statement is the statement in the first returned data that has the highest similarity to the first local area.

16. The method according to any one of claims 11-15, characterized in that, The method further includes: If the first returned data includes an end identifier, the data of the unread text in the first returned data will be determined as the data of the first pre-read text; The unread text is the data corresponding to the sentence after the read position in the first returned data; the end marker is used to indicate the end of the paragraph.

17. The method according to any one of claims 11-16, characterized in that, If the first returned data does not include an end marker but includes one or more punctuation characters, the method further includes: According to the first rule, the unread text is divided into pre-read text and pre-reserved text, and the unread text corresponds to the unread data; The data of the pre-read text is determined to be the data of the first pre-read text.

18. The method according to claim 17, characterized in that, The first rule includes: when the number of sentences in the unread text is greater than a first threshold, configuring the last N sentences in the unread text as the pre-reserved text, and determining the sentences in the unread text that are different from the pre-reserved text as the pre-read text; N is a preset positive integer.

19. The method according to claim 17 or 18, characterized in that, Before determining the data of the pre-read text as the data of the first pre-read text, the method further includes: The data length of the pre-read text is determined to be less than a second threshold, and the number of sentences in the unread text is greater than a fourth threshold; or, The data length of the pre-reserved text is determined to be less than a third threshold, and the number of sentences in the unread text is greater than a fourth threshold; or, The data length of the pre-read text is determined to be greater than a second threshold, and the number of sentences in the pre-reserved text is greater than a third threshold.

20. An electronic device, characterized in that, The electronic device includes: a memory and one or more processors; the memory and the processors are coupled. The memory is used to store computer program code, which includes computer instructions. When the processor executes the computer instructions, the electronic device performs the method as described in any one of claims 1-19.

21. A chip system, characterized in that, The chip system is applied to an electronic device; the chip system includes one or more interface circuits and one or more processors; the interface circuits and the processors are interconnected via lines; the interface circuits are used to receive signals from the memory of the electronic device and send the signals to the processors, the signals including computer instructions stored in the memory; when the processor executes the computer instructions, the electronic device performs the method as described in any one of claims 1-19.

22. A computer-readable storage medium, characterized in that, Includes computer instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1-19.

23. A communication system, characterized in that, The communication system includes the electronic device as described in claim 20, and a cloud server; The cloud server is used to generate a first audio reading based on the first pre-read text and speech rate parameters sent by the electronic device; The cloud server is also used to send the first audio reading to the electronic device.

Citation Information

Patent Citations

  • Voice outputting method, terminal and computer readable storage medium

    CN109686359A

  • Network-based simultaneous interpretation method and system

    CN110677406A

  • Machine simultaneous interpretation output audio dynamic synthesis method, device and equipment

    CN112233649A

  • Processing method, mobile terminal and storage medium

    CN113314095A

  • Electronic device and method of controlling speech recognition by electronic device

    US20200302913A1