Processing method, intelligent terminal and storage medium

By using pre-translated content and preset data sets to assist translation, the problems of low accuracy and efficiency in information processing are solved, and the user experience is improved.

CN119294407BActive Publication Date: 2025-10-10SHENZHEN TECNO TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411825348.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-10-10
Estimated Expiration
2044-12-12

AI Technical Summary

Technical Problem

Existing information processing functions have low accuracy and/or low efficiency in some usage scenarios, resulting in a poor user experience.

Method used

By generating or determining the second information, using pre-translated content, a preset data set and a first preset model for auxiliary translation, and combining the processing of voice, text, image and gesture information, the accuracy and efficiency of information processing are improved.

Benefits of technology

It improves the accuracy and efficiency of information processing and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119294407B_ABST
    Figure CN119294407B_ABST
Patent Text Reader

Abstract

The application provides a processing method, an intelligent terminal and a storage medium. The processing method can be applied to a first device and includes the following steps: S10: generating or determining second information based on first content and first information. Through the technical solution, the accuracy and / or efficiency of information processing can be improved, and the use experience of a user can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of terminal application technology, and in particular to a processing method, an intelligent terminal and a storage medium. Background Art

[0002] In the context of globalization, the information processing function of smart terminals plays an important role in the transmission of information between different languages ​​and cultures.

[0003] During the process of conceiving and implementing this application, the inventors discovered that there are at least the following problems: the current information processing function has low accuracy and / or low efficiency in some usage scenarios (such as instant translation scenarios), which can easily lead to a poor user experience.

[0004] The preceding description is intended to provide general background information and does not necessarily constitute prior art. Summary of the Invention

[0005] In response to the above technical problems, the present application provides a processing method, an intelligent terminal and a storage medium, aiming to improve the accuracy and / or efficiency of information processing and improve the user experience.

[0006] This application provides a processing method, which can be applied to a first device, including the steps of:

[0007] S10: Generate or determine second information based on the first content and the first information.

[0008] Optionally, the method further comprises at least one of the following:

[0009] The first content includes at least one of pre-translated content, a preset data set, and a first preset model;

[0010] The first information corresponds to the source language;

[0011] The first information includes at least one of voice, text, image and gesture;

[0012] The second information corresponds to the target language;

[0013] The second information is output by the first device and / or the second device;

[0014] The output mode of the second information includes at least one of audio broadcast, text display, image display and video display.

[0015] Optionally, the method for obtaining or determining the first content includes at least one of the following:

[0016] In response to receiving the translation-related content, selecting or determining pre-translated content corresponding to the translation-related content;

[0017] Acquire historical voice information, and construct a preset data set and / or a first preset model based on the historical voice information.

[0018] Optionally, the method further comprises at least one of the following:

[0019] The translation-related content includes at least one of the document content, the subject name, and the keyword information;

[0020] Pre-translated content includes specialized terminology and / or context-specific terms;

[0021] The preset data set includes pronunciation habit information corresponding to at least one language;

[0022] Pre-translated content is provided through the first language model;

[0023] The second information is generated or determined by the second language model and / or the first preset model;

[0024] The first content is output to the second device, so that the second device generates or determines second information corresponding to the first information based on the first content.

[0025] Optionally, step S10 includes:

[0026] A first process is performed on the first information according to the first content to obtain second information.

[0027] Optionally, the method further comprises at least one of the following:

[0028] In response to receiving the preset instruction, selecting or determining a target language, and generating or determining second information in the target language;

[0029] In response to receiving the first translation request from the second device, selecting or determining a target language, and generating or determining second information in the target language;

[0030] outputting second information;

[0031] outputting the second information to the second device so that the second device outputs the second information;

[0032] In response to receiving a second translation request from the second device, sending the first content to the second device according to the second translation request, so that the second device generates or determines second information corresponding to the first information based on the first content;

[0033] The first processing includes at least one of translation, supplementation, summarization, deletion, replacement and word order adjustment.

[0034] The present application also provides a processing method, which can be applied to a second device, comprising the steps of:

[0035] S20: outputting second information, the second information being generated or determined based on the first content and the first information.

[0036] Optionally, the method further comprises at least one of:

[0037] At least one of the first information, the second information and the first content is provided by the first device;

[0038] The first content comprises at least one of pre-translation content, a preset data set and a first preset model;

[0039] The first information corresponds to a source language;

[0040] The first information comprises at least one of voice, text, image and gesture;

[0041] The second information corresponds to a target language;

[0042] The output mode of the second information comprises at least one of audio broadcast, text display, image display and video display.

[0043] The application further provides an intelligent terminal, comprising a memory and a processor, wherein the memory stores a processing program, and the processing program is executed by the processor to implement the steps of the above method.

[0044] The application further provides a storage medium storing a computer program, and the computer program is executed by the processor to implement the steps of the above method.

[0045] As described above, the processing method of the application can be applied to the first device, comprising the step of: S10: generating or determining second information based on the first content and the first information. The application further provides a processing method which can be applied to the second device (such as a mobile phone), comprising the step of: S20: outputting second information, the second information being generated or determined based on the first content and the first information. The processing method provided by the application can improve the accuracy and efficiency of information processing and improve the user's experience. BRIEF DESCRIPTION OF DRAWINGS

[0046] The accompanying drawings incorporated in and forming a part of the specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the application. In order to more clearly illustrate the technical solutions of the embodiments of the application, the drawings needed to be used in the embodiment description will be briefly introduced as follows. Obviously, for those of ordinary skill in the art, without paying creative labor, other drawings can also be obtained based on these drawings.

[0047] Figure 1 A hardware structure schematic diagram of an intelligent terminal for implementing various embodiments of the application;

[0048] Figure 2 A communication network system architecture diagram provided in an embodiment of the present application;

[0049] Figure 3 is a flowchart of a processing method according to the first embodiment;

[0050] Figure 4 is a flowchart of a processing method according to the second embodiment;

[0051] Figure 5 is a flowchart of a processing method according to a third embodiment;

[0052] Figure 6 is a flowchart of a processing method according to a fourth embodiment;

[0053] Figure 7 is a schematic diagram of an interaction sequence according to a fourth embodiment;

[0054] Figure 8 is a schematic diagram of an interface according to a fourth embodiment;

[0055] Figure 9 Schematic diagram of the structure of the processing device provided in the embodiment of the present application Figure 1 ;

[0056] Figure 10 Schematic diagram of the structure of the processing device provided in the embodiment of the present application Figure 2 .

[0057] The purpose of this application, its features, and advantages will be further described in conjunction with the embodiments and with reference to the accompanying drawings. The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and the accompanying text are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of this application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION

[0058] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.

[0059] It should be noted that, in this document, the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the sentence "comprising a ..." does not exclude the presence of other identical elements in the process, method, article or device comprising the element. In addition, components, features, and elements with the same name in different embodiments of the present application may have the same meaning or different meanings, and their specific meanings need to be determined by their explanation in the specific embodiment or further combined with the context of the specific embodiment.

[0060] It should be understood that although the terms "first," "second," "third," etc. may be used herein to describe various information, such information should not be limited to these terms. These terms are used solely to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the term "if," as used herein, may be interpreted as "upon," "when," or "in response to a determination." Furthermore, as used herein, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context indicates otherwise. It should be further understood that the terms "comprising" and "including" indicate the presence of the recited features, steps, operations, elements, components, items, types, and / or groups, but do not preclude the presence, occurrence, or addition of one or more other features, steps, operations, elements, components, items, types, and / or groups. The terms "or," "and / or," "including at least one of the following," etc., as used herein, may be interpreted as inclusive, meaning any one or any combination. For example, “comprising at least one of the following: A, B, C” means “any of the following: A; B; C; A and B; A and C; B and C; A and B and C”; and for another example, “A, B or C” or “A, B and / or C” means “any of the following: A; B; C; A and B; A and C; B and C; A and B and C”. An exception to this definition will occur only when a combination of elements, functions, steps or operations are inherently mutually exclusive in some manner.

[0061] It should be understood that, although the various steps in the flowchart in the embodiment of the present application are shown in sequence according to the indication of the arrows, these steps are not necessarily performed in sequence in the order indicated by the arrows. Unless clearly stated herein, the execution of these steps is not strictly limited in order, and they can be performed in other orders. Moreover, at least a portion of the steps in the figure may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but can be performed at different times, and their execution order is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of other steps or sub-steps or stages of other steps.

[0062] As used herein, the words "if" and "if" may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.

[0063] It should be noted that in this article, step codes such as S10 and S20 are used for the purpose of expressing the corresponding content more clearly and concisely, and do not constitute a substantial limitation on the order. When implementing the step, those skilled in the art may execute S20 first and then S10, etc., but these should all be within the scope of protection of this application.

[0064] It should be understood that the specific embodiments described herein are only used to explain the present application and are not intended to limit the present application.

[0065] In the subsequent description, the use of suffixes such as "module", "component" or "unit" to represent elements is only for the purpose of facilitating the description of the present application and has no specific meaning. Therefore, "module", "component" or "unit" can be used interchangeably.

[0066] Smart terminals can be implemented in various forms. For example, the smart terminals described in this application may include smart terminals such as mobile phones, tablet computers, laptop computers, PDAs, portable media players (PMPs), navigation devices, wearable devices, smart bracelets, pedometers, etc., as well as fixed terminals such as digital TVs and desktop computers.

[0067] The subsequent description will be made by taking a mobile terminal as an example. It will be understood by those skilled in the art that, in addition to components specifically used for mobile purposes, the configuration according to the embodiments of the present application can also be applied to fixed-type terminals.

[0068] See also Figure 1 , which is a schematic diagram of the hardware structure of a mobile terminal for implementing various embodiments of the present application. The mobile terminal 100 may include: an RF (Radio Frequency) unit 101, a WiFi module 102, an audio output unit 103, an A / V (Audio / Video) input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, a processor 110, and a power supply 111. Those skilled in the art will understand that Figure 1 The structure of the mobile terminal shown in the figure does not constitute a limitation to the mobile terminal. The mobile terminal may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0069] The following combination Figure 1 A detailed introduction to the various components of the mobile terminal:

[0070] The RF unit 101 can be used to send and receive information or receive signals during calls. Specifically, it receives downlink information from the base station and transmits it to the processor 110 for processing. It also transmits uplink data to the base station. Typically, the RF unit 101 includes, but is not limited to, an antenna, at least one amplifier, a transceiver, a coupler, a low-noise amplifier, a duplexer, and more. Furthermore, the RF unit 101 can communicate with the network and other devices via wireless communication. The above-mentioned wireless communications may use any communication standard or protocol, including but not limited to GSM (Global System of Mobile communication), GPRS (General Packet Radio Service), CDMA2000 (Code Division Multiple Access 2000), WCDMA (Wideband Code Division Multiple Access), TD-SCDMA (Time Division-Synchronous Code Division Multiple Access), FDD-LTE (Frequency Division Duplexing- Long Term Evolution), TDD-LTE (Time Division Duplexing- Long Term Evolution) and 5G, etc.

[0071] WiFi is a short-range wireless transmission technology. Mobile terminals can help users send and receive emails, browse web pages, and access streaming media through the WiFi module 102. It provides users with wireless broadband Internet access. Figure 1 The WiFi module 102 is shown, but it is understandable that it is not an essential component of the mobile terminal and can be omitted as needed without changing the essence of the invention.

[0072] The audio output unit 103 can convert audio data received by the RF unit 101 or the WiFi module 102 or stored in the memory 109 into an audio signal and output it as sound when the mobile terminal 100 is in a call signal reception mode, a talk mode, a recording mode, a voice recognition mode, a broadcast reception mode, or the like. Furthermore, the audio output unit 103 can also provide audio output related to a specific function performed by the mobile terminal 100 (e.g., a call signal reception sound, a message reception sound, etc.). The audio output unit 103 may include a speaker, a buzzer, or the like.

[0073] The A / V input unit 104 is used to receive audio or video signals. The A / V input unit 104 may include a graphics processing unit (GPU) 1041 and a microphone 1042. The GPU 1041 processes image data from still images or videos captured by an image capture device (such as a camera) in video capture mode or image capture mode. The processed image frames may be displayed on the display unit 106. The image frames processed by the GPU 1041 may be stored in the memory 109 (or other storage medium) or transmitted via the RF unit 101 or the WiFi module 102. The microphone 1042 can receive sound (audio data) in various operating modes, such as phone call mode, recording mode, and voice recognition mode, and process such sound into audio data. In phone call mode, the processed audio (voice) data may be converted into a format that can be transmitted to a mobile communication base station via the RF unit 101. The microphone 1042 may implement various noise cancellation (or suppression) algorithms to eliminate (or suppress) noise or interference generated during the reception and transmission of audio signals.

[0074] The mobile terminal 100 also includes at least one sensor 105, such as a light sensor, a motion sensor, or other sensors. Optionally, the light sensor includes an ambient light sensor and a proximity sensor. Optionally, the ambient light sensor can adjust the brightness of the display panel 1061 based on the brightness of the ambient light, and the proximity sensor can turn off the display panel 1061 and / or the backlight when the mobile terminal 100 is brought to your ear. An accelerometer, a type of motion sensor, can detect acceleration in all directions (typically three axes) and, when stationary, can detect the magnitude and direction of gravity. This can be used for applications that recognize the phone's posture (e.g., switching between landscape and portrait modes, related games, magnetometer posture calibration), vibration recognition-related functions (e.g., pedometer, tapping), and other functions. Other sensors that may be configured on a mobile phone, such as a fingerprint sensor, pressure sensor, iris sensor, molecular sensor, gyroscope, barometer, hygrometer, thermometer, infrared sensor, etc., are not described here.

[0075] The display unit 106 is used to display information input by the user or information provided to the user. The display unit 106 may include a display panel 1061, which may be configured in the form of a liquid crystal display (LCD), an organic light-emitting diode (OLED), or the like.

[0076] The user input unit 107 can be used to receive input digital or character information and generate key signal input related to user settings and function control of the mobile terminal. Optionally, the user input unit 107 may include a touch panel 1071 and other input devices 1072. The touch panel 1071, also known as a touch screen, can detect user touch operations on or near it (for example, operations performed on or near the touch panel 1071 using a finger, stylus, or any other suitable object or accessory) and drive corresponding connected devices according to pre-set programs. The touch panel 1071 may include a touch detection device and a touch controller. Optionally, the touch detection device detects the user's touch position and detects signals generated by the touch operation, transmitting the signals to the touch controller. The touch controller receives the touch information from the touch detection device, converts it into touch point coordinates, and then sends it to the processor 110. The touch controller can also receive and execute commands from the processor 110. Furthermore, the touch panel 1071 can be implemented using various types of sensors, including resistive, capacitive, infrared, and surface acoustic wave. In addition to the touch panel 1071, the user input unit 107 may further include other input devices 1072. Optionally, the other input devices 1072 may include, but are not limited to, one or more of a physical keyboard, function keys (such as volume control keys, power keys, etc.), a trackball, a mouse, a joystick, etc., and the specifics are not limited here.

[0077] Optionally, the touch panel 1071 may cover the display panel 1061. When the touch panel 1071 detects a touch operation on or near it, it transmits the information to the processor 110 to determine the type of touch event. The processor 110 then provides a corresponding visual output on the display panel 1061 according to the type of touch event. Figure 1 In the embodiment, the touch panel 1071 and the display panel 1061 are two independent components to realize the input and output functions of the mobile terminal. However, in some embodiments, the touch panel 1071 and the display panel 1061 can be integrated to realize the input and output functions of the mobile terminal, which is not limited here.

[0078] The interface unit 108 serves as an interface through which at least one external device can be connected to the mobile terminal 100. For example, the external device may include a wired or wireless headset port, an external power supply (or battery charger) port, a wired or wireless data port, a memory card port, a port for connecting a device with an identification module, an audio input / output (I / O) port, a video I / O port, a headphone port, etc. The interface unit 108 may be used to receive input (e.g., data information, power, etc.) from an external device and transmit the received input to one or more elements within the mobile terminal 100 or may be used to transmit data between the mobile terminal 100 and an external device.

[0079] The memory 109 can be used to store software programs and various data. The memory 109 can mainly include a program storage area and a data storage area, and the program storage area can store an operating system, application programs required by at least one function (such as a sound playing function, an image playing function, etc.), and the like; and the data storage area can store data created according to the use of the mobile phone (such as audio data, a phone book, etc.), and the like. In addition, the memory 109 can include a high-speed random access memory, and can also include a nonvolatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state memory device.

[0080] The processor 110 is the control center of the mobile terminal, connects all parts of the mobile terminal through various interfaces and lines, executes various functions of the mobile terminal and processes data by running or executing software programs and / or modules stored in the memory 109 and calling data stored in the memory 109, and thus monitors the mobile terminal as a whole. The processor 110 can include one or more processing units; preferably, the processor 110 can integrate an application processor and a modem processor, and the application processor can mainly process an operating system, a user interface, and application programs, and the like, and the modem processor can mainly process wireless communication. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 110.

[0081] The mobile terminal 100 can also include a power supply 111 (such as a battery) for supplying power to various components; preferably, the power supply 111 can be logically connected to the processor 110 through a power management system, so as to realize the functions of managing charging, discharging, and power consumption management, and the like through the power management system.

[0082] Although Figure 1 The mobile terminal 100 can also include a Bluetooth module and the like, which are not described here.

[0083] In order to facilitate the understanding of the embodiments of the present application, the communication network system based on the mobile terminal of the present application is described below.

[0084] Please refer to Figure 2 , Figure 2 A communication network system architecture diagram provided by the embodiments of the present application, the communication network system is a LTE system of general mobile communication technology, the LTE system includes a UE (User Equipment, user equipment) 201, an E-UTRAN (Evolved UMTS Terrestrial Radio Access Network, evolved UMTS terrestrial radio access network) 202, an EPC (Evolved Packet Core, evolved packet core network) 203 and an operator's IP service 204 which are sequentially connected in communication.

[0085] Optionally, UE201 may be the above-mentioned terminal 100, which will not be described in detail here.

[0086] E-UTRAN 202 includes eNodeB 2021 and other eNodeBs 2022 . Optionally, eNodeB 2021 may be connected to other eNodeBs 2022 via a backhaul (eg, an X2 interface). eNodeB 2021 is connected to EPC 203 , and eNodeB 2021 may provide UE 201 with access to EPC 203 .

[0087] The EPC 203 may include an MME (Mobility Management Entity) 2031, an HSS (Home Subscriber Server) 2032, other MMEs 2033, an SGW (Serving GateWay) 2034, a PGW (PDN GateWay) 2035, and a PCRF (Policy and Charging Rules Function) 2036. Optionally, the MME 2031 is a control node that processes signaling between the UE 201 and the EPC 203, providing bearer and connection management. The HSS 2032 provides registers for managing functions such as the Home Location Register (not shown) and stores user-specific information such as service features and data rates. All user data can be sent through SGW2034. PGW2035 can provide IP address allocation and other functions for UE 201. PCRF2036 is the policy and charging control policy decision point for service data flows and IP bearer resources. It selects and provides available policy and charging control decisions for the policy and charging execution function unit (not shown in the figure).

[0088] The IP service 204 may include the Internet, an intranet, an IMS (IP Multimedia Subsystem), or other IP services.

[0089] Although the above introduction takes the LTE system as an example, those skilled in the art should know that this application is not only applicable to the LTE system, but also to other wireless communication systems, such as GSM, CDMA2000, WCDMA, TD-SCDMA, 5G and future new network systems (such as 6G), etc., which are not limited here.

[0090] Based on the above-mentioned mobile terminal hardware structure and communication network system, various embodiments of the present application are proposed.

[0091] First embodiment

[0092] Reference Figure 3 , Figure 3 : is a flowchart of a processing method according to the first embodiment. The processing method of the embodiment of the present application can be applied to a first device (such as a mobile phone), including the following steps:

[0093] S10: Generate or determine second information based on the first content and the first information.

[0094] Optionally, the processing method in the embodiment of the present application is considered from different perspectives and is applicable to a variety of application scenarios. In this embodiment, the number of execution devices, information type and processing method are used as examples for explanation.

[0095] (1) Considering the number of execution devices, the application scenario includes at least one of the following:

[0096] Scenario A: The first device obtains first information, generates or determines second information based on the first content and the first information, and outputs the second information;

[0097] Scenario B: The first device obtains the first information, generates or determines the second information based on the first content and the first information, and sends the second information to the second device so that the second device outputs the second information.

[0098] Optionally, in scenario A, the first user and the second user communicate using the first device. For example, the first device displays the information collected from different users on a split screen and is used as a translator, which is suitable for face-to-face communication scenarios.

[0099] Optionally, in scenario B, a first user uses a first device to communicate with a second user holding a second device, which can be applicable to an instant translation scenario of face-to-face communication or a remote call scenario.

[0100] Optionally, the method for acquiring the first information includes collecting the first information through a sensor and / or receiving the first information sent by an associated device. Optionally, the associated device includes a second device or other devices.

[0101] (2) From the perspective of information type, the application scenario includes at least one of the following:

[0102] Scenario C: The first information is voice. The first device collects the voice of the first user through a sensor such as a microphone and / or receives voice sent by an associated device to determine the first information.

[0103] Scenario D: The first information is at least one of text, image, and gesture. When it is inconvenient for the first user to speak, text can be input into the first device, and / or an image can be captured or selected through the first device, and / or the user's gesture information can be collected through the first device to determine the first information.

[0104] Optionally, in scenario C, the first information needs to be processed by the first device to generate or determine the second information to conform to the language habits of the first user and / or the second user. The first processing includes at least one of translation, supplementation, summarization, deletion, replacement and word order adjustment. The second information can be voice, or at least one of text, image or video.

[0105] Optionally, in scenario D, the first user can only input text content such as keywords that he wants to express, and the first device performs semantic understanding and / or content supplement based on the text input or selected by the user to form a complete sentence, and / or converts the complete sentence into voice and outputs it to the second user as second information. This can be repeated to assist the first user in communicating with the second user.

[0106] Optionally, in scenario D, the first user can also take a photo or select an image, and the first device can understand the user's intention through image recognition, such as recognizing information such as text and objects contained in the image taken by the user, or understanding the meaning represented by images such as emoticons selected by the user, and then generate a complete sentence, and / or convert the complete sentence into voice and output it to the second user as the second information. This process can be repeated to assist the first user in communicating with the second user.

[0107] Optionally, in scenario D, the first device can also use sensors to identify information such as the first user's facial expressions, lip shapes, movements, and gestures. For example, when the first user is a person with a speech impaired person and needs to communicate with the second user through the first device, the first device can identify and process the collected first information such as the first user's facial expressions, lip shapes, movements, and gestures, and then generate or determine second information to output to the second user. This process can be repeated to assist the first user in communicating with the second user. Similarly, when the second user is a person with a speech impaired person, the second information can also be output through the first device and / or the second device in the form of at least one of audio broadcast, text display, image display, and video display. Optionally, the video display includes synthesized sign language video, etc.

[0108] (3) From the perspective of processing methods, the application scenarios include at least one of the following:

[0109] Scenario E: The first device translates the first information to obtain second information, and outputs the second information to the second user via the first device and / or the second device, so that the second information conforms to the first user's expression habits and / or the second user's language habits.

[0110] Scenario F: The first device performs at least one of supplementation, summarization, deletion, replacement, and word order adjustment on the first information to obtain second information, and outputs the second information to the second user through the first device and / or the second device, so that the second information conforms to the expression habits of the first user and / or the language habits of the second user.

[0111] Optionally, in scenario E, if the first information contains professional terms and / or specific context terms, and / or the first user has a heavy accent, it may increase the difficulty of translation, reduce the accuracy of the translation result, and / or extend the translation time. In the embodiment of the present application, auxiliary translation is performed by providing the first content to reduce the difficulty of translation and / or improve the accuracy of the translation result and / or translation efficiency. Optionally, the first content includes at least one of pre-translated content, a preset data set, and a first preset model.

[0112] Optionally, in scenario F, if there is missing information in the first information, the first information can be supplemented, for example, the missing content in the first information can be supplemented in combination with the context content to obtain the second information; if there is redundant information in the first information, the first information can be summarized and / or deleted, for example, the information can be summarized according to the degree of conciseness required by the user, and / or the repeated content in the first information can be deleted to obtain the second information, thereby optimizing the first information so that the obtained second information conforms to the language habits of the second user; the translated first information can also be processed by vocabulary replacement, word order adjustment, and timbre simulation to obtain the second information, so that the second information fits the expression habits of the first user.

[0113] Optionally, the application scenarios of the processing method in the embodiments of the present application include but are not limited to the above scenarios, and the scenarios can be implemented in any combination. For example, the combined implementation of Scenario B, Scenario C, and Scenario E is as follows: a first device obtains first information, the first information being speech in a source language, translates the first information based on the first content to obtain second information in a target language, and sends the second information to a second device, so that the second device outputs the second information. Such combinations are not described one by one in the embodiments of the present application.

[0114] Through the technical solution of this embodiment, specifically by generating or determining the second information based on the first content and the first information, the accuracy and / or efficiency of information processing can be improved, and the user experience can be improved.

[0115] Second embodiment

[0116] Based on the above Figure 3 The first embodiment shown in FIG. 1 is a second embodiment of the present application, which proposes a processing method, see Figure 4 , Figure 4 FIG. 1 is a flow chart of a processing method according to a second embodiment. The method includes a method for obtaining or determining first content, specifically including:

[0117] Step S01: In response to receiving translation-related content, selecting or determining pre-translation content corresponding to the translation-related content;

[0118] Optionally, when the language used in the first information is different from the target language, that is, the source language used in the first information (for example, Chinese) and the required target language (for example, English) are different languages, the first information needs to be translated and processed by the first device and / or the second device to obtain the second information in the target language.

[0119] Alternatively, when the first message contains professional terminology and / or context-specific terms (such as slang, idioms, and two-part allegorical sayings), if the first device directly uses a local model for instant translation, the translation accuracy may be low and / or the translation time may be long due to translation difficulty and / or insufficient computing power, thus affecting the effectiveness of the instant translation. If the first message is directly translated instantly using an associated device (such as a server), there is a risk of privacy leakage. The embodiments of this application propose a solution for assisting translation through pre-translated content.

[0120] Optionally, the translation-related content includes at least one of document content, subject name, and keyword information.

[0121] Optionally, the pre-translated content includes professional terms and / or context-specific terms.

[0122] Optionally, pre-translated content is provided via a first language model.

[0123] Optionally, before starting a conversation and / or during the conversation, the user can provide translation-related content related to the conversation, such as relevant documents, topic names, keywords and other information. The first device and / or associated devices can directly translate the professional terms and / or specific context terms involved in the translation-related content to obtain pre-translated content; and / or, the first device and / or associated devices can determine the relevant fields based on the translation-related content, and / or, call the relevant database, select or determine relevant background knowledge information as pre-translation content, thereby reducing the translation delay in the subsequent instant translation process, and / or, can use the computing power of the associated device to provide auxiliary support for the subsequent execution of instant translation by the local model.

[0124] Optionally, when pre-translation is performed on an associated device (e.g., a server), the pre-translation can be performed using a large language model (e.g., a first language model) deployed on the associated device to assist in the subsequent real-time translation performed by the first device. Alternatively, if the translation-related content still requires strong privacy, the translation-related content can be encrypted before being sent to the associated device. For example, perturbation processing and / or differential privacy processing can be performed. The content returned by the associated device can then be decrypted to obtain the pre-translated content, thereby further protecting data privacy during the pre-translation process.

[0125] Step S02: Acquire historical voice information, and construct a preset data set and / or a first preset model based on the historical voice information.

[0126] Alternatively, if the first user has a strong accent or pronunciation problems, if the first device and / or directly performs instant translation of the first information corresponding to the first user, the translation result may be less accurate and / or take longer due to reasons such as greater translation difficulty, thereby affecting the effectiveness of the instant translation. In the embodiments of the present application, a solution for assisting translation using a preset data set and / or a first preset model is proposed.

[0127] Optionally, the preset data set includes pronunciation habit information corresponding to at least one language.

[0128] Optionally, the first device may collect historical voice data with the user's permission. Optionally, the historical voice data may come from the user's voice assistant, call recordings, and voice messages. Optionally, by collecting voice data under different environmental noise conditions, different emotional states, different speaking speeds, different conversation goals, and different conversation topics, the diversity and / or robustness of the dataset can be increased.

[0129] Optionally, data preprocessing and / or feature extraction may be performed on the collected historical speech data.

[0130] Optionally, the data preprocessing process includes at least one of the following:

[0131] Perform noise reduction on speech data, using filters and / or noise reduction algorithms to reduce background noise;

[0132] Perform speech segmentation to divide the continuous speech stream into single word or syllable segments, and cut the continuous speech signal into shorter frames, for example, 25 milliseconds per frame, with 10 milliseconds overlap between frames to reduce information loss;

[0133] Normalize the speech signal to ensure that the speech segments have a uniform volume level;

[0134] The high-frequency signal is enhanced by pre-emphasis through a first-order high-pass filter to compensate for the attenuation of the high-frequency signal during the recording process.

[0135] Optionally, the feature extraction method includes at least one of the following:

[0136] Through the Short-Time Fourier Transform (STFT), the signal is converted from the time domain to the frequency domain to analyze the spectral characteristics of the signal;

[0137] Calculate the logarithmic energy spectrum of the magnitude spectrum, convert the linear frequency into nonlinear Mel frequency using a Mel filter bank to simulate the human auditory perception mechanism, and perform a discrete cosine transform (DCT) on the output of the filter bank to obtain Mel frequency cepstral coefficients (MFCCs).

[0138] Describe the vocal tract response of the speech signal through linear prediction and extract the linear prediction coefficients;

[0139] Analyze the frequency and duration of audio signals to extract pitch and rhythm information.

[0140] Optionally, the extracted features are organized into a preset data set, where each sample corresponds to a pronunciation of the user. The data set may also contain rich annotation information, such as pronunciation accuracy, speaking speed, emotion, etc. The data set may be divided into a training set, a validation set, and a test set to facilitate the use of the preset data set for model training and / or evaluation.

[0141] Optionally, based on the constructed preset data set, a first preset model can be further trained to obtain a traditional machine learning model, such as a random forest, a support vector machine, or a deep learning model, such as a recurrent neural network (RNN) and / or a convolutional neural network (CNN).

[0142] Optionally, the first information of the first user may be directly and quickly identified and / or processed through a preset data set and / or a first preset model predetermined by the first device to generate or determine the second information.

[0143] Optionally, when the received first information is identified and / or processed by the second device, for the second device, if the identification and / or processing is temporarily performed based on the first information during the call, the generated second information is difficult to simulate the tone that fits the first user who generated the first information, that is, it is difficult to simulate the pronunciation habits of the first user, and voice recognition also has a certain degree of difficulty. Therefore, in an embodiment of the present application, the preset data set and / or the first preset model predetermined by the first device is provided to the second device, so that the second device can generate or determine the second information corresponding to the first information based on the preset data set and / or the first preset model.

[0144] Optionally, when the first device sends the preset data set and / or the first preset model to the second device, the data privacy of the first user needs to be ensured. Data desensitization and other methods can be used before sending. Privacy++, an emotional privacy protection framework mechanism based on a cyclic adversarial generative neural network, can also be used to protect the privacy information contained in the user's voice.

[0145] Optionally, the pronunciation habits of the user may change over time and other factors, so the preset data set and / or the first preset model may be updated according to user settings and / or periodically.

[0146] Through the technical solution of this embodiment, specifically by selecting or determining the pre-translation content corresponding to the translation-related content in response to receiving the translation-related content; and / or, obtaining historical voice information, constructing a preset data set and / or a first preset model based on the historical voice information, at least one of the pre-translation content, the preset data set and the first preset model can be determined to provide auxiliary support for the instant translation process, which can reduce the difficulty of instant translation, and / or improve the accuracy of the translation results, and / or reduce the delay of the translation process.

[0147] Third embodiment

[0148] Based on any of the above embodiments, the third embodiment of this application is proposed. This embodiment proposes a processing method, see Figure 5 , Figure 5 FIG. 1 is a flow chart of a processing method according to the third embodiment. Step S10 of the method includes:

[0149] S101: Perform a first process on first information according to first content to obtain second information.

[0150] Optionally, the first processing includes at least one of translation, supplementation, summarization, deletion, replacement and word order adjustment.

[0151] Optionally, a triggering method for performing translation processing on the first information includes:

[0152] In response to receiving a preset instruction, selecting or determining a target language, and generating or determining second information in the target language; and / or,

[0153] In response to receiving the first translation request from the second device, a target language is selected or determined, and second information in the target language is generated or determined.

[0154] Optionally, the first device may trigger an instant translation function according to a preset instruction.

[0155] Optionally, the preset instruction can be initiated by user operation or automatically by the first device. For example, when it is identified that the current call object (second user) is in the preset list, and / or the remark information of the current call object (second user) contains the target language, the translation instruction can be triggered to translate the first information according to the selected or determined target language to obtain the second information in the target language.

[0156] Optionally, the first device may trigger an instant translation function according to the first translation request of the second device.

[0157] Optionally, when a second user holding a second device starts a call with a first user holding a first device, if the second device recognizes that the source language used by the first user does not match the target language of the second user, the second device can send a first translation request to the first device. The first translation request includes at least one of the target language, information type, and output method. After the first device receives the first translation request, it can call the first content related to the target language to perform instant translation of the first information. Optionally, the first content includes at least one of pre-translated content, a preset data set, and a first preset model.

[0158] Optionally, when the first device generates or determines the second information based on the first information, if the first information is speech, the speech signal may be converted into text data using speech recognition technology. For example, deep learning models such as recurrent neural networks (RNNs) and / or convolutional neural networks (CNNs) may be used in conjunction with an attention mechanism to generate the converted text data. Optionally, the converted text data is input into a natural language processing (NLP) module. This module can perform tasks such as language detection, word segmentation, part-of-speech tagging, and named entity recognition, providing a foundation for subsequent translation and / or other text processing. Neural network-based translation models can translate source language text into target language text. These translation models can be pre-trained using large-scale bilingual corpora to generate fluent and accurate translation results.

[0159] Optionally, for text that needs to be supplemented or summarized, a text generation model, such as a Transformer-based model, can be used to generate additional information or summarize key points. The text generation model can generate coherent and relevant text content based on the context. Optionally, the text generation model used to supplement or summarize the first information is the first preset model in the first content.

[0160] Optionally, for text that needs to be deleted or replaced, a rule-based approach or machine learning model can be used to identify and modify specific portions of the text. For example, a text processing model can be trained to identify and replace sensitive information and / or inappropriate content. Optionally, the text processing model used to delete or replace the first information is a first preset model in the first content.

[0161] Optionally, to adapt to the grammatical structure of the target language and / or the language habits desired by the user, the translated text may need to undergo word order adjustment. This can be achieved through dependency parsing and syntactic analysis to ensure that the text is grammatically and semantically correct. Optionally, the word order adjustment process can be performed based on a preset dataset in the first content, or the first preset model in the first content can be used as a word order adjustment model to further optimize and adjust the translated text.

[0162] Alternatively, if the second information required by the second user is speech, the processed text can be converted back into a speech signal and output as speech using text-to-speech (TTS) technology. This can be achieved by using a vocoder and a speech synthesis model to generate natural and fluent speech output.

[0163] Optionally, the first device can also use an end-to-end speech model to directly convert the speech of the first information into the speech output of the second information, and use a unified model structure to directly map the original speech signal to the target output. In the generation stage, the decoder is responsible for converting the probability distribution of the model output into understandable text or speech signals, eliminating intermediate steps such as feature extraction, achieving real-time conversion, and further reducing the delay of instant translation, and / or, it can be combined with a language model to improve the fluency and accuracy of the generated results, and / or, adopt anti-noise technology and multimodal input (for example, combined with visual information) to improve recognition accuracy.

[0164] Through the technical solution of this embodiment, specifically by performing a first processing on the first information based on the first content to obtain the second information, and performing at least one of translation, supplementation, summarization, deletion, replacement and word order adjustment, it is ensured that the generated second information conforms to the expression habits of the first user and / or conforms to the language habits of the second user, thereby improving the user experience.

[0165] Fourth embodiment

[0166] Referring to Figure 6 , Figure 6 is a flowchart of a processing method according to the fourth embodiment. The processing method of the embodiments of the present application can be applied to a second device (such as a mobile phone), and includes the following steps:

[0167] S20: outputting second information, the second information being generated or determined based on the first content and the first information.

[0168] Optionally, at least one of the first information, the second information, and the first content is provided by the first device.

[0169] Optionally, the first content includes at least one of pre-translation content, a preset data set, and a first preset model.

[0170] Optionally, the first information corresponds to a source language.

[0171] Optionally, the first information includes at least one of voice, text, images, and gestures.

[0172] Optionally, the second information corresponds to a target language.

[0173] Optionally, the output mode of the second information includes at least one of audio broadcast, text display, image display, and video display.

[0174] Optionally, in a scenario where the first user uses the first device to communicate with a second user holding the second device, the first device can send the generated or determined second information to the second device to make the second device output the second information. Optionally, if the second device identifies that the source language used by the first user does not match the target language of the second user, the second device can send a first translation request to the first device, the first translation request including at least one of the target language, the information type, and the output mode. After receiving the first translation request, the first device can call the first content related to the target language to perform instant translation on the first information. Optionally, the first content includes at least one of pre-translation content, a preset data set, and a first preset model.

[0175] Optionally, the process of generating or determining the second information based on the first content and the first information in the embodiments of the present application can also be performed by an associated device (such as a network device).

[0176] Referring to Figure 7 , Figure 7 is an interaction timing diagram according to the fourth embodiment, such as Figure 7As shown, in a scenario where a first user uses a first device to communicate with a second user holding a second device, the first device and / or the second device are associated with a network device. Optionally, if the second device recognizes that the source language used by the first user does not match the target language of the second user, the second device can send a first processing request to the network device. The first processing request includes at least one of the target language, information type, and output method. After the network device receives the first processing request, it can send relevant request information to the first device, such as a first translation request. In response to receiving the request information sent by the network device, the first device sends the first content and / or the first information to the network device, so that the network device generates or determines the second information based on the first content and the first information. Optionally, the first content includes at least one of pre-translated content, a preset data set, and a first preset model. Optionally, the network device sends the generated or determined second information to the second device, so that the second device outputs the second information.

[0177] Optionally, the second device selects or determines the target language in a manner including at least one of the following:

[0178] Select or determine the target language according to system settings;

[0179] selecting or determining a target language based on the third information;

[0180] Select or determine the target language based on user operations.

[0181] Optionally, the third information includes voice information currently sent by the second user.

[0182] Optionally, the second device may perform a first process on the received first information according to the needs of the second user to obtain the second information. Optionally, the first process includes at least one of translation, supplementation, summarization, deletion, replacement, and word order adjustment.

[0183] Optionally, in a scenario where a first user uses a first device to communicate with a second user holding a second device, the second device may also generate or determine second information based on the first information and output it. Optionally, if the second device recognizes that the source language used by the first user does not match the target language of the second user, a second translation request may be sent to the first device. In response to receiving the second translation request from the second device, the first device sends the first content to the second device according to the second translation request, so that the second device generates or determines the second information corresponding to the first information based on the first content. Optionally, the first content includes at least one of pre-translated content, a preset data set, and a first preset model. Optionally, the second device outputs the second information. Optionally, the second information is output in the form of at least one of audio broadcast, text display, image display, and video display.

[0184] Optionally, for the second device, if it is necessary to learn information such as the pronunciation habits of the other user separately during calls with different devices and / or users, the real-time performance of information processing will be poor and / or the learning cost will be high. If at least one of the pre-translated content, preset data set and the first preset model provided by the other user can be obtained, the speed of information processing can be increased and / or the learning cost in the information processing process can be reduced and / or the effect of information processing can be improved.

[0185] Optionally, in a scenario where a first user and a second user communicate using the first device, the first device may directly output the generated or determined second information, with the second information being output in at least one of an audio broadcast, a text display, an image display, and a video display. Optionally, the video display includes a synthesized sign language video. Optionally, the first device may also send the second information to the second device, causing the second device to output the second information.

[0186] Optionally, the processing method further includes:

[0187] generating or determining target content based on at least one of the first information, the second information, and the third information;

[0188] generating or determining fourth information according to the target content;

[0189] The second process is performed on the fourth information.

[0190] Optionally, the step of generating or determining fourth information according to the target content includes:

[0191] Calling a target application and / or target file according to the target content;

[0192] Fourth information is generated or determined based on the target application and / or the target file.

[0193] Optionally, performing a second processing on the fourth information includes at least one of the following:

[0194] displaying fourth information on the first device;

[0195] displaying fourth information on the second device;

[0196] A call summary and / or a to-do list is generated based on the fourth information.

[0197] Reference Figure 8 , Figure 8 is a schematic diagram of an interface according to the fourth embodiment, as shown in FIG. Figure 8As shown, the first device and / or the second device can identify the target content in the first message and / or the second message in real time during the call, and then call the relevant target application and / or target file accordingly. For example, the third message sent by the second user contains keywords such as "tomorrow's meeting time". The first user does not need to exit the call interface to manually search for relevant information. The first device can automatically call the meeting appointment information of the conference software and display it on the interface of the first device, and / or, after confirmation by the first user, send the meeting appointment information to the second device so that the meeting appointment information is displayed on the interface of the second device. This can provide the user with the target information needed in a timely manner, assist the user in performing related operations, improve the efficiency of the voice call, and / or improve the intelligence of the call process to improve the user's experience.

[0198] Optionally, the first device and / or the second device can generate a call summary and / or a to-do list based on the target information involved in the call. Optionally, the call summary and / or the to-do list can be synchronized between the user's multiple devices, such as at least one of a smartphone, a tablet, a smartwatch, and a computer, so that the user can view and update them anytime and anywhere. The call summary can help the user quickly review and organize key information in the call to avoid missing important content. Especially when the call content is extensive or involves multiple matters, the summary can provide a clear and concise overview of the information; by automatically generating a to-do list, it can assist the user in quickly converting action points in the call into specific tasks, facilitating subsequent execution and tracking, and helping to ensure that tasks are completed on time, thereby improving the user experience.

[0199] Through the technical solution of this embodiment, specifically by outputting the second information, which is generated or determined based on the first content and the first information, the accuracy and / or efficiency of information processing can be improved, and / or, by combining the fourth information, the efficiency of voice calls can be improved, and / or, the intelligence of the call process can be improved to improve the user experience.

[0200] Fifth embodiment

[0201] See Figure 9 , Figure 9 Schematic diagram of the structure of the processing device provided in the embodiment of the present application Figure 1 , the device can be mounted on or be the first device in the above method embodiment. Figure 9 The processing device shown can be used to perform part or all of the functions of the method embodiment described in the above embodiment. Figure 9 As shown, the processing device 1100 includes:

[0202] The processing module 1101 is configured to generate or determine second information based on the first content and the first information.

[0203] Optionally, the device further comprises at least one of the following:

[0204] The first content includes at least one of pre-translated content, a preset data set, and a first preset model;

[0205] The first information corresponds to the source language;

[0206] The first information includes at least one of voice, text, image and gesture;

[0207] The second information corresponds to the target language;

[0208] The second information is output by the first device and / or the second device;

[0209] The output mode of the second information includes at least one of audio broadcast, text display, image display and video display.

[0210] Optionally, the method for obtaining or determining the first content includes at least one of the following:

[0211] In response to receiving the translation-related content, selecting or determining pre-translated content corresponding to the translation-related content;

[0212] Acquire historical voice information, and construct a preset data set and / or a first preset model based on the historical voice information.

[0213] Optionally, the device further comprises at least one of the following:

[0214] The translation-related content includes at least one of the document content, the subject name, and the keyword information;

[0215] Pre-translated content includes specialized terminology and / or context-specific terms;

[0216] The preset data set includes pronunciation habit information corresponding to at least one language;

[0217] Pre-translated content is provided through the first language model;

[0218] The second information is generated or determined by the second language model and / or the first preset model;

[0219] The first content is output to the second device, so that the second device generates or determines second information corresponding to the first information based on a preset data set and / or a first preset model.

[0220] Optionally, generating or determining the second information based on the first content and the first information includes:

[0221] A first process is performed on the first information according to the first content to obtain second information.

[0222] Optionally, the device further comprises at least one of the following:

[0223] In response to receiving the preset instruction, selecting or determining a target language, and generating or determining second information in the target language;

[0224] In response to receiving the first translation request from the second device, selecting or determining a target language, and generating or determining second information in the target language;

[0225] outputting second information;

[0226] outputting the second information to the second device so that the second device outputs the second information;

[0227] In response to receiving a second translation request from the second device, sending the first content to the second device according to the second translation request, so that the second device generates or determines second information corresponding to the first information based on the first content;

[0228] The first processing includes at least one of translation, supplementation, summarization, deletion, replacement and word order adjustment.

[0229] The processing device provided in the embodiment of the present application has similar implementation principles and beneficial effects to the technical solutions shown in the above-mentioned corresponding method embodiments, and will not be described in detail here.

[0230] Sixth embodiment

[0231] See Figure 10 , Figure 10 Schematic diagram of the structure of the processing device provided in the embodiment of the present application Figure 2 , the device can be mounted on or is the second device in the above method embodiment. Figure 10 The processing device shown can be used to perform part or all of the functions of the method embodiment described in the above embodiment. Figure 10 As shown, the processing device 1200 includes:

[0232] The output module 1201 is configured to output second information, where the second information is generated or determined based on the first content and the first information.

[0233] Optionally, the device further comprises at least one of the following:

[0234] At least one of the first information, the second information, and the first content is provided by the first device;

[0235] The first content includes at least one of pre-translated content, a preset data set, and a first preset model;

[0236] The first information corresponds to the source language;

[0237] The first information includes at least one of voice, text, image and gesture;

[0238] The second information corresponds to the target language;

[0239] The output mode of the second information includes at least one of audio broadcast, text display, image display and video display.

[0240] The processing device provided in the embodiment of the present application has similar implementation principles and beneficial effects to the technical solutions shown in the above-mentioned corresponding method embodiments, and will not be described in detail here.

[0241] An embodiment of the present application further provides an intelligent terminal, comprising a memory and a processor, wherein a processing program is stored in the memory, and when the processing program is executed by the processor, the steps of the processing method in any of the above embodiments are implemented.

[0242] An embodiment of the present application further provides a storage medium having a processing program stored thereon. When the processing program is executed by a processor, the steps of the processing method in any of the above embodiments are implemented.

[0243] In the embodiments of the smart terminal and storage medium provided in this application, all technical features of any of the above-mentioned processing method embodiments may be included. The expanded and explained contents of the specification are basically the same as those of the embodiments of the above-mentioned methods and will not be repeated here.

[0244] An embodiment of the present application further provides a computer program product, which includes computer program code. When the computer program code runs on a computer, the computer executes the methods in the various possible implementation modes described above.

[0245] An embodiment of the present application also provides a chip, including a memory and a processor, wherein the memory is used to store computer programs, and the processor is used to call and run the computer programs from the memory, so that a device equipped with the chip executes the methods in the various possible implementation modes as described above.

[0246] It is understood that the above scenarios are merely examples and do not limit the application scenarios of the technical solutions provided in the embodiments of this application. The technical solutions of this application can also be applied to other scenarios. For example, those skilled in the art will appreciate that with the evolution of system architecture and the emergence of new business scenarios, the technical solutions provided in the embodiments of this application will also be applicable to similar technical problems.

[0247] The serial numbers of the above embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0248] The steps in the method of the embodiment of the present application can be adjusted in order, combined and deleted according to actual needs.

[0249] The units in the device of the embodiment of the present application can be merged, divided and deleted according to actual needs.

[0250] In this application, the same or similar terminology, technical solutions and / or application scenario descriptions are generally only described in detail the first time they appear. When they appear again later, they are generally not repeated for the sake of brevity. When understanding the technical solutions and other contents of this application, for the same or similar terminology, technical solutions and / or application scenario descriptions that are not described in detail later, you can refer to the previous relevant detailed descriptions.

[0251] In this application, the description of each embodiment has its own focus. For parts that are not described or recorded in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0252] The various technical features of the technical solution of this application can be combined arbitrarily. In order to make the description concise, not all possible combinations of the various technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0253] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) as mentioned above, and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, controlled terminal, or network device, etc.) to execute the method of each embodiment of the present application.

[0254] In the above embodiments, all or part of the embodiments can be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments can be implemented in the form of a computer program product. A computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions according to the embodiments of the present application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a storage medium or transmitted from one storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) means. The storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that integrates one or more available media. Available media can be magnetic media (e.g., floppy disks, storage disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media (e.g., solid-state disks (SSDs)).

[0255] The above are only preferred embodiments of the present application and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. A processing method, applied to a first device, characterized in that: Including steps: generating or determining second information based on the first content and the first information, and / or, Sending the first information and the first content to the second device, so that the second device generates or determines second information corresponding to the first information based on the first content; The first content includes pre-translated content, a preset data set, and a first preset model; The method for determining the first content includes: in response to receiving translation-related content, selecting or determining pre-translation content corresponding to the translation-related content; before and / or during the conversation, sending the translation-related content related to the conversation to the associated device, so that the associated device determines a relevant field based on the translation-related content; and / or, calling a relevant database to select or determine relevant background knowledge information as pre-translation content, where the translation-related content includes at least one of document content, subject name, and keyword information; when the associated device is used for pre-translation, the pre-translation is performed with the help of a large language model deployed in the associated device to provide auxiliary support for the first device to perform instant translation; the preset data set and the first preset model are constructed by the first device based on the acquired historical voice information, the preset data set includes pronunciation habit information corresponding to at least one language, and the first preset model is trained based on the preset data set; The first information corresponds to a source language, and the second information corresponds to a target language.

2. The processing method according to claim 1, characterized in that Also include at least one of the following: The first information includes at least one of voice, text, image and gesture; The second information is output by the first device and / or the second device; The output mode of the second information includes at least one of audio broadcast, text display, image display and video display.

3. The processing method according to claim 1, characterized in that Also include at least one of the following: Pre-translated content includes specialized terminology and / or context-specific terms; The second information is generated or determined by a second language model.

4. The processing method according to claim 2, characterized in that The method further comprises: A first process is performed on the first information according to the first content to obtain second information.

5. The processing method according to claim 4, characterized in that Also include at least one of the following: In response to receiving the preset instruction, selecting or determining a target language, and generating or determining second information in the target language; In response to receiving the first translation request from the second device, selecting or determining a target language, and generating or determining second information in the target language; outputting the second information to the second device so that the second device outputs the second information; In response to receiving a second translation request from the second device, sending the first content to the second device according to the second translation request, so that the second device generates or determines second information corresponding to the first information based on the first content; The first processing includes at least one of translation, supplementation, summarization, deletion, replacement and word order adjustment.

6. A processing method, applied to a second device, characterized in that: Including steps: outputting second information, where the second information is generated or determined by the first device and / or the second device based on the first content and the first information; The first content includes pre-translated content, a preset data set, and a first preset model; The method for determining the first content includes: in response to receiving the translation-related content, the first device selects or determines pre-translation content corresponding to the translation-related content; before and / or during the conversation, the translation-related content related to the conversation is sent to the associated device, so that the associated device determines the relevant field based on the translation-related content, and / or calls a relevant database to select or determine relevant background knowledge information as pre-translation content, where the translation-related content includes at least one of document content, subject name, and keyword information; when the associated device is used for pre-translation, the pre-translation is performed with the help of a large language model deployed in the associated device to provide auxiliary support for the first device to perform instant translation; the preset data set and the first preset model are constructed by the first device based on the acquired historical voice information, the preset data set includes pronunciation habit information corresponding to at least one language, and the first preset model is trained based on the preset data set; The first information corresponds to a source language, and the second information corresponds to a target language.

7. The processing method according to claim 6, characterized in that Also include at least one of the following: The first information includes at least one of voice, text, image and gesture; The output mode of the second information includes at least one of audio broadcast, text display, image display and video display.

8. An intelligent terminal, characterized in that: include: A memory and a processor, wherein a processing program is stored in the memory, and when the processing program is executed by the processor, the steps of the processing method according to any one of claims 1 to 7 are implemented.

9. A storage medium, characterized in that: The storage medium stores a computer program, which, when executed by a processor, implements the steps of the processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Voice conversion method and device, voice conversion system and storage medium

    CN116935851A

  • Machine simultaneous interpretation system and method, test method and device and related equipment

    CN116935853A