Information processing method and electronic equipment
By collecting and classifying audio information in real time, automatically determining the language and translating it, the problem of cumbersome operation of existing translation software is solved, and efficient and convenient cross-language communication is achieved.
Patent Information
- Application Number
- CN202510238287.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-02
- Publication Date
- 2025-06-10
AI Technical Summary
The existing translation software is complicated to operate, and users need to frequently enter the corresponding language of the language to be translated, which leads to inconvenience in operation.
By collecting audio information in real time, classifying and determining at least two categories of languages, and automatically converting them to the target language, realizing the translation process without the need for users to manually select languages.
It improves the efficiency and convenience of translation, reduces the steps of user operations, and realizes real-time and automatic cross-language communication.
Smart Images

Figure CN120124645A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information technology, and in particular, to an information processing method and an electronic device. Background Art
[0002] As an important tool software, translation software can break down language barriers in real time, enabling smooth communication and rapid and accurate information transmission whether in cross-border business negotiations, international academic exchanges, overseas travel, or online multicultural social interactions. It has become a powerful assistant for people to cross the language gap and expand the boundaries of communication.
[0003] Since translation software usually requires users to frequently input the language to be translated and the target language for translation, the operation is cumbersome. Therefore, there is an urgent need for a processing method to solve the above problems. Summary of the Invention
[0004] The first aspect of this application provides an information processing method, including:
[0005] Classify the real-time collected audio information to obtain at least two categories corresponding to the audio information;
[0006] Determine the source language corresponding to each category in the at least two categories;
[0007] Based on the source language corresponding to each category, convert the audio information into information corresponding to a target language, where the target language is different from the source language.
[0008] In a possible implementation, the at least two categories at least include a first category;
[0009] Determining the source language corresponding to each category in the at least two categories includes:
[0010] Query whether there is a first correspondence relationship, where the first correspondence relationship represents the source language corresponding to the first category;
[0011] If the first correspondence relationship exists, determine that the source language corresponding to the first category is the first language according to the first language in the first correspondence relationship;
[0012] If the first correspondence relationship does not exist, determine a language from the candidate languages as the source language corresponding to the first category, where the candidate languages are languages for which the correspondence relationship has not been determined.
[0013] In a possible implementation, the at least two categories further include a second category;
[0014] The determining a language from the candidate languages as the source language corresponding to the first category includes:
[0015] If there is a second correspondence relationship, determine the remaining first language among the two preset languages as the source language corresponding to the first category; the second correspondence relationship indicates that the source language corresponding to the second category is the second language, the second language is different from the first language, and the determination time of the second category is earlier than the determination time of the first category.
[0016] If there is no such second correspondence relationship, according to the preset language priority, select the first language among the two preset languages as the source language corresponding to the first category, and the second language as the target language corresponding to the first category, and the priority of the first language is higher than that of the second language; establish a first correspondence relationship, and the first correspondence relationship indicates that the source language corresponding to the first category is the first language.
[0017] In a possible implementation, it further includes:
[0018] If there is no such first correspondence relationship and both of the two preset languages have corresponding relationships, according to the preset language priority, select the first language among the two preset languages as the source language corresponding to the first category, and the priority of the first language is higher than that of the second language;
[0019] Establish a first correspondence relationship, and the first correspondence relationship indicates that the source language corresponding to the first category is the first language;
[0020] Based on the first language as the source language corresponding to the first category, convert the audio information into information corresponding to the target language, and the second language is the target language corresponding to the first category.
[0021] In a possible implementation, after converting the audio information into information corresponding to the target language based on the first language as the source language corresponding to the first category, it further includes:
[0022] Analyze the information corresponding to the second language;
[0023] If the information corresponding to the second language does not meet the language rules of the second language, determine that the source language corresponding to the first category is the second language;
[0024] Update the source language corresponding to the first category in the first correspondence relationship according to the second language;
[0025] Based on the second language as the source language corresponding to the first category, convert the audio information into information corresponding to the first language, and the first language is the target language corresponding to the first category.
[0026] In a possible implementation, after selecting the first language among the two preset languages as the source language corresponding to the first category according to the preset language priority, it further includes:
[0027] Analyze the audio information based on the first language.
[0028] If the audio information does not meet the language rules of the first language, determine that the source language corresponding to the first category is the second language; establish a third correspondence relationship, where the third correspondence relationship represents that the source language corresponding to the first category is the second language; based on the second language as the source language corresponding to the first category, convert the audio information into information corresponding to the first language, with the first language as the target language corresponding to the first category.
[0029] If the audio information meets the language rules of the first language, execute the step of establishing the first correspondence relationship.
[0030] In a possible implementation, the classifying the real-time collected audio information to obtain at least two categories corresponding to the audio information includes at least one of the following:
[0031] In response to a first instruction, analyze the collected audio information to obtain at least two voiceprint features; based on each voiceprint feature, determine at least two categories corresponding to the audio information.
[0032] In response to a first instruction, analyze the collected audio information to obtain at least two timbre features; based on each timbre feature, determine at least two categories corresponding to the audio information.
[0033] In response to a first instruction, analyze the collected audio information to obtain at least two sound source directions; based on each sound source direction, determine at least two categories corresponding to the audio information.
[0034] In a possible implementation, it further includes:
[0035] Display the information corresponding to the first language converted from the audio information in the first area of the display screen.
[0036] Display the information corresponding to the second language converted from the audio information in the second area of the display screen, where the first area and the second area are different.
[0037] In a possible implementation, before classifying the real-time collected audio information to obtain at least two categories corresponding to the audio information, it further includes:
[0038] Receive a first instruction, where the first instruction is used to indicate the conversion of the language of the audio information.
[0039] Collect audio information in real time.
[0040] Correspondingly, it further includes:
[0041] Based on receiving the second instruction, collecting audio information is prohibited, and the second instruction is used to indicate to stop converting the language of the audio information;
[0042] Release the correspondence between various categories and languages.
[0043] The second aspect of the present application provides an electronic device, including:
[0044] An audio collection device for real-time collection of audio information;
[0045] A processing device for classifying the real-time collected audio information to obtain at least two categories corresponding to the audio information; determining the source language corresponding to each category in the at least two categories; based on the source language corresponding to each category, converting the audio information into information corresponding to the target language, and the target language is different from the source language.
[0046] The third aspect of the present application provides a computer program product, including computer-readable instructions, when the computer-readable instructions run on an electronic device, enabling the electronic device to implement the information processing method of the first aspect or any implementation manner of the first aspect.
[0047] The fourth aspect of the present application provides an electronic device, including at least one processor and a memory connected to the processor, wherein:
[0048] The memory is used to store a computer program;
[0049] The processor is used to execute the computer program so that the electronic device can implement the information processing method of the first aspect or any implementation manner of the first aspect.
[0050] The fifth aspect of the present application provides a computer storage medium, and the storage medium carries one or more computer programs, when the one or more computer programs are executed by an electronic device, enabling the electronic device to implement the information processing method of the first aspect or any implementation manner of the first aspect. Description of the Drawings
[0051] Combined with the drawings and referring to the following specific implementation manners, the above and other features, advantages and aspects of the embodiments of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale.
[0052] Figure 1 It is a schematic diagram of the current translation software interface;
[0053] Figure 2 It is a flowchart of an information processing method provided by an embodiment of the present application;
[0054] Figure 3 It is a schematic diagram of an interface provided by an embodiment of the present application;
[0055] Figure 4 It is a schematic diagram of a translation scenario provided by an embodiment of the present application;
[0056] Figure 5 It is a schematic diagram of a process for determining the source language corresponding to each of at least two categories provided by an embodiment of the present application;
[0057] Figure 6 It is a schematic diagram of a process for determining the source language corresponding to each of at least two categories provided by an embodiment of the present application;
[0058] Figure 7 It is a schematic diagram of a process for determining the corresponding source language for the first category provided by an embodiment of the present application;
[0059] Figure 8 It is another schematic diagram of a process for determining the corresponding source language for the first category provided by an embodiment of the present application;
[0060] Figure 9 It is yet another schematic diagram of a process for determining the corresponding source language for the first category provided by an embodiment of the present application;
[0061] Figure 10 It is still another schematic diagram of a process for determining the corresponding source language for the first category provided by an embodiment of the present application;
[0062] Figure 11 It is a schematic diagram of a process for displaying and outputting the converted information provided by an embodiment of the present application;
[0063] Figure 12 It is a schematic diagram of a display screen provided by an embodiment of the present application Figure 1 ;
[0064] Figure 13 It is a schematic diagram of a display screen provided by an embodiment of the present application Figure 2 ;
[0065] Figure 14 It is a schematic diagram of a display screen provided by an embodiment of the present application Figure 3 ;
[0066] Figure 15 It is a schematic diagram of a display screen provided by an embodiment of the present application Figure 4 ;
[0067] Figure 16 It is a schematic diagram of a display screen provided by an embodiment of the present application Figure 5 ;
[0068] Figure 17It is another schematic flowchart for displaying and outputting the converted information provided by the embodiments of the present application;
[0069] Figure 18 It is a schematic diagram of the display screen provided by the embodiments of the present application Figure 6 ;
[0070] Figure 19 It is a schematic diagram of the display screen provided by the embodiments of the present application Figure 7 ;
[0071] Figure 20 It is a schematic structural diagram of an electronic device provided by the embodiments of the present application
[0072] Figure 21 It is another schematic structural diagram of an electronic device provided by the embodiments of the present application. Detailed implementation manners
[0073] The embodiments of the present application will be described below with reference to the accompanying drawings in the embodiments of the present application. The terms used in the implementation part of the present application are only used to explain the specific embodiments of the present application, rather than to limit the present application.
[0074] The embodiments of the present application will be described below with reference to the accompanying drawings. Those of ordinary skill in the art will know that with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0075] The terms "first", "second", etc. in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing when describing objects with the same attributes in the embodiments of the present application. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device including a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.
[0076] An information processing method provided by the embodiments of the present application can be applied to an electronic device with audio acquisition ability and audio processing ability, and the specific structural form of the electronic device is not limited in the present application.
[0077] Before the present application, the large models adopted by translation software often need to input the language to be recognized, and corresponding options will be set for the input language. The user selects the corresponding language option and inputs the audio of the corresponding language, and the large model uses the input language to translate the audio into text.
[0078] Figure 1It is a schematic diagram of the interface of the current translation software. This interface contains two buttons in Chinese and English. When user A who uses Chinese wants to perform translation, press the Chinese button. After collecting the voice of user A, the large model translates the voice into text according to the Chinese language, and then the text can be translated into English; the translation for English is similar. During the communication between the two parties, a click operation needs to be performed for each translated segment, and the operation is cumbersome.
[0079] This application proposes an information processing method to solve the above problems. Refer to Figure 2 , Figure 2 It is a schematic flowchart of an information processing method provided by an embodiment of this application. As shown in Figure 2 , an information processing method provided by an embodiment of this application may include steps 201 to 203, and the following will describe these steps in detail.
[0080] 201. Classify the real-time collected audio information to obtain at least two categories corresponding to the audio information;
[0081] Among them, in the translation scenario, the on-site audio is collected in real time to obtain audio information, and the audio information contains the audio of both parties of the language to be translated.
[0082] Among them, in the translation scenario, two or more speakers speak, and the multiple speakers speak in turn to communicate, and the multiple speakers use two or more languages.
[0083] Among them, the electronic device collects the sound in the scene in real time and performs real-time classification.
[0084] As an example, at the scene where the speaker is speaking, the on-site audio is collected in real time to perform real-time classification on the collected audio information and determine its corresponding categories.
[0085] Among them, the category can be determined according to the voiceprint feature, timbre feature, sound source direction, etc. in the audio information.
[0086] In a possible implementation, it may be that after receiving the first instruction, in response to the first instruction, the real-time collected audio information is classified to obtain at least two corresponding classifications.
[0087] Among them, the first instruction is an instruction used to indicate the start of translation. The first instruction can be generated based on the triggering of a specific button set on the electronic device, or can be generated based on detecting that the audio information contains content indicating the start of translation, or can be triggered based on a detected specified gesture. This application does not make a limitation here.
[0088] Figure 3This is a schematic diagram of an interface provided by an embodiment of the present application. There is a button "Start Translation" on this interface. When this button is triggered, a first instruction is generated, and the electronic device starts to process the subsequently real-time collected audio information. The collected audio information can be the continuous audio information of both parties in a conversation. In a complete conversation translation scenario, triggering this button once generates a first instruction, and it is not necessary for each speaker to trigger separately before speaking.
[0089] In a possible implementation, in response to the first instruction, the collected audio information is analyzed to obtain at least two voiceprint features; based on each voiceprint feature, at least two categories corresponding to the audio information are determined.
[0090] Among them, each person's voiceprint is unique and serves as an effective basis for distinguishing different people.
[0091] Therefore, voiceprints can be used to distinguish different speakers.
[0092] Among them, by analyzing the audio information, the voiceprint features contained therein are obtained, and based on the voiceprint features therein, it can be determined whether any two segments of audio are from the same speaker.
[0093] Correspondingly, based on each voiceprint feature, the category corresponding to the audio information can be determined. Generally, one voiceprint feature corresponds to one speaker, and one speaker can correspond to one category of audio information.
[0094] For example, in the case where it is determined according to the voiceprint features that there are two speakers, it can be determined that the audio information corresponds to two categories; in the case where it is determined according to the voiceprint features that there are three speakers, it can be determined that the audio information corresponds to three categories, and so on.
[0095] In a possible implementation, in response to the first instruction, the collected audio information is analyzed to obtain at least two timbre features; based on each timbre feature, at least two categories corresponding to the audio information are determined.
[0096] Among them, timbre is the characteristic of sound. Different people have different lengths, widths, thicknesses of the vocal cords and different shapes and sizes of the resonance cavities (such as the oral cavity, nasal cavity, etc.), resulting in different combinations of overtones when speaking, thus forming different timbres.
[0097] Therefore, timbre can be used to distinguish different speakers.
[0098] Among them, by analyzing the audio information, the timbre features contained therein are obtained, and based on the timbre features therein, it can be determined whether any two segments of audio are from the same speaker.
[0099] Correspondingly, according to each timbre feature, the category corresponding to the audio information can be determined. Generally, one timbre feature corresponds to one speaker, and one speaker can correspond to one category of audio information.
[0100] In a possible implementation, in response to the first instruction, the collected audio information is analyzed to obtain at least two sound source directions; according to each sound source direction, at least two categories corresponding to the audio information are determined.
[0101] Wherein, the sound source direction is the direction of the speaker relative to the electronic device.
[0102] Wherein, in the translation scenario, each speaker in the conversation is located in a different direction relative to the electronic device. The direction of the sound source in the audio information can be analyzed and determined, and different sound source directions correspond to different categories.
[0103] Wherein, compared with the speaker at a farther distance, the speaker at a closer distance may be a more familiar person. Generally, the speaker at a closer distance speaks in the same language. Therefore, the category can be determined by the sound source direction, and different sound source directions correspond to different categories.
[0104] Figure 4 FIG. is a schematic diagram of the translation scenario provided in the embodiment of the present application. There are scenarios with two and three people in this schematic diagram. Among them, (a) is a schematic diagram of a scenario with two speakers, including user A, user B, and the electronic device. User A and user B are on both sides of the electronic device respectively, and user A and user B are each used as a sound source. (b) is a schematic diagram of a scenario with three speakers, including user A, user B, user C, and the electronic device. User A is on one side of the electronic device and is used as a sound source; user B and user C are on the other side of the electronic device, and the directions of user B and user C relative to the electronic device are close. User A and user B are used as a sound source. (c) is another schematic diagram of a scenario with three speakers, including user A, user B, user C, and the electronic device. User A is on one side of the electronic device, and user B and user C are on the other side of the electronic device. User B and user C are farther away, and user A, user B, and user C are each used as a sound source. The relative position relationship between each sound source and the electronic device is marked with a dotted line in the figure.
[0105] Among them, for the category obtained by classifying the audio information, the category can be recorded by assigning a speaker identifier.
[0106] For example, through voiceprint feature analysis, the voiceprint feature of a speaker is determined. If the speaker is the first speaker, label it as speaker 1, and so on, and label each speaker accordingly.
[0107] In a possible implementation, the audio information collected in real time can be processed in slices to determine the category of the audio segment, and the audio segments that are temporally continuous and belong to the same category are combined in chronological order to obtain a piece of audio information belonging to the same category, and subsequent steps are performed based on this piece of audio information.
[0108] 202. Determine the source language corresponding to each category among the at least two categories.
[0109] Among them, after determining the at least two categories corresponding to the audio information, determine the source language corresponding to each category.
[0110] Among them, the source language is the language used in a part of the audio information corresponding to each category, and the target language corresponding to the source language is the specific language that one wants to convert to.
[0111] Among them, one category corresponds to one language. Correspondingly, for each category corresponding to the audio information, determine the corresponding source language.
[0112] Among them, the source language corresponding to the category can be determined from among a preset plurality of languages.
[0113] Among them, in a continuous conversation, different methods are used to determine the source language for the first obtained classification and the non-first obtained classification in the audio information.
[0114] Among them, in a continuous conversation, for the audio information collected in real time, when a category is first analyzed, the source language corresponding to the category can be determined from among the preset languages, and the corresponding relationship between the source language and the category is established. When the category is analyzed again during this conversation, the source language can be determined using this corresponding relationship.
[0115] Among them, the language corresponding to each category can be any one of a preset plurality of languages.
[0116] Among them, the preset languages can be pre-set by the user of the electronic device, and when entering the translation scenario, the corresponding language information can be obtained from the settings.
[0117] As an example, the configuration information of the electronic device can be detected, such as the language environment configuration, to obtain the language used by the electronic device, which is generally the language used by the owner of the electronic device; it is also possible to select the language that will be used in this scenario in the electronic device interface when entering the translation scenario, and use this language and the language used by the electronic device as the preset plurality of languages.
[0118] 103. Based on the source language corresponding to each category, convert the audio information into information corresponding to the target language, where the target language is different from the source language.
[0119] Among them, after determining the source language corresponding to each category, the audio corresponding to the corresponding part in the audio information can be converted into the information corresponding to the target language.
[0120] The information corresponding to the target language can be text information or audio information, which is not limited in this application.
[0121] Among them, the target language is the language opposite to the source language.
[0122] Among them, the selection process of the target language can follow the determination process of the source language. When the determined categories are two categories, the languages corresponding to the two categories are respectively used as the source language of one's own side and the target language of the other side.
[0123] As an example, for the source language and the target language, any one of Chinese, English, German, etc. can be adopted, and the target language and the source language adopt different languages.
[0124] As an example, the audio information corresponds to two categories. It is determined that the source language corresponding to category 1 is Chinese, and the source language corresponding to category 2 is English. Correspondingly, when converting the audio information part corresponding to category 1, English is used as the target language; when converting the audio information part corresponding to category 2, Chinese is used as the target language.
[0125] In an application scenario, user A and user B communicate across languages through an electronic device. For example, user A uses Chinese and user B uses English. After user A presses the "Start Translation" button on the interface, user A starts to speak to user B in Chinese. After the electronic device collects and classifies the audio information of user A in real time, it translates the Chinese spoken by user A into English and outputs it. After user B receives the English, user B speaks to user A in English. The electronic device collects and classifies the audio of user B in real time, determines that user B uses English, and translates the audio of user B into Chinese and outputs it. In this process, only the first speaking user (user A in this scenario) needs to click the start button before speaking, and no further clicks are required during the whole process. The electronic device continuously collects, classifies the language, and translates the audio during the communication process between user A and user B in real time.
[0126] It should be noted that this application does not make specific restrictions on the output form of the translated information. For example, the translated information can be output by displaying it in the display area of the electronic device, or it can be output in the form of audio playback, and the output in the form of audio playback does not affect the real-time audio collection described above (including in S201).
[0127] In a possible implementation, a language model is used to process the audio information. The language model can translate the input audio information using a determined source language and automatically translate it into text or speech in the target language.
[0128] Among them, in the case where the converted information is text information, the language model includes an STT (Speech to Text) model and a translation model. The STT model converts the audio information into text format. The identifier of the source language and the audio information are input into the STT model, and the STT model converts the audio information into text format according to the source language. The translation model converts the information in text format into information in text format in the target language.
[0129] Among them, in the case where the converted information is speech information, the language model includes an STT model, a translation model, and a TTS (Text To Speech) model. The STT model converts the audio information into text format. The identifier of the source language and the audio information are input into the STT model, and the STT model converts the audio information into text format according to the source language. The translation model converts the information in text format into information in text format in the target language. The TTS model generates speech based on the target language, converts the information in text format in the target language into speech, and outputs it.
[0130] In a possible implementation, a language detection model can also be set. The real-time collected audio information is input into the language detection model. The language detection model detects the language used in the current audio information, and this language can be used as the source language to convert the corresponding audio information into information corresponding to the target language.
[0131] In this embodiment, the real-time collected audio information is classified to obtain at least two categories corresponding to the audio information; the source language corresponding to each category in the at least two categories is determined; based on the source language corresponding to each category, the audio information is converted into information corresponding to the target language, and the target language is different from the source language. By determining at least two categories corresponding to the real-time collected audio information and determining the source language corresponding to each category, it is realized to automatically analyze and determine the source language of the real-time collected audio information without manually annotating the source language for each collected audio information. It can be realized to automatically convert the real-time collected audio information into information corresponding to the target language according to the source language. During the communication process between the two speakers, there is no need to interrupt multiple times to select the language used, which improves the communication efficiency.
[0132] In a possible implementation, the at least two categories at least include a first category, and the source language corresponding to it can be determined for the first category. The specific implementation process is as follows Figure 5 shown.
[0133] Figure 5 FIG. 0 is a schematic flowchart for determining the source language corresponding to each category among at least two categories provided by an embodiment of the present application, which may include steps 501 to 503. The following will describe these steps in detail.
[0134] 501. Query whether there is a first correspondence relationship, which characterizes the source language corresponding to the first category.
[0135] Among them, after analyzing the category of the real-time collected audio information to determine the category, query whether there is a correspondence relationship of this category.
[0136] In this embodiment, one category is taken as an example for illustration, and the source language of other categories is also analyzed and determined by this method.
[0137] Among them, after determining the category of the audio information, query whether there is a first correspondence relationship related to this category, and this first correspondence relationship can characterize the source language corresponding to the first category.
[0138] As an example, this category is the speaker 1, and the first correspondence relationship characterizes the source language corresponding to this speaker 1, and this source language can be any language.
[0139] Among them, if the query finds the first correspondence relationship, execute the subsequent step 502; otherwise, execute step 503.
[0140] 502. If there is the first correspondence relationship, determine that the source language corresponding to the first category is the first language according to the first language in the first correspondence relationship.
[0141] Among them, if there is the first correspondence relationship, it means that in the previous audio information, the source language corresponding to the first category has been determined.
[0142] Among them, the first correspondence relationship records that the first category corresponds to the first language, then, using this first correspondence relationship, determine that the source language corresponding to the first category is this first language.
[0143] As an example, this first category is the speaker 1, and the first correspondence relationship characterizes that the corresponding language of this speaker 1 is Chinese, then determine that the source language corresponding to the first category is Chinese.
[0144] As an example, this first category is the speaker 2, and the first correspondence relationship characterizes that the corresponding language of this speaker 2 is English, then determine that the source language corresponding to the first category is English.
[0145] 503. If there is no such first correspondence relationship, determine a language from the candidate languages as the source language corresponding to the first category, and this candidate language is a language for which the correspondence relationship has not been determined.
[0146] Among them, if the first corresponding relationship does not exist, it indicates that the source language corresponding to the first category has not been determined in the previous audio information.
[0147] Correspondingly, in the candidate languages, determine one as the source language corresponding to the first category.
[0148] Among them, multiple languages are preset, and one can be selected from the preset multiple languages as the source language of the first category.
[0149] In a possible implementation, the candidate languages can be the remaining languages after being selected by other categories, or all of the multiple languages that have never been selected.
[0150] As an example, the candidate language can be unique, and the unique candidate language can be directly used as the source language corresponding to the first category; if the candidate language is not unique, one can be selected from the multiple candidate languages according to the set selection conditions as the source language corresponding to the first category.
[0151] Among them, after determining the source language corresponding to the first category, establish the first corresponding relationship between the language used by the source language and the first category for subsequent use when converting the audio information corresponding to the first category.
[0152] In an application scenario, user A and user B communicate in different languages. During the communication process, user A speaks first, user B answers, user B makes a supplementary answer, user A speaks again, and user B answers again. The audio information during the communication between the two users is collected in real time, and the collected audio information in real time is classified and the source language is determined.
[0153] First, the audio information when user A speaks is collected and analyzed. The voiceprint feature of user A who makes the sound is obtained through the analysis, and this voiceprint feature is labeled as speaker 1. The first corresponding relationship of this speaker 1 is not found. Among the candidate languages (Chinese and English), Chinese is determined as the source language of this speaker 1. Using Chinese as the source language, the audio of this speaker 1 is converted into English, and it is recorded that the source language of this speaker 1 is Chinese. Then, the audio information when user B answers is collected and analyzed. The voiceprint feature of user B who makes the sound is obtained. The voiceprint feature of this user B is different from that of user A, and it is determined that this audio information is not from the same speaker as before. This voiceprint feature is labeled as speaker 2, and this user B makes a sound for the first time. The first corresponding relationship of this speaker 2 is not found. Among the candidate languages (Chinese and English), English is determined as the source language of this speaker 2. Using English as the source language, the audio of this speaker 2 is converted into Chinese, and it is recorded that the source language of this second speaker is English. Then, the audio information of user B's supplementary answer is collected in real time and analyzed. The voiceprint feature of user B who makes the sound is obtained through the analysis. The voiceprint feature of this user B is consistent with that of speaker 2, and it is determined that this audio information corresponds to speaker 2. The first corresponding relationship of this speaker 2 is found, and using the first corresponding relationship of this speaker 2, English is determined as the source language of this speaker 2. Using English as the source language, the audio of this speaker 2 is converted into Chinese. Subsequently, the audio information of user A speaking again and user B answering again are collected in real time. By adopting the above process, the source languages corresponding to the audio information of this speaking again and the answer again can be determined, and then they are respectively converted into the corresponding target languages for output.
[0154] In this embodiment, after determining the category corresponding to the audio information, it can be queried whether there is a first corresponding relationship representing the source language corresponding to the first category. If there is a first corresponding relationship, according to the first language in the first corresponding relationship, it is determined that the source language corresponding to the first category is the first language; if there is no first corresponding relationship, a language is determined from the candidate languages as the source language corresponding to the first category. The candidate languages are the languages for which the corresponding relationships have not been determined. The source language determined historically can be used as the source language of the current audio information, without having to select from multiple preset languages each time, reducing the data processing volume.
[0155] In a possible implementation, the at least two categories further include a second category, and the source language corresponding to the first category can be determined according to whether the second category has been selected. The specific implementation process is as follows Figure 6 as shown.
[0156] Figure 6 It is a flowchart of the process for determining the source language corresponding to each of the at least two categories provided by an embodiment of the present application, which may include steps 601 to 603. The following will describe these steps in detail.
[0157] 601. If there is a second correspondence relationship, determine the remaining first language among the two preset languages as the source language corresponding to this first category; this second correspondence relationship indicates that the source language corresponding to this second category is the second language, which is different from the first language, and the determination time of this second category is earlier than the determination time of this first category;
[0158] Among them, if there is no first correspondence relationship corresponding to this first category, it indicates that the source language has not been determined for this first category, and select one from the candidate languages as the source language corresponding to this first category.
[0159] Among them, in the case of having two preset languages, if the source language corresponding to the second language has been determined to be the second language, the remaining first language is used as the candidate language. When selecting the corresponding source language for the first category, this first language can be directly used as the source language corresponding to the first category.
[0160] As an example, the two preset languages are Chinese and English respectively. The categories are represented by speakers. The determination time of speaker 1 is earlier than that of speaker 2. The second category is speaker 1, and this first category is speaker 2. It has been determined that the corresponding source language of speaker 1 is Chinese, and a second correspondence relationship between English and speaker 1 has been established. Then, when determining the corresponding source language for speaker 2, directly use the remaining Chinese as the corresponding source language.
[0161] 602. If there is no such second correspondence relationship, according to the preset language priority, select the first language among the two preset languages as the source language corresponding to this first category, and the second language as the target language corresponding to this first category. The priority of this first language is higher than that of the second language;
[0162] Among them, if there is no second correspondence relationship, it indicates that both of the two preset languages have been selected. It is necessary to select one from these two preset languages as the source language corresponding to this first category, and the remaining one as the target language corresponding to this first category.
[0163] Among them, for the preset language priority, preferably select the language with a higher priority as the source language corresponding to this first category.
[0164] Among them, one of the two preset languages can be determined from the language environment of the electronic device, and the other is determined according to the user's selection. Determining the language from the language environment of the electronic device can be obtained from the language configuration information of the electronic device.
[0165] In general, in the scenario of using an electronic device for translation, the owner of the electronic device speaks first. Therefore, the priority of the language configured for the electronic device can be set higher than that of other languages. When classifying audio information and determining the corresponding source language for the classification, the source language corresponding to the category of the audio information of the first speaker (usually the owner) is made consistent with the language of the electronic device.
[0166] 603. Establish a first correspondence relationship, which represents that the source language corresponding to the first category is the first language.
[0167] Among them, after determining that the source language corresponding to the first category is the first language, establish the first correspondence relationship between the first language and the first category.
[0168] Correspondingly, after subsequently collecting audio information in real time again, if it is determined that the audio information is of the first category, directly using this first correspondence relationship, it can be determined that the source language corresponding to the first category is the first language.
[0169] In this embodiment, if there is a second correspondence relationship indicating that the source language corresponding to the second category is the second language, determine the remaining first language among the two preset languages as the source language corresponding to the first category; the second language is different from the first language, and the determination time of the second category is earlier than that of the first category; if there is no such second correspondence relationship, according to the preset language priority, select the first language among the two preset languages as the source language corresponding to the first category, and the second language as the target language corresponding to the first category, and the priority of the first language is higher than that of the second language; establish a first correspondence relationship, which represents that the source language corresponding to the first category is the first language, realizing the selection of one of the candidate languages as the source language corresponding to the second category, enabling the automatic conversion of the corresponding audio information of the first determined category into the target language using this source language, realizing the process of automatic translation, without the need for manual selection of the source language operation, and improving the translation efficiency.
[0170] In a possible implementation, if corresponding relationships already exist for both of the two preset languages, when determining a new category again, one of the two preset languages can be selected as the source language corresponding to the first category. The specific implementation process is as follows Figure 7 as shown.
[0171] Figure 7 is a flowchart diagram for determining the corresponding source language for the first category provided by an embodiment of the present application, which may include steps 701 to 703, and the following will describe these steps in detail.
[0172] 701. If the first corresponding relationship does not exist and corresponding relationships exist for both of the two preset languages, select the first language as the source language corresponding to the first category from the two preset languages according to the preset language priority, where the priority of the first language is higher than that of the second language.
[0173] Among them, if the first corresponding relationship corresponding to the first category is not queried and corresponding relationships already exist for both of the two preset languages, it indicates that in addition to the two determined categories, a third category of audio information has appeared, and the audio information of this third category may use any one of the two preset languages.
[0174] As an example, the two preset languages are Chinese and English, the source language corresponding to speaker 1 is English, and the source language corresponding to speaker 2 is Chinese. After speaker 3 is determined, when selecting the source language corresponding to speaker 3, according to the preset language priority, select one of English and Chinese as the source language corresponding to speaker 3. If the priority of Chinese is higher, select Chinese as the source language corresponding to speaker 3. If the priority of English is higher, select English as the source language corresponding to the speaker.
[0175] In a possible implementation, when speaker 4 appears later, to determine the source language corresponding to speaker 4, the Figure 7 same process is also adopted.
[0176] 702. Establish a first corresponding relationship, where this first corresponding relationship indicates that the source language corresponding to the first category is the first language.
[0177] Among them, after determining that the source language corresponding to the first category is the first language, establish the first corresponding relationship between the first language and the first category.
[0178] Correspondingly, after subsequently collecting audio information in real time again, if it is determined that the audio information is of the first category, directly using this first corresponding relationship, it can be determined that the source language corresponding to the first category is the first language.
[0179] 703. Based on the first language as the source language corresponding to the first category, convert the audio information into information corresponding to the target language, where the second language is the target language corresponding to the first category.
[0180] Among them, after determining that the source language corresponding to the first category is the first language, use the second language as the target language corresponding to the first category, and convert the part of the audio information corresponding to the first category into information corresponding to the second language.
[0181] In an application scenario, User A, User B, and User C communicate in different languages. During the communication process, User A speaks first, User B replies, User C makes a supplementary reply, User A speaks again, and User B replies again. The audio information during the communication between two users is collected in real time, and the collected audio information in real time is classified and the source language is determined. Among them, for the process of the source languages corresponding to the audio information of User A and User B respectively, refer to the corresponding explanations in the foregoing embodiments, which will not be elaborated here. It has been determined that Chinese is the source language of Speaker 1, and English is determined as the source language of Speaker 2. After collecting a piece of audio information (the audio information when User C is speaking), the audio information is analyzed, and the voiceprint feature of the speaking User C is obtained through the analysis. This voiceprint feature is inconsistent with the voiceprint features of Speaker 1 and Speaker 2. This voiceprint feature is marked as Speaker 3, and no first corresponding relationship of this Speaker 3 is found. Moreover, there are corresponding relationships for all candidate languages. Since the priority of Chinese is higher than that of English, Chinese is determined as the source language of this Speaker 3 among the candidate languages (Chinese and English), and the audio of this Speaker 3 is converted into English using Chinese as the source language, and it is recorded that the source language of this Speaker 3 is Chinese.
[0182] In this embodiment, if there is no such first corresponding relationship and there are corresponding relationships for both of the two preset languages, it means that after the corresponding relationships of the two preset languages have been determined and then a new category is obtained, according to the priority of the preset languages, the first language is selected as the source language corresponding to this first category, and the priority of the first language is higher than the priority of the second language; establish the first corresponding relationship, where the first corresponding relationship indicates that the source language corresponding to the first category is the first language; based on the first language as the source language corresponding to the first category, convert the audio information into the information corresponding to the target language, and the second language is the target language corresponding to the first category, which realizes the automatic translation of the audio information of the third or subsequent users participating in the communication process for the case of two preset languages, without the need for manual selection of the source language operation, improving the translation efficiency.
[0183] In a possible implementation, it is also necessary to determine whether the source language selected for the first category is accurate, and if it is not accurate, the source language corresponding to the first category needs to be adjusted. The specific implementation process is as follows Figure 8 as shown.
[0184] Figure 8 It is another schematic flowchart for determining the corresponding source language for the first category provided by the embodiments of the present application, which may include steps 801 to 804. The following will describe these steps in detail.
[0185] 801. Analyze the information corresponding to the second language;
[0186] Among them, after converting the audio information into information corresponding to the target language (the second language) based on the first language as the source language corresponding to the first category, the information corresponding to the second language is used for analysis to reverse-infer whether the source language determined for the first category is accurate.
[0187] Among them, according to the language rules of the second language, the information of the second language obtained by conversion is analyzed.
[0188] As an example, the target language is English. According to the language rules of English, the information obtained by conversion is analyzed to obtain an analysis result.
[0189] As an example, the target language is Chinese. According to the language rules of Chinese, the information obtained by conversion is analyzed to obtain an analysis result.
[0190] Among them, if the information corresponding to the second language meets the language rules of the second language, it can be inferred that the source language determined for the first category is accurate as the first language, and then the first corresponding relationship is continued to be used to determine the source language of the audio information.
[0191] 802. If the information corresponding to the second language does not meet the language rules of the second language, it is determined that the source language corresponding to the first category is the second language;
[0192] Among them, if the information corresponding to the second language obtained by conversion does not meet the language rules of the second language, it can be inferred that the source language determined for the first category is incorrect as the first language, and the source language corresponding to the first category is adjusted to the second language.
[0193] As an example, for the first category determined by classifying the audio information, it is determined that the source language corresponding to the first category is Chinese. Correspondingly, the target language is English. According to the language rules of English, the information obtained by conversion is analyzed. If the information obtained by conversion does not meet the language rules of English, it is determined that the source language corresponding to the first category is English. Correspondingly, the target language is Chinese.
[0194] As an example, for the first category determined by classifying the audio information, it is determined that the source language corresponding to the first category is English. Correspondingly, the target language is German. According to the language rules of German, the information obtained by conversion is analyzed. If the information obtained by conversion does not meet the language rules of German, it is determined that the source language corresponding to the first category is German. Correspondingly, the target language is English.
[0195] 803. Update the source language corresponding to the first category in the first corresponding relationship according to the second language;
[0196] Among them, after determining the second language as the source language of the first category, the first corresponding relationship of the first category is updated.
[0197] Among them, the source language corresponding to the first category in the updated first correspondence relationship is the second language.
[0198] 804. Based on the second language as the source language corresponding to the first category, convert the audio information into information corresponding to the first language, and the first language is used as the target language corresponding to the first category.
[0199] Among them, according to the updated source language (the second language), re-convert the corresponding audio information and convert it into information corresponding to the first language.
[0200] In a possible implementation, use the second language as the source language, use the first language as the target language, and convert the audio information from the second language to the first language to obtain information corresponding to the first language.
[0201] In an application scenario, user A makes a sound. Based on the collected audio information of user A's voice, first determine that the source language of the audio information is English, and convert the audio information into Chinese, but the converted information does not meet the language rules of Chinese; modify the source language of the audio information to Chinese, and convert the audio information into English to obtain the converted English information.
[0202] In this embodiment, after converting the audio information into information corresponding to the second language based on the first language as the source language corresponding to the first category, analyze the information corresponding to the second language; if the information corresponding to the second language does not meet the language rules of the second language, determine that the source language corresponding to the first category is the second language; update the source language corresponding to the first category in the first correspondence relationship according to the second language; based on the second language as the source language corresponding to the first category, convert the audio information into information corresponding to the first language, and the first language is used as the target language corresponding to the first category. After initially determining that the source language is the first language and converting the audio information into the second language as the target language, if the information converted into the second language does not meet the language rules of the second language, adjust the source language of the audio information to the second language, and convert the audio information into the first language as the target language, which realizes the judgment of whether the determined source language and target language are accurate and improves the translation accuracy.
[0203] In a possible implementation, it is also necessary to determine whether the selected source language for the first type is accurate, and if it is not accurate, adjust the source language corresponding to the first type. The specific implementation process is as follows Figure 9 as shown.
[0204] Figure 9 It is another process schematic diagram for determining the corresponding source language for the first category provided by the embodiments of the present application, which may include steps 901 to 904, and the following will describe these steps in detail.
[0205] 901. Analyze the audio information based on the first language.
[0206] Among them, after selecting the first language as the source language corresponding to the first category from two preset languages according to the preset language priority, it is necessary to analyze whether the source language determined for the audio information is the first language accurately.
[0207] Among them, analyze the audio information according to the language rules of the first language.
[0208] As an example, if the source language is English, analyze the converted information according to the language rules of English to obtain the analysis result.
[0209] As an example, if the source language is Chinese, analyze the converted information according to the language rules of Chinese to obtain the analysis result.
[0210] Among them, if the information corresponding to the first language meets the language rules of the first language, it can be determined that the source language determined for the first category as the first language is accurate, and then continue to use the first corresponding relationship to determine the source language of the audio information.
[0211] Among them, if the audio information does not meet the language rules of the first language, execute steps 902 - 904; if the audio information meets the language rules of the first language, execute step 905.
[0212] 902. If the audio information does not meet the language rules of the first language, determine that the source language corresponding to the first category is the second language.
[0213] Among them, if the audio information does not meet the language rules of the first language, it can be determined that the previously determined source language for the first category as the first language is incorrect. Therefore, adjust the source language corresponding to the first category to the second language.
[0214] As an example, for the first category determined by classifying the audio information, if it is determined that the source language corresponding to the first category is Chinese and the target language is English, correspondingly, analyze the audio information according to the language rules of Chinese. If the audio information does not meet the language rules of English, determine that the source language corresponding to the first category is English, and correspondingly, the target language is Chinese.
[0215] As an example, for the first category determined by classifying the audio information, if it is determined that the source language corresponding to the first category is English and the target language is German, correspondingly, analyze the audio information according to the language rules of English. If the audio information does not meet the language rules of English, determine that the source language corresponding to the first category is German, and correspondingly, the target language is English.
[0216] In a possible implementation, if there are only two preset languages, it can be directly determined that the other language is the source language of the audio information of the first category.
[0217] 903. Establish a third correspondence relationship, which represents that the source language corresponding to the first category is the second language;
[0218] Among them, after re-determining the second language as the source language of the first category, a third correspondence relationship is established for the first category.
[0219] Among them, since it is the first time to determine the source language corresponding to the first category and there is no corresponding relationship for the first category before, a corresponding relationship (the third correspondence relationship) is established for the first category.
[0220] Among them, the source language corresponding to the first category in the third correspondence relationship is the second language.
[0221] 904. Based on the second language as the source language corresponding to the first category, convert the audio information into information corresponding to the first language, and the first language is the target language corresponding to the first category;
[0222] Among them, according to the updated source language (the second language), the corresponding audio information is re-converted and converted into information corresponding to the first language.
[0223] In a possible implementation, the second language is used as the source language, the first language is used as the target language, and the audio information is converted from the second language to the first language to obtain information corresponding to the first language.
[0224] In an application scenario, user A makes a sound. Based on the collected audio information of user A's sound, it is first determined that the source language of the audio information is Chinese, and the corresponding target language is English. The audio information is converted into Chinese, but the converted information does not meet the language rules of Chinese; the source language of the audio information is modified to Chinese, the corresponding target language is English, and the audio information is converted into English to obtain the converted English information.
[0225] 905. If the audio information meets the language rules of the first language, execute the step of establishing the first correspondence relationship.
[0226] Among them, if the audio information meets the language rules of the first language, it can be determined that the source language determined for the first category as the first language is correct, and a first correspondence relationship is established for the first category. The source language corresponding to the first category in the first correspondence relationship is the second language.
[0227] The subsequent execution steps can refer to Figure 7 Steps 702 - 703 in, which will not be elaborated in this embodiment.
[0228] In a possible implementation, it is also possible to convert the audio information into text information according to the first language determined as the source language, and determine whether the text information in the first language satisfies the language rules of the first language, so as to determine whether the determination of the source language is accurate. When the determination of the source language is accurate, continue to translate the text information in the first language with the first language as the source language and the second language as the target language to obtain the information in the second language. When the determination of the source language is incorrect, determine the second language as the source language and the first language as the target language.
[0229] In this embodiment, after selecting the first language as the source language corresponding to the first category according to the preset language priority among the two preset languages, the corresponding audio information is analyzed based on the first language; if the audio information does not satisfy the language rules of the first language, it is determined that the source language corresponding to the first category is the second language; a third correspondence is established, and the third correspondence indicates that the source language corresponding to the first category is the second language; based on the second language as the source language corresponding to the first category, the audio information is converted into the information corresponding to the first language, and the first language is the target language corresponding to the first category; if the audio information satisfies the language rules of the first language, perform subsequent steps such as establishing the first correspondence. Before converting the audio information according to the initially determined source language being the first language, first determine whether the audio information satisfies the language rules of the first language. If not, adjust the source language of the audio information to the second language, and convert the audio information into the first language as the target language, realizing the judgment of whether the determined source language and target language are accurate and improving the translation accuracy.
[0230] In a possible implementation, if there are more than two preset languages and a new first type appears, the language with a lower priority can be selected according to the priority as the source language corresponding to the first category. The specific implementation process is as follows Figure 10 as shown.
[0231] Figure 10 FIG. is another schematic flowchart for determining the corresponding source language for the first category provided by the embodiments of the present application, which may include steps 1001 to 1003, and the following will describe these steps in detail.
[0232] 1001. If the first correspondence does not exist and both the first language and the second language have corresponding relationships, according to the preset language priority, select the third language as the source language corresponding to the first category, and the priority of the third language is lower than the priority of the first language and the priority of the second language;
[0233] Among them, the preset language priority can be three or more preset ones.
[0234] In a general translation scenario, there may be two or three languages. When there are three preset languages, they can be selected among the three languages according to the priority and in the order of the determination category time. The one with the earliest determination category time has the highest selection priority, the second in the determination category time order selects the second-level priority language, and the third in the determination category time order selects the third-level priority language, and so on.
[0235] Among them, if no first corresponding relationship corresponding to the first category is found, and corresponding relationships already exist in the first two preset languages with higher priority rankings, indicating that in addition to the two determined categories, a third category of audio information appears, then according to the priority of the preset languages, the third language with a lower level is selected as the source language of the audio information of this third category.
[0236] As an example, the three preset languages are Chinese, English, and German. The source language corresponding to speaker 1 is English, and the source language corresponding to speaker 2 is Chinese. After determining speaker 3, when selecting the source language corresponding to speaker 3, according to the priority of the preset languages, German with the third priority ranking is selected as the source language corresponding to speaker 3.
[0237] In a possible implementation, when there are multiple preset languages, if the source language corresponding to a certain category is determined to be the first language, the second language is defaulted to be the target language of this category. When it is subsequently determined that the source language corresponding to the first category is the third language, the audio information in the first language can be converted into the information in the second language and the third language, and the information in the second language and the information in the third language are respectively output.
[0238] 1002. Establish a fourth corresponding relationship, which represents that the source language corresponding to the first category is the third language;
[0239] Among them, after determining the third language as the source language of the first category, a fourth corresponding relationship is established for the first category.
[0240] Among them, since this is the first time to determine the source language corresponding to the first category and there was no corresponding relationship for the first category before, therefore, a corresponding relationship (the fourth corresponding relationship) is established for the first category.
[0241] Among them, the source language corresponding to the first category in the fourth corresponding relationship is the third language.
[0242] 1003. Based on the third language as the source language corresponding to the first category, convert the audio information into the information corresponding to the target language. The first language is used as the target language corresponding to the first category, and the priority of the first language is higher than the priority of the second language and the third language.
[0243] Among them, according to the source language (the third language) corresponding to the first category, the corresponding audio information is converted into the information corresponding to the first language.
[0244] Generally, since the language with the highest priority (the first language) is the language used by the owner of the electronic device, it is preferable to translate the information in other languages into the language used by the holder of the electronic device.
[0245] Of course, after determining that the source language corresponding to the first category is the third language, the third language is used to analyze the audio information to determine whether the selection is accurate. If it is not accurate, according to the priority, one of the first language and the second language is selected as the source language for this first category.
[0246] Because in the communication process, it is easier to join the communication in the same language. Therefore, the probability of selecting the same language as the holder of the electronic device is greater. Generally, when reselecting the language for the first category, the language with the highest priority in the priority order is selected.
[0247] In a possible implementation, the third language is used as the source language, the first language is used as the target language, and the audio information is converted from the third language to the first language to obtain the information corresponding to the first language.
[0248] In an application scenario, user A makes a sound. Based on the collected audio information of user A's voice, it is first determined that the source language of this audio information is Chinese, and the corresponding target language is English. The audio information is converted into Chinese, but the converted information does not meet the language rules of Chinese; the source language of the audio information is modified to Chinese, the corresponding target language is English, and the audio information is converted into English to obtain the converted English information.
[0249] As an example, the three preset languages are Chinese, English, and German. The source language corresponding to speaker 1 is English, and the source language corresponding to speaker 2 is Chinese. After determining that the source language corresponding to speaker 3 is German, the German audio information is converted into Chinese.
[0250] In this embodiment, if there is no first corresponding relationship and there are corresponding relationships for both the first language and the second language, it indicates that after determining the corresponding relationships of two preset languages and then obtaining a new category, according to the preset language priority, the third language ranked after the first language and the second language in the priority order is used as the source language corresponding to the first category; a fourth corresponding relationship is established, and the fourth corresponding relationship indicates that the source language corresponding to the first category is the third language; based on the third language being the source language corresponding to the first category, the audio information is converted into information corresponding to the target language, the first language is the target language corresponding to the first category, and the priority of the first language is higher than the priorities of the second language and the third language. This realizes automatic translation of the audio information of the third user or the user ranked further back who joins the communication process in the case of multiple preset languages, without the need for manual selection of the source language, improving the translation efficiency.
[0251] In a possible implementation, it is also necessary to display and output the converted information. The specific implementation process is as follows Figure 11 as shown.
[0252] Figure 11 FIG. is a schematic flowchart of displaying and outputting the converted information provided by an embodiment of the present application, which may include steps 1101 to 1102, and the following will describe these steps in detail.
[0253] 1101. Convert the audio information into information corresponding to the first language and display it in the first area of the display screen;
[0254] 1102. Convert the audio information into information corresponding to the second language and display it in the second area of the display screen, where the first area and the second area are different.
[0255] Among them, in the case of two languages, the information corresponding to different languages can be displayed in different areas of the display screen.
[0256] Among them, the first area and the second area of the display screen can be preset areas or areas divided according to the sound source direction.
[0257] In a possible implementation, the translation interface on the display screen is divided into two areas, one area is on the side close to the bottom of the screen, and one area is close to the top of the display screen.
[0258] Since the converted information is for the other party to view, therefore, the converted information is in a display manner that matches the perspective of the party using the target language. The display contents of the two areas face different directions, and the directions are both with the middle of the display screen as the top and the edge of the display screen as the bottom.
[0259] In a possible implementation, the translation interface on the display screen is divided into two regions, the left and the right. The first region displays information corresponding to the first language, and the second region displays information corresponding to the second language.
[0260] In a possible implementation, during the communication corresponding to the displayed information, there may be two speakers using different languages, or there may be multiple speakers using two different languages.
[0261] Figure 12 It is a schematic diagram of the display screen provided by the embodiments of the present application Figure 1 , in the application scenario of this schematic diagram, the first language is Chinese and the second language is English. There are two speakers in this schematic diagram, the first language is Chinese and the second language is English. Speaker 1 corresponds to the source language of Chinese, and Speaker 2 corresponds to the source language of English. The display screen is divided into two regions, the left and the right. The first region 1201 is used to display Chinese, and the second region 1202 is used to display English. In this application scenario, after the previous simple communication, users A and B cannot communicate directly, but it has been determined that Speaker 1 (user A) corresponds to the source language of Chinese, and Speaker 2 (user B) corresponds to the source language of English. User A asks in Chinese "Hello! I want to go to the supermarket. Could you tell me how to get there?" After the audio information is collected in real time, it is translated into "Hello! I want to go to the supermarket. Could you tell me how to get there?" The translated text is displayed in the second region 1202. Subsequently, user B (Speaker 2) replies in English "Go straight ahead and turn left at the next intersection. The supermarket is on your left." After the audio information of this user B is collected in real time, it is translated into "Go straight ahead and turn left at the next intersection. The supermarket is on your left.", and the translated text is displayed in the first region 1201.
[0262] Figure 13 It is a schematic diagram of the display screen provided by the embodiments of the present application Figure 2, in the application scenario of this schematic diagram, the first language is Chinese and the second language is English. There are two speakers in this schematic diagram. The first language of the first speaker is Chinese and the second language of the second speaker is English. The source language corresponding to Speaker 1 is Chinese, and the source language corresponding to Speaker 2 is English. The display screen is divided into upper and lower regions. The first region 1301 is used to display Chinese, and the second region 1302 is used to display English. In this application scenario, after a previous simple conversation, users A and B cannot communicate directly, but it has been determined that the source language corresponding to Speaker 1 (user A) is Chinese, and the source language corresponding to Speaker 2 (user B) is English. User A asks in Chinese, "Hello! I want to go to the supermarket. Could you tell me how to get there?" After the audio information is collected in real time, it is translated into "Hello! I want to go to the supermarket. Could you tell me how to get there?" The translated text is displayed in the second region 1302. Subsequently, user B replies in English, "Go straight ahead and turn left at the next intersection. The supermarket is on your left." After the audio information of user B is collected in real time, it is translated into "Go straight ahead and turn left at the next intersection. The supermarket is on your left." The translated text is displayed in the first region 1301.
[0263] Figure 14 Schematic diagram of the display screen provided by the embodiment of the present application Figure 3, in the application scenario of this schematic diagram, the first language is Chinese and the second language is English. There are two speakers in this schematic diagram. The first language of the first speaker is Chinese and the second language of the second speaker is English. The source language corresponding to Speaker 1 is Chinese, and the source language corresponding to Speaker 2 is English. The display screen is divided into upper and lower regions. The first region 1401 is used to display Chinese, and the second region 1402 is used to display English. The display contents of the two regions face different directions, and the top of both directions is the middle of the display screen, and the bottom is the edge of the display screen. In this application scenario, after a previous simple conversation, users A and B cannot communicate directly, but it has been determined that the source language corresponding to Speaker 1 (user A) is Chinese, and the source language corresponding to Speaker 2 (user B) is English. User A asks in Chinese, "Hello! I want to go to the supermarket. Could you tell me how to get there?" After the audio information is collected in real time, it is translated into "Hello! I want to go to the supermarket. Could you tell me how to get there?" The translated text is displayed in the second region 1402. Subsequently, user B replies in English, "Go straight ahead and turn left at the next intersection. The supermarket is on your left." After the audio information of user B is collected in real time, it is translated into "Go straight ahead and turn left at the next intersection. The supermarket is on your left." The translated text is displayed in the first region 1401.
[0264] Figure 15 This is a schematic diagram of the display screen provided by an embodiment of the present application Figure 4, in the application scenario of this schematic diagram, the first language is Chinese and the second language is English. There are three speakers in this schematic diagram. The first language is Chinese and the second language is English. The source language corresponding to Speaker 1 is Chinese, and the source languages corresponding to Speaker 2 and Speaker 3 are English. The display screen is divided into two left and right areas. The first area 1501 is used to display Chinese, and the second area 1502 is used to display English. In this application scenario, after a previous simple conversation, users ABC cannot communicate directly, but it has been determined that the source language corresponding to Speaker 1 (User A) is Chinese, and the source languages corresponding to Speaker 2 (User B) and Speaker 3 (User C) are English. User A asks in Chinese "Hello! I want to go to the supermarket. Could you tell me how to get there?" After collecting this audio information in real time, it is translated into "Hello! I want to go to the supermarket. Could you tell me how to get there?" The translated text is displayed in the second area 1502. Subsequently, User B replies in English "Go straight ahead and turn left at the next intersection. The supermarket is on your left." After collecting the audio information of this User B in real time, it is translated into "Go straight ahead and turn left at the next intersection. The supermarket is on your left.", and the translated text is displayed in the first area 1501. Subsequently, User C adds in English "The supermarket isn't open yet. It will open at 9 o'clock." After collecting this audio information in real time, it is translated into "The supermarket isn't open yet. It will open at 9 o'clock.", and the translated text is displayed in the first area 1501.
[0265] Figure 16 This is a schematic diagram of the display screen provided by an embodiment of the present application Figure 5, in the application scenario of this schematic diagram, the first language is Chinese and the second language is English. There are three speakers in this schematic diagram, with the first language being Chinese and the second language being English. The source languages corresponding to Speaker 1 and Speaker 3 are Chinese, and the source language corresponding to Speaker 2 is English. The display screen is divided into upper and lower regions. The first region 1601 is used to display Chinese, and the second region 1602 is used to display English. In this application scenario, after some simple exchanges before, users ABC cannot communicate directly, but it has been determined that the source languages corresponding to Speaker 1 (User A) and Speaker 3 (User C) are Chinese, and the source language corresponding to Speaker 2 (User B) is English. User A asks in Chinese "Hello! I want to go to the supermarket. Could you tell me how to get there?" After the audio information is collected in real time, it is translated into "Hello! I want to go to the supermarket. Could you tell me how to get there?" The translated text is displayed in the second region 1502. Subsequently, User B replies in English "Go straight ahead and turn left at the next intersection. The supermarket is on your left." After the audio information of User B is collected in real time, it is translated into "Go straight ahead for one intersection and turn left. The supermarket is on your left." The translated text is displayed in the first region 1601. Subsequently, User C asks in Chinese "What time does this supermarket close at night?" After the audio information is collected in real time, it is translated into "What time does this supermarket close at night?" The translated text is displayed in the first region 1602.
[0266] In this embodiment, the information corresponding to the first language converted from the audio information is displayed in the first region of the display screen; the information corresponding to the second language converted from the audio information is displayed in the second region of the display screen, and the first region and the second region are different. The audio information corresponding to each language obtained by audio conversion is displayed in different regions of the display screen, so that communication participants can quickly view the text content in the corresponding language and improve the communication efficiency.
[0267] In a possible implementation, it is also necessary to display and output the converted information. The specific implementation process is as follows Figure 17 as shown.
[0268] Figure 17 It is another process schematic diagram for displaying and outputting the converted information provided by the embodiment of the present application, which may include steps 1701 to 1703. The following will describe these steps in detail.
[0269] 1701. Convert the audio information into information corresponding to the first language and display it in the first area of the display screen;
[0270] 1702. Convert the audio information into information corresponding to the second language and display it in the second area of the display screen;
[0271] 1703. Convert the audio information into information corresponding to the third language and display it in the third area of the display screen; the first area, the second area, and the third area are different from each other.
[0272] Among them, in the case of three languages, the information corresponding to different languages can be displayed in different areas of the display screen.
[0273] Among them, the first area, the second area, and the third area of the display screen can be preset areas or areas divided according to the sound source direction.
[0274] In a possible implementation, divide the translation interface on the display screen into three areas, one area is on the side close to the bottom of the screen, one area is close to the top of the display screen, and one area is in the middle area of the screen.
[0275] Since the converted information is for the other party to view, therefore, the converted information is in a display mode that matches the perspective of the party using the target language. The second area and the third area are arranged in the left-right direction, and the display contents of the second area and the third area face the same direction. The display contents of the first area and the second area face different directions, and the directions are both with the middle of the display screen as the top and the edge of the display screen as the bottom.
[0276] In a possible implementation, divide the translation interface on the display screen into two left and right areas, the first area displays the information corresponding to the first language, and the second area displays the information corresponding to the second language.
[0277] In a possible implementation, during the communication corresponding to the displayed information, there may be three speakers using different languages, or multiple speakers using three different languages.
[0278] Figure 18 It is a schematic diagram of the display screen provided by the embodiments of the present application Figure 6, in the application scenario of this schematic diagram, there are three speakers in the schematic diagram. The first language is Chinese, the second language is English, and the third language is Russian. Speaker 1 corresponds to the source language of Chinese, Speaker 2 corresponds to the source language of English, and Speaker 3 corresponds to the source language of Russian. The display screen is divided into three regions: upper, middle, and lower. The first region 1801 is used to display Chinese, the second region 1802 is used to display English, and the third region 1803 is used to display Russian. In this application scenario, after previous simple communication, users A, B, and C cannot communicate directly, but it has been determined that Speaker 1 corresponds to the source language of Chinese, Speaker 2 corresponds to the source language of English, and Speaker 3 corresponds to the source language of Russian. User A (Speaker 1) asks in Chinese, "Hello! I want to go to the supermarket. Could you tell me how to get there?" After the audio information is collected in real time, it is translated into English as "Hello! I want to go to the supermarket. Could you tell me how to get there?" The translated text is displayed in the second region 1802, and it is also translated into Russian as "Привет! Я хочу пойти в супермаркет. Пожалуйста, скажите, как туда попасть?" The translated text is displayed in the third region 1802. Subsequently, User B (Speaker 2) replies in English, "Go straight ahead and turn left at the next intersection. The supermarket is on your left." After the audio information of User B is collected in real time, it is translated into "Go straight ahead at an intersection and turn left. The supermarket is on your left." The translated text is displayed in the first region 1801. User C (Speaker 3) says in Russian, "Супермаркет еще не открыт. Он откроется в 9 часов." After the audio information is collected in real time, it is translated into "The supermarket is not open yet. It will open at 9 o'clock." The translated text is displayed in the first region 1801.
[0279] Figure 19 This is a schematic diagram of the display screen provided by the embodiments of the present application Figure 7, in the application scenario of this schematic diagram, there are three speakers in the schematic diagram. The first language is Chinese, the second language is English, and the third language is Russian. Speaker 1 corresponds to the source language Chinese, Speaker 2 corresponds to the source language English, and Speaker 3 corresponds to the source language Russian. The display screen is divided into three upper and lower regions, and the upper half region is divided into left and right regions. The first region 1901 is used to display Chinese, the second region 1902 is used to display English, and the third region 1903 is used to display Russian. In this application scenario, after a previous simple conversation, users A and B cannot communicate directly, but it has been determined that Speaker 1 corresponds to the source language Chinese and Speaker 2 corresponds to the source language English. User A (Speaker 1) asks in Chinese "Hello! I want to go to the supermarket. Could you tell me how to get there?" After the audio information is collected in real time, it is translated into English "Hello! I want to go to the supermarket. Could you tell me how to get there?" The translated text is displayed in the second region 1902. Subsequently, User B (Speaker 2) replies in English "Go straight ahead and turn left at the next intersection. The supermarket is on your left." After the audio information of this User B is collected in real time, it is translated into "Go straight ahead at the next intersection and turn left. The supermarket is on your left." The translated text is displayed in the first region 1901. User C (Speaker 3) says in Russian "Супермаркет еще не открыт. Он откроется в 9 часов." After the audio information is collected in real time, it is translated into "The supermarket is not open yet. It will open at 9 o'clock." The translated text is displayed in the first region 1901. User A (Speaker 1) replies in Chinese "Thank you!" After the audio information is collected in real time, it is translated into English "Thank you!" The translated text is displayed in the second region 1902, and is also translated into Russian "Спасибо вам!" The translated text is displayed in the third region 1903.
[0280] In this embodiment, the information corresponding to the first language converted from the audio information is displayed in the first region of the display screen; the information corresponding to the second language converted from the audio information is displayed in the second region of the display screen; the information corresponding to the third language converted from the audio information is displayed in the third region of the display screen; the first region, the second region, and the third region are different from each other. The audio information corresponding to each language obtained by audio conversion is displayed in different regions of the display screen according to the language, so that communication participants can quickly view the text content in the corresponding language and improve the communication efficiency.
[0281] In a possible implementation, different instructions can be used to start or stop the current information processing flow.
[0282] Before classifying the real-time collected audio information to obtain at least two categories corresponding to the audio information, it further includes:
[0283] Step 1: Receive a first instruction, which is used to indicate the conversion of the language of the audio information.
[0284] Among them, the first instruction is an instruction used to indicate the start of translation. The first instruction can be generated based on the triggering of a specific button set on the electronic device, or can be generated based on detecting that the audio information contains content indicating the start of translation.
[0285] Among them, after receiving the first instruction, it is determined that the language of the audio information needs to be converted, and translation starts.
[0286] Step 2: Collect audio information in real time;
[0287] Among them, start collecting audio information in real time, and perform subsequent processes of determining the category of the audio information, determining the corresponding language, and converting it into the target language.
[0288] In a possible implementation, the electronic device collects audio information in real time without processing the audio information. After receiving the first instruction, it starts to perform the processes of determining the category of the audio information, determining the corresponding language, and converting it into the target language.
[0289] Among them, during a translation communication process, only the first instruction needs to be triggered once before the start of translation, and there is no need to trigger a translation operation every time a paragraph is finished speaking during the translation communication process.
[0290] Correspondingly, the information processing method further includes:
[0291] Step 3: Based on receiving a second instruction, prohibit the collection of audio information. The second instruction is used to indicate the stop of the conversion of the language of the audio information.
[0292] Among them, the second instruction is used to indicate the stop of the conversion of the language of the audio information and is an instruction indicating the end of translation.
[0293] Among them, the second instruction can be generated by triggering the same button as the first instruction
[0294] The second instruction can be generated based on the triggering of a specific button set on the electronic device, or can be generated based on detecting that the audio information contains content indicating the stop of translation.
[0295] Among them, upon receiving the second instruction, the acquisition of audio information is stopped, and the translation process is stopped.
[0296] Step 4: Release the correspondence between each category and each language.
[0297] Among them, after stopping the translation, release the correspondence established between each language and each category during this translation process. So that when translating again, it will not be affected by this translation.
[0298] In this embodiment, when starting the translation, it starts by receiving a first instruction for instructing the language conversion of audio information, and acquires audio information in real time. During the translation process, the real-time acquired audio information is translated in real time. Based on receiving a second instruction indicating to stop the language conversion of audio information, the acquisition of audio information is prohibited and no more translation is performed. Moreover, release the correspondence between each category and each language established during this translation process. In this process, there is no need to trigger a translation operation every time a paragraph is finished, improving the communication efficiency.
[0299] The above introduces an information processing method provided by an embodiment of the present application. Next, an electronic device for executing the above information processing method will be introduced.
[0300] Please refer to Figure 20 , Figure 20 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As Figure 20 shown, the electronic device 2000 includes:
[0301] An audio acquisition device 2001, configured to acquire audio information in real time;
[0302] A processing device 2002, configured to classify the real-time acquired audio information to obtain at least two categories corresponding to the audio information; determine the source language corresponding to each of the at least two categories; and based on the source language corresponding to each category, convert the audio information into information corresponding to a target language, where the target language is different from the source language.
[0303] The processing device includes:
[0304] A classification module, configured to classify the real-time acquired audio information to obtain at least two categories corresponding to the audio information;
[0305] A first determination module, configured to determine the source language corresponding to each of the at least two categories;
[0306] A conversion module, configured to convert the audio information into information corresponding to a target language based on the source language corresponding to each category, where the target language is different from the source language.
[0307] In a possible implementation, the at least two categories at least include a first category;
[0308] The first determination module includes:
[0309] A query unit for querying whether there is a first corresponding relationship, where the first corresponding relationship represents the source language corresponding to the first category;
[0310] A first determination unit for, if there is the first corresponding relationship, determining that the source language corresponding to the first category is the first language according to the first language in the first corresponding relationship;
[0311] A second determination unit for, if the first corresponding relationship does not exist, determining a language from the candidate languages as the source language corresponding to the first category, where the candidate languages are languages for which the corresponding relationship has not been determined.
[0312] In a possible implementation, the at least two categories further include a second category;
[0313] The second determination unit is specifically configured to:
[0314] If there is a second corresponding relationship, determining the remaining first language among the two preset languages as the source language corresponding to the first category; the second corresponding relationship represents that the source language corresponding to the second category is the second language, the second language is different from the first language, and the determination time of the second category is earlier than the determination time of the first category;
[0315] If the second corresponding relationship does not exist, selecting the first language from the two preset languages as the source language corresponding to the first category according to the preset language priority, the second language as the target language corresponding to the first category, and the priority of the first language is higher than that of the second language; establishing a first corresponding relationship, where the first corresponding relationship represents that the source language corresponding to the first category is the first language.
[0316] In a possible implementation, the second determination unit is further configured to:
[0317] If the first corresponding relationship does not exist and corresponding relationships exist for both of the two preset languages, selecting the first language from the two preset languages as the source language corresponding to the first category according to the preset language priority, and the priority of the first language is higher than that of the second language;
[0318] Establishing a first corresponding relationship, where the first corresponding relationship represents that the source language corresponding to the first category is the first language;
[0319] Based on the first language as the source language corresponding to the first category, converting the audio information into information corresponding to the target language, and the second language as the target language corresponding to the first category.
[0320] In a possible implementation, the second determination unit is further configured to:
[0321] Based on the first language as the source language corresponding to the first category, after converting the audio information into information corresponding to the target language, analyze the information corresponding to the second language;
[0322] If the information corresponding to the second language does not meet the language rules of the second language, determine that the source language corresponding to the first category is the second language;
[0323] Update the source language corresponding to the first category in the first correspondence relationship according to the second language;
[0324] Based on the second language as the source language corresponding to the first category, convert the audio information into information corresponding to the first language, and the first language is used as the target language corresponding to the first category.
[0325] In a possible implementation, the second determination unit is further configured to:
[0326] After selecting the first language as the source language corresponding to the first category according to the preset language priority, analyze the audio information based on the first language;
[0327] If the audio information does not meet the language rules of the first language, determine that the source language corresponding to the first category is the second language; establish a third correspondence relationship, where the third correspondence relationship represents that the source language corresponding to the first category is the second language; based on the second language as the source language corresponding to the first category, convert the audio information into information corresponding to the first language, and the first language is used as the target language corresponding to the first category;
[0328] If the audio information meets the language rules of the first language, execute the step of establishing the first correspondence relationship.
[0329] In a possible implementation, the classification module is configured to perform at least one of the following:
[0330] In response to a first instruction, analyze the collected audio information to obtain at least two voiceprint features; determine at least two categories corresponding to the audio information according to each voiceprint feature;
[0331] In response to a first instruction, analyze the collected audio information to obtain at least two timbre features; determine at least two categories corresponding to the audio information according to each timbre feature;
[0332] In response to a first instruction, analyze the collected audio information to obtain at least two sound source directions; determine at least two categories corresponding to the audio information according to each sound source direction.
[0333] In a possible implementation, it further includes:
[0334] A display module, configured to convert the audio information into information corresponding to a first language and display it in a first area of a display screen; convert the audio information into information corresponding to a second language and display it in a second area of the display screen, where the first area and the second area are different.
[0335] In a possible implementation, it further includes:
[0336] A receiving module, configured to receive a first instruction before classifying the real-time collected audio information to obtain at least two categories corresponding to the audio information, where the first instruction is used to indicate to convert the language of the audio information;
[0337] A triggering module, configured to trigger an audio collection device to collect audio information in real time according to the first instruction;
[0338] Correspondingly, it further includes:
[0339] A prohibiting module, configured to prohibit the collection device from collecting audio information based on receiving a second instruction, where the second instruction is used to indicate to stop converting the language of the audio information;
[0340] A releasing module, configured to release the corresponding relationship between each category and each language.
[0341] It should be noted that for the structural functions of the electronic device applying the information processing method provided in the embodiments of the present application, please refer to the explanations in the foregoing method embodiments, and no further details will be provided in this embodiment.
[0342] In this embodiment, an audio collection device collects audio information in real time; a processing device classifies the real-time collected audio information to obtain at least two categories corresponding to the audio information; determines the source language corresponding to each category among the at least two categories; based on the source language corresponding to each category, converts the audio information into information corresponding to a target language, where the target language is different from the source language. By determining at least two categories corresponding to the real-time collected audio information and determining the source language corresponding to each category, it realizes automatic analysis and determination of the source language of the real-time collected audio information, without the need for manual annotation of the source language for each collected audio information, and can automatically convert the real-time collected audio information into information corresponding to the target language according to the source language. During the communication process between the two speakers, there is no need to interrupt multiple times to select the language to be used, improving the communication efficiency.
[0343] In the embodiments of the present application, an electronic device is further provided. Refer to Figure 21 As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the information processing method in the embodiments of the present application. The electronic device in the embodiments of the present application may include, but is not limited to, fixed terminals such as mobile phones, laptop computers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), desktop computers, and the like.Figure 21 The illustrated electronic device is merely an example and should not impose any limitation on the functions and scope of use of the embodiments of the present application.
[0344] As Figure 21 shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 2101, which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 2102 or a program loaded from a storage device 2108 into a random access memory (RAM) 2103. When the electronic device is powered on, various programs and data required for the operation of the electronic device are also stored in the RAM 2103. The processing device 2101, the ROM 2102, and the RAM 2103 are connected to each other through a bus 2104. An input / output (I / O) interface 2105 is also connected to the bus 2104.
[0345] Generally, the following devices may be connected to the I / O interface 2105: an input device 2106 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 2107 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 2108 including, for example, a memory card, a hard disk, etc.; and a communication device 2109. The communication device 2109 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 21 an electronic device with various devices is shown, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0346] The embodiments of the present application also provide a computer program product including computer-readable instructions, which, when running on an electronic device, enable the electronic device to implement any information processing method provided by the embodiments of the present application.
[0347] The embodiments of the present application also provide a computer-readable storage medium carrying one or more computer programs, which, when executed by an electronic device, can enable the electronic device to implement any information processing method provided by the embodiments of the present application.
[0348] It should be further noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationships between the modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.
[0349] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits or dedicated circuits. However, for this application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods of various embodiments of this application.
[0350] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0351] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partly generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another, for example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a training device or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
Claims
1. An information processing method, comprising: Classifying the audio information collected in real time to obtain at least two categories corresponding to the audio information; Determining a source language corresponding to each of the at least two categories; Based on a source language corresponding to each category, the audio information is converted into information corresponding to a target language, where the target language is different from the source language.
2. The information processing method according to claim 1, wherein the at least two categories include at least a first category; Determining the source language corresponding to each of the at least two categories includes: querying whether a first corresponding relationship exists, where the first corresponding relationship represents a source language corresponding to the first category; If the first corresponding relationship exists, determining, according to the first language in the first corresponding relationship, that the source language corresponding to the first category is the first language; If the first corresponding relationship does not exist, a language is determined from among the languages to be selected as the source language corresponding to the first category, and the language to be selected is a language for which the corresponding relationship has not been determined.
3. The information processing method according to claim 2, wherein the at least two categories further include a second category; The step of determining a language from the candidate languages as the source language corresponding to the first category includes: If the second corresponding relationship exists, determining the remaining first language among the two preset languages as the source language corresponding to the first category; The second corresponding relationship indicates that the source language corresponding to the second category is a second language, the second language is different from the first language, and the determination time of the second category is earlier than the determination time of the first category; If the second corresponding relationship does not exist, according to the preset language priority, selecting a first language from two preset languages as the source language corresponding to the first category, and the second language as the target language corresponding to the first category, and the priority of the first language is higher than the priority of the second language; A first corresponding relationship is established, wherein the first corresponding relationship indicates that the source language corresponding to the first category is a first language.
4. The information processing method according to claim 3, further comprising: If the first corresponding relationship does not exist and both preset languages have corresponding relationships, selecting the first language from the two preset languages as the source language corresponding to the first category according to the preset language priorities, and the priority of the first language is higher than the priority of the second language; Establishing a first corresponding relationship, wherein the first corresponding relationship indicates that the source language corresponding to the first category is a first language; Based on the first language being used as a source language corresponding to the first category, the audio information is converted into information corresponding to a target language, and the second language is used as a target language corresponding to the first category.
5. The information processing method according to claim 4, after converting the audio information into information corresponding to a target language based on the first language as the source language corresponding to the first category, further comprising: analyzing the information corresponding to the second language; If the information corresponding to the second language does not satisfy the language rules of the second language, determining that the source language corresponding to the first category is the second language; updating the source language corresponding to the first category in the first corresponding relationship according to the second language; Based on the second language being the source language corresponding to the first category, the audio information is converted into information corresponding to the first language, and the first language is used as the target language corresponding to the first category.
6. The information processing method according to claim 4, after selecting the first language from two preset languages as the source language corresponding to the first category according to the preset language priority, further comprising: analyzing the audio information based on the first language; If the audio information does not satisfy the language rule of the first language, determining that the source language corresponding to the first category is the second language; Establishing a third correspondence relationship, wherein the third correspondence relationship indicates that the source language corresponding to the first category is the second language; based on the second language being the source language corresponding to the first category, converting the audio information into information corresponding to the first language, wherein the first language is the target language corresponding to the first category; If the audio information satisfies the language rule of the first language, the step of establishing the first corresponding relationship is performed.
7. The information processing method according to claim 1, wherein the step of classifying the audio information collected in real time to obtain at least two categories corresponding to the audio information comprises at least one of the following: In response to the first instruction, analyzing the collected audio information to obtain at least two voiceprint features; Determine at least two categories corresponding to the audio information according to each voiceprint feature; In response to the first instruction, analyzing the collected audio information to obtain at least two timbre features; Determine at least two categories corresponding to the audio information according to each timbre feature; In response to the first instruction, analyzing the collected audio information to obtain at least two sound source directions; At least two categories corresponding to the audio information are determined according to each sound source direction.
8. The information processing method according to any one of claims 1 to 6, further comprising: Converting the audio information into information corresponding to the first language and displaying it in a first area of the display screen; The audio information is converted into information corresponding to a second language and displayed in a second area of the display screen, wherein the first area and the second area are different.
9. The information processing method according to any one of claims 2 to 6, before classifying the audio information collected in real time to obtain at least two categories corresponding to the audio information, further comprising: receiving a first instruction, wherein the first instruction is used to instruct to convert the language of the audio information; Collect audio information in real time; Correspondingly, it also includes: Based on receiving a second instruction, prohibiting the collection of audio information, wherein the second instruction is used to instruct to stop converting the language of the audio information; Release the correspondence between each category and each language.
10. An electronic device, comprising: An audio acquisition device, used for acquiring audio information in real time; A processing device, used to classify the audio information collected in real time to obtain at least two categories corresponding to the audio information; Determine a source language corresponding to each of the at least two categories; and based on the source language corresponding to each category, convert the audio information into information corresponding to a target language, where the target language is different from the source language.