Interaction method and device, electronic equipment, medium and program product

By locally verbal recognition and generating voice reply messages on smart wearable devices, the problem of too long time spent in the interaction process is solved and the user experience is improved.

CN120067262APending Publication Date: 2025-05-30THE FOURTH PARADIGM BEIJING TECH CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510152239.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-11
Publication Date
2025-05-30

Smart Images

  • Figure CN120067262A_ABST
    Figure CN120067262A_ABST
Patent Text Reader

Abstract

The invention provides an interaction method and device, electronic equipment, a medium and a program product, the interaction method is applied to wearable equipment, and the method comprises the following steps: under the condition that an input message of a user is received, obtaining a text reply message corresponding to the input message, carrying out language identification on the text reply message, and sending the text reply message to the wearable equipment; obtaining language information corresponding to the text reply message; based on the text reply message and the language information, a voice reply message is generated, and the voice reply message is an audio message obtained after the text reply message is converted into a target language indicated by the language information; and outputting the voice reply message. According to the invention, the time consumption in the voice reply message generation process is reduced, and the user experience in the interaction process is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of wearable devices, and in particular to an interaction method, device, electronic device, medium, and program product. Background Art

[0002] With the improvement of people's living standards, smart wearable devices are favored by consumers due to their advantages such as small size and portability. However, due to the small size of smart wearable devices, their own computing power is generally weak. In the process of interacting with users, some smart wearable devices often need to call third-party services multiple times to generate interactive information for interacting with users. Each call to a third-party service requires waiting for a certain period of time, and usually it is necessary to wait for the third-party service called last time to return the result before the next call to the third-party service can be made, which makes the process relatively time-consuming, resulting in the phenomenon of slow response of smart wearable devices in the process of interacting with users, resulting in a poor user experience. It can be seen that in the related art, when smart wearable devices interact with users, it is easy for the process of generating interactive information to take too long. Summary of the Invention

[0003] The embodiments of the present disclosure provide an interaction method, apparatus, electronic device, medium, and program product to solve the problem in related technologies that, when a smart wearable device interacts with a user, the generation process of interaction information tends to take too long.

[0004] To solve the above problems, the present disclosure is implemented as follows:

[0005] In a first aspect, the present disclosure provides an interaction method, applied to a wearable device, the method comprising:

[0006] Upon receiving an input message from a user, obtaining a text reply message corresponding to the input message, and performing language recognition on the text reply message to obtain language information corresponding to the text reply message;

[0007] generating a voice reply message based on the text reply message and the language information, wherein the voice reply message is an audio message obtained by converting the text reply message into a target language indicated by the language information;

[0008] Output the voice reply message.

[0009] In a second aspect, the present disclosure provides an interactive device, applied to a wearable device, comprising:

[0010] an identification module configured to, upon receiving an input message from a user, obtain a text reply message corresponding to the input message, and perform language identification on the text reply message to obtain language information corresponding to the text reply message;

[0011] a generating module configured to generate a voice reply message based on the text reply message and the language information, wherein the voice reply message is an audio message obtained by converting the text reply message into a target language indicated by the language information;

[0012] An output module is used to output the voice reply message.

[0013] In a third aspect, the present disclosure provides an electronic device comprising: a memory, a processor, and a program stored in the memory and executable on the processor; the processor is configured to read the program in the memory to implement the steps of the interaction method described in the first aspect.

[0014] In a fourth aspect, the present disclosure provides a readable storage medium for storing a program, which, when executed by a processor, implements the steps in the interaction method described in the first aspect.

[0015] In a fifth aspect, the present disclosure provides a computer program product, which is stored in a storage medium and executed by at least one processor to implement the steps in the interaction method as described in the first aspect.

[0016] In the disclosed embodiment, during the process of interaction between the wearable device and the user, the language of the text reply message can be identified locally on the wearable device to obtain the language information corresponding to the text reply message without calling a third-party service for language identification. This helps reduce the number of times the third-party service is called during the interaction process, and further helps reduce the time spent in generating voice reply messages, thereby improving the user experience during the interaction process. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the description of the embodiments of the present disclosure. Obviously, the drawings described below are only some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0018] Figure 1 A flowchart of an interactive method provided in an embodiment of the present disclosure;

[0019] Figure 2 A schematic diagram of the structure of an interactive device provided in an embodiment of the present disclosure;

[0020] Figure 3 A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present disclosure in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0022] The terms "first", "second", etc. in the embodiments of the present disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or that are inherent to these processes, methods, products or devices. In addition, "and / or" is used in the present disclosure to represent at least one of the connected objects, such as A and / or B and / or C, which means seven situations including A alone, B alone, C alone, and both A and B exist, both B and C exist, both A and C exist, and all A, B and C exist.

[0023] See Figure 1 , Figure 1 A flowchart of an interactive method provided in an embodiment of the present disclosure, the method comprising the following steps:

[0024] Step 101: Upon receiving an input message from a user, obtaining a text reply message corresponding to the input message, and performing language recognition on the text reply message to obtain language information corresponding to the text reply message;

[0025] Step 102: Generate a voice reply message based on the text reply message and the language information, wherein the voice reply message is an audio message obtained by converting the text reply message into the target language indicated by the language information;

[0026] Step 103: Output the voice reply message.

[0027] The above-mentioned wearable devices can be various types of smart wearable devices, such as smart watches, smart bracelets, smart rings, smart glasses, VR devices, etc.

[0028] The above-mentioned input message can be a message input by the user through the wearable device in various interaction scenarios with the wearable device. For example, the input message can be query information input through the wearable device when the user needs to search for relevant information about a certain object. For another example, the input message can also be a voice message input by the user in a dialogue scenario with the wearable device. For another example, the input message can also be a message input to the wearable device when the user queries the answer to a question based on the wearable device. The input form of the input message can be various input forms supported by the wearable device, for example, voice input form or text input form.

[0029] Among them, the above-mentioned text reply message can be a reply message of the wearable device to the input message input by the user. Specifically, the above-mentioned text reply message corresponding to the input message can be a text reply message obtained through a related artificial intelligence (AI) service. For example, the input message can be sent to a related AI big model, and the AI ​​big model replies to the input message to obtain the text reply message. In this way, the generation speed of text reply messages can be increased without occupying too much computing power of the wearable device, which is conducive to improving the user experience.

[0030] For users of wearable devices, when interacting with AI on the wearable device, in some scenarios, they usually want to obtain the information output by the wearable device by listening to audio, rather than staring at the screen every time they interact, such as in a conversation scenario. In addition, if the screen size of the wearable device is small and it is worn in a fixed position, such as on the user's hand, if the user needs to query the text information on the wearable device, the user needs to raise his hand to view it, which will result in a poor user interaction experience. Based on this, in the embodiment of the present disclosure, after obtaining the text reply message, the text reply message is further converted into a voice reply message, and the voice reply message is output via audio to improve the user's interaction experience.

[0031] Because before converting a text reply message into a voice reply message, it is usually necessary to determine the language to be converted, wherein the language may be Chinese, English, Japanese, etc. Based on this, the wearable device may perform language recognition on the text reply message to obtain the language information corresponding to the text reply message. Specifically, the language information may be determined by the language of the text used in the text reply message. For example, when the text in the text reply message is in Chinese, the language information may be Chinese. For another example, when the text in the text reply message is in English, the language information may be English. For another example, when the text in the text reply message is in Japanese, the language information may be Japanese.

[0032] The above-mentioned generation of a voice reply message based on the text reply message and the language information can be achieved by converting the text reply message into audio of the target language indicated by the language information through various text-to-audio methods, and using the converted audio as the voice reply message.

[0033] The above-mentioned outputting of the voice reply message may specifically be broadcasting the voice reply message through the wearable device.

[0034] In some embodiments of the present disclosure, while outputting the voice reply message, the text reply message can also be displayed through the screen of the wearable device. Specifically, the text reply message can be displayed in a related interaction window, wherein the interaction window can display the interaction record between the user and the wearable device, for example, displaying the chat record between the user and the wearable device, so that the user can view historical interaction messages according to the interaction window.

[0035] In this embodiment, during the interaction between the wearable device and the user, the language of the text reply message can be identified locally to obtain the language information corresponding to the text reply message without calling a third-party service for language identification. This helps reduce the number of times the third-party service is called during the interaction process, and further helps reduce the time spent in generating voice reply messages, thereby improving the user experience during the interaction process.

[0036] Optionally, the acquiring a text reply message corresponding to the input message and performing language identification on the text reply message to obtain language information corresponding to the text reply message includes:

[0037] Sending the input message to the first server;

[0038] The text reply message sent by the first server is received, and language identification is performed on the text reply message to obtain language information corresponding to the text reply message.

[0039] After receiving the input message, the first server may convert the input message into a text format to obtain a text message, and generate a reply message corresponding to the text message based on the relevant services in the first server. For example, in some embodiments of the present disclosure, an AI large model may be deployed in the first server, and the first server may input the text message as a prompt word of the AI ​​large model into the AI ​​large model, and use the result output by the AI ​​large model based on the prompt word as the text reply message. Then, the text reply message is sent to the wearable device.

[0040] It is understandable that after receiving the text reply message sent by the first server, the wearable device can store the received text reply message in a cache, so that in the subsequent language recognition process, the text can be directly obtained from the cache for recognition to improve the efficiency of text recognition.

[0041] In this implementation, the process of obtaining the text reply message can be achieved by sending the input message to the first server and receiving the text reply message sent by the first server.

[0042] Optionally, the receiving the text reply message sent by the first server and performing language recognition on the text reply message to obtain language information corresponding to the text reply message includes:

[0043] In the process of receiving the text reply message sent by the first server, language recognition is performed on the text in the received text reply message to obtain language information corresponding to the text reply message.

[0044] Specifically, since it takes a certain amount of time for the wearable device to receive the text reply message, the wearable device can start two processes during the process of receiving the text reply message. One process is used to receive the text reply message, and the other process can call an algorithm to perform language recognition on the text in the received text reply message. In this way, when the text reply message is received, the language recognition process is basically completed, which is conducive to further reducing the time consumption of the voice reply message generation process.

[0045] It can be understood that the method for performing language recognition on the text in the received text reply message to obtain the language information corresponding to the text reply message can adopt the recognition method in the following embodiments, for example, performing text processing on the text in the received text reply message to obtain text processing information; based on the text processing information, determining the language information corresponding to the text in the text reply message.

[0046] In this embodiment, by performing language recognition on the text in the received text reply message during the process of receiving the text reply message sent by the first server, the language information corresponding to the text reply message is obtained, that is, the text reply message reception process and the language recognition process can be carried out simultaneously, which is conducive to further reducing the time consumption of the voice reply message generation process.

[0047] Optionally, performing language identification on the text reply message to obtain language information corresponding to the text reply message includes:

[0048] Performing text processing on the text reply message to obtain text processing information, wherein the text processing information includes text feature information corresponding to each character in the text reply message;

[0049] Based on the text feature information, the language information corresponding to the text reply message is determined.

[0050] The text processing information may include multiple text feature information corresponding to the multiple characters in the text reply message, that is, each character in the text reply message has a corresponding text feature information. The text feature information may include various attribute features of the characters, for example, the language family to which the characters belong, the glyph structure of the characters, and other features. Since the language to which the characters belong can usually be determined based on the characteristics of the characters, the language to which the corresponding characters belong can be determined based on the text feature information.

[0051] In this embodiment, text processing information is obtained by performing text processing on the text reply message, and the text processing information includes text feature information corresponding to each character in the text reply message. Therefore, the language of each character in the text reply message can be determined based on the text feature information, and then the language of the entire text reply message can be determined, thereby realizing the process of determining the language of the text reply message.

[0052] Optionally, determining the language information corresponding to the text reply message based on the text feature information includes:

[0053] Determining the language of each character in the text reply message based on the character feature information;

[0054] The language information is determined based on the language of each character in the text reply message, wherein the language information includes the target language, the target language being the language of the characters in the text reply message, corresponding to the largest number of characters in the text reply message, and the target language being the language of the voice reply message.

[0055] In related technologies, because the wearable device system cannot identify the encoding format of text information, it is also impossible to directly determine the language of the text in the text reply message. Based on this, in some embodiments of the present disclosure, feature information of text in various languages ​​can be pre-acquired and stored in the wearable device. In this way, the wearable device can match the stored feature information of text in various languages ​​with the text feature information of the text in the text reply message to determine the language of each text in the text reply message.

[0056] The target language is the language of the text reply message, that is, the language indicated by the language information.

[0057] All the characters in the above-mentioned text reply message may belong to the same language. In this case, the language to which all the characters belong can be determined as the target language, wherein the language of the text reply message is the target language. In addition, the language to which the characters in the text reply message belong may also include more than two languages. For example, in the process of communicating in Chinese, some English words may be mixed in with Chinese. Based on this, the language that appears the most times among the languages ​​to which the characters in the text reply message belong can be determined as the target language. For example, if the text reply message includes 10 characters, of which 6 characters belong to language 1, 3 characters belong to language 2, and 1 character belongs to language 3, at this time, since language 1 appears the most times, that is, the number of characters in the text reply message corresponding to language 1 is the largest, therefore, language 1 can be determined as the target language.

[0058] In this embodiment, the language to which each character in the text reply message belongs is determined based on the text feature information, and the language to which the characters in the text reply message belong, which corresponds to the largest number of characters in the text reply message, is determined as the target language to which the text reply message belongs. This is conducive to improving the accuracy of the determined target language.

[0059] Optionally, the text feature information includes a language identifier and glyph information, the language identifier being used to identify the language to which the corresponding text belongs, wherein different language identifiers correspond to different language families. Determining the language of each text in the text reply message based on the text feature information includes:

[0060] Determining at least two candidate languages ​​based on the language family identifier corresponding to a first character, wherein the first character is any character in the text reply message, and the at least two candidate languages ​​include: a language in the language family indicated by the language family identifier corresponding to the first character, and the glyph information corresponding to the first character includes k glyphs, where k is an integer greater than 0;

[0061] Determining, from the at least two candidate languages, a recognition language corresponding to each of the k glyphs to obtain k recognition languages, wherein the recognition languages ​​are languages ​​from the at least two candidate languages, and the k recognition languages ​​correspond one-to-one to the k glyphs;

[0062] The language that appears most frequently among the k recognized languages ​​is determined as the language to which the first character belongs.

[0063] It can be understood that each character in the above-mentioned text reply message can determine the language to which each character belongs according to the above-mentioned first character recognition method. The embodiment of this disclosure only takes the process of determining the language to which the first character belongs as an example to explain the process of determining the language to which each character in the text reply message belongs.

[0064] Related art has proposed classifying languages ​​worldwide according to their genealogical family. For example, the resulting language families include: 1) Sino-Tibetan, 2) Indo-European, 3) Altaic, 4) Semitic-Hamitic, 5) Uralic, 6) Caucasian (Iberian-Caucasian), 7) Austronesian (Malayo-Polynesian), 8) Austroasiatic, and 9) Dravidian. Each language family typically includes more than two languages. For example, the Sino-Tibetan family includes the following: Chinese, Tibeto-Burman, Miao-Yao, and Zhuang-Dong. Another example is the Indo-European family, which encompasses over 400 languages, including the Indo-Iranian, Germanic, Romance, Celtic, Slavic, Greek, Baltic, Albanian, and Armenian.

[0065] Based on this, in the process of determining the language of the first character, the language family of the first character can be determined in advance based on the language family identifier of the first character, and then the language of the first character can be determined within the language family of the first character based on the glyph information of the first character. In this process, determining the language family of the first character helps narrow the matching range for determining the language based on glyph information, thereby reducing the time required for the matching process and reducing the computing resources required for the matching process.

[0066] Specifically, the at least two candidate languages ​​may include all languages ​​in the language family to which the first character belongs.

[0067] The electronic device may store font information, which includes fonts in multiple languages. The font of each language may include font information of a large number of commonly used characters in the language.

[0068] The glyph information may include glyph names, glyph indexes, code points, and other information. The glyph names may be named using commonly used glyph naming methods, such as horizontal, vertical, left-falling, and right-falling strokes. The glyph index may include the index code of the corresponding glyph in the font library, and the code points may record indexing rules.

[0069] It can be understood that a character may include at least one glyph. For example, the character "十" includes two glyphs, namely "一" and "|". Another example is that the word "and" includes three glyphs, namely "a", "n", and "d". In some embodiments of the present disclosure, after obtaining the glyph information of all glyphs of the first character, at least two candidate font libraries corresponding to the at least two candidate languages can be determined from the font library information included in the wearable device. Then, the glyph information of each glyph is respectively matched in each candidate font library to determine the font library corresponding to each glyph, and the language of the font library to which the glyph corresponds is determined as the recognized language corresponding to the glyph. Among them, when a certain font library includes a specific glyph, it is determined that the font library corresponds to the specific glyph, and the language of the font library is determined as the recognized language corresponding to the glyph.

[0070] In this embodiment, in the process of determining the language to which the first character belongs, by first determining at least two candidate languages based on the language family identifier corresponding to the first character, it is beneficial to narrow the candidate matching range. Then, among the at least two candidate languages, the recognized language corresponding to each of the k glyphs is determined to obtain k recognized languages, and the language with the most occurrences among the k recognized languages is determined as the language to which the first character belongs. In this way, it is beneficial to improve the accuracy of the determined language to which the first character belongs.

[0071] Optionally, the wearable device stores font library information, and the font library information includes font libraries in multiple languages. The processing the text reply message to obtain text processing information includes:

[0072] Decompose the characters in the text reply message to obtain text decomposition information, where the text decomposition information includes: each character in the text reply message, and the glyph information corresponding to each character in the text reply message;

[0073] Based on the font library information, identify each character in the text reply message to obtain identification information, where the identification information includes the language family identifier of each character in the text reply message;

[0074] Take the glyph information and language family identifier corresponding to each character as the character feature information corresponding to each character, and take the text decomposition information and the identification information as the text processing information.

[0075] The above-mentioned glyph information corresponding to each character and the language family identifier as the character feature information corresponding to each character may mean that for any character in the text reply message, the glyph information and the language family identifier of the character can be used as the character feature information corresponding to the character. For example, the glyph information of the second character in the text reply message and the language family identifier of the second character are used as the corresponding character feature information of the second character, where the second character is any character in the text reply message.

[0076] The above-mentioned multiple languages may include all languages in the world, or include the more commonly used languages in the world. It should be noted that in a wearable device, the font library information is usually自带 so that users can select the voice of the wearable device, thus facilitating people in different countries to use the wearable device. Therefore, the above-mentioned font library information may be the font library information自带 by the wearable device.

[0077] The above-mentioned disassembling of the characters in the text reply message may specifically include: disassembling the characters in the text reply message character by character, and performing glyph disassembling on the disassembled characters. For example, the character "十" can be disassembled into two glyphs "一" and "|". Another example is that "and" can be disassembled into three glyphs "a", "n", and "d". After the disassembling is completed, the language family to which each character belongs can be determined based on the font library information. The font library information may include a large number of commonly used characters in each commonly used language. For example, by matching the characters with the font library information, the language family to which the characters belong can be determined, and a unique language family identifier can be set for each language family. The language family identifier may be a Unicode encoding. It should be noted that during the process of character processing, the order of the characters in the text reply message will not be disrupted, and the order of the language family identifiers of the obtained characters is the same as the order of the corresponding characters in the text reply message. In this way, the problem of reduced accuracy of subsequent recognition results caused by the disrupted order can be avoided.

[0078] In this embodiment, by disassembling the characters in the text reply message, character disassembling information is obtained, and by recognizing each character in the text reply message based on the font library information, recognition information is obtained, so as to obtain the language family identifier and glyph information of each character in the text reply message, which is convenient for subsequent language recognition of the text reply message based on the language family identifier and glyph information.

[0079] Optionally, the generating of the voice reply message based on the text reply message and the language information includes:

[0080] Sending the text reply message and the language information to a second server;

[0081] Receive the voice reply message sent by the second server.

[0082] After receiving the text reply message and the language information, the second server can convert the text reply message into audio in the target language according to the target language indicated by the language information. It should be noted that when the text reply message includes text in more than two languages, during the audio conversion process, only the text corresponding to the target language can be converted into audio in the target language, and text in languages ​​other than the target language may not be processed. For example, when there are three national languages ​​A, B, and C in the text reply message, and A appears the most times, the audio is requested in the national language A, and B and C are not processed.

[0083] The second server can be deployed with a large AI model. The second server can input the text reply message and the language information into the large AI model, so that the second server can convert the text reply message into audio in the target language according to the target language indicated by the language information. This can increase the speed of generating voice reply messages without requiring excessive computing power from the wearable device, thereby improving the user experience.

[0084] The first server and the second server may be the same server or different servers. When the first server and the second server are the same server, the AI ​​large model deployed in the first server and the AI ​​large model deployed in the second server in the above embodiment may be the same AI large model. For example, the AI ​​large model may be a large language model. Furthermore, when the first server and the second server are the same server, the AI ​​large model deployed in the first server and the AI ​​large model deployed in the second server in the above embodiment may be two different AI large models. This can be configured as needed.

[0085] In this embodiment, by sending the text reply message and the language information to a second server and receiving the voice reply message sent by the second server, the text reply message can be converted into a voice reply message, which is conducive to the subsequent output of the message content of the text reply message in the form of voice broadcast.

[0086] See Figure 2 , Figure 2 This is a schematic diagram of the structure of an interactive device 200 provided in an embodiment of the present disclosure, wherein the interactive device 200 includes:

[0087] The recognition module 201 is configured to, upon receiving an input message from a user, obtain a text reply message corresponding to the input message, and perform language recognition on the text reply message to obtain language information corresponding to the text reply message;

[0088] A generating module 202 is configured to generate a voice reply message based on the text reply message and the language information, wherein the voice reply message is an audio message obtained by converting the text reply message into a target language indicated by the language information;

[0089] The output module 203 is configured to output the voice reply message.

[0090] Optionally, the identification module 201 includes:

[0091] A first sending submodule, configured to send the input message to a first server;

[0092] The identification submodule is used to receive the text reply message sent by the first server, and perform language identification on the text reply message to obtain language information corresponding to the text reply message.

[0093] Optionally, the identification submodule is specifically used to perform language identification on the text in the received text reply message during the process of receiving the text reply message sent by the first server, so as to obtain language information corresponding to the text reply message.

[0094] Optionally, the identification module 201 further includes:

[0095] a processing submodule, configured to perform text processing on the text reply message to obtain text processing information, wherein the text processing information includes text feature information corresponding to each character in the text reply message;

[0096] The determination submodule is used to determine the language information corresponding to the text reply message based on the text feature information.

[0097] Optionally, the determination submodule is specifically configured to determine the language of each character in the text reply message based on the character feature information;

[0098] The determination submodule is further specifically configured to determine the language information based on the language of each character in the text reply message, wherein the language information includes the target language, and the target language is the language of the characters in the text reply message that corresponds to the largest number of characters in the text reply message.

[0099] Optionally, the text feature information includes a language family identifier and glyph information, the language family identifier being used to identify the language family to which the corresponding text belongs, wherein different language families correspond to different language family identifiers, and the determination submodule is further configured to determine at least two candidate languages ​​based on the language family identifier corresponding to the first text, wherein the first text is any text in the text reply message, and the at least two candidate languages ​​include: a language in the language family indicated by the language family identifier corresponding to the first text, and the glyph information corresponding to the first text includes k glyphs, where k is an integer greater than 0;

[0100] The determination submodule is further configured to determine, from the at least two candidate languages, a recognition language corresponding to each of the k glyphs, to obtain k recognition languages, wherein the recognition languages ​​are languages ​​from the at least two candidate languages, and the k recognition languages ​​correspond one-to-one to the k glyphs;

[0101] The determination submodule is further configured to determine the language that appears the most times among the k recognized languages ​​as the language to which the first character belongs.

[0102] Optionally, the wearable device stores font information, the font information including fonts in multiple languages, and the processing submodule is specifically configured to decompose the characters in the text reply message to obtain text decomposition information, wherein the text decomposition information includes: each character in the text reply message, and font information corresponding to each character in the text reply message;

[0103] The processing submodule is further configured to identify each character in the text reply message based on the character library information to obtain identification information, wherein the identification information includes a language identification of each character in the text reply message;

[0104] The processing submodule is further configured to use the glyph information and language identification corresponding to each character as the character feature information corresponding to each character, and use the character decomposition information and the recognition information as the character processing information.

[0105] Optionally, the generating module 202 includes:

[0106] A second sending submodule, configured to send the text reply message and the language information to a second server;

[0107] A receiving submodule is used to receive the voice reply message sent by the second server.

[0108] In this embodiment, during the interaction between the wearable device and the user, the language of the text reply message can be identified locally to obtain the language information corresponding to the text reply message without calling a third-party service for language identification. This helps reduce the number of times the third-party service is called during the interaction process, and further helps reduce the time spent in generating voice reply messages, thereby improving the user experience during the interaction process.

[0109] The present disclosure also provides an electronic device. Figure 3 , the electronic device may include a processor 301, a memory 302, and a program 3021 stored in the memory 302 and executable on the processor 301.

[0110] When the program 3021 is executed by the processor 301, it can achieve Figure 1 Any step in the corresponding method embodiment. Specifically, when program 3021 is executed by processor 301, the following steps can be implemented:

[0111] Upon receiving an input message from a user, obtaining a text reply message corresponding to the input message, and performing language recognition on the text reply message to obtain language information corresponding to the text reply message;

[0112] generating a voice reply message based on the text reply message and the language information, wherein the voice reply message is an audio message obtained by converting the text reply message into a target language indicated by the language information;

[0113] Output the voice reply message.

[0114] Optionally, the acquiring a text reply message corresponding to the input message and performing language identification on the text reply message to obtain language information corresponding to the text reply message includes:

[0115] Sending the input message to the first server;

[0116] The text reply message sent by the first server is received, and language identification is performed on the text reply message to obtain language information corresponding to the text reply message.

[0117] Optionally, the receiving the text reply message sent by the first server and performing language recognition on the text reply message to obtain language information corresponding to the text reply message includes:

[0118] In the process of receiving the text reply message sent by the first server, language recognition is performed on the text in the received text reply message to obtain language information corresponding to the text reply message.

[0119] Optionally, performing language identification on the text reply message to obtain language information corresponding to the text reply message includes:

[0120] Performing text processing on the text reply message to obtain text processing information, wherein the text processing information includes text feature information corresponding to each character in the text reply message;

[0121] Based on the text feature information, the language information corresponding to the text reply message is determined.

[0122] Optionally, determining the language information corresponding to the text reply message based on the text feature information includes:

[0123] Determining the language of each character in the text reply message based on the character feature information;

[0124] The language information is determined based on the language of each character in the text reply message, wherein the language information includes the target language, and the target language is the language of the characters in the text reply message that has the largest number of corresponding characters in the text reply message.

[0125] Optionally, the text feature information includes a language identifier and glyph information, the language identifier being used to identify the language to which the corresponding text belongs, wherein different language identifiers correspond to different language families. Determining the language of each text in the text reply message based on the text feature information includes:

[0126] Determining at least two candidate languages ​​based on the language family identifier corresponding to a first character, wherein the first character is any character in the text reply message, and the at least two candidate languages ​​include: a language in the language family indicated by the language family identifier corresponding to the first character, and the glyph information corresponding to the first character includes k glyphs, where k is an integer greater than 0;

[0127] Determining, from the at least two candidate languages, a recognition language corresponding to each of the k glyphs to obtain k recognition languages, wherein the recognition languages ​​are languages ​​from the at least two candidate languages, and the k recognition languages ​​correspond one-to-one to the k glyphs;

[0128] The language that appears most frequently among the k recognized languages ​​is determined as the language to which the first character belongs.

[0129] Optionally, the wearable device stores character library information, the character library information including character libraries in multiple languages, and the text processing of the text reply message to obtain the text processing information includes:

[0130] Decomposing the characters in the text reply message to obtain character decomposition information, wherein the character decomposition information includes: each character in the text reply message and font information corresponding to each character in the text reply message;

[0131] Recognize each character in the text reply message based on the character library information to obtain recognition information, wherein the recognition information includes a language identifier of each character in the text reply message;

[0132] The glyph information and language identification corresponding to each character are used as the character feature information corresponding to each character, and the character decomposition information and the recognition information are used as the character processing information.

[0133] Optionally, generating a voice reply message based on the text reply message and the language information includes:

[0134] Sending the text reply message and the language information to a second server;

[0135] Receive the voice reply message sent by the second server.

[0136] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the computer program implements the various processes of the above-mentioned interactive method embodiment and can achieve the same technical effect. To avoid repetition, the details are not described here. The computer-readable storage medium is, for example, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0137] The embodiments of the present disclosure further provide a computer program product, which is stored in a storage medium. The computer program product is executed by at least one processor to implement the various processes of the above-mentioned interaction method embodiment and can achieve the same technical effect. To avoid repetition, it will not be repeated here.

[0138] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0139] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better embodiment. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in each embodiment of the present disclosure.

[0140] The embodiments of the present disclosure are described above in conjunction with the accompanying drawings, but the present disclosure is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present disclosure, ordinary technicians in this field can also make many forms without departing from the scope of protection of the purpose of the present disclosure and the claims, all of which are protected by the present disclosure.

Claims

1. An interaction method, applied to a wearable device, characterized in that: The method comprises: When receiving an input message from a user, obtaining a text reply message corresponding to the input message, and performing language recognition on the text reply message to obtain language information corresponding to the text reply message; Based on the text reply message and the language information, generating a voice reply message, wherein the voice reply message is an audio message obtained by converting the text reply message into a target language indicated by the language information; The voice reply message is output.

2. The method according to claim 1, characterized in that: The acquiring of a text reply message corresponding to the input message, and performing language recognition on the text reply message to obtain language information corresponding to the text reply message includes: Sending the input message to the first server; The text reply message sent by the first server is received, and language identification is performed on the text reply message to obtain language information corresponding to the text reply message.

3. The method according to claim 2, characterized in that The receiving the text reply message sent by the first server and performing language recognition on the text reply message to obtain language information corresponding to the text reply message includes: In the process of receiving the text reply message sent by the first server, language recognition is performed on the text in the received text reply message to obtain language information corresponding to the text reply message.

4. The method according to claim 1, characterized in that: The performing language identification on the text reply message to obtain language information corresponding to the text reply message includes: Performing text processing on the text reply message to obtain text processing information, wherein the text processing information includes text feature information corresponding to each character in the text reply message; Based on the text feature information, the language information corresponding to the text reply message is determined.

5. The method according to claim 4, characterized in that The determining, based on the text feature information, the language information corresponding to the text reply message includes: Based on the text feature information, determining the language of each character in the text reply message; The language information is determined based on the language to which each character in the text reply message belongs, wherein the language information includes the target language, and the target language is the language to which the characters in the text reply message belong and which corresponds to the largest number of characters in the text reply message.

6. The method according to claim 5, characterized in that The text feature information includes a language identifier and font information, wherein the language identifier is used to identify the language to which the corresponding text belongs, wherein different language identifiers correspond to different languages, and determining the language to which each text in the text reply message belongs based on the text feature information includes: Determine at least two candidate languages ​​based on the language family identifier corresponding to the first character, wherein the first character is any character in the text reply message, and the at least two candidate languages ​​include: a language in the language family indicated by the language family identifier corresponding to the first character, and the glyph information corresponding to the first character includes k glyphs, where k is an integer greater than 0; Determine the recognition language corresponding to each of the k glyphs among the at least two candidate languages, and obtain k recognition languages, wherein the recognition languages ​​are languages ​​among the at least two candidate languages, and the k recognition languages ​​correspond to the k glyphs one-to-one; The language that appears most frequently among the k recognized languages ​​is determined as the language to which the first character belongs.

7. The method according to claim 4, characterized in that The wearable device stores character library information, the character library information includes character libraries in multiple languages, and the text reply message is subjected to character processing to obtain the text processing information, including: Decomposing the characters in the text reply message to obtain character decomposition information, wherein the character decomposition information includes: each character in the text reply message and font information corresponding to each character in the text reply message; Recognize each character in the text reply message based on the character library information to obtain recognition information, wherein the recognition information includes a language identification of each character in the text reply message; The glyph information and language identification corresponding to each character are used as the character feature information corresponding to each character, and the character decomposition information and the recognition information are used as the character processing information.

8. An interactive device, applied to a wearable device, characterized in that: The device comprises: The recognition module is used to obtain a text reply message corresponding to the input message when receiving the input message of the user, and perform language recognition on the text reply message to obtain language information corresponding to the text reply message; A generating module, configured to generate a voice reply message based on the text reply message and the language information, wherein the voice reply message is an audio message obtained by converting the text reply message into a target language indicated by the language information; An output module is used to output the voice reply message.

9. An electronic device, comprising: A memory, a processor, and a program stored in the memory and executable on the processor; wherein the processor is used to read the program in the memory to implement the steps of the interactive method as described in any one of claims 1 to 7.

10. A computer program product, characterized in that The computer program product is stored in a storage medium, and the computer program product is executed by at least one processor to implement the steps in the interaction method according to any one of claims 1 to 7.