An input method vocabulary determination method and device, vehicle, and storage medium

By collecting lip and facial features of people inside the vehicle and analyzing music playback statistics, the dialect type of the input method is automatically determined, solving the problem of dialect word matching in the input method relying on the user, thus improving user experience and driving safety.

CN115392224BActive Publication Date: 2025-11-11ECARX (HUBEI) TECHCO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210998159.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-19
Publication Date
2025-11-11
Estimated Expiration
2042-08-19

AI Technical Summary

Technical Problem

The existing input methods rely too heavily on manual selection by users to match dialect words, resulting in a poor user experience, which in particular affects driving safety when used by drivers.

Method used

By collecting data on the lip movements, facial expressions, and music playback of people inside the vehicle, and combining this with preset probability thresholds, the system automatically determines the target dialect and matches it to the input method's dictionary, thus avoiding manual operation by the user.

Benefits of technology

It eliminates the need for users to manually select dialect dictionaries, improving the intelligence and user experience of the input method, and especially enhancing safety in driving environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115392224B_ABST
    Figure CN115392224B_ABST
Patent Text Reader

Abstract

This invention discloses a method, apparatus, vehicle, and storage medium for determining the dictionary of an input method. The method includes: if the decibel value of a human voice inside the vehicle falls within a first preset range, then determining a target dialect result based on at least two of a first dialect result, a second dialect result, and a third dialect result. The first dialect result is determined based on the lip movements of a preset person inside the vehicle; the second dialect result is determined based on the facial features of a preset person; and the third dialect result is determined based on a music playback statistics database inside the vehicle. The dialect result includes a correspondence between dialect types and their probabilities. The target dialect type is determined based on the target dialect result and a preset probability threshold. Finally, the dictionary of the vehicle's input method is determined based on the target dialect type. This invention solves the problem of excessive reliance on manual user selection for dialect dictionary matching in input methods, intelligently determining the dictionary for the user and improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of input method technology, and in particular to a method, apparatus, vehicle, and storage medium for determining the dictionary of an input method. Background Technology

[0002] Input methods refer to the encoding methods used to input various symbols into electronic information devices (such as computers and mobile phones). The encoding methods for Chinese character input are basically based on sound, form, and meaning.

[0003] Currently, input methods are basically divided into two categories: 1) those based on the shape and meaning of characters, which offer fast and accurate input but require specialized learning; and 2) those based on pronunciation, which have lower accuracy, with keyboards typically printed with pinyin letters, requiring no special learning for inputting commonly used words. Among Chinese input methods, Pinyin or Wubi are commonly used as encoding methods, while some regions widely use Cangjie or Zhuyin Fuhao.

[0004] For users who don't speak Mandarin or are accustomed to using their dialect for speech-to-text input, the accuracy of the dialect dictionary matching between the input method and the user is particularly important. Inaccurate matching will lead to errors in speech-to-text conversion. However, the matching between the user and the input method's dialect dictionary mainly relies on the user's manual selection. This method is overly dependent on the user, lacks intelligence, and results in a poor user experience. For example, for drivers, manually selecting a dialect dictionary while driving can affect driving safety. Summary of the Invention

[0005] This invention provides a method, apparatus, vehicle, and storage medium for determining the dictionary of an input method, in order to solve the problem of over-reliance on the user when matching the dialect dictionary of an input method.

[0006] In a first aspect, embodiments of the present invention provide a method for determining the dictionary of an input method, including:

[0007] If the decibel value of the human voice inside the vehicle falls within a first preset range, then a target dialect result is determined based on at least two of the first dialect result, the second dialect result, and the third dialect result. The first dialect result is determined based on the lip movements of a preset person inside the vehicle, the second dialect result is determined based on the facial features of the preset person, and the third dialect result is determined based on the music playback statistics database inside the vehicle. The dialect result includes the correspondence between dialect types and the probability of each dialect type.

[0008] Based on the target dialect results and a preset probability threshold, the target dialect type is determined;

[0009] The vocabulary of the vehicle's input method is determined based on the target dialect.

[0010] Secondly, embodiments of the present invention provide a dictionary determination device for an input method, comprising:

[0011] The target dialect result determination module is used to determine the target dialect result based on at least two of the first dialect result, the second dialect result, and the third dialect result if the decibel value of the human voice in the vehicle is within a first preset range. The first dialect result is determined based on the lip movements of a preset person in the vehicle, the second dialect result is determined based on the facial features of the preset person, and the third dialect result is determined based on the music playback statistics database in the vehicle. The dialect result includes the correspondence between dialect types and the probability of the dialect types.

[0012] The target dialect type determination module is used to determine the target dialect type based on the target dialect result and a preset probability threshold.

[0013] The first lexicon determination module is used to determine the lexicon of the vehicle's input method based on the target dialect type.

[0014] Thirdly, embodiments of the present invention provide a vehicle, the vehicle comprising:

[0015] At least one processor;

[0016] and memory that is communicatively connected to at least one processor;

[0017] The memory stores a computer program that can be executed by at least one processor, which is executed by at least one processor to enable the at least one processor to perform the dictionary determination method of the input method described in the first aspect.

[0018] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing computer instructions for causing a processor to execute the dictionary determination method of the input method described in the first aspect.

[0019] The input method dictionary determination scheme provided in this embodiment of the invention, if the decibel value of human voices inside the vehicle falls within a first preset range, then a target dialect result is determined based on at least two of the first dialect result, the second dialect result, and the third dialect result. The first dialect result is determined based on the lip movements of a preset person inside the vehicle, the second dialect result is determined based on the facial features of the preset person, and the third dialect result is determined based on a music playback statistics database inside the vehicle. The dialect result includes a correspondence between dialect types and the probabilities of those dialect types. Based on the target dialect result and a preset probability threshold, the target dialect type is determined, and based on the target dialect type, the dictionary of the vehicle's input method is determined. By adopting the above technical solution, when the dialect of a person's voice inside the vehicle cannot be determined through speech recognition (i.e., the decibel value falls within the first preset range), the target dialect result can be determined based on at least two items from the preset facial features, lip movements, and music playback statistics database. Then, based on the target dialect result and a preset probability threshold, the target dialect type is obtained. Finally, based on the target dialect type, the vocabulary of the vehicle's input method can be determined. The entire process requires no manual operation from the user, solving the problem that matching the dialect vocabulary of the input method relies too heavily on manual selection by the user. It intelligently determines the vocabulary of the input method for the user, improving the user experience.

[0020] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a flowchart of a dictionary determination method for an input method according to Embodiment 1 of the present invention;

[0023] Figure 2 This is a flowchart of a dictionary determination method for an input method according to Embodiment 2 of the present invention;

[0024] Figure 3 This is a flowchart of a dictionary determination method for an input method according to Embodiment 3 of the present invention;

[0025] Figure 4 This is a flowchart of a dictionary determination method for an input method according to Embodiment 4 of the present invention;

[0026] Figure 5 This is a flowchart of a dictionary determination method for an input method according to Embodiment 5 of the present invention;

[0027] Figure 6 This is a schematic diagram of the structure of a dictionary determination device for an input method according to Embodiment Six of the present invention;

[0028] Figure 7 This is a structural schematic diagram of a vehicle provided according to Embodiment Seven of the present invention. Detailed Implementation

[0029] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0030] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. In the description of this invention, unless otherwise stated, "a plurality of" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or devices.

[0031] Example 1

[0032] Figure 1The flowchart of an input method dictionary determination method provided in Embodiment 1 of the present invention is applicable to the determination of dialect dictionary of input method. The method can be executed by input method dictionary determination device, which can be implemented in hardware and / or software. The input method dictionary determination device can be configured in a vehicle and can be implemented in hardware and / or software.

[0033] like Figure 1 As shown in the figure, the word database determination method for an input method provided in Embodiment 1 of the present invention specifically includes the following steps:

[0034] S101. If the decibel value of human voices inside the vehicle falls within the first preset range, then the target dialect result is determined based on at least two of the first dialect result, the second dialect result, and the third dialect result.

[0035] The first dialect result is determined based on the lip movements of the preset personnel in the vehicle, the second dialect result is determined based on the facial features of the preset personnel, and the third dialect result is determined based on the music playback statistics database in the vehicle. The dialect results include the correspondence between dialect types and the probability of dialect types.

[0036] In this embodiment, a microphone inside the vehicle can be used to collect human voices inside the vehicle. When the volume of the human voices inside the vehicle is within a first preset range, such as a decibel value greater than 10 decibels and less than 40 decibels, the purpose of setting the first preset range is to distinguish whether the human voices inside the vehicle have clearly identifiable semantics and tones, and can accurately determine the dialect type. If the decibel value of the human voices inside the vehicle is not within the first preset range, there is a possibility of directly determining the dialect type based on the human voice. If the decibel value of the human voices inside the vehicle is within the first preset range, such as when people inside the vehicle are whispering, the volume of the human voice is too low, and it is impossible to directly determine the dialect type based on the human voice. Generally, people's lip movements differ when speaking different dialects, and these patterns follow certain rules. These rules can be used to determine the dialect spoken by the people in the vehicle and its probability. People from different regions also exhibit significant differences in facial features; for example, the facial features of people from Shandong differ from those from people from Guangdong. Based on the facial features of people from different regions, the region to which the people in the vehicle belong can be inferred. Then, based on the dialect spoken in that region, the dialect spoken by the people in the vehicle and its probability can be determined. Furthermore, analyzing the music playback statistics database within the vehicle can determine the music preferences of the people in the vehicle. The preferences of people speaking different dialects follow certain patterns; for example, Cantonese speakers prefer Cantonese songs. Therefore, based on music preferences, the dialect spoken by the people in the vehicle and its probability can be inferred. In summary, based on the lip movements, facial features, and the music playback statistics database within the vehicle, three dialect results containing the dialect type and its probability can be obtained. Using at least two of these results, the target dialect can be determined with relatively high accuracy. The lip movements of the preset personnel can be captured by in-vehicle cameras. A music playback statistics database can be used; first, the microphones in the vehicle capture music played inside, then relevant speech recognition algorithms are used to identify the music track and language. The human voices inside the vehicle can be those of the preset personnel. The distinction between human and non-human voices can be made using a voice recognition algorithm. Preset personnel include the driver and / or front passenger. The target dialect result can include one or more dialect types and their probabilities. Dialect types include Cantonese, Minnan, and Sichuanese, among others, without limitation. The dialect probability represents the likelihood that the preset personnel will use the corresponding dialect type. For example, if the dialect result includes Cantonese with a probability of 10%, it means that after identification using the dialect recognition method corresponding to that result, the probability of the preset personnel using Cantonese is 10%, meaning the probability of needing to use a Cantonese input method dictionary is 10%.

[0037] S102. Determine the target dialect type based on the target dialect results and the preset probability threshold.

[0038] In this embodiment, the probabilities in the target dialect results obtained in the above steps can be compared with a preset probability threshold. Probabilities greater than or equal to the preset probability threshold can be filtered out, and the dialect types corresponding to these filtered probabilities are the target dialect types. The number of dialect types corresponding to the filtered probabilities can be one or more, thus the number of target dialect types can also be one or more.

[0039] S103. Determine the vocabulary of the vehicle's input method based on the target dialect.

[0040] In this embodiment, when the target dialect is a single dialect, the vocabulary of the vehicle's input method can be determined as the vocabulary corresponding to that single dialect. When the target dialect is multiple dialects, the vocabulary of the vehicle's input method can be kept unchanged, that is, the originally selected vocabulary can be maintained without any changes.

[0041] For example, if the single dialect is Cantonese, then the vehicle's input method's dictionary can be set to the Cantonese dictionary. If the single dialect is both Cantonese and Hokkien, and the vehicle's input method is currently using the Mandarin dictionary, then the vehicle's input method's dictionary will remain the Mandarin dictionary and will not be changed.

[0042] The input method vocabulary determination method provided in this embodiment of the invention, if the decibel value of human voices inside the vehicle falls within a first preset range, then a target dialect result is determined based on at least two of a first dialect result, a second dialect result, and a third dialect result. The first dialect result is determined based on the lip movements of a preset person inside the vehicle, the second dialect result is determined based on the facial features of the preset person, and the third dialect result is determined based on a music playback statistics database inside the vehicle. The dialect result includes a correspondence between dialect types and the probabilities of those dialect types. The target dialect type is determined based on the target dialect result and a preset probability threshold. Finally, the vocabulary of the input method for the vehicle is determined based on the target dialect type. The technical solution of this invention addresses the issue that when the dialect of a person's voice inside a vehicle cannot be determined through speech recognition (i.e., the decibel value falls within a first preset range), the target dialect can be determined based on at least two factors from a preset database of facial features, lip movements, and music playback statistics. Then, based on the target dialect result and a preset probability threshold, the target dialect type is obtained. Finally, the dictionary for the vehicle's input method can be determined based on this target dialect type. This entire process requires no manual user intervention, solving the problem of excessive reliance on manual selection for dialect dictionary matching in input methods. It intelligently determines the dictionary for the user's input method, improving the user experience.

[0043] Example 2

[0044] Figure 2This is a flowchart of a method for determining the dictionary of an input method according to Embodiment 2 of the present invention. The technical solution of the present invention is further optimized based on the above optional technical solutions, and a specific method for determining the dictionary of an input method is given.

[0045] Optionally, the method further includes: acquiring the lip movements of a preset person inside the vehicle; matching the lip movements with sample lip movements in a preset lip movement library to obtain a similarity result; wherein the preset lip movement library contains a correspondence between dialect types and lip movement samples; the similarity result includes the lip movement similarity between the lip movement and the sample lip movement, as well as the dialect type corresponding to the lip movement similarity; and determining the dialect types and their corresponding lip movement similarities that meet a first preset requirement in the similarity result as the probability of the dialect type in the first dialect result. The advantage of this setup is that it allows for the inference of the preset person's dialect type and its corresponding probability based solely on their lip movements, without requiring the preset person's voice.

[0046] Optionally, the method further includes: collecting the facial features of a preset person inside the vehicle; identifying the region to which the facial features belong using a preset facial feature database; and obtaining a region result, wherein the preset facial feature database contains a correspondence between facial feature samples and their respective regions; and the region result includes the facial feature similarity between the facial feature and the sample facial features, as well as the region corresponding to the facial feature similarity. Based on the region result, a preset regional dialect database, and a second preset requirement, a second dialect result is determined, wherein the preset regional dialect database contains a correspondence between the region, the dialect type, and the probability of using the dialect type; and the probability of the dialect type in the second dialect result is determined based on the facial feature similarity and the probability of use. The advantage of this approach is that it uncovers the connection between facial features, region, and dialect type, ensuring the accuracy of the inferred dialect type and its corresponding probability.

[0047] Optionally, determining the target dialect type based on the target dialect results and a preset probability threshold includes: filtering the probability of satisfying preset conditions and the corresponding dialect type from the target dialect results to obtain a target probability result, wherein the target probability result includes a target probability and the target dialect type corresponding to the target probability; and determining the target dialect type based on the preset probability threshold and the target probability result. The advantage of this setting is that by using preset conditions and a preset probability threshold to filter the target dialect type, the accuracy of the target dialect type is improved, and the probability of a target dialect type having multiple dialect types is significantly reduced.

[0048] like Figure 2 As shown in Embodiment 2 of the present invention, a method for determining the dictionary of an input method specifically includes the following steps:

[0049] S201. If the decibel value of the human voice inside the vehicle is within the first preset range, then obtain the lip movements of the preset people inside the vehicle, match the lip movements with the sample lip movements in the preset lip movement library, and obtain the similarity result.

[0050] The preset lip shape database contains the correspondence between dialect types and lip shape samples. The similarity results include the lip shape similarity between the lip shape and the sample lip shape, as well as the dialect type corresponding to the lip shape similarity.

[0051] Specifically, the similarity result can include one or more dialect types, and there is a one-to-one correspondence between dialect types and lip shape similarity. The process of matching lip shapes with sample lip shapes in a preset lip shape library and obtaining lip shape similarity can be achieved using relevant lip shape matching algorithms. Sample lip shapes can be understood as sample lip shape features.

[0052] For example, if the first preset range is greater than 10 decibels and less than 40 decibels, the preset person is a driver, and the preset lip-reading database contains lip-reading samples corresponding to dialects such as Cantonese, Guangxi Bai, Hakka, and Sichuanese, and the lip-reading samples correspond one-to-one with the dialects, then when the decibel value of the voice inside the vehicle is 30 decibels, that is, the decibel value of the voice inside the vehicle is within the first preset range, the driver's lip-reading can be captured by the in-vehicle camera, and the lip-reading can be matched with the sample lip-readings in the preset lip-reading database. The lip-reading similarity between the lip-reading and the sample lip-readings corresponding to dialects such as Cantonese, Guangxi Bai, Hakka, and Sichuanese can be obtained. Based on this result, a similarity result composed of the lip-reading similarity and the dialect corresponding to the lip-reading similarity can be obtained.

[0053] S202. The dialect types and their corresponding lip shape similarities that meet the first preset requirements in the similarity results are determined as the dialect types and their probabilities in the first dialect results.

[0054] Specifically, the first dialect result may contain one or more dialect types, and there is a one-to-one correspondence between the dialect types and their probabilities.

[0055] For example, if the first preset requirement is that the lip shape similarity is greater than 60%, in the similarity results, the lip shape similarity is arranged in descending order as follows: 80% similarity with Cantonese, 75% similarity with Guangxi Bai, 70% similarity with Hakka, 50% similarity with Sichuanese, and so on. Then, Cantonese, Guangxi Bai, and Hakka, and their corresponding probabilities, can be determined as the dialect type and the probability of the dialect type in the first dialect result.

[0056] S203. Collect the facial features of preset personnel inside the vehicle, use the preset facial feature library to identify the region to which the facial features belong, and obtain the region result.

[0057] The preset face database contains the correspondence between face samples and their respective regions. The region results include the face similarity between the face and the sample face, as well as the region corresponding to the face similarity.

[0058] Specifically, the similarity results can include one or more dialect types, and there is a one-to-one correspondence between dialect types and facial similarity. The process of matching facial shapes with sample facial shapes in a preset facial shape library and obtaining facial similarity can be achieved using relevant facial shape matching algorithms. Sample facial shapes can be understood as sample facial shape features.

[0059] For example, if the preset person is a driver, and the preset face database contains face shapes of residents from Guangdong, Guangxi, Fujian, and Sichuan, and each face shape corresponds one-to-one with the region, then the driver's face shape can be captured by the in-vehicle camera and matched with sample faces in the preset face database. This will give the face shape similarity to the sample faces corresponding to the regions of Guangdong, Guangxi, Fujian, and Sichuan. Based on this result, a similarity result consisting of the face shape similarity and the corresponding region can be obtained.

[0060] S204. Based on the results of the region, the preset regional dialect database, and the second preset requirements, determine the second dialect result.

[0061] The preset regional dialect database contains the correspondence between the region, dialect type, and the probability of using the dialect type. The probability of the dialect type in the second dialect result is determined based on facial similarity and usage probability.

[0062] Specifically, the process begins by filtering out regions from the regional results that match those in the preset regional dialect database. Then, using a preset calculation method (such as product calculation), the facial similarity of these regions is compared to the probability of using a particular dialect type in the preset regional dialect database. Finally, this is compared to a second preset requirement. If the result meets the second preset requirement, the probability of that result being identified as a dialect type in the second dialect results is determined, and the corresponding dialect type is officially recognized as a dialect type in the second dialect results. The preset calculation method can be product calculation or weighted product calculation, and the second preset requirement can be a requirement regarding probability magnitude. The second dialect results can contain one or more dialect types, with a one-to-one correspondence between dialect types and their probabilities.

[0063] For example, if the second preset requirement is a probability greater than 30%, the preset regional dialect database contains dialects used in Guangdong, Guangxi, Fujian, and Sichuan, along with the probability of each dialect. For instance, the probability of Cantonese being used in Guangdong is 50%, Hakka is 10%, and Hokkien is 5%. The regional result would be a facial similarity of 90% with residents of Guangdong, 87% with residents of Guangxi, 86% with residents of Fujian, and 84% with residents of Sichuan. The preset calculation method is multiplication. Therefore, the regional result corresponding to the dialect spoken in Guangdong would be... The similarity is 90%. The probability of using Cantonese in Guangdong is calculated by multiplying the similarity with the pre-defined regional dialect database. The probability of using Cantonese is 50%, Hakka is 10%, and Hokkien is 5%. This process is repeated for Guangxi, Fujian, and Sichuan. Finally, from the product results, languages ​​with a probability greater than 30% are selected. For example, if the product of 90% (similarity to facial features of Guangdong residents) * 50% (probability of Guangdong residents using Cantonese) is 45%, and 45% is greater than 30%, then Cantonese is identified as the language in the second dialect result, and 45% is identified as the probability of Cantonese in the second dialect result.

[0064] S205. Based on the results of the first dialect and the second dialect, determine the result of the target dialect.

[0065] Specifically, the results from the first dialect and the second dialect can be identified as the results from the target dialect.

[0066] S206. From the target dialect results, filter out the probabilities that meet the preset conditions and the corresponding dialect types to obtain the target probability results.

[0067] The target probability result includes the target probability and the target dialect type corresponding to the target probability.

[0068] For example, if the preset condition is to filter the highest probability and its corresponding language type in the first dialect result and to filter the highest probability and its corresponding language type in the second dialect result, if the filtering result of the first dialect result is Cantonese 80% and the filtering result of the second dialect result is Cantonese 45%, then the target probability results are: Cantonese 80% and Cantonese 45%.

[0069] S207. Based on the preset probability threshold and the target probability result, determine the target dialect type.

[0070] Specifically, the probabilities in the target probability results can be processed, such as summing the probabilities of the same language types. If the processed probability is greater than a preset probability threshold, such as 90%, then the language type corresponding to that probability can be determined as the target dialect type.

[0071] Optionally, determining the target dialect type based on the preset probability threshold and the target probability result includes: determining the dialect types in the target probability result as candidate dialect types; for each candidate dialect type, using a preset weighted summation method, performing a weighted summation calculation on the corresponding probabilities in the target probability result to obtain the candidate dialect result, wherein the candidate dialect result includes the correspondence between the candidate dialect type and the probability of the candidate dialect type; and determining the candidate dialect type corresponding to the probability greater than or equal to the preset probability threshold in the candidate dialect result as the target dialect type.

[0072] Specifically, if the target probability result contains one or more correspondences between probabilities and dialect types, then all dialect types in the target probability result can be identified as candidate dialect types. The probability of each dialect type in the target probability result is calculated by a preset weighted summation to obtain the correspondence between candidate dialect types and their probabilities, i.e., the candidate dialect result. Then, from the candidate dialect result, the probability greater than or equal to the preset probability threshold is selected. The candidate dialect type corresponding to the selected probability is the target dialect type.

[0073] For example, if the target probability results are Cantonese 80% and Cantonese 45%, the preset probability threshold is 0.5, and the preset weighted summation method is to add a weighting coefficient of 0.6 to the probability selected from the first dialect result and a weighting coefficient of 0.2 to the probability selected from the second dialect result, and sum the probabilities of the same dialect type, then the probability calculation process of the candidate dialect result is 0.6*0.8+0.2*0.45=0.57, the probability of the candidate dialect result is 0.57, the candidate dialect type is Cantonese, and since 0.57 is greater than 0.5, the target dialect type is Cantonese.

[0074] For example, if the target probability results are Cantonese 80%, Hokkien 60%, Hokkien 45%, and Cantonese 40%, with a preset probability threshold of 0.5 and a preset weighted summation method of adding a weighting coefficient of 0.6 to the probabilities selected from the first dialect results (Cantonese 80% and Hokkien 60%), adding a weighting coefficient of 0.2 to the probabilities selected from the second dialect results (Hokkien 45% and Cantonese 40%), and summing the probabilities of the same dialect type, then the probability calculation process for the candidate dialect results is: 0.6*0.8+0.2*0.4=0.56, 0.6*0.6+0.45*0.2=0.45. The probability of the candidate dialect results is Cantonese 56% and Hokkien 45%, and the candidate dialect type is Cantonese. Since 0.56 is greater than 0.5, the target dialect type is Cantonese.

[0075] The input method vocabulary determination method provided in this embodiment of the invention involves the following steps: If the decibel value of the human voice inside the vehicle falls within a first preset range, the lip movements of a preset person inside the vehicle are acquired. Based on the similarity between this lip movement and the sample lip movements in a preset lip movement library, a first dialect result is determined. The facial features of the preset person inside the vehicle are also collected, and the regional result is determined based on this facial feature. Based on the regional result, a preset regional dialect library, and a second preset requirement, a second dialect result is determined. Based on the first and second dialect results, a target dialect result is determined. From this target dialect result, the probability of meeting preset conditions and the corresponding dialect type are selected to obtain a target probability result. Finally, based on a preset probability threshold and the target probability result, the target dialect type is determined. This allows the determination of the input method vocabulary for the vehicle. By collecting the lip and facial features of people in the vehicle, the dialect type used by the user can be inferred before the user uses the input method, enabling timely and accurate recommendation of the input method vocabulary, significantly improving the matching efficiency between the user and the input method vocabulary.

[0076] Based on the above embodiments, the method may further include:

[0077] 1) If the decibel value of the human voice inside the vehicle is within the second preset range, the dialect type of the human voice is identified to obtain the fourth dialect result. The fourth dialect result includes the correspondence between the dialect type and the probability of the dialect type. The upper limit of the first preset range is less than the lower limit of the second preset range.

[0078] Specifically, when the decibel level of a person's voice inside the vehicle falls within a second preset range (e.g., greater than 41 decibels but less than 70 decibels), the purpose of setting this second preset range is to distinguish whether the voice inside the vehicle can be clearly identified in terms of semantics and tone. If the decibel level of the voice inside the vehicle is within the second preset range, relevant speech recognition algorithms can be used to identify the semantics and tone of the voice, thereby identifying the dialect type of the voice and the probability that the voice belongs to that dialect type, thus obtaining a fourth dialect result. The fourth dialect result can include a correspondence between one or more dialect types and the probability of each dialect type.

[0079] 2) Determine the vocabulary of the vehicle's input method based on the dialect type with the highest probability in the fourth dialect result.

[0080] Specifically, if the dialect with the highest probability in the fourth dialect result is Cantonese, then it can be determined that the vehicle's input method's dictionary is a Cantonese dictionary.

[0081] Example 3

[0082] Figure 3 This is a flowchart of a method for determining the dictionary of an input method according to Embodiment 2 of the present invention. The technical solution of the present invention is further optimized based on the above optional technical solutions, and a specific method for determining the dictionary of an input method is given.

[0083] Optionally, the method further includes: determining music dialect type results based on a music playback statistics database within the vehicle and a third preset requirement, wherein the music playback statistics database contains a correspondence between music tracks, playback counts, and dialect types, and the music dialect type results include a correspondence between the dialect types of music tracks and the proportion of each dialect type; and determining a third dialect result based on the music dialect type results, a preset music dialect database, and a fourth preset requirement, wherein the preset music dialect database includes a correspondence between the dialect types of music tracks, the dialect type used by the listener, and the probability of the dialect type used by the listener, and the probability of the dialect type in the third dialect result is determined based on the proportion of each dialect type and the probability of the dialect type used by the listener. The advantage of this setup is that it uncovers the connection between the language type of the music listened to by the user and the dialect type used by the user, and based on this connection, accurately infers the dialect used by the user and the probability of using that dialect.

[0084] like Figure 3 As shown in Embodiment 3 of the present invention, a method for determining the dictionary of an input method specifically includes the following steps:

[0085] S301. If the decibel value of the human voice inside the vehicle is within the first preset range, then obtain the lip movements of the preset people inside the vehicle, match the lip movements with the sample lip movements in the preset lip movement library, and obtain the similarity result.

[0086] S302. The dialect types and their corresponding lip shape similarities that meet the first preset requirements in the similarity results are determined as the dialect types and their probabilities in the first dialect results.

[0087] S303. Based on the music playback statistics database in the vehicle and the third preset requirements, determine the music dialect category results.

[0088] The music playback statistics database includes the correspondence between music tracks, number of plays, and dialect types. The music dialect type results include the correspondence between the dialect types of music tracks and the proportion of each dialect type.

[0089] For example, if the third preset requirement is to filter out a preset number (such as three) of dialect types with the highest percentage of music playbacks, then the music tracks in the music playback statistics library can be classified according to dialect type, and the number of playbacks of music tracks in each dialect type can be calculated. Then, the percentage of playbacks of music tracks in each dialect type can be calculated, and the three dialect types with the highest percentage of music playbacks can be filtered out to obtain the music dialect type results, such as Cantonese 70%, Hokkien 20%, Mandarin 5%, and English 5%.

[0090] S304. Based on the results of music dialect types, the preset music dialect database, and the fourth preset requirement, determine the third dialect result.

[0091] The preset music dialect database includes the dialect types of music tracks, the dialect types used by listeners, and the probability of the dialect types used by listeners. The probability of dialect types in the third dialect results is determined based on the proportion of dialect types and the probability of the dialect types used by listeners.

[0092] Specifically, the process begins by filtering out dialects from the results of music dialect categories that match the dialect categories of music tracks in a preset music dialect database. Then, using a preset calculation method, such as multiplication, the probability of that dialect category being used is calculated, along with the probability of the listener using that dialect category for the same music track in the preset music dialect database. Finally, this is compared with a fourth preset requirement. If the calculation result meets the fourth preset requirement, then that result is determined as the probability of a dialect category in the third dialect result, and the corresponding dialect category is identified as a dialect category in the third dialect result. The preset calculation method can be multiplication or weighted multiplication, etc. The fourth preset requirement can be a requirement for the probability magnitude, such as a probability value greater than a preset probability value. The third dialect result can contain one or more dialect categories, and there is a one-to-one correspondence between dialect categories and their probabilities.

[0093] For example, if the fourth preset requirement is a probability greater than 10%, the preset person is a driver, and the preset music dialect database contains a correspondence between the dialect types of music tracks, the probability of listeners using Mandarin being 50%, Cantonese being 30%, and Hokkien being 2%, and the probability of each dialect type used by the listener, and the corresponding probability of each dialect type used by the listener. The music dialect type results include a correspondence between the dialect types of music tracks and the percentage of each dialect type, such as Cantonese being 70%, Hokkien being 20%, and Mandarin being 5%. Then, the probability of Cantonese can be calculated as 0.3 * 0.7 = 0.21. Similarly, the probabilities of different dialect types such as Hokkien and Mandarin are calculated separately. Finally, from the product results, language types with a probability greater than 30%, such as Cantonese with a probability of 21%, are selected. Cantonese is then determined as the language type in the third dialect result, and 21% is determined as the probability of Cantonese in the third dialect result.

[0094] S305. Based on the results of the first dialect and the third dialect, determine the result of the target dialect.

[0095] Specifically, the results from the first dialect and the third dialect can be identified as the results from the target dialect.

[0096] S306. From the target dialect results, filter out the probabilities that meet the preset conditions and the corresponding dialect types to obtain the target probability results.

[0097] The target probability result includes the target probability and the target dialect type corresponding to the target probability.

[0098] For example, the preset conditions can be to filter the highest probability of the first dialect result and its corresponding language type, and to filter the highest probability of the third dialect result and its corresponding language type. If the filtering result of the first dialect result is Cantonese 80% and the filtering result of the third dialect result is Cantonese 21%, then the target probability results are: Cantonese 80% and Cantonese 21%.

[0099] S307. Based on the preset probability threshold and the target probability result, determine the target dialect type.

[0100] Optionally, determining the target dialect type based on the preset probability threshold and the target probability result includes: determining the dialect types in the target probability result as candidate dialect types; for each candidate dialect type, using a preset weighted summation method, performing a weighted summation calculation on the corresponding probabilities in the target probability result to obtain the candidate dialect result, wherein the candidate dialect result includes the correspondence between the candidate dialect type and the probability of the candidate dialect type; and determining the candidate dialect type corresponding to the probability greater than or equal to the preset probability threshold in the candidate dialect result as the target dialect type.

[0101] For example, if the target probability results are Cantonese 80% and Cantonese 21%, the preset probability threshold is 0.5, and the preset weighted summation method is to add a weighting coefficient of 0.6 to the probability selected from the first dialect result and a weighting coefficient of 0.2 to the probability selected from the third dialect result, and sum the probabilities of the same dialect type, then the probability calculation process of the candidate dialect result is 0.6*0.8+0.2*0.21=0.522. The probability of the candidate dialect result is 0.522, and the candidate dialect type is Cantonese. Since 0.522 is greater than 0.5, the target dialect type is Cantonese.

[0102] The input method vocabulary determination method provided in this embodiment of the invention involves obtaining the lip movements of a preset person in the vehicle if the decibel value of the voice inside the vehicle falls within a first preset range. Based on the similarity between this lip movement and the sample lip movements in a preset lip movement library, a first dialect result is determined. Then, based on a music playback statistics library inside the vehicle, a music dialect type result is determined. Finally, based on the music dialect type result, the preset music dialect library, and a fourth preset requirement, a third dialect result is determined. Based on the first and third dialect results, a target dialect result is determined. From this target dialect result, the probability and corresponding dialect type that meet preset conditions are selected to obtain a target probability result. Finally, based on a preset probability threshold and the target probability result, the target dialect type is determined, thus determining the vocabulary of the vehicle's input method. By collecting the lip movement features of people in the vehicle and the features of the music played inside, the dialect type used by the user can be inferred before the user uses the input method, allowing for timely and accurate recommendation of the input method's vocabulary, significantly improving the matching efficiency between the user and the input method's vocabulary.

[0103] Based on the above embodiments, the method may further include: if the decibel value of the human voice inside the vehicle is within a second preset range, then identify the dialect type of the human voice to obtain a fourth dialect result, wherein the fourth dialect result includes the correspondence between the dialect type and the probability of the dialect type, and the upper limit of the first preset range is less than the lower limit of the second preset range; and determine the vocabulary of the vehicle's input method based on the dialect type with the highest probability in the fourth dialect result.

[0104] Example 4

[0105] Figure 4 This is a flowchart of a method for determining the dictionary of an input method according to Embodiment 2 of the present invention. The technical solution of the present invention is further optimized based on the above optional technical solutions, and a specific method for determining the dictionary of an input method is given.

[0106] like Figure 4 As shown in Embodiment 4 of the present invention, a method for determining the dictionary of an input method specifically includes the following steps:

[0107] S401. If the decibel value of the human voice inside the vehicle is within the first preset range, then collect the facial features of the preset people inside the vehicle, use the preset facial feature library to identify the region to which the facial features belong, and obtain the region result.

[0108] S402. Based on the regional results, the preset regional dialect database, and the second preset requirements, determine the second dialect result.

[0109] S403. Based on the music playback statistics database in the vehicle and the third preset requirements, determine the music dialect category results.

[0110] S404. Based on the results of music dialect types, the preset music dialect database, and the fourth preset requirement, determine the third dialect result.

[0111] S405. Based on the results of the second dialect and the third dialect, determine the result of the target dialect.

[0112] Specifically, the results of the second dialect and the third dialect can be identified as the results of the target dialect.

[0113] S406. From the target dialect results, filter out the probabilities that meet the preset conditions and the corresponding dialect types to obtain the target probability results.

[0114] The target probability result includes the target probability and the target dialect type corresponding to the target probability.

[0115] For example, if the preset condition is to filter the highest probability of the second dialect result and its corresponding language type, and to filter the highest probability of the third dialect result and its corresponding language type, if the filtering result of the second dialect result is Cantonese 45% and the filtering result of the third dialect result is Cantonese 21%, then the target probability results are: Cantonese 45% and Cantonese 21%.

[0116] S407. Based on the preset probability threshold and the target probability result, determine the target dialect type.

[0117] Optionally, determining the target dialect type based on the preset probability threshold and the target probability result includes: determining the dialect types in the target probability result as candidate dialect types; for each candidate dialect type, using a preset weighted summation method, performing a weighted summation calculation on the corresponding probabilities in the target probability result to obtain the candidate dialect result, wherein the candidate dialect result includes the correspondence between the candidate dialect type and the probability of the candidate dialect type; and determining the candidate dialect type corresponding to the probability greater than or equal to the preset probability threshold in the candidate dialect result as the target dialect type.

[0118] For example, if the target probability results are Cantonese 45% and Cantonese 21%, the preset probability threshold is 0.15, and the preset weighted summation method is to add a weighting coefficient of 0.3 to the probability selected from the second dialect results and a weighting coefficient of 0.2 to the probability selected from the second dialect results, and sum the probabilities of the same dialect type, then the probability calculation process of the candidate dialect result is 0.3*0.45+0.2*0.21=0.177. The probability of the candidate dialect result is 0.177, and the candidate dialect type is Cantonese. Since 0.177 is greater than 0.15, the target dialect type is Cantonese.

[0119] The input method vocabulary determination method provided in this embodiment of the invention involves the following steps: If the decibel value of the human voice inside the vehicle falls within a first preset range, the facial features of preset individuals inside the vehicle are collected. Based on these facial features, the region of origin is determined. A second dialect result is determined based on the region of origin, a preset regional dialect database, and a second preset requirement. A music dialect type result is determined based on a music playback statistics database inside the vehicle. A third dialect result is determined based on the music dialect type result, the preset music dialect database, and a fourth preset requirement. A target dialect result is determined based on the second and third dialect results. From this target dialect result, the probability and corresponding dialect type that meet preset conditions are selected to obtain a target probability result. Finally, based on a preset probability threshold and the target probability result, the target dialect type is determined. This allows the determination of the input method vocabulary for the vehicle. By collecting facial features of individuals inside the vehicle and the characteristics of the music played inside, the dialect type used by the user can be inferred before the user uses the input method. This enables timely and accurate recommendation of the input method vocabulary, significantly improving the matching efficiency between the user and the input method vocabulary.

[0120] Based on the above embodiments, the method may further include: if the decibel value of the human voice inside the vehicle is within a second preset range, then identify the dialect type of the human voice to obtain a fourth dialect result, wherein the fourth dialect result includes the correspondence between the dialect type and the probability of the dialect type, and the upper limit of the first preset range is less than the lower limit of the second preset range; and determine the vocabulary of the vehicle's input method based on the dialect type with the highest probability in the fourth dialect result.

[0121] Example 5

[0122] Figure 5 This is a flowchart of a method for determining the dictionary of an input method according to Embodiment 2 of the present invention. The technical solution of the present invention is further optimized based on the above optional technical solutions, and a specific method for determining the dictionary of an input method is given.

[0123] like Figure 5 As shown in Embodiment 5 of the present invention, a method for determining the dictionary of an input method specifically includes the following steps:

[0124] S501. If the decibel value of the human voice inside the vehicle is within the first preset range, then obtain the lip movements of the preset people inside the vehicle, match the lip movements with the sample lip movements in the preset lip movement library, and obtain the similarity result.

[0125] S502. In the similarity results, the dialect types that meet the first preset requirements and the lip shape similarity corresponding to the dialect types are determined as the dialect types and the probabilities of the dialect types in the first dialect results.

[0126] S503. Collect the facial features of preset personnel inside the vehicle, use the preset facial feature library to identify the region to which the facial features belong, and obtain the region result.

[0127] S504. Based on the results of the region, the preset regional dialect database, and the second preset requirements, determine the second dialect result.

[0128] S505. Based on the music playback statistics database in the vehicle and the third preset requirements, determine the music dialect category results.

[0129] S506. Based on the results of music dialect types, the preset music dialect database, and the fourth preset requirement, determine the third dialect result.

[0130] S507. Based on the results of the first dialect, the second dialect, and the third dialect, determine the target dialect result.

[0131] Specifically, the results from the first dialect, the second dialect, and the third dialect can be identified as the target dialect results.

[0132] S508. From the target dialect results, filter out the probabilities that meet the preset conditions and the dialect types corresponding to the probabilities to obtain the target probability results.

[0133] The target probability result includes the target probability and the target dialect type corresponding to the target probability.

[0134] For example, if the preset condition is to filter the results of the first dialect, the second dialect, and the third dialect respectively, and the maximum probability and the corresponding language, if the filtering result of the first dialect is Cantonese 80%, the filtering result of the second dialect is Cantonese 45%, and the filtering result of the third dialect is Hokkien 21%, then the target probability results are: Cantonese 80%, Cantonese 45%, and Hokkien 21%.

[0135] S509. Based on the preset probability threshold and the target probability result, determine the target dialect type.

[0136] Optionally, determining the target dialect type based on the preset probability threshold and the target probability result includes: determining the dialect types in the target probability result as candidate dialect types; for each candidate dialect type, using a preset weighted summation method, performing a weighted summation calculation on the corresponding probabilities in the target probability result to obtain the candidate dialect result, wherein the candidate dialect result includes the correspondence between the candidate dialect type and the probability of the candidate dialect type; and determining the candidate dialect type corresponding to the probability greater than or equal to the preset probability threshold in the candidate dialect result as the target dialect type.

[0137] For example, if the target probability results are Cantonese 80%, Cantonese 45%, and Hokkien 21%, the preset probability threshold is 0.5, and the preset weighted summation method is to add a weighting coefficient of 0.6 to the probability selected from the first dialect result, a weighting coefficient of 0.2 to the probability selected from the second dialect result, and a weighting coefficient of 0.2 to the probability selected from the third dialect result, and sum the probabilities of the same dialect type, then the probability calculation process of the candidate dialect result is: 0.6*0.8+0.45*0.2=0.57, 0.2*0.21=0.042. The probabilities of the candidate dialect results are 0.57 and 0.042, and the candidate dialect types are Cantonese and Hokkien. Since 0.57 is greater than 0.5, the target dialect type is Cantonese.

[0138] The dictionary determination method for the input method provided in this embodiment of the invention, if the decibel value of the human voice inside the vehicle falls within a first preset range, then the lip movements of a preset person inside the vehicle are acquired. Based on the similarity between the lip movements and sample lip movements in a preset lip movement library, a first dialect result is determined. The facial features of the preset person inside the vehicle are also acquired, and their regional origin is determined based on these facial features. Based on the regional origin result, a preset regional dialect library, and a second preset requirement, a second dialect result is determined. Then, based on a music playback statistics library inside the vehicle, a music dialect category result is determined. Finally, based on the music dialect category result, a preset music dialect library, and a fourth preset requirement, a third dialect result is determined. Based on the results of the first dialect, the second dialect, and the third dialect, the target dialect result is determined. From this target dialect result, the probability of meeting preset conditions and the corresponding dialect type are selected to obtain the target probability result. Finally, based on the preset probability threshold and the target probability result, the target dialect type is determined, thus determining the vocabulary of the vehicle's input method. By collecting the lip and facial features of people in the vehicle and the features of the music played in the car, the dialect type used by the user can be inferred before the user uses the input method, and the vocabulary of the input method can be recommended to the user in a timely and accurate manner, which greatly improves the matching efficiency between the user and the vocabulary of the input method.

[0139] Based on the above embodiments, the method may further include: if the decibel value of the human voice inside the vehicle is within a second preset range, then identify the dialect type of the human voice to obtain a fourth dialect result, wherein the fourth dialect result includes the correspondence between the dialect type and the probability of the dialect type, and the upper limit of the first preset range is less than the lower limit of the second preset range; and determine the vocabulary of the vehicle's input method based on the dialect type with the highest probability in the fourth dialect result.

[0140] Example 6

[0141] Figure 6 This is a schematic diagram of a dictionary determination device for an input method provided in Embodiment 3 of the present invention. Figure 6As shown, the device includes: a target dialect result determination module 601, a target dialect type determination module 602, and a first lexicon determination module 603, wherein:

[0142] The target dialect result determination module is used to determine the target dialect result based on at least two of the first dialect result, the second dialect result, and the third dialect result if the decibel value of the human voice in the vehicle is within a first preset range. The first dialect result is determined based on the lip movements of a preset person in the vehicle, the second dialect result is determined based on the facial features of the preset person, and the third dialect result is determined based on the music playback statistics database in the vehicle. The dialect result includes the correspondence between dialect types and the probability of the dialect types.

[0143] The target dialect type determination module is used to determine the target dialect type based on the target dialect result and a preset probability threshold.

[0144] The first lexicon determination module is used to determine the lexicon of the vehicle's input method based on the target dialect type.

[0145] The input method dictionary determination device provided in this embodiment of the invention can determine the target dialect result based on at least two items from a preset database of facial features, lip movements, and music playback statistics when the dialect of a person in a vehicle cannot be determined by speech recognition (i.e., the decibel value is within a first preset range). Then, based on the target dialect result and a preset probability threshold, the target dialect type is obtained. Finally, the dictionary of the vehicle's input method can be determined based on the target dialect type. The entire process does not require manual operation by the user, solving the problem that the matching of dialect dictionaries in input methods relies too much on manual selection by the user. It intelligently determines the dictionary of the input method for the user, improving the user experience.

[0146] Optionally, the device may also include:

[0147] The fourth dialect result determination module is used to identify the dialect type of the human voice when the decibel value of the human voice in the vehicle is within the second preset range, and obtain the fourth dialect result. The fourth dialect result includes the correspondence between the dialect type and the probability of the dialect type, and the upper limit of the first preset range is less than the lower limit of the second preset range.

[0148] The second dictionary determination module is used to determine the dictionary of the vehicle's input method based on the dialect type with the highest probability in the fourth dialect results.

[0149] Optionally, the device may also include:

[0150] The similarity result determination module is used to obtain the lip movements of a preset person in the vehicle, match the lip movements with the sample lip movements in a preset lip movement library, and obtain a similarity result. The preset lip movement library contains the correspondence between dialect types and lip movement samples. The similarity result includes the lip movement similarity between the lip movement and the sample lip movement, as well as the dialect type corresponding to the lip movement similarity.

[0151] The first dialect result determination module is used to determine the dialect type and the lip shape similarity corresponding to the dialect type in the similarity result as the probability of the dialect type in the first dialect result.

[0152] Optionally, the device may also include:

[0153] The region determination module is used to collect the facial features of preset people in the vehicle, identify the region of the facial features using a preset facial feature library, and obtain the region result. The preset facial feature library contains the correspondence between facial feature samples and their respective regions. The region result includes the facial feature similarity between the facial feature and the sample facial features and the region corresponding to the facial feature similarity.

[0154] The second dialect result determination module is used to determine the second dialect result based on the regional result, the preset regional dialect database, and the second preset requirements. The preset regional dialect database contains the correspondence between the regional name, the dialect type, and the probability of using the dialect type. The probability of the dialect type in the second dialect result is determined based on the facial similarity and the probability of use.

[0155] Optionally, the device may also include:

[0156] The music dialect category determination module is used to determine the music dialect category result based on the music playback statistics library in the vehicle and the third preset requirements. The music playback statistics library contains the correspondence between music tracks, playback times and dialect categories. The music dialect category result includes the correspondence between the dialect category of the music track and the proportion of the dialect category.

[0157] The third-party dialect result determination module is used to determine the third-party dialect result based on the music dialect type result, the preset music dialect library, and the fourth preset requirement. The preset music dialect library includes the correspondence between the dialect type of the music track, the dialect type used by the listener, and the probability of the dialect type used by the listener. The probability of the dialect type in the third-party dialect result is determined according to the proportion of the dialect type and the probability of the dialect type used by the listener.

[0158] Optional, the target dialect type determination module includes:

[0159] The target probability result determination unit is used to filter the probability that meets the preset conditions and the dialect type corresponding to the probability from the target dialect result to obtain the target probability result, wherein the target probability result includes the target probability and the target dialect type corresponding to the target probability;

[0160] The target dialect type determination unit is used to determine the target dialect type based on a preset probability threshold and the target probability result.

[0161] Furthermore, determining the target dialect type based on the preset probability threshold and the target probability result includes: determining the dialect types in the target probability result as candidate dialect types; for each candidate dialect type, using a preset weighted summation method, performing a weighted summation calculation on the corresponding probabilities in the target probability result to obtain the candidate dialect result, wherein the candidate dialect result includes the correspondence between the candidate dialect type and the probability of the candidate dialect type; and determining the candidate dialect type corresponding to the probability greater than or equal to the preset probability threshold in the candidate dialect result as the target dialect type.

[0162] The dictionary determination device for input methods provided in this embodiment of the invention can execute the dictionary determination method for input methods provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0163] Example 7

[0164] Figure 7 A schematic diagram of the structure of a vehicle 70 that can be used to implement embodiments of the present invention is shown. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the invention described and / or claimed herein.

[0165] like Figure 7 As shown, the vehicle 70 includes at least one processor 71 and a memory, such as a read-only memory (ROM) 72 and a random access memory (RAM) 73, communicatively connected to the at least one processor 71. The memory stores computer programs executable by the at least one processor. The processor 71 can perform various appropriate actions and processes based on the computer program stored in the ROM 72 or loaded from storage unit 78 into the RAM 73. The RAM 73 can also store various programs and data required for the operation of the vehicle 70. The processor 71, ROM 72, and RAM 73 are interconnected via a bus 74. An input / output (I / O) interface 75 is also connected to the bus 74.

[0166] Multiple components in vehicle 70 are connected to I / O interface 75, including: input unit 76, such as keyboard, mouse, etc.; output unit 77, such as various types of displays, speakers, etc.; storage unit 78, such as disk, optical disk, etc.; and communication unit 79, such as network card, modem, wireless transceiver, etc. Communication unit 79 allows vehicle 70 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0167] Processor 71 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 71 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Processor 71 performs the various methods and processes described above, such as the dictionary determination method for an input method.

[0168] In some embodiments, the input method's dictionary determination method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 78. In some embodiments, part or all of the computer program may be loaded and / or installed on vehicle 70 via ROM 72 and / or communication unit 79. When the computer program is loaded into RAM 73 and executed by processor 71, one or more steps of the input method's dictionary determination method described above may be performed. Alternatively, in other embodiments, processor 71 may be configured to perform the input method's dictionary determination method by any other suitable means (e.g., by means of firmware).

[0169] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0170] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0171] The computer equipment provided above can be used to execute the dictionary determination method of the input method provided in any of the above embodiments, and has corresponding functions and beneficial effects.

[0172] Example 8

[0173] In the context of this invention, the computer-readable storage medium may be a tangible medium, and the computer-executable instructions, when executed by a computer processor, are used to perform a dictionary determination method for an input method, the method comprising:

[0174] If the decibel value of the human voice inside the vehicle falls within a first preset range, then a target dialect result is determined based on at least two of the first dialect result, the second dialect result, and the third dialect result. The first dialect result is determined based on the lip movements of a preset person inside the vehicle, the second dialect result is determined based on the facial features of the preset person, and the third dialect result is determined based on the music playback statistics database inside the vehicle. The dialect result includes the correspondence between dialect types and the probability of each dialect type.

[0175] Based on the target dialect results and a preset probability threshold, the target dialect type is determined;

[0176] The vocabulary of the vehicle's input method is determined based on the target dialect.

[0177] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by, or in conjunction with, an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0178] The computer equipment provided above can be used to execute the dictionary determination method of the input method provided in any of the above embodiments, and has corresponding functions and beneficial effects.

[0179] It is worth noting that in the embodiments of the above-mentioned input method word database determination device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the scope of protection of the present invention.

[0180] Note that the above description is merely a preferred embodiment of the present invention and the technical principles employed. Those skilled in the art will understand that the present invention is not limited to the specific embodiments described herein, and various obvious changes, readjustments, and substitutions can be made without departing from the scope of protection of the present invention. Therefore, although the present invention has been described in detail through the above embodiments, the present invention is not limited to the above embodiments, and may include many other equivalent embodiments without departing from the concept of the present invention, the scope of which is determined by the scope of the appended claims.

Claims

1. A method for determining the dictionary of an input method, characterized in that, Applied to vehicles, the method includes: If the decibel value of the human voice inside the vehicle falls within a first preset range, then a target dialect result is determined based on at least two of the first dialect result, the second dialect result, and the third dialect result. The first dialect result is determined based on the lip movements of a preset person inside the vehicle, the second dialect result is determined based on the facial features of the preset person, and the third dialect result is determined based on the music playback statistics database inside the vehicle. The dialect result includes the correspondence between dialect types and the probability of each dialect type. Based on the target dialect results and a preset probability threshold, the target dialect type is determined; Based on the target dialect type, determine the vocabulary of the vehicle's input method; If the decibel value of the human voice inside the vehicle is within the second preset range, the dialect type of the human voice is identified to obtain a fourth dialect result. The fourth dialect result includes the correspondence between the dialect type and the probability of the dialect type. The upper limit of the first preset range is less than the lower limit of the second preset range. The vocabulary of the vehicle's input method is determined based on the dialect type with the highest probability in the fourth dialect result.

2. The method according to claim 1, further comprising: The lip movements of a preset person inside the vehicle are obtained, and the lip movements are matched with sample lip movements in a preset lip movement library to obtain a similarity result. The preset lip movement library contains the correspondence between dialect types and sample lip movements. The similarity result includes the lip movement similarity between the lip movement and the sample lip movement, as well as the dialect type corresponding to the lip movement similarity. The dialect types that meet the first preset requirements and the lip shape similarity corresponding to the dialect types in the similarity results are determined as the probability of the dialect type in the first dialect result and the dialect type.

3. The method according to claim 1, further comprising: The system collects the facial features of a preset person inside the vehicle, uses a preset facial feature library to identify the region to which the facial features belong, and obtains the region result. The preset facial feature library contains the correspondence between sample facial features and their respective regions. The region result includes the facial feature similarity between the facial feature and the sample facial features, as well as the region to which the facial feature similarity belongs. Based on the regional results, the preset regional dialect database, and the second preset requirements, a second dialect result is determined. The preset regional dialect database contains the correspondence between the region, the dialect type, and the probability of using the dialect type. The probability of the dialect type in the second dialect result is determined based on the facial similarity and the probability of use.

4. The method according to claim 1, further comprising: Based on the music playback statistics database in the vehicle and the third preset requirement, the music dialect category results are determined. The music playback statistics database contains the correspondence between music tracks, playback times, and dialect categories. The music dialect category results include the correspondence between the dialect category of the music track and the proportion of the dialect category. Based on the results of the music dialect types, the preset music dialect database, and the fourth preset requirement, a third dialect result is determined. The preset music dialect database includes the correspondence between the dialect types of the music tracks, the dialect types used by the listeners, and the probability of the dialect types used by the listeners. The probability of the dialect types in the third dialect result is determined based on the proportion of the dialect types and the probability of the dialect types used by the listeners.

5. The method according to any one of claims 1-4, characterized in that, The step of determining the target dialect type based on the target dialect result and a preset probability threshold includes: From the target dialect results, the probabilities that meet preset conditions and the dialect types corresponding to the probabilities are filtered to obtain the target probability results, wherein the target probability results include the target probability and the target dialect type corresponding to the target probability; The target dialect type is determined based on the preset probability threshold and the target probability result.

6. The method according to claim 5, characterized in that, The determination of the target dialect type based on the preset probability threshold and the target probability result includes: The dialect types in the target probability results are determined as candidate dialect types. For each candidate dialect type, the corresponding probabilities in the target probability results are weighted and summed using a preset weighted summation method to obtain the candidate dialect results. The candidate dialect results include the correspondence between the candidate dialect types and the probabilities of the candidate dialect types. The candidate dialect types whose probabilities are greater than or equal to a preset probability threshold are identified as the target dialect types.

7. A dictionary determination device for an input method, characterized in that, include: The target dialect result determination module is used to determine the target dialect result based on at least two of the first dialect result, the second dialect result, and the third dialect result if the decibel value of the human voice in the vehicle is within a first preset range. The first dialect result is determined based on the lip movements of a preset person in the vehicle, the second dialect result is determined based on the facial features of the preset person, and the third dialect result is determined based on the music playback statistics database in the vehicle. The dialect result includes the correspondence between dialect types and the probability of the dialect types. The target dialect type determination module is used to determine the target dialect type based on the target dialect result and a preset probability threshold. The first lexicon determination module is used to determine the lexicon of the vehicle's input method based on the target dialect type; The fourth dialect result determination module is used to identify the dialect type of the human voice when the decibel value of the human voice in the vehicle is within the second preset range, and obtain the fourth dialect result. The fourth dialect result includes the correspondence between the dialect type and the probability of the dialect type, and the upper limit of the first preset range is less than the lower limit of the second preset range. The second dictionary determination module is used to determine the dictionary of the vehicle's input method based on the dialect type with the highest probability in the fourth dialect results.

8. A vehicle, characterized in that, The vehicles include: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the dictionary determination method of the input method according to any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that, when executed by a processor, implement the dictionary determination method of the input method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Input method switch method and device

    CN106774976A

  • Language identification method and device, language identification training method and device, medium and terminal

    CN108389573A