Voice broadcasting method and device, electronic equipment and storage medium
By logging in to the associated voice information account in the device, and automatically switching the language type of voice broadcast based on the number of usages of the language type and user feedback, the problem that the device cannot switch according to user habits is solved, and the user experience and device adaptability are improved.
Patent Information
- Application Number
- CN202510578391.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-06
- Publication Date
- 2025-08-26
AI Technical Summary
Existing devices require users to manually switch language types when broadcasting voice, resulting in cumbersome dialect broadcast function and the inability to automatically switch according to the language type used by users, reducing users' interest and experience.
By logging in to the associated voice information account in response to the user's login operation, the target language type of voice broadcast is automatically determined according to the number of times used by each language type in the account, and the number of times used by the voice control command is iteratively updated, and the number of times used by the language type is combined with the preset threshold and user feedback, the automatic switching of language type is realized.
It realizes automatic switching of voice broadcasts based on the language type used by users, which improves user experience and interest, simplifies operational processes, and improves equipment adaptability.
Smart Images

Figure CN120544556A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of voice technology, and specifically to a voice broadcasting method, device, electronic device and storage medium. Background Art
[0002] With the development of artificial intelligence, human-computer interaction through voice is becoming increasingly common. This process utilizes speech recognition and voice broadcast technologies. Speech recognition technology allows devices to convert voice signals into understandable commands, while voice broadcast technology converts text or other information into understandable speech. In other words, devices can "understand" human speech through speech recognition technology, and can also "respond" in human language through voice broadcast technology.
[0003] However, influenced by various factors, including geography and culture, human language types vary greatly. For example, Mandarin and local dialects (such as Shandong and Sichuan dialects) differ significantly in pronunciation, intonation, and vocabulary. In other words, people in different regions use different types of language. Furthermore, with the rise of new media, young people are increasingly interested in various language types and are imitating them.
[0004] Currently, devices can automatically understand a variety of languages. However, when the device performs voice announcements, users are often required to manually switch the language. This makes the dialect announcement function cumbersome to operate and unfavorable for use in dialects. In addition, each user has their own preferred language. If the user does not manually switch the voice announcement language, the default language is often used for announcements, which is likely to be inconsistent with the user's accustomed language, thus reducing the user's interest in and experience with the voice announcement function.
[0005] It should be pointed out that the information disclosed in the background technology section of this application is only intended to deepen the understanding of the general background technology of this application, and should not be regarded as an admission or any form of implication that the information constitutes prior art already known to those skilled in the art. Summary of the Invention
[0006] In view of this, the present application provides a voice broadcast method, device, electronic device and storage medium to help solve the problem that when the user performs human-computer interaction through voice, the device cannot automatically switch the language type of voice broadcast according to the language type that the user is accustomed to using.
[0007] In a first aspect, an embodiment of the present application provides a voice broadcast method, applied to a voice broadcast device, the method comprising: In response to a login operation, logging into a corresponding voice information account, the voice information account being associated with N language type information, each piece of language type information being used to represent a language type and a usage count corresponding to the language type, each piece of the N language type information corresponding to a different language type, where N ≥ 1; The target language type of the voice broadcast is determined according to the number of times each of the language types in the N language type information is used.
[0008] In a possible implementation, after determining the target language type for voice broadcast based on the number of times each language type in the N pieces of language type information is used, the method further includes: Determining, according to the voice control instruction, a language type corresponding to the voice control instruction, wherein the voice control instruction is used to control the voice broadcast device to perform voice broadcast; Iteratively updating the number of times the language type corresponding to the voice control instruction is used to the voice information account according to the language type corresponding to the voice control instruction; The target language type for voice broadcast is determined according to the number of times each of the N language types is used after iterative update.
[0009] In a possible implementation, determining the target language type of the voice announcement according to the number of times each language type in the N pieces of language type information is used includes: comparing the usage counts of each of the N language types; The language type that is used the most times among the N language types is determined as the target language type for voice broadcast.
[0010] In a possible implementation, determining the target language type of the voice announcement according to the number of times each language type in the N pieces of language type information is used includes: Determining, based on the number of times each of the language types in the N language type information is used, a single type usage ratio corresponding to each of the N language types, where the single type usage ratio is a ratio of the number of times corresponding to the language type to the number of times all language types corresponding to the voice information account are used; When there are M language types whose corresponding single type usage ratios are greater than or equal to the corresponding preset threshold, the language type corresponding to the maximum single type usage ratio is determined as the target language type for voice broadcast, wherein the language type corresponding to the maximum single type usage ratio is the language type among the M language types, wherein 1≤M≤N.
[0011] In a possible implementation, determining the target language type of the voice announcement according to the number of times each language type in the N pieces of language type information is used includes: Determining, based on the number of times each of the language types in the N language type information is used, a single type usage ratio corresponding to each of the N language types, where the single type usage ratio is a ratio of the number of times corresponding to the language type to the number of times all language types corresponding to the voice information account are used; When the single-type usage ratio corresponding to each of the language types is less than the corresponding preset threshold, the preset language type is determined as the target language type for voice broadcast.
[0012] In one possible implementation, when there are M language types corresponding to the single type usage proportions being greater than or equal to the corresponding preset threshold, determining the language type corresponding to the largest single type usage proportion as the target language type for voice broadcast includes: When the usage ratio of the single type corresponding to the M language types is greater than or equal to the corresponding preset threshold, a prompt message is output, where the prompt message is used to remind the user whether to switch the target language type of the voice broadcast; If the first response information input by the user is received, the language type corresponding to the largest single type usage ratio is determined as the target language type of the voice broadcast, and the first response information is used to indicate that the user allows the target language type of the voice broadcast to be switched.
[0013] In one possible implementation, when there are M language types corresponding to the single type usage proportions being greater than or equal to the corresponding preset threshold, determining the language type corresponding to the largest single type usage proportion as the target language type for voice broadcast includes: When the usage ratio of the single type corresponding to the M language types is greater than or equal to the corresponding preset threshold, a prompt message is output, where the prompt message is used to remind the user whether to switch the target language type of the voice broadcast; If the second response information input by the user is received, it is determined that the target language type of the voice broadcast is the voice type currently used for the voice broadcast, and the second response information is used to indicate that the user does not allow the target language type of the voice broadcast to be switched.
[0014] In one possible implementation, when there are M language types corresponding to the single type usage proportions being greater than or equal to the corresponding preset threshold, determining the language type corresponding to the largest single type usage proportion as the target language type for voice broadcast includes: When the usage ratio of the single type corresponding to the M language types is greater than or equal to the corresponding preset threshold, a prompt message is output, where the prompt message is used to remind the user whether to switch the target language type of the voice broadcast; If a second response message input by the user is received, the target language type of the voice broadcast is determined to be the language type corresponding to the response message input by the user last time, and the second response message is used to indicate that the user does not allow the target language type of the voice broadcast to be switched.
[0015] In a possible implementation, when there are M language types whose corresponding single-type usage proportions are greater than or equal to a corresponding preset threshold, outputting a prompt message includes: When there are M language types corresponding to the single-type usage ratios greater than or equal to the corresponding preset threshold, determining whether the voice information account has outputted prompt information within a preset time period, where the time period is a time interval from a first moment to a current moment, and the first moment is earlier than the current moment; If the voice information account outputs prompt information within a preset time period, determining the target language type of the voice broadcast to be the language type currently used for the voice broadcast; If the voice information account does not output prompt information within a preset time period, the prompt information is output.
[0016] In one possible implementation, in response to the login operation, logging into the corresponding voice information account includes: In response to a voice wake-up instruction, a corresponding voice information account is logged in, the voice information account is associated with the voiceprint corresponding to the voice wake-up instruction, and the voice wake-up instruction is used to activate the voice broadcast function of the voice broadcast device.
[0017] In a possible implementation, the iteratively updating, according to the language type corresponding to the voice control instruction, the usage count of the language type corresponding to the voice control instruction includes: Determine, according to the voice control instruction, a target voice information account corresponding to the voice control instruction, wherein the target voice information account is a voice information account corresponding to the voiceprint of the voice control instruction; According to the language type corresponding to the voice control instruction, the number of times the language type corresponding to the voice control instruction is used is iteratively updated to the target voice information account.
[0018] In one possible implementation, in response to the login operation, logging into the corresponding voice information account includes: If the adaptive voice broadcast has been set, then in response to the login operation, the corresponding voice information account is logged in, and the adaptive voice broadcast is used to indicate that the voice broadcast device automatically switches the language type of the voice broadcast; If the adaptive voice broadcast is not set, the target language type of the voice broadcast is determined to be the language type currently used for the voice broadcast.
[0019] In a second aspect, an embodiment of the present application provides a voice broadcast device, which is applied to a voice broadcast device, and the device includes: a voice information account login module, in response to a login operation, logging into a corresponding voice information account, wherein the voice information account is associated with N language type information, each piece of language type information being used to represent a language type and a usage count corresponding to the language type, each of the N pieces of language type information corresponding to a different language type, wherein N ≥ 1; The target language type determination module is used to determine the target language type of the voice broadcast according to the number of times each language type in the N language type information is used.
[0020] In a third aspect, an embodiment of the present application provides an electronic device, including: processor; Memory; and a computer program, wherein the computer program is stored in the memory, and when the computer program is executed by the processor, the electronic device executes the method described in any one of the first aspects.
[0021] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which includes a stored program, wherein when the program is running, the device where the computer-readable storage medium is located is controlled to execute the method described in any one of claims 1 to 11.
[0022] In an embodiment of the present application, in response to a user's login operation, a corresponding voice information account is logged in. The voice information account is associated with at least one piece of language type information, wherein each piece of language type information represents a language type and the number of times the language type is used. Based on the number of times each language type in the at least one piece of language type information is used, the target language type for voice broadcast is determined. This solves the problem of the device being unable to automatically switch the language type for voice broadcast based on the user's accustomed language type when the user is interacting with the device via voice. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0024] Figure 1 A schematic diagram of an application scenario provided for an embodiment of the present application.
[0025] Figure 2 A flowchart of a voice broadcast method provided in an embodiment of the present application.
[0026] Figure 3 A schematic diagram of the association relationship between a voice information account and language type information provided in an embodiment of the present application.
[0027] Figure 4 A schematic diagram of logging into a voice information account using a voice wake-up command provided in an embodiment of the present application.
[0028] Figure 5 A schematic diagram of iterative updating of language type information provided in an embodiment of the present application.
[0029] Figure 6 Another schematic diagram of iterative updating of language type information provided in an embodiment of the present application.
[0030] Figure 7 A structural diagram of a voice broadcasting device provided in an embodiment of the present application.
[0031] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0032] In order to better understand the technical solution of the present application, the embodiments of the present application are described in detail below with reference to the accompanying drawings.
[0033] It should be clear that the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.
[0034] The terms used in the embodiments of the present application are for the purpose of describing specific embodiments only and are not intended to limit the present application. The singular forms "a", "an", "the" and "the" used in the embodiments of the present application and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.
[0035] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. Furthermore, the character " / " in this document generally indicates an "or" relationship between the associated objects.
[0036] With the development of artificial intelligence, human-computer interaction through voice is becoming increasingly common. This process utilizes speech recognition and voice broadcast technologies. Speech recognition technology allows devices to convert voice signals into understandable commands, while voice broadcast technology converts text or other information into understandable speech. In other words, devices can "understand" human speech through speech recognition technology, and can also "respond" in human language through voice broadcast technology.
[0037] See also Figure 1 , provides a schematic diagram of an application scenario for the embodiment of this application. Figure 1 As shown, the voice interaction device 10 includes: a controller 101, a voice recognition device 102 and a voice broadcast device 103. The voice interaction device can be a smart speaker, a car voice interaction assistant, a smart home assistant, etc., which is not specifically limited in the embodiments of the present application.
[0038] Specifically, the controller 101 can control the voice recognition device 102 to receive and recognize voice commands, and can generate response information based on the voice commands. For example, if the voice command received and recognized by the voice recognition device 102 is "Please open the window", the controller 101 can generate a text type or other type of response information of "OK, I will open the window for you right away" based on the voice command "Please open the window". The controller 101 can control the voice broadcast device 103 to perform voice broadcast based on the response information. The voice recognition device 102 can receive voice commands and "translate" the voice command recognition into a text type or other type of machine language that the controller 101 can understand. The voice broadcast device 103 can "translate" the response information generated by the controller 101 into a voice that humans can understand and broadcast it.
[0039] in addition, Figure 1 The voice interaction device shown in the figure is only an exemplary description and should not be regarded as limiting the scope of protection of this application.
[0040] However, influenced by various factors, including geography and culture, human language types vary greatly. For example, Mandarin and local dialects (such as Shandong and Sichuan dialects) differ significantly in pronunciation, intonation, and vocabulary. In other words, people in different regions use different types of language. Furthermore, with the rise of new media, young people are increasingly interested in various language types and are imitating them.
[0041] Currently, devices can automatically understand a variety of languages. However, when the device performs voice announcements, users are often required to manually switch the language. This makes the dialect announcement function cumbersome to operate and unfavorable for its use. Furthermore, each user has their own preferred language when speaking. If the user does not manually switch the voice announcement language, the default language is often used, which is likely inconsistent with the user's conversational habits, thus reducing the user's interest in and experience with the voice announcement function.
[0042] To address the above issues, in an embodiment of the present application, in response to a user's login operation, a corresponding voice information account is logged in. The voice information account is associated with at least one piece of language type information, where each piece of language type information represents a language type and the number of times that language type is used. Based on the number of times each language type in the at least one piece of language type information is used, the target language type for voice broadcast is determined. This solves the problem of the device being unable to automatically switch the language type for voice broadcast based on the user's accustomed language type when the user is interacting with the device via voice.
[0043] Specifically, a detailed description is given below with reference to the accompanying drawings and specific embodiments.
[0044] See also Figure 2 The embodiment provides a flow chart of a voice broadcast method. The method can be applied to Figure 1 The controller shown, such as Figure 2 As shown, it specifically includes the following steps.
[0045] Step S201: In response to the login operation, log in to the corresponding voice information account.
[0046] It should be noted that the login operation can be a direct response to the user's login operation, such as the user manually entering the account and password, fingerprint recognition, voiceprint recognition, etc., which is not limited in the embodiments of the present application. In addition, the login operation can also be an indirect response to the user's login operation. For example, the user logs in on the first device, the first device exchanges information with the second device, and the first device sends the user's login information to the second device. Then, the second device responds to the user's login operation and logs in to the user's corresponding account.
[0047] In an embodiment of the present application, a voice information account is associated with N language type information, each language type information represents a language type and the number of uses corresponding to the language type, and each of the N language type information corresponds to a different language type, where N≥1.
[0048] The language type may be a standard language type, such as Mandarin, or a non-standard language type, such as Shandong dialect, Sichuan dialect, and other local dialects in various regions.
[0049] When a user uses the voice interaction feature, they enter a voice command. Each voice command has a corresponding language type. Each time a user enters a voice command, the corresponding language type is used once. The usage count for the corresponding language type in the user's corresponding voice messaging account increments by 1. The usage count refers to the number of times the user has used each language type in their voice messaging account from the time the account was created to the current time.
[0050] See also Figure 3 , provides a schematic diagram of the association relationship between voice information account and language type information for the embodiment of this application. Figure 3 As shown, database 30 includes multiple data tables 301, including a voice information account table 3011 and a language type information table 3012. Voice information account table 3011 includes fields such as user account ID, user account, and password. Language type information table 3012 includes fields such as user account ID, language type, and usage count.
[0051] The database can be installed and deployed locally or on a remote server, which is not limited in the embodiments of the present application. In addition, the database is a relational database, which can be SQL Server, MySQL, Oracle, etc., which is not limited in the embodiments of the present application.
[0052] It's important to note that a user account ID is a unique code throughout the system. For example, the user account ID "38f34aa7-bebc-0375-c4d3-4817162296b6" will not exist in the entire system. In other words, a user account ID can be used to uniquely identify a user account.
[0053] For example, Figure 3As shown, in voice information account table 3011, the user account "11111" corresponds to the user account ID "38f34aa7-bebc-0375-c4d3-4817162296b6." This user account ID can be used to query the language type information table 3012, finding three pieces of information: language type 1, usage count 400; language type 2, usage count 600; and language type 3, usage count 10.
[0054] It should be pointed out that Figure 3 The schematic diagram of the association relationship between a voice information account and language type information is merely an exemplary description of the association relationship between a voice information account and language type information.
[0055] In addition, the number of uses in the language type information table 3012 can be recorded directly by setting the "number of uses" field in the table, or the number of uses of each language type corresponding to each user account ID can be determined through a virtual table (view). Among them, view refers to the view technology in the database, which is a virtual table. Like a real table, a view contains a series of named row and column data. However, some fields in the view may not exist in the various data tables of the database, but are dynamically generated based on existing fields through query technologies such as aggregate functions. Table 1 provides a language type information table for an embodiment of the present application.
[0056] Table 1:
[0057] It should be noted that Table 1 is only an exemplary description of the language type information table.
[0058] For example, through the view technology of the database, Table 2 can be determined based on Table 1. Table 2 provides another language type information table for the embodiment of the present application.
[0059] Table 2:
[0060] Table 2 is not a real data table in the database, but is generated through view technology based on the dynamic query of Table 1.
[0061] In other words, after a user logs in, a voice messaging account and its unique authentication information can be uniquely identified based on the user's login information. Using the voice messaging account's unique authentication information, such as the user account ID, the language type information corresponding to that user account can be linked and queried from other data tables. It is understood that similarly, other information corresponding to the user account, such as fingerprint information and voiceprint information, can also be linked and queried based on the user account.
[0062] As described above, there are multiple ways for a user to log in to a corresponding voice information account.
[0063] In one possible implementation, in response to a voice wake-up instruction, a corresponding voice information account is logged in, and the voice information account is associated with the voiceprint corresponding to the voice wake-up instruction, wherein the voice wake-up instruction is used to activate the voice broadcast function of the voice broadcast device.
[0064] The voice wake-up command is the first voice command a user needs to enter when using the voice interaction feature. Using the voiceprint corresponding to the voice wake-up command to log in to the corresponding voice message account reduces the number of steps required to use the voice interaction feature, facilitating operation and improving the user experience.
[0065] A voiceprint refers to the sound wave spectrum revealed by electroacoustic instrumentation when analyzing each person's voice. It's important to note that the size and shape of the vocal organs (such as lips, teeth, larynx, tongue, and fangs) used when speaking vary from person to person. Consequently, the sound wave spectrum corresponding to each person's speech will also vary. In other words, each person has a unique voiceprint, which can be used to identify a unique user.
[0066] Reference Figure 4 , provides a schematic diagram of a voice wake-up command to log in to a voice information account in the embodiment of the present application. Figure 4 As shown, user A enters the voice wake-up command "Xiao H." After receiving the voice wake-up command, voice broadcast device 401 identifies the voiceprint information corresponding to the voice wake-up command. The voiceprint information is then sent to database 402, which contains multiple data tables 4021: a voiceprint information table, a voice information account table, and a language type information table. As described above, the voiceprint information table is linked to the voice information account table via the user account ID, and the language type information table is linked to the voice information account table via the user account ID. In other words, the user's account ID can be determined from the voiceprint information in the voiceprint information table. Based on the user account ID, the user's login information (e.g., username, password) can be retrieved from the voice information account table, enabling login to the voice information account using the voiceprint information.
[0067] It is understood that the user's voiceprint information can be determined based on the voice wake-up command input by the user. Based on the voiceprint information, the user's corresponding voice information account can be associated and queried, thereby realizing the login operation of the voice information account.
[0068] It should be noted that when using voiceprint information to log in to a voice messaging account, the voiceprint information must be pre-registered or associated with an existing voice messaging account. In an embodiment of the present application, when the voiceprint information is not registered or associated with an existing voice messaging account, the voiceprint information can be associated with a public voice messaging account. The public voice messaging account is a special voice messaging account that allows multiple different voiceprint information to complete the login operation. Through the public voice messaging account, users who do not have registered voiceprint information can experience the automatic switching function of the voice broadcast language type.
[0069] Step S202: determining a target language type for voice announcement according to the usage count of each language type in the at least one language type information.
[0070] In a possible implementation, the usage counts of each language type in at least one language type are compared; and the language type in the at least one language type that is used the most times is determined as the target language type for voice broadcast.
[0071] In actual applications, each user has a language type that they are familiar with or accustomed to using. Therefore, during voice interaction, it is generally considered that the language type used most frequently by a user is the language type that the user is accustomed to using.
[0072] For example, Table 3 provides another language type information table for an embodiment of the present application.
[0073] Table 3:
[0074] It should be noted that Table 3 is only an exemplary description of the language type information associated with the voice information account corresponding to the user account ID "38f34aa7-bebc-0375-c4d3-4817162296b6".
[0075] As shown in Table 3, the language type information table associated with this user contains three language type records: the first language type has been used 200 times, the second language type has been used 700 times, and the third language type has been used 100 times. Since the second language type has been used the most times, when this user uses the voice broadcast function, the target language for the voice broadcast is determined to be the second language type.
[0076] The language most frequently used in voice messaging accounts is determined as the target language for voice broadcasts. This allows users to choose the language they prefer based on their preferred language, making voice broadcasts more familiar and enhancing their experience.
[0077] After a user enters a voice command, a language recognition model (hereinafter referred to as the model) needs to identify the language type corresponding to the voice command. However, in practice, different language types present varying degrees of difficulty for the model to recognize. For example, in real life, Beijing dialect and Mandarin are very similar, making Beijing dialect more challenging for the model to recognize. Cantonese, on the other hand, has distinct characteristics and may be easier to recognize.
[0078] After the model identifies the language type corresponding to the voice command, the number of times the corresponding language type is used in the user's corresponding voice information account is incremented by 1. However, for some languages that are more difficult to recognize, the recognition results may be inaccurate. Therefore, determining the target language type for voice broadcast based solely on the most frequently used language type in the user's corresponding voice information account may deviate from the user's actual habits.
[0079] For example, when the user actually inputs voice commands, the first language type, such as Mandarin, is used 100 times, and the second language type, such as Beijing dialect, is used 200 times. However, affected by the recognition accuracy of the model, in the voice information account corresponding to the user, the first language type is used 160 times, and the second language type is used 140 times. According to the language type that the user is actually accustomed to, the second language type should be used for voice broadcast. However, the language type used most times in the actual voice account is the first language type, so the language type of the voice broadcast is determined to be the first language type. This makes the language type of the voice broadcast mismatch with the language type that the user is actually accustomed to, reducing the user's experience of using the voice broadcast function.
[0080] In one possible implementation, based on the number of times each language type is used in N language type information, the single type usage ratio corresponding to each language type in the N language types is determined, wherein the single type usage ratio is the ratio of the number of times the language type is used to the number of times all language types corresponding to the voice information account are used; when there are M language types whose corresponding single type usage ratios are greater than or equal to the corresponding preset threshold, the language type corresponding to the maximum single type usage ratio is determined as the target language type for voice broadcast, wherein the language type corresponding to the maximum single type usage ratio is the language type among the M language types, wherein N≥1, 1≤M≤N.
[0081] For example, Table 4 provides another language type information table for an embodiment of the present application.
[0082] Table 4:
[0083] It should be noted that Table 4 is only an exemplary description of the language type information associated with the voice information account corresponding to the user account ID "38f34aa7-bebc-0375-c4d3-4817162296b6".
[0084] Table 4 shows that for the voicemail account corresponding to user ID "38f34aa7-bebc-0375-c4d3-4817162296b6," the first language type was used 160 times, corresponding to a preset threshold of 60%, and the second language type was used 140 times, corresponding to a preset threshold of 40%. Note that for the same voicemail account, the sum of the preset thresholds for multiple language types can be 1 or different.
[0085] In addition, the first language type is similar to the second language type. The model can more easily recognize the first language type and has difficulty recognizing the second language type. It is easy to recognize the second language type as the first language type.
[0086] When the user actually input voice commands, they used the first language type 100 times and the second language type 200 times. Due to recognition errors in the model, the user's language type information account shows that the first language type was used 160 times and the second language type was used 140 times.
[0087] The second language type is more difficult to identify and is more likely to be identified as the first language type. Understandably, the number of times the second language type is used in the voice message account is less than the actual number of times the user inputs voice commands, while the number of times the first language type is used is more than the actual number of times the user inputs voice commands. Therefore, when the target language type for voice broadcast is determined based on the most frequently used language type, there is a deviation from the language type that the user actually uses.
[0088] Table 4 shows that the percentage of single-type usage in the first language is 53%, while the percentage of single-type usage in the second language is 47%. Considering the model's recognition errors for both language types, the threshold for the first language is set at 60%, and the threshold for the second language is set at 40%. At this point, the percentage of single-type usage in the second language exceeds the corresponding threshold, while the percentage of single-type usage in the first language does not meet the threshold. Therefore, the target language for voice broadcast is determined to be the second language.
[0089] It can be understood that by setting thresholds for different language types respectively, the recognition errors existing in the model and the errors in recording the number of times a language type is used can be corrected, so that the language type used in the voice broadcast is more in line with the user's actual habits, further improving the user's experience of using the voice broadcast function.
[0090] It should be noted that in the same voice information account, there may be multiple language types whose single-type usage ratio is greater than or equal to the corresponding preset threshold. At this time, the language type corresponding to the largest single-type usage ratio is determined as the target language type for voice broadcast.
[0091] For example, Table 5 provides another language type information table for an embodiment of the present application.
[0092] Table 5:
[0093] It should be noted that Table 5 is only an exemplary description of the language type information associated with the voice information account corresponding to the user account ID "38f34aa7-bebc-0375-c4d3-4817162296b6".
[0094] According to Table 4, in the voice message account corresponding to the user account ID "38f34aa7-bebc-0375-c4d3-4817162296b6", the first language type was used 160 times, with a corresponding single type usage ratio of 27%, and a corresponding preset threshold of 25%, which reached the preset threshold; the second language type was used 140 times, with a corresponding single type usage ratio of 40%, and a corresponding preset threshold of 20%, which reached the preset threshold; the third language type was used 200 times, with a corresponding single type usage ratio of 33%, and a corresponding preset threshold of 50%, which did not reach the preset threshold.
[0095] It is understood that in this voice message account, both the first language type and the second language type have reached the preset threshold. However, the usage percentage of the second language type is higher than the usage percentage of the first language type. Therefore, the target language type for voice broadcast is determined to be the second language type.
[0096] Although the model has errors in recognizing some language types, in actual applications, the recognition errors are relatively small. In addition, setting a threshold for each language type can compensate for the problems caused by recognition errors to a certain extent. Therefore, when the usage ratio of a single type of a certain language type is greater than or equal to the corresponding preset threshold, the model is considered to have accurately recognized the language type. At this time, among the multiple language types that reach the preset threshold, the language type corresponding to the largest single type usage ratio is considered to be more in line with the user's actual habits. Therefore, the language type corresponding to the largest single type usage ratio is determined as the target language type for voice broadcast.
[0097] It should be noted that within the same voice message account, when the usage percentage of each language type is less than the corresponding preset threshold, the preset language type is determined as the target language type for voice broadcast. The preset language type is usually a standard language type, such as Mandarin.
[0098] If the usage percentage of each language type in a voice message account is less than the corresponding preset threshold, the model is considered to have potential errors in identifying each language type in that voice message account. To avoid incorrectly determining the language type used for voice broadcasts, a preset language type is used for voice broadcasts. The preset language type is typically a standard language type, such as Mandarin, to avoid degrading the user experience of the voice broadcast function.
[0099] During voice interaction, if the conditions for automatic language switching are met, the target language of the voice announcement will automatically switch. However, a sudden switch from one language to another may cause confusion among users, leading them to suspect a malfunction in the voice announcement device, thus reducing their experience with the voice announcement feature.
[0100] In an embodiment of the present application, when there are M language types corresponding to a single type usage ratio greater than or equal to the corresponding preset threshold, a prompt message is output, wherein the prompt message is used to remind the user whether to switch the target language type of the voice broadcast; if the first response information input by the user is received, the language type corresponding to the largest single type usage ratio is determined as the target language type of the voice broadcast, wherein the first response information is used to indicate that the user allows switching the target type of the voice broadcast.
[0101] The prompt information may be in the form of voice, text, video, etc., which is not limited in the embodiments of the present application.
[0102] In other words, when the conditions for switching the voice announcement language are met, a prompt should be displayed to ask the user for their consent. Once the user approves the switch, the voice announcement language should be switched. While respecting the user, avoid sudden language changes that could degrade the user experience.
[0103] In a possible implementation, if the second response information input by the user is received, the target language type of the voice broadcast is determined to be the voice type currently used for the voice broadcast, wherein the second response information is used to indicate that the user does not allow the target language type of the voice broadcast to be switched.
[0104] In other words, when a user disagrees with switching the voice announcement language, it may indicate that the user is more accustomed to the current language than the language to be switched. Alternatively, the current language is acceptable to the user, thus avoiding a deterioration in the user's experience with the voice announcement feature.
[0105] In actual applications, when the user does not allow the language type of the voice broadcast to be switched, it is possible that the user does not like the current language type of the voice broadcast and only dislikes the language type to be switched even more, so the switching is not allowed.
[0106] In one possible implementation, if a second response message input by the user is received, the target language type of the voice broadcast is determined to be the language type corresponding to the response message input by the user last time, wherein the second response message is used to indicate that the user does not allow switching to the target language type of the voice broadcast.
[0107] The language type that the user last selected according to the prompt information may be one that the user prefers or is accustomed to. If the user does not allow the language type to be switched for voice announcement according to the prompt information, the language type for voice announcement can be determined to be the language type that the user last selected according to the prompt information, which may be more in line with the user's interests or habits, thereby further improving the user's experience of using the voice announcement function.
[0108] In actual applications, if users frequently use voice interaction features, input a large number of voice commands, and often use voice commands in different languages, prompt messages may be output frequently. Frequent prompt messages may annoy users and reduce their experience with the voice announcement feature.
[0109] In an embodiment of the present application, when there are M language types corresponding to a single type usage ratio greater than or equal to the corresponding preset threshold, it is determined whether the voice information account has output prompt information within a preset time period, wherein the time period is the time interval from the first moment to the current moment, and the first moment is earlier than the current moment; if the voice information account has output prompt information within the preset time period, it is determined that the target language type of the voice broadcast is the language type currently used in the voice broadcast; if the voice information account does not output prompt information within the preset time period, the prompt information is output.
[0110] When the language type switching conditions of the voice broadcast are met, before outputting the prompt information, it is determined whether the voice information account has output prompt information within a preset time period, which can avoid the problem of frequent output of prompt information and thus avoid the problem of user disgust.
[0111] According to the method described above, the language type of the voice broadcast is determined based on the language type information in the user's corresponding voice message account, thereby realizing the function of automatically switching the voice broadcast language type based on the user's actual habits. In other words, the language type information data in the user's corresponding voice message account needs to be continuously updated based on the voice commands input by the user.
[0112] It should be noted that voice commands include voice wake-up commands and voice control commands. In the embodiment of the present application, the language type information in the voice information account is the language type information of the voice control command. Among them, the voice control command is used to control the voice broadcast device to perform voice broadcast.
[0113] In an embodiment of the present application, after determining the target language type of the voice broadcast based on the number of times each language type in at least one language type information is used, it also includes: determining the language type corresponding to the voice control instruction based on the voice control instruction; iteratively updating the number of times the language type corresponding to the voice control instruction is used to the voice information account based on the language type corresponding to the voice control instruction; and determining the target language type of the voice broadcast based on the number of times each language type in the at least one language type after the iterative update is used.
[0114] See also Figure 5 , provides a schematic diagram of iterative update of language type information for the embodiment of this application. Figure 5As shown, after user A logs in, their user account ID is "38f34aa7-bebc-0375-c4d3-4817162296b6." When user A enters the voice control command "Open the car window," voice broadcast device 501 uses the model to identify the language type corresponding to the voice control command as the first language type. Based on user A's user account ID and the first language type, the usage count in the language type information table in database 502 is iteratively updated accordingly.
[0115] It should be noted that if the voice control command input by user A is recognized as the fifth language type, user A's voice information account does not have the fifth language type in the language type information table. In this case, a data entry is added to the language type information table, indicating that the user account ID is "38f34aa7-bebc-0375-c4d3-4817162296b6", the language type is the fifth language type, and the number of uses is 1.
[0116] In actual applications, voice control commands typically don't undergo voiceprint recognition. That is, when user A logs in and enters a first voice control command, user B might also enter a second voice control command into the same voice broadcast device. Since voiceprint recognition isn't performed on voice control commands, the second voice command will also be considered input by user A. In other words, the language type information corresponding to the first and second voice control commands will be associated and recorded in user A's corresponding voice information account. Therefore, user A's corresponding voice information account may be associated with other users' language type information, potentially interfering with the determination of user A's language usage habits.
[0117] In an embodiment of the present application, based on the voice control instruction, the target voice information account corresponding to the voice control instruction is determined, wherein the target voice information account is the voice information account corresponding to the voiceprint of the voice control instruction; based on the language type corresponding to the voice control instruction, the number of times the language type corresponding to the voice control instruction is used is iteratively updated to the target voice information account.
[0118] See also Figure 6 , provides another schematic diagram of iterative update of language type information for the embodiment of this application. Figure 6As shown, after user A logs in, their user account ID is "38f34aa7-bebc-0375-c4d3-4817162296b6." When user A enters the voice control command "open the car window," the user account ID can be linked to the voiceprint information corresponding to the voice control command in database 602. Voice broadcast device 601 uses the model to identify the language type corresponding to the voice control command as the first language type. Based on user A's user account ID "38f34aa7-bebc-0375-c4d3-4817162296b6" and the first language type, the usage count in the language type information table can be iteratively updated accordingly. At this point, if user B enters the voice control command "Turn on the air conditioner," database 602 can correlate the voiceprint information associated with this voice control command with user B's account ID, "f03e72d8-aa31-5549-36df-edb1839673f8." Voice broadcast device 601 uses the model to identify the language type corresponding to this voice control command as the fourth language type. Based on user B's account ID, "f03e72d8-aa31-5549-36df-edb1839673f8," and the fourth language type, the usage count in the language type information table in database 602 can be iteratively updated accordingly.
[0119] Voiceprint recognition is performed on each voice control command, accurately identifying the user corresponding to each voice control command and then accurately matching the user's corresponding voice information account. After the user logs into the voice information account, voice control commands entered by others are prevented from interfering with the iterative updates of the language type information corresponding to the logged-in user. This ensures that the language type of the voice broadcast is more consistent with the logged-in user's actual habits, further improving the user experience of the voice broadcast function.
[0120] It is understandable that the automatic switching of voice broadcast language types based on user habits requires continuous identification of the language type corresponding to each voice command through the language type recognition model. In addition, the language type information corresponding to each voice command needs to be iteratively updated to the corresponding voice information account. These operations will place a burden on the controller, so whether to use this function should respect the user's wishes. In other words, users can choose to use or not use this function.
[0121] Therefore, in the embodiment of the present application, before logging into the corresponding voice information account in response to a login operation, it is necessary to determine whether adaptive voice broadcast has been set, where adaptive voice broadcast means that the voice broadcast device automatically switches the language type of the voice broadcast. If adaptive voice broadcast has been set, the corresponding voice information account is logged in in response to the login operation; if adaptive voice broadcast has not been set, the target language type of the voice broadcast is determined to be the language type currently used by the voice broadcast.
[0122] Corresponding to the above method embodiment, the present application also provides a voice broadcast device. Specifically, see Figure 7 , is a structural diagram of a voice broadcasting device provided in an embodiment of the present application. Figure 7 As shown, the figure shows a voice broadcast device 70. The voice broadcast device 70 includes: a voice information account login module 701, and a target language type determination module 702. Specifically, the voice information account login module 701, in response to a login operation, logs in to the corresponding voice information account, wherein the voice information account is associated with N language type information, each language type information is used to represent a language type and the number of times the language type is used, and each of the N language type information corresponds to a different language type, where N ≥ 1; the target language type determination module 702 is used to determine the target language type for the voice broadcast based on the number of times each of the N language type information is used.
[0123] Corresponding to the above embodiment, an embodiment of the present application further provides an electronic device.
[0124] See also Figure 8 , is a structural diagram of an electronic device provided in an embodiment of the present application. Figure 8 As shown, the electronic device 80 may include: a processor 801, a memory 802, and a communication unit 803. These components communicate via one or more buses. Those skilled in the art will appreciate that the electronic device structure shown in the figure does not limit the embodiments of the present application. It may be a bus structure or a star structure, and may include more or fewer components than shown, or combine certain components, or arrange the components differently.
[0125] The communication unit 803 is used to establish a communication channel so that the electronic device can communicate with other devices.
[0126] The processor 801 is the control center of the electronic device. It uses various interfaces and lines to connect various parts of the entire electronic device. It runs or executes software programs and / or modules stored in the memory 802, and calls data stored in the memory to perform various functions of the electronic device and / or process data. The processor can be composed of an integrated circuit (IC), for example, it can be composed of a single packaged IC, or it can be composed of multiple packaged ICs with the same or different functions. For example, the processor 801 can include only a central processing unit (CPU). In the embodiment of the present application, the CPU can be a single computing core or multiple computing cores.
[0127] The memory 802 is used to store execution instructions of the processor 801. The memory 802 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0128] When the execution instructions in the memory 802 are executed by the processor 801 , the electronic device 80 is enabled to execute part or all of the steps in the above method embodiment.
[0129] Corresponding to the above embodiment, embodiments of the present application further provide a computer-readable storage medium, wherein the computer-readable storage medium may store a program. When the program is executed, the program may control the device containing the computer-readable storage medium to execute some or all of the steps of the above method embodiments. In a specific implementation, the computer-readable storage medium may be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0130] Corresponding to the above embodiment, an embodiment of the present application further provides a computer program product, which includes executable instructions. When the executable instructions are executed on a computer, the computer executes some or all of the steps in the above method embodiment.
[0131] In the embodiments of the present application, "at least one" refers to one or more, and "more" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can be represented by: a, b, c, ab, ac, bc, or abc, where a, b, c can be single or multiple.
[0132] Those skilled in the art will appreciate that the various units and algorithm steps described in the embodiments disclosed herein can be implemented using a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0133] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0134] In the several embodiments provided in this application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to the existing technology, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of this application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk, and other media that can store program code.
[0135] The above description is merely a specific embodiment of the present application. Any person skilled in the art may easily conceive of variations or substitutions within the technical scope disclosed in this application, and such variations or substitutions shall be within the scope of protection of this application. The scope of protection of this application shall be subject to the scope of protection of the claims.
Claims
1. A voice broadcasting method, characterized in that: Applied to a voice broadcast device, the method includes: In response to a login operation, logging into a corresponding voice information account, the voice information account being associated with N language type information, each piece of language type information being used to represent a language type and a usage count corresponding to the language type, each piece of the N language type information corresponding to a different language type, where N ≥ 1; The target language type of the voice broadcast is determined according to the number of times each of the language types in the N language type information is used.
2. The method according to claim 1, characterized in that After determining the target language type for voice broadcast according to the number of times each language type in the N language type information is used, the method further includes: Determining, according to the voice control instruction, a language type corresponding to the voice control instruction, wherein the voice control instruction is used to control the voice broadcast device to perform voice broadcast; Iteratively updating the number of times the language type corresponding to the voice control instruction is used to the voice information account according to the language type corresponding to the voice control instruction; The target language type for voice broadcast is determined according to the number of times each of the N language types is used after iterative update.
3. The method according to claim 1, characterized in that The step of determining the target language type for voice broadcast according to the number of times each language type in the N language type information is used includes: comparing the usage counts of each of the N language types; The language type that is used the most times among the N language types is determined as the target language type for voice broadcast.
4. The method according to claim 1, wherein The step of determining the target language type for voice broadcast according to the number of times each language type in the N language type information is used includes: Determining, based on the number of times each of the language types in the N language type information is used, a single type usage ratio corresponding to each of the N language types, where the single type usage ratio is a ratio of the number of times corresponding to the language type to the number of times all language types corresponding to the voice information account are used; When there are M language types whose corresponding single type usage ratios are greater than or equal to the corresponding preset threshold, the language type corresponding to the maximum single type usage ratio is determined as the target language type for voice broadcast, wherein the language type corresponding to the maximum single type usage ratio is the language type among the M language types, wherein 1≤M≤N.
5. The method according to claim 4, characterized in that Also includes: When the single-type usage ratio corresponding to each of the language types is less than the corresponding preset threshold, the preset language type is determined as the target language type for voice broadcast.
6. The method according to claim 4, characterized in that When there are M language types corresponding to the single type usage ratio being greater than or equal to the corresponding preset threshold, determining the language type corresponding to the largest single type usage ratio as the target language type for voice broadcast includes: When the usage ratio of the single type corresponding to the M language types is greater than or equal to the corresponding preset threshold, a prompt message is output, where the prompt message is used to remind the user whether to switch the target language type of the voice broadcast; If the first response information input by the user is received, the language type corresponding to the largest single type usage ratio is determined as the target language type of the voice broadcast, and the first response information is used to indicate that the user allows the target language type of the voice broadcast to be switched.
7. The method according to claim 6, characterized in that Also includes: If the second response information input by the user is received, it is determined that the target language type of the voice broadcast is the voice type currently used for the voice broadcast, and the second response information is used to indicate that the user does not allow the target language type of the voice broadcast to be switched.
8. The method according to claim 6, characterized in that Also includes: If a second response message input by the user is received, the target language type of the voice broadcast is determined to be the language type corresponding to the response message input by the user last time, and the second response message is used to indicate that the user does not allow the target language type of the voice broadcast to be switched.
9. The method according to claim 6, characterized in that When the usage ratio of the single type corresponding to the M language types is greater than or equal to the corresponding preset threshold, outputting a prompt message includes: When there are M language types corresponding to the single-type usage ratios greater than or equal to the corresponding preset threshold, determining whether the voice information account has outputted prompt information within a preset time period, where the time period is a time interval from a first moment to a current moment, and the first moment is earlier than the current moment; If the voice information account outputs prompt information within a preset time period, determining the target language type of the voice broadcast to be the language type currently used for the voice broadcast; If the voice information account does not output prompt information within a preset time period, the prompt information is output.
10. The method according to claim 1, characterized in that The step of logging into the corresponding voice information account in response to the login operation includes: In response to a voice wake-up instruction, a corresponding voice information account is logged in, the voice information account is associated with the voiceprint corresponding to the voice wake-up instruction, and the voice wake-up instruction is used to activate the voice broadcast function of the voice broadcast device.
11. The method according to claim 2, characterized in that The iteratively updating the usage count of the language type corresponding to the voice control instruction according to the language type corresponding to the voice control instruction includes: Determine, according to the voice control instruction, a target voice information account corresponding to the voice control instruction, wherein the target voice information account is a voice information account corresponding to the voiceprint of the voice control instruction; According to the language type corresponding to the voice control instruction, the number of times the language type corresponding to the voice control instruction is used is iteratively updated to the target voice information account.
12. The method according to claim 1, characterized in that The step of logging into the corresponding voice information account in response to the login operation includes: If the adaptive voice broadcast has been set, then in response to the login operation, the corresponding voice information account is logged in, and the adaptive voice broadcast is used to indicate that the voice broadcast device automatically switches the language type of the voice broadcast; If the adaptive voice broadcast is not set, the target language type of the voice broadcast is determined to be the language type currently used for the voice broadcast.
13. A voice broadcasting device, characterized in that: Applied to a voice broadcast device, the device comprises: a voice information account login module, in response to a login operation, logging into a corresponding voice information account, wherein the voice information account is associated with N language type information, each piece of language type information being used to represent a language type and a usage count corresponding to the language type, each of the N pieces of language type information corresponding to a different language type, wherein N ≥ 1; The target language type determination module is used to determine the target language type of the voice broadcast according to the number of times each language type in the N language type information is used.
14. An electronic device, characterized in that: include: processor; Memory; and a computer program, wherein the computer program is stored in the memory, and when the computer program is executed by the processor, causes the electronic device to perform the method according to any one of claims 1 to 11.
15. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 11.