Voice database updating method and related device
By acquiring user voice data, filtering and converting it into text data, the problem of low voice database update efficiency is solved, and efficient database update and expansion is achieved.
Patent Information
- Application Number
- CN202510991851.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-18
- Publication Date
- 2025-09-30
AI Technical Summary
The updating efficiency of existing voice databases is low, mainly due to the slow efficiency of manual collection, which makes it difficult to update and expand quickly.
By obtaining user voice data, filtering voice data of non-official language types, converting it into text data, determining the content that is not included in the basic voice database, and adding it to the database for update training.
Improves the efficiency of voice database updates, ensures that the data for each update is maximized, and enhances the coverage and adaptability of the database.
Smart Images

Figure CN120726992A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of database updating, and in particular to a method for updating a speech database and related devices. Background Art
[0002] When developing AI model voice functions, it is usually necessary to build a voice database for training. Currently, voice databases are generally collected by staff based on existing text, or customer feedback is collected and included in a general voice database. The efficiency of manual collection is slow, resulting in low update efficiency of the voice database. Therefore, how to improve the update efficiency of the voice database has become a problem that needs to be solved urgently. Summary of the Invention
[0003] The embodiments of the present application provide a method for updating a voice database and related devices, which improve the updating efficiency of the voice database.
[0004] In a first aspect, embodiments of the present application provide a method for updating a speech database, which is applied to an electronic device having a preset AI model, including:
[0005] Obtaining a basic speech database corresponding to the preset AI model;
[0006] During the use of the preset AI model, voice data of different users are collected to obtain k pieces of voice data; k is a positive integer;
[0007] Filtering out speech data of non-official languages from the k pieces of speech data to obtain m pieces of speech data, where m is a positive integer less than or equal to k;
[0008] Determine m dialect types corresponding to the m pieces of speech data;
[0009] Convert each of the m pieces of voice data into text data based on the m dialect types to obtain m pieces of text data;
[0010] Determine, based on the m pieces of text data, the speech data that are not included in the basic speech database among the m pieces of speech data, to obtain n pieces of speech data; n is a natural number smaller than m;
[0011] The n pieces of voice data are added to the basic voice database to obtain a target voice database; and the preset AI model is updated and trained using the target voice database.
[0012] In a second aspect, an embodiment of the present application provides a device for updating a speech database, which is applied to an electronic device, wherein a preset AI model is provided in the electronic device, and the device includes: an acquisition unit and a database update unit, wherein:
[0013] The acquisition unit is configured to acquire a basic voice database corresponding to the preset AI model; during the use of the preset AI model, voice data of different users are collected to obtain k pieces of voice data; k is a positive integer;
[0014] The database updating unit is configured to filter out speech data of non-official language types from the k pieces of speech data to obtain m pieces of speech data; m is a positive integer less than or equal to k; determine the m dialect types corresponding to the m pieces of speech data; convert each piece of speech data from the m pieces of speech data into text data based on the m dialect types to obtain m pieces of text data; determine the speech data from the m pieces of speech data that is not included in the basic speech database based on the m pieces of text data to obtain n pieces of speech data; n is a natural number less than m; add the n pieces of speech data to the basic speech database to obtain a target speech database; and update and train the preset AI model using the target speech database.
[0015] In a third aspect, an embodiment of the present application provides an electronic device comprising: a processor, a memory, a communication interface, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the processor, and the program includes instructions for executing the steps in the first aspect of the embodiment of the present application.
[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the above-mentioned computer-readable storage medium stores a computer program for electronic data exchange, wherein the above-mentioned computer program enables a computer to execute some or all of the steps described in the first aspect of the embodiment of the present application.
[0017] In a fifth aspect, embodiments of the present application provide a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is operable to cause a computer to perform some or all of the steps described in the first aspect of the embodiments of the present application. The computer program product may be a software installation package.
[0018] The implementation of this application has the following beneficial effects:
[0019] It can be seen that the voice database updating method described in the present application converts the collected m voice data into text data to obtain m text data. The m text data are compared with the basic voice database to directly locate the unrecorded content, thereby obtaining n unrecorded voice data. Compared with audio feature comparison, text comparison is more efficient. Then, these n voice data are stored in the basic voice database to obtain the target voice database. The newly added n voice data are all the missing content in the basic voice database, which maximizes the effective data of each update, thereby improving the update efficiency of the voice database. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the background technology, the drawings required for use in the embodiments of the present application or the background technology will be described below.
[0021] Figure 1 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application;
[0022] Figure 2 This is a schematic diagram of a scenario of an electronic device provided in an embodiment of the present application;
[0023] Figure 3 This is a flow chart of a method for updating a voice database provided in an embodiment of the present application;
[0024] Figure 4 This is a user task flow chart of a vehicle diagnostic scenario provided by an embodiment of the present application;
[0025] Figure 5 is a segmented graph of second voice data provided in an embodiment of the present application;
[0026] Figure 6 This is a segmented graph after segmented time points are adjusted, provided in an embodiment of the present application;
[0027] Figure 7 This is a block diagram of the functional units of a speech database updating device provided in an embodiment of the present application;
[0028] Figure 8 It is a structural diagram of another electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0029] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0030] The terms "first," "second," and the like in the specification and claims of this application and the accompanying drawings are used to distinguish between different objects, not to describe a particular order. Furthermore, the terms "including," "having," and any variations thereof, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or apparatus comprising a series of steps or elements is not limited to the listed steps or elements but may optionally include steps or elements not listed, or may optionally include other steps or elements inherent to the process, method, product, or apparatus.
[0031] It should be understood that the term "and / or" in this document is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this document indicates that the associated objects are in an "or" relationship. The "plurality" appearing in the embodiments of this application refers to two or more.
[0032] In the embodiments of the present application, "at least one item" or similar expressions refers to any combination of these items, including any combination of single items or plural items, and refers to one or more, and multiple refers to two or more. For example, at least one item (item) of a, b, or c can represent the following seven situations: a, b, c, a and b, a and c, b and c, a, b, and c. Among them, each of a, b, and c can be an element or a set containing one or more elements.
[0033] The "connection" appearing in the embodiments of the present application refers to various connection methods such as direct connection or indirect connection to achieve communication between devices, and the embodiments of the present application do not impose any limitations on this.
[0034] References herein to "embodiments" mean that a particular feature, structure, or characteristic described in connection with the embodiments may be included in at least one embodiment of the present application. The appearance of this phrase in various places in the specification does not necessarily refer to the same embodiment, nor does it constitute an independent or alternative embodiment that is mutually exclusive of other embodiments. It is understood, both explicitly and implicitly, by those skilled in the art that the embodiments described herein may be combined with other embodiments.
[0035] The electronic devices described in the embodiments of the present application may include smart phones (such as Android phones, iOS phones, Windows Phone phones, etc.), tablet computers, PDAs, laptops, video matrices, monitoring platforms, mobile internet devices (MIDs) or wearable devices, etc. The above are only examples and not exhaustive, including but not limited to the above devices.
[0036] Of course, the above-mentioned electronic device can also be a server, for example, a cloud server.
[0037] The following describes the relevant contents, concepts, meanings, technical issues, technical solutions, beneficial effects, etc. involved in the embodiments of this application.
[0038] First, some professional terms involved in this application are explained:
[0039] Dialect type: This refers to the classification of sub-linguistic forms of natural languages with unique phonetic, lexical, and grammatical features based on their regional, social, or functional variations. In the context of voice database updates, this specifically refers to regional language variants of non-official standard languages, such as Sichuanese, Cantonese, and Wu dialects in Chinese, and American and British English dialects in English.
[0040] Pre-trained language model: A neural network model pre-trained based on large-scale text data. It learns the universal semantic representation of language through unsupervised learning and can be fine-tuned for specific tasks (such as text classification and semantic understanding). For example, the BERT model and the RoBERTa model.
[0041] The wave module in Python is used to operate audio files in WAV format, and supports reading, writing, and modifying basic parameters of audio data (such as sampling rate, number of channels, and frame data).
[0042] spaCy is an open-source Python natural language processing library that provides efficient text processing capabilities, including word segmentation, part-of-speech tagging, syntactic analysis, and named entity recognition. It is particularly well-suited for processing large amounts of text data. It can be used for punctuation supplementation, grammatical analysis, keyword extraction, and dialect vocabulary calibration.
[0043] Dialect ASR Model: A speech-to-text model specifically trained for specific dialects (such as Cantonese, Sichuan dialect, Wu dialect, etc.), which can accurately convert dialect speech into text. It is trained based on dialect speech data, optimizing the pronunciation features unique to dialects (such as the nine tones and six pitches of Cantonese, the confusion of "n / l" in Sichuan dialect); it includes a dialect-specific vocabulary dictionary (such as "啱", "要得") to improve the recognition accuracy; it can be constructed through transfer learning (such as fine-tuning based on a Mandarin ASR model) or training from scratch. It can be used in scenarios such as dialect speech-to-text, dialect customer service systems, and data collection for dialect culture protection.
[0044] Please refer to Figure 1 , Figure 1 Figure [ID] is a schematic structural diagram of an electronic device provided by an embodiment of the present application; it can be seen that the electronic device may include: a voice acquisition module, a control module, a database update module, etc., which are not limited herein. Among them:
[0045] The voice acquisition module is usually composed of devices such as a microphone array, an audio preamplifier, and an analog-to-digital converter. The microphone array is used to collect voice signals in the surrounding environment, which can improve the sensitivity and directivity of voice acquisition; the audio preamplifier amplifies the weak voice electrical signal to make it reach an appropriate level; the analog-to-digital converter converts the analog voice signal into a digital signal for subsequent processing. The voice acquisition module is responsible for collecting voice data of different users during the use of the preset AI model, obtaining multiple voice data. It is the source of obtaining the original voice data and can provide basic materials for subsequent data screening, processing, and database update.
[0046] The control module generally consists of a microprocessor (such as a single-chip microcomputer, an ARM chip, etc.), a memory (including a random access memory RAM and a read-only memory ROM), etc. The microprocessor is responsible for executing various instructions and coordinating the work of each module; the memory is used to store program codes and temporary data, etc. The control module plays a role of overall coordination in the electronic device. On the one hand, it can control when the voice acquisition module starts and stops collecting voice data; on the other hand, during the data processing process, it can control the execution order of operations such as data screening, dialect type determination, and speech-to-text, ensuring that the entire voice database update process proceeds in an orderly manner according to the established logic.
[0047] The database update module can be composed of storage devices (such as hard disks, solid-state drives, etc.), database management system software, etc. The storage device is used to actually store the data of the voice database; the database management system is responsible for creating, querying, updating, deleting and other operations on the database. The database update module is used to add the voice data that has not been included in the basic voice database after screening, conversion and other processing to the voice database according to the instructions of the control module to obtain an updated voice database. In addition, the database update module can also update and train the preset AI model through the updated voice database, thereby realizing the update of the voice database and the optimization of the AI model.
[0048] See also Figure 2 , Figure 2 This is a schematic diagram of a scenario of an electronic device provided in an embodiment of the present application. It can be seen that the target user, as the operating subject, can operate the electronic device. Two options are displayed on the electronic device, one is "Start Preset AI Model" and the other is "Update Voice Database". The target user clicks "Start Preset AI Model" to start the preset AI model in the electronic device and interact with the preset AI model. Clicking "Update Voice Database" allows the electronic device to update the voice database in the electronic device based on the built-in algorithm. Specifically, the electronic device can execute the voice database update method provided in the embodiment of the present application to update the voice database. The specific steps are as follows:
[0049] Obtaining a basic speech database corresponding to the preset AI model;
[0050] During the use of the preset AI model, voice data of different users are collected to obtain k pieces of voice data; k is a positive integer;
[0051] Filtering out speech data of non-official languages from the k pieces of speech data to obtain m pieces of speech data, where m is a positive integer less than or equal to k;
[0052] Determine m dialect types corresponding to the m pieces of speech data;
[0053] Convert each of the m pieces of voice data into text data based on the m dialect types to obtain m pieces of text data;
[0054] Determine, based on the m pieces of text data, the speech data that are not included in the basic speech database among the m pieces of speech data, to obtain n pieces of speech data; n is a natural number smaller than m;
[0055] The n pieces of voice data are added to the basic voice database to obtain a target voice database; and the preset AI model is updated and trained using the target voice database.
[0056] It should be explained that the above electronic device can execute part or all of the steps of the method for updating the multi-language database provided in the embodiment of the present application.
[0057] See also Figure 3 , Figure 3 This is a flowchart of a method for updating a voice database provided in an embodiment of the present application; the method can be applied to an electronic device having a preset AI model provided therein, and the method specifically includes but is not limited to the following steps:
[0058] S301. Obtain a basic speech database corresponding to the preset AI model.
[0059] In an embodiment of the present application, the preset AI model can be preset or defaulted in advance. For example, the preset AI model can be a deep learning model.
[0060] Optionally, step S301, obtaining a basic speech database corresponding to the preset AI model, may include the following steps:
[0061] A1. Obtain the target application scenario of the preset AI model;
[0062] A2. Determine a professional voice database corresponding to the target application scenario;
[0063] A3. Determine the user type corresponding to the target application scenario, and obtain a user type; a is a positive integer;
[0064] A4. Determine the usage habits corresponding to each of the a user types to obtain a usage habits;
[0065] A5. Adjust the professional speech database based on the a usage habits to obtain the basic speech database.
[0066] In the embodiments of the present application, the target application scenario may include one of the following: image recognition scenario, vehicle diagnosis scenario, financial investment scenario, education and learning scenario, etc., which are not limited here.
[0067] In a specific embodiment, the target application scenario of the preset AI model can be obtained first. Specifically, the algorithm principle, technical architecture and capabilities of the preset AI model can be studied in depth. The target application scenario can be determined based on the capabilities of the model. For example, if the preset AI model is a model with strong image recognition capabilities, it can be used in image recognition scenarios such as security monitoring to identify suspects, industrial quality inspection to detect product defects, and medical image analysis to assist in disease diagnosis. If the preset AI model is a model that is good at vehicle fault identification, it can establish a vehicle-related knowledge graph covering knowledge such as vehicle structure, component functions, fault cases, and maintenance methods. Through the knowledge graph, the model can quickly associate and query relevant information to assist in fault diagnosis and generate maintenance suggestions, and its target application scenario can be determined to be a vehicle diagnosis scenario.
[0068] Next, the professional voice database corresponding to the target application scenario can be determined. Specifically, the mapping relationship between the preset application scenario and the voice database can be pre-stored, and the professional voice database corresponding to the target application scenario can be determined based on the mapping relationship; then, the user type corresponding to the target application scenario can be determined, and a user types can be obtained. Specifically, a user task flow chart for the target application scenario can be drawn, and all participating roles in the user task flow chart can be identified to obtain at least one participating role. Then, the user type corresponding to each participating role in the at least one participating role can be determined to obtain a user types. For example, the mapping relationship between the preset participating roles and the user type can be pre-stored, and the a user types corresponding to the a participating roles can be determined based on the mapping relationship.
[0069] It should be explained that a participating role may correspond to one or more user types. For example, assuming that a participating role is a car owner, the user type corresponding to the car owner role may include at least one of the following: novice car owner, experienced car owner, car enthusiast, etc., which is not limited here.
[0070] For example, assuming the target application scenario is a vehicle diagnosis scenario, please refer to Figure 4 , Figure 4 This is a user task flow chart of a vehicle diagnosis scenario provided by an embodiment of the present application. It can be seen that the user task flow chart includes the following steps:
[0071] The owner discovers a vehicle fault: This is the starting point of the entire process, which means that an abnormal condition has occurred in the vehicle during use and has been noticed by the owner. For example, the owner finds that the fault light on the vehicle dashboard is on, or that there are abnormal noises or vibrations when the vehicle is driving.
[0072] Send fault information to the preset AI model of the electronic device: After the car owner discovers the fault, he or she will use specific electronic devices (for example, mobile phone APP, on-board diagnostic system, etc.) to pass the vehicle fault-related information (for example, fault phenomenon description, vehicle performance, etc.) to the preset AI model for subsequent analysis.
[0073] The preset AI model provides possible faults and maintenance plans: After receiving the fault information sent by the car owner, the preset AI model uses its own algorithms, data and knowledge reserves to analyze and diagnose the fault information, determine the possible faults, and formulate a corresponding maintenance plan. For example, if it determines that a certain component of the engine is damaged, it will provide maintenance suggestions such as replacing the component and specific operating steps.
[0074] The car owner or maintenance technician repairs the vehicle according to the maintenance plan: Based on the maintenance plan given by the preset AI model, the car owner himself (if it is a simple and feasible maintenance task and the car owner has the relevant capabilities) or a professional maintenance technician (in most cases it is performed by a maintenance technician) repairs the vehicle and solves the vehicle failure problem according to the steps and requirements of the maintenance plan.
[0075] The owner confirms the repair results: After the repair work is completed, the owner will check and confirm the condition of the vehicle after repair to see whether the fault has been eliminated and whether the vehicle has returned to normal operation.
[0076] Providing customer service with feedback on the user experience and / or suggestions for the preset AI model: Throughout the entire process, the customer service staff of the preset AI model may communicate with the car owner to understand the car owner's experience using the preset AI model for vehicle diagnosis, collect opinions and suggestions from the car owner, and then feed this information back to relevant parties for optimization and improvement of the preset AI model.
[0077] It can be seen that the participating roles in the user task flow chart include: car owner, maintenance technician, and customer service staff.
[0078] Next, the usage habits corresponding to each of the a user types can be determined to obtain a usage habits. Specifically, a common language questionnaire can be designed. For users of different user types, the common language questionnaire can be used to ask them about common vocabulary, sentence patterns, dialect types, etc. in scenarios such as describing vehicle failures, communicating with maintenance personnel, and using relevant diagnostic tools. For example, a novice car owner can be asked "When you find a vehicle failure, how would you describe it to the maintenance personnel" and provide multiple options, while setting open questions for them to fill in freely. For another example, text information from users of different user types in car forums, maintenance record documents, customer service communication records, etc. can be collected. For example, on a car forum, the language characteristics in posts where car enthusiasts share their troubleshooting experiences can be analyzed, and the diagnostic description terms of professional maintenance technicians can be viewed from maintenance record documents. In this way, a usage habits can be obtained; finally, the voice in the professional voice database can be adjusted according to the a usage habits, thereby obtaining a basic voice database.
[0079] By identifying the user types and language habits corresponding to target application scenarios, the voice database and AI model can be tailored to the language expressions of different users. For example, in vehicle diagnosis, the terminology used by novice car owners and professional repair technicians differs significantly: novice car owners may express themselves in vague and colloquial language, while repair technicians use professional terminology. Adjusting the voice database based on different language habits allows the model to better understand different user expressions, improving the interactive experience and diagnostic efficiency.
[0080] Optionally, each usage habit includes: a dialect type and personal usage preference information; step A5, adjusting the professional speech database based on the a usage habits to obtain the basic speech database, may include the following steps:
[0081] B1. Obtain a first professional voice; the first professional voice is any professional voice data in the professional voice database;
[0082] B2. adjusting the first professional voice according to the dialect type of each of the a usage habits to obtain a dialect voice;
[0083] B3. Determine the preferred voice corresponding to the personal language preference information of each of the a language habits, and obtain i preferred voices, where i is an integer greater than or equal to a;
[0084] B4. Determine b basic voices based on the i preferred voices and the a dialect voices; b is an integer greater than a; the b basic voices are all basic voices in the basic voice database that are associated with the first professional voice.
[0085] In the embodiments of the present application, each dialect type may include one of the following: Cantonese, Minnan dialect, Wu dialect, Gan dialect, Xiang dialect, Hakka dialect, etc., which are not limited herein.
[0086] In a specific embodiment, the first professional voice can be obtained from a professional voice database; then, according to the dialect type of each usage habit in a usage habits, the first professional voice can be adjusted to obtain a dialect voices. Specifically, by consulting dialect research materials, language databases, and interviewing and conducting questionnaires on people using different dialects, etc., the voice, vocabulary, and grammar features of various dialects can be comprehensively collected. For example, for Cantonese, it is necessary to clarify its unique nine tones and six tones, unique vocabulary (such as "ngaam" meaning "suitable"), and some special sentence patterns (such as "Have you had your meal?"), and then, according to these voice, vocabulary, and grammar features, the first professional voice can be converted into the corresponding dialect, thus obtaining a dialect voices. For example, assuming the first professional voice is the Mandarin "The vehicle engine has a problem and it may be necessary to check whether there is a problem with the spark plug", and it needs to be converted into Cantonese, then the converted dialect voice is "The engine of the vehicle has a problem and it may be necessary to check whether there is a problem with the spark plug".
[0087] Then, the preference voices corresponding to the personal usage preference information of each usage habit in a usage habits can be determined to obtain i preference voices. For example, assuming that a certain personal usage preference information is: frequently using "Oh my god" and "awesome", then the corresponding preference voices can be generated: the voice of "Oh my god" and the voice of "awesome", thus obtaining i preference voices; finally, b basic voices can be determined according to the i preference voices and the a dialect voices. Specifically, the i preference voices can be randomly fused with the a dialect voices to obtain b basic voices.
[0088] In this way, by adjusting the professional voice according to the dialect type, the preset AI model can be adapted to the language habits of different regions. For example, in a medical consultation scenario, explaining professional terms in the local dialect (such as "penicillin" may be called "penicillin" in Sichuan dialect) can lower the user's understanding threshold, improve the information transmission efficiency, and avoid communication barriers caused by dialect differences, especially suitable for regions where dialects are frequently used (such as towns and townships, ethnic minority concentrated areas), and expand the geographical coverage of AI applications.
[0089] Optionally, in step B4, the determining b basic voices according to the i preference voices and the a dialect voices includes:
[0090] C1. Identifying the i preference voices and the a dialect voices through semantic recognition technology to obtain i preference semantic information and a dialect semantic information;
[0091] C2. determining a first semantic similarity between each dialect semantic information in the a dialect semantic information and the i preference semantic information, to obtain a first semantic similarity sets; each first semantic similarity set includes i first semantic similarities;
[0092] C3, filtering out the first semantic similarity in each of the a first semantic similarity sets that is greater than a preset semantic similarity, to obtain c second semantic similarity sets; c is a positive integer less than or equal to a;
[0093] C4. Determining c preferred speech sets corresponding to the c second semantic similarity sets based on the i preferred speech sets;
[0094] C5. Randomly fuse the preferred voices in the c preferred voice sets into the a dialect voices through a preset voice fusion method to obtain the b basic voices.
[0095] In the embodiment of the present application, the preset semantic similarity and the preset speech fusion method can be preset or defaulted in advance.
[0096] In a specific embodiment, i preferred voices and a dialect voices can be recognized through semantic recognition technology to obtain i preference semantic information and a dialect semantic information. Specifically, the ambient noise can be eliminated by using a noise suppression algorithm (such as WaveNet noise reduction) to address possible accent interference in the voice (such as the nine tones and six tones of Cantonese and the erhua sound of Sichuan dialect). Then, volume normalization is performed to ensure the consistency of audio features of different dialect voices and preferred voices. Then, the i preferred voices and a dialect voices can be converted into text data through a preset voice recognition model to obtain i preference text data and a dialect text data. Then, semantic recognition technology is used to recognize these i preference text data and a dialect text data to obtain i preference semantic information and a dialect semantic information.
[0097] Furthermore, the first semantic similarity between each of the a dialect semantic information and the i preference semantic information can be determined to obtain a set of a first semantic similarities. Specifically, for each dialect semantic information, the regional vocabulary therein can be first converted into a standard semantic equivalent expression (e.g., "搞莫子" → "做什么"), and a pre-trained language model (e.g., BERT model, RoBERTa model) can be used to generate text embedding vectors. For example, assuming that the dialect semantic information includes "夜儿个下雨咯" (Henan dialect "昨天下雨了") → standardized text "昨天下雨了" → corresponding dialect semantic vector is generated through the BERT model. Similarly, each preference semantic information can also be converted into a preference semantic vector to obtain i preference semantic vectors. Then, a preset similarity calculation method (e.g., cosine similarity) can be used to calculate the first semantic similarity between each dialect semantic vector and the i preference semantic vectors, thereby obtaining a set of a first semantic similarities.
[0098] Then, the first semantic similarities greater than the preset semantic similarity in each semantic similarity set of the a first semantic similarity sets can be screened out to obtain c second semantic similarity sets. Specifically, for each first semantic similarity set, all the first semantic similarities greater than the preset semantic similarity therein can be found, and a second semantic similarity set can be formed by these first semantic similarities. Thus, c second semantic similarity sets can be obtained.
[0099] Furthermore, c preference voice sets corresponding to the c second semantic similarity sets can be determined based on the i preference voices. Specifically, for each second semantic similarity set, the preference voice corresponding to the second semantic similarity therein can be first determined to obtain at least one preference voice, and then a preference voice set can be formed by these at least one preference voices. Thus, c preference voice sets can be obtained.
[0100] Finally, the preference voices in the c preference voice sets can be randomly fused into the a dialect voices through a preset voice fusion method to obtain b basic voices. Specifically, the preset voice fusion method can include one of the following: feature-level fusion method, waveform-level fusion method, semantic-driven fusion method, which is not limited herein. For example, assuming the waveform-level fusion method, a preference voice is randomly selected from the c preference voice sets, and the waveform segment of the preference voice (e.g., the catchphrase "你知道吧" in the preference voice) is directly inserted into an appropriate position (e.g., the beginning / end of a sentence) of a certain dialect voice to obtain a basic voice. Since the preset voice fusion method is a conventional technology, the specific fusion process is not elaborated herein. Thus, b basic voices can be obtained. For example, for a certain dialect voice "你吃了饭没得" (Sichuan dialect), the catchphrase "对吧" in the preference voice can be randomly fused, and the fused voice "你吃了饭没得,对吧" can be obtained after fusion.
[0101] In this way, through semantic recognition technology, the preferred speech and dialect speech are converted into semantic information, realizing the abstraction from "voice signal" to "semantic content". For example, when the user speaks a dialect, the system can recognize the core meaning expressed by the user (such as "tomorrow morning" may be expressed as "ming zao chen" in Sichuan dialect), and the semantic information in the preferred speech (such as the colloquial expressions commonly used by the user) can be synchronously extracted, providing a semantic-level consistency basis for subsequent matching. In addition, by calculating the similarity between the dialect semantics and the preferred semantics, the personal language habits that best match the user's dialect expression can be screened out. For example, when the user uses Cantonese to express "eat" ("sik faan"), the system can match the more commonly used colloquial expression (such as "have dinner") in the user's preferred speech through similarity calculation, so as to integrate the user's more habitual expression into the dialect speech and enhance the intimacy of communication.
[0102] S302. During the use of the preset AI model, collect the voice data of different users to obtain k pieces of voice data; k is a positive integer.
[0103] In the embodiment of the present application, a recording module can be built into the application end (such as APP, applet) of the preset AI model. When the user uses the preset AI model, the recording module can be started to record the user's voice to obtain multiple pieces of voice data. Or, the application end can also prompt and guide the user to participate through pop-up windows or tasks. For example, the voice assistant APP prompts "Participate in the voice optimization plan, recording 10 daily expressions can get integral rewards" when the user uses it. After the preset AI model is started by the user, the user starts to use the preset AI model and can record through the recording module to collect the voice data of different users to obtain k pieces of voice data. Or, the voice collection task can also be released on the crowdsourcing platform or social platform, and screening conditions such as region and age are set. For example, the activity of "recording local dialect tongue twisters" is released in the dialect interest community, and users are encouraged to participate through red envelopes, guiding the users to have a conversation with the preset AI model and recording the users' voices in real time, so as to obtain k pieces of voice data.
[0104] S303. Screen out the voice data of non-official language types from the k pieces of voice data to obtain m pieces of voice data; m is a positive integer less than or equal to k.
[0105] In the embodiment of the present application, the official language type refers to the general formal language type stipulated by the official. For example, the official language type can be Mandarin.
[0106] In specific embodiments, Mandarin ASR can be used to convert each of the k voice data into text data, obtaining k text data. Then, for each text data, its corresponding transcription error rate can be determined. If the transcription error rate is greater than a preset error rate (for example, Cantonese "shifan" is transcribed as "chifan" but the context semantics is incorrect), the corresponding voice data can be marked as a non-official language type. Thus, m voice data can be obtained.
[0107] S304. Determine the m dialect types corresponding to the m voice data.
[0108] In the embodiments of the present application, standard voice samples of various local dialects (for example, Cantonese, Sichuan dialect, Wu dialect, etc.) can be collected first, and the corresponding dialect types can be marked to form a reference voice database. Then, for each of the m voice data, acoustic feature extraction can be performed, for example, spectral features (such as Mel Frequency Cepstral Coefficients), pitch, duration, formants, etc. The extracted acoustic features are compared with the acoustic features of the samples in the reference voice database to determine the dialect type corresponding to each voice data, obtaining m dialect types. For example, the method of template matching can be used for comparison, calculating the similarity between each voice data and the templates of various local dialects, and selecting the dialect type with the highest similarity as its corresponding dialect type.
[0109] S305. Based on the m dialect types, convert each of the m voice data into text data, obtaining m text data.
[0110] In the embodiments of the present application, for each voice data, text conversion can be performed based on the characteristics of its corresponding dialect type to obtain text data. Thus, m text data can be obtained.
[0111] Optionally, in step S305, the converting each of the m voice data into text data based on the m dialect types to obtain m text data may include the following steps:
[0112] D1. Obtain the first voice data and its corresponding first dialect type; the first voice data is any one of the m voice data;
[0113] ]>D2. Perform data preprocessing on the first voice data to obtain second voice data;
[0114] D3. Determine the target voice length corresponding to the second voice data;
[0115] D4. When the target voice length is greater than or equal to a preset voice length, segment the second voice data according to the first dialect type and the target voice length to obtain d partial voice data; d is an integer greater than 1;
[0116] D5. Convert the d segments of speech data into text data in parallel according to the first dialect type to obtain d segments of text data;
[0117] D6. Combine the d segments of partial text data in chronological order to obtain reference text data;
[0118] D7. Perform text calibration on the reference text data to obtain text data corresponding to the first voice data.
[0119] In an embodiment of the present application, data preprocessing may include at least one of the following: format conversion operation, noise reduction operation, normalization operation, data enhancement operation, etc., which are not limited here; the preset speech length can be preset or defaulted in advance.
[0120] In a specific embodiment, first speech data and its corresponding first dialect type can be obtained first; then, data preprocessing can be performed on the first speech data to obtain second speech data. For example, data preprocessing can include format conversion operations and noise reduction operations. Speech in different formats (for example, MP3, WAV, FLAC, etc.) can be first converted into a unified format (usually WAV is selected, which is lossless and easy to process). Then, a noise reduction operation can be performed to obtain second speech data. For example, spectral subtraction can be used for noise reduction to estimate the noise spectrum and subtract it from the speech spectrum.
[0121] Then, the target speech length corresponding to the second speech data can be determined. Specifically, the storage file of the second speech data can be read through a language processing tool (for example, the wave module in Python) to obtain its audio information. The duration can be calculated based on the total number of audio frames and the sampling rate. For example, if the total number of audio frames is 16000 and the sampling rate is 1kHz (1000Hz), the calculation formula is as follows:
[0122] t=16000 / 1000(Hz)=16(s);
[0123] Let \(t\) be the duration. According to the above formula, the duration, that is, the target speech length, can be obtained. When the target speech length is greater than or equal to the preset speech length, the second speech data can be segmented according to the first dialect type and the target speech length to obtain \(d\) segments of partial speech data. Then, the \(d\) segments of partial speech data can be concurrently converted into text data according to the first dialect type to obtain \(d\) segments of partial text data. Specifically, the first dialect ASR model corresponding to the first dialect type can be determined. For example, the mapping relationship between the preset dialect types and the dialect ASR models can be stored in advance, and the first dialect ASR model corresponding to the first dialect type can be determined based on this mapping relationship. Then, the \(d\) segments of partial speech data can be input into the first dialect ASR model simultaneously for text conversion to obtain \(d\) segments of partial text data.
[0124] Furthermore, these \(d\) segments of partial text data can be combined in chronological order to obtain the reference text data. Specifically, when each segment of partial speech data is converted into text, the time interval (start time, end time) in the original audio (i.e., the second speech data) needs to be recorded. For example, assume that the first segment of speech corresponds to 0 - 3 seconds of the original audio, the second segment corresponds to 3 - 6 seconds, and so on. Then, these \(d\) segments of text data can be sorted in ascending order according to their corresponding time intervals (from the smallest start time), and in the sorted order, the content of each segment of partial text data can be concatenated in sequence to form the complete reference text data. Finally, the reference text data can be calibrated for text to obtain the text data corresponding to the first speech data. Specifically, text calibration can include text error correction and / or adding punctuation, etc., which are not limited here. A dialect homophone dictionary can be established, and the reference text data can be corrected for homophones and common misspelled words according to the dialect homophone dictionary. For example, in Cantonese, "佢" is miswritten as "渠", and in Sichuan dialect, "巴适" is miswritten as "巴适得板". Then, tools such as spaCy can also be used to analyze the sentence structure and add punctuation to the text (applicable to text transcriptions without punctuation), or the reference text data can be calibrated by the target user and / or professionals. In this way, the text data corresponding to the first speech data can be obtained.
[0125] When the target speech length is less than the preset speech length, the second speech data does not need to be segmented, and the second speech data can be directly converted into text data according to the first dialect type to obtain the reference text data. Then, the reference text data can be calibrated for text to obtain the text data corresponding to the first speech data.
[0126] In this way, by triggering segmented parallel processing for long audio and directly single-threaded processing for short audio, memory overflow or computing resource exhaustion caused by processing large files at one time can be avoided. For example, 1 minute of audio is processed by a single core, and 10 minutes of audio is automatically allocated to 10 cores for parallel processing, making full use of hardware resources and improving data processing efficiency.
[0127] Optionally, segmenting the second speech data according to the first dialect type and the target speech length to obtain d segments of partial speech data may include the following steps:
[0128] E1. Determine the target professional field dictionary corresponding to the target application scenario;
[0129] E2. Determine the time intervals corresponding to all professional keywords in the second voice data according to the target professional field dictionary, and obtain e time intervals; e is a natural number;
[0130] E3. Determine a target time window according to the first dialect type and the target speech length;
[0131] E4. Determine segmented time points in the second speech data according to the target time window to obtain f segmented time points, where f is a positive integer;
[0132] E5. Determine whether there is a segmented time point in the f segmented time points that is within the e time intervals;
[0133] E6. If so, determine the segmented time points and their corresponding time intervals among the f segmented time points, and obtain h segmented time points and h time intervals; h is a natural number less than e;
[0134] E7. Adjust the h segmentation time points according to the h time intervals so that each professional keyword is completely retained in the same segment of partial voice data, obtaining h target segmentation time points;
[0135] E8. Segment the second voice data according to the h target segmentation time points and fh segmentation time points other than the h segmentation time points among the f segmentation time points to obtain the d segments of partial voice data;
[0136] E9. If not, segment the second voice data according to the f segmentation time points to obtain the d segments of partial voice data.
[0137] In an embodiment of the present application, a target professional field dictionary corresponding to a target application scenario is determined. For example, a mapping relationship between a preset application scenario and a professional field dictionary can be pre-stored, and the target professional field dictionary corresponding to the target application scenario is determined based on the mapping relationship; then, the time intervals corresponding to all professional keywords in the second voice data can be determined according to the target professional field dictionary, and e time intervals are obtained. Specifically, the keywords in the second voice data can be identified by voice recognition technology to obtain multiple keywords, and it is determined whether the multiple keywords are included in the target professional field dictionary to obtain e keywords included in the target professional field dictionary, that is, e professional keywords. Then, the starting time and ending time of these e professional keywords in the second voice data can be determined to obtain e starting times and e ending times. The corresponding ending times of these e starting times and e ending times are combined to obtain e time intervals.
[0138] Next, the target time window can be determined based on the first dialect type and the target speech length; then, the segmentation time points in the second speech data can be determined based on the target time window to obtain f segmentation time points. Specifically, the window length corresponding to the target time window can be determined, and then the second speech data can be divided into multiple segments of speech data based on the window length, thereby obtaining f segmentation time points. For example, if the total length of the second speech data is 60 seconds and the target time window is set to 15 seconds, the segmentation points are 15s, 30s, and 45s (i.e., f=3).
[0139] Next, it can be determined whether there is a segmented time point in the e time intervals among the f segmented time points. Specifically, for each segmented time point, it can be determined in turn whether it falls in the e time intervals. For example, assuming f=3, the f segmented time points are 15s, 30s, and 45s respectively, and e=2, the e time intervals are [3, 5] and [28, 31] respectively. It can be seen that the segmented time point 30s falls in the time interval [28, 31], then it can be determined that there is a segmented time point in the e time intervals among the f segmented time points. Conversely, if no segmented time point falls in the e time intervals, it can be determined that there is no segmented time point in the e time intervals among the f segmented time points.
[0140] If there is at least one segmented time point that falls within e time intervals, the segmented time point and its corresponding time interval in the f segmented time points that are within the e time intervals can be found to obtain h segmented time points and h time intervals. Then, the h segmented time points can be adjusted according to the h time intervals so that each professional keyword is completely retained in the same segment of partial voice data to obtain h target segmented time points. For example, taking the first segmented time point and its corresponding first time interval as an example, the first segmented time point is any segmented time point among the h segmented time points, and the first time interval is the h time intervals corresponding to the first segmented time point in the h time intervals. The starting time and ending time of the first time interval can be determined first, and then the first difference between the first segmented time point and the starting time, as well as the first segmented time interval, can be calculated. The second difference between the starting point and the end time, if the first difference is greater than or equal to the second difference, the first segment time point can be moved toward the end time by the distance of the second difference plus the preset difference to obtain the target segment time point corresponding to the first segment time point, wherein the preset difference can be preset or defaulted in advance. For example, assuming that the first segment time point is the 30th second, it falls in the time interval [28, 31], the first difference is 2, the second difference is 1, and the preset difference can be preset to 1. Then the first segment time point can be moved toward the end time by a distance of 2s to obtain the target segment time point corresponding to the first segment time point (i.e., the 32nd second); if the first difference is less than the second difference, the first segment time point can be moved toward the start time by the distance of the first difference plus the preset difference to obtain the target segment time point corresponding to the first segment time point.
[0141] For an example, see Figure 5 , Figure 5 It is a segmentation diagram of the second voice data provided by an embodiment of the present application; it can be seen that f=5, and the f segmentation time points are: the first segmentation time point s1, the second segmentation time point s2, the third segmentation time point s3, the fourth segmentation time point s4, and the fifth segmentation time point s5; among them, the second segmentation time point s2 falls into the time interval [p1, p2], dividing the time interval [p1, p2] from the inside, that is, the professional keywords corresponding to the time interval [p1, p2] are separated.
[0142] See also Figure 6 , Figure 6 This is a segmented diagram after segmented time point adjustment provided by an embodiment of the present application. It can be seen that Figure 6 and Figure 5 The difference is that the second segment time point s2 is moved to the left of p1, and the positions of other segment time points remain unchanged; in this way, the professional keywords corresponding to the time interval [p1, p2] are not divided, but are completely in the same segment.
[0143] Furthermore, the second voice data can be segmented according to h target segmentation time points and fh segmentation time points other than the h segmentation time points among the f segmentation time points to obtain d segments of partial voice data; if they do not exist, the second voice data can be directly segmented according to the f segmentation time points to obtain d segments of partial voice data.
[0144] In this way, by judging whether the segmentation time point is within the professional keyword time interval, on-demand adjustment can be achieved. If there is a conflict (the segmentation point cuts off the keyword), h segmentation points are adjusted to ensure the integrity of the keyword; if there is no conflict, it is directly processed according to the original segmentation point to avoid excessive adjustment affecting efficiency.
[0145] Optionally, determining the target time window according to the first dialect type and the target speech length may include the following steps:
[0146] F1. Obtaining pronunciation feature information corresponding to the first dialect type;
[0147] F2. Determine a reference time window based on the pronunciation feature information;
[0148] F3. Adjust the reference time window according to the target speech length to obtain the target time window.
[0149] In an embodiment of the present application, the pronunciation feature information may include at least one of the following: prosodic features (for example, average syllable duration, tone duration, pause duration between syllables, etc.), acoustic features (for example, resonance peak frequency, phoneme duration, etc.), pronunciation habits (for example, connected reading frequency, tail sound omission ratio, special pronunciation rules, etc.), etc., which are not limited here.
[0150] In a specific embodiment, pronunciation feature information corresponding to the first dialect type can be obtained. Specifically, a mapping relationship between a preset dialect type and feature information can be pre-stored, and the pronunciation feature information corresponding to the first dialect type can be determined based on the mapping relationship. Then, a reference time window can be determined based on the pronunciation feature information. For example, the pronunciation feature information can be an average syllable duration, and the reference time window can be determined based on the average syllable duration, as follows:
[0151] Average pause coefficient = (total speech duration - total silence duration) / total speech duration;
[0152] Reference time window = average syllable duration × average pause coefficient;
[0153] Among them, the total voice duration is the total duration of the second voice data; the total silence duration is the silence duration of the second voice data; according to the above formula, the reference time window can be obtained. Finally, the reference time window can be adjusted according to the target voice length to obtain the target time window. Specifically, the target voice length can be divided by the length of the reference time window to obtain a first value. If the first value is an integer, the reference time window can be directly used as the target time window without adjusting. If the first value is not an integer, the size of the reference time window can be appropriately increased or decreased to obtain the target time window. Among them, the target voice length can be divided by the target time window. Specifically, the target voice length can be divided by the second value to obtain the target time window. For example, assuming that the target voice length is 10 and the reference time window is 3, the first value is approximately equal to 3.33. The reference time window can be reduced to 2.5 to obtain the target time window, or the reference time window can be increased to 5 to obtain the target time window.
[0154] Because pronunciation varies significantly across dialects, such as the four tones of Mandarin, the nine tones and six tones of Cantonese, and the voiced consonant characteristics of Wu, obtaining information on pronunciation characteristics specific to each dialect (such as pitch variation, syllable duration, and liaison rules) can help align the reference time window more closely with the natural rhythm of that dialect. For example, in Cantonese, where entering-tone characters are pronounced briefly, the reference time window can be narrowed to avoid segmenting short syllables. In Minnan dialect, where many connected words are present, the reference time window can be appropriately widened to preserve complete connected units.
[0155] S306. Determine, based on the m pieces of text data, the speech data that are not included in the basic speech database among the m pieces of speech data, to obtain n pieces of speech data; n is a natural number smaller than m.
[0156] In the embodiment of the present application, the basic voice database may include multiple voice data and text data corresponding to each voice data.
[0157] In a specific embodiment, each text data in the m text data can be compared with the text data in the basic voice database to find the text data not stored in the basic voice database, and obtain n text data. Then, the n voice data corresponding to these n text data in the m voice data can be determined.
[0158] S307: Add the n pieces of voice data to the basic voice database to obtain a target voice database; and update and train the preset AI model using the target voice database.
[0159] In an embodiment of the present application, the basic voice database can update these n voice data into its own database, thereby obtaining a target voice database; then, the preset AI model can also be updated and trained using the data in the target voice database.
[0160] It can be seen that the voice database updating method described in the present application converts the collected m voice data into text data to obtain m text data. The m text data are compared with the basic voice database to directly locate the unrecorded content, thereby obtaining n unrecorded voice data. Compared with audio feature comparison, text comparison is more efficient. Then, these n voice data are stored in the basic voice database to obtain the target voice database. The newly added n voice data are all the missing content in the basic voice database, which maximizes the effective data of each update, thereby improving the update efficiency of the voice database.
[0161] See also Figure 7 , Figure 7 This is a block diagram of the functional units of a speech database updating device 700 provided in an embodiment of the present application; the speech database updating device 700 can be applied to an electronic device having a preset AI model. The speech database updating device 700 includes: an acquisition unit 701 and a database updating unit 702, wherein:
[0162] The acquisition unit 701 is configured to acquire a basic voice database corresponding to the preset AI model; during the use of the preset AI model, voice data of different users are collected to obtain k pieces of voice data; k is a positive integer;
[0163] The database updating unit 702 is configured to filter out speech data of non-official language types from the k pieces of speech data to obtain m pieces of speech data, where m is a positive integer less than or equal to k; determine the m dialect types corresponding to the m pieces of speech data; convert each piece of speech data from the m pieces of speech data into text data based on the m dialect types to obtain m pieces of text data; determine the speech data from the m pieces of speech data that is not included in the basic speech database based on the m pieces of text data to obtain n pieces of speech data; where n is a natural number less than m; add the n pieces of speech data to the basic speech database to obtain a target speech database; and update and train the preset AI model using the target speech database.
[0164] In a specific implementation, the speech database updating apparatus 700 described in the embodiment of the present invention may also execute other implementations described in the speech database updating method provided in the above embodiment of the present invention, which will not be described in detail here.
[0165] See also Figure 8 , Figure 8 This is a structural diagram of an electronic device provided in an embodiment of the present application. The electronic device may include a processor, a memory, a communication interface, and one or more programs. The processor, memory, and communication interface may be interconnected via a bus; the one or more programs are stored in the memory and configured to be executed by the processor; in the embodiment of the present application, the one or more programs include instructions for executing other implementations described in the method for updating the voice database provided in the embodiment of the present invention, which will not be repeated here.
[0166] An embodiment of the present application also provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program for electronic data exchange, and the computer program enables a computer to execute part or all of the steps of any method described in the above method embodiments, and the above computer includes an electronic device.
[0167] The present application also provides a computer program product comprising a non-transitory computer-readable storage medium storing a computer program, wherein the computer program is operable to cause a computer to perform some or all of the steps of any of the methods described in the above method embodiments. The computer program product may be a software installation package, and the computer may comprise an electronic device.
[0168] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.
[0169] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0170] In the several embodiments provided in this application, it should be understood that the disclosed devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, and the indirect coupling or communication connection of devices or units can be electrical or other forms.
[0171] Those skilled in the art will appreciate that all or part of the process steps in the above-described method embodiments can be implemented by a computer program instructing the relevant hardware. The program can be stored in a computer-readable storage medium, and when executed, the program can include the process steps in the above-described method embodiments. The aforementioned storage medium includes various media capable of storing program code, such as ROM or random access memory (RAM), magnetic disks, or optical disks.
[0172] The steps of the method or algorithm described in the embodiments of the present application can be implemented in hardware or by a processor executing software instructions. The software instructions can be composed of corresponding software modules, which can be stored in RAM, flash memory, ROM, EPROM, electrically erasable programmable read-only memory (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium well known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and storage medium can be located in an ASIC. In addition, the ASIC can be located in a terminal device or a management device. Of course, the processor and storage medium can also be present in a terminal device or a management device as discrete components.
[0173] Those skilled in the art will appreciate that in one or more of the above examples, the functions described in the embodiments of the present application can be implemented in whole or in part through software, hardware, firmware, or any combination thereof. When implemented using software, they can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part.
[0174] The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more available media integrated therein.
[0175] The available media may be magnetic media (eg, floppy disks, hard disks, magnetic tapes), optical media (eg, digital video discs (DVDs)), or semiconductor media (eg, solid state disks (SSDs)).
[0176] The modules / units included in the devices and products described in the above embodiments may be software modules / units, hardware modules / units, or partly software modules / units and partly hardware modules / units. For example, for the devices and products applied to or integrated in the chip, the modules / units included therein may all be implemented in the form of hardware such as circuits, or at least part of the modules / units may be implemented in the form of software programs, which run on the processor integrated inside the chip, and the remaining (if any) modules / units may be implemented in the form of hardware such as circuits; for the devices and products applied to or integrated in the chip module, the modules / units included therein may all be implemented in the form of hardware such as circuits, and different modules / units may be located in the same component (such as chip, circuit module, etc.) or different components of the chip module, or at least part of the modules / units may be It is implemented in the form of a software program, which runs on the processor integrated inside the chip module, and the remaining (if any) modules / units can be implemented in the form of hardware such as circuits; for various devices and products applied to or integrated in the terminal equipment, the various modules / units contained therein can be implemented in the form of hardware such as circuits, and different modules / units can be located in the same component (for example, chip, circuit module, etc.) or different components in the terminal equipment, or, at least some modules / units can be implemented in the form of a software program, which runs on the processor integrated inside the terminal equipment, and the remaining (if any) modules / units can be implemented in the form of hardware such as circuits.
[0177] The specific implementation methods described above further illustrate the purpose, technical solutions and beneficial effects of the embodiments of the present application. It should be understood that the above description is only a specific implementation method of the embodiments of the present application and is not intended to limit the scope of protection of the embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the embodiments of the present application should be included in the scope of protection of the embodiments of the present application.
Claims
1. A method for updating a speech database, characterized in that: Applied to electronic devices, wherein the electronic devices are provided with a preset AI model, including: Obtaining a basic speech database corresponding to the preset AI model; During the use of the preset AI model, voice data of different users are collected to obtain k pieces of voice data; k is a positive integer; Filtering out speech data of non-official languages from the k pieces of speech data to obtain m pieces of speech data, where m is a positive integer less than or equal to k; Determine m dialect types corresponding to the m pieces of speech data; Convert each of the m pieces of voice data into text data based on the m dialect types to obtain m pieces of text data; Determine, based on the m pieces of text data, the speech data that are not included in the basic speech database among the m pieces of speech data, to obtain n pieces of speech data; n is a natural number smaller than m; The n pieces of voice data are added to the basic voice database to obtain a target voice database; and the preset AI model is updated and trained using the target voice database.
2. The method according to claim 1, wherein The obtaining of a basic speech database corresponding to the preset AI model includes: Obtaining the target application scenario of the preset AI model; Determine a professional voice database corresponding to the target application scenario; Determine the user type corresponding to the target application scenario, and obtain a user type; a is a positive integer; Determining the usage habits corresponding to each of the a user types to obtain a usage habits; The professional speech database is adjusted based on the a usage habits to obtain the basic speech database.
3. The method according to claim 2, wherein Each usage habit includes: a dialect type and personal usage preference information; the professional speech database is adjusted based on the a usage habits to obtain the basic speech database, including: Acquire a first professional voice; the first professional voice is any professional voice data in the professional voice database; Adjusting the first professional voice according to the dialect type of each of the a usage habits to obtain a dialect voices; Determine the preferred voice corresponding to the personal language preference information of each of the a language habits, and obtain i preferred voices, where i is an integer greater than or equal to a; b basic voices are determined according to the i preferred voices and the a dialect voices; b is an integer greater than a; the b basic voices are all basic voices in the basic voice database that are associated with the first professional voice.
4. The method according to claim 3, wherein The determining b basic voices according to the i preferred voices and the a dialect voices includes: Recognize the i preferred voices and the a dialect voices using semantic recognition technology to obtain i pieces of preferred semantic information and a pieces of dialect semantic information; Determining a first semantic similarity between each dialect semantic information in the a dialect semantic information and the i preference semantic information to obtain a first semantic similarity sets; each first semantic similarity set includes i first semantic similarities; Filtering out the first semantic similarity in each of the a first semantic similarity sets that is greater than a preset semantic similarity, to obtain c second semantic similarity sets; c is a positive integer less than or equal to a; Determining c preferred speech sets corresponding to the c second semantic similarity sets based on the i preferred speech sets; The preferred voices in the c preferred voice sets are randomly fused into the a dialect voices through a preset voice fusion method to obtain the b basic voices.
5. The method according to any one of claims 2 to 4, characterized in that The converting each of the m pieces of voice data into text data based on the m dialect types to obtain m pieces of text data includes: Acquire first voice data and its corresponding first dialect type; the first voice data is any voice data among the m voice data; performing data preprocessing on the first voice data to obtain second voice data; Determining a target speech length corresponding to the second speech data; When the target speech length is greater than or equal to the preset speech length, segmenting the second speech data according to the first dialect type and the target speech length to obtain d segments of partial speech data; d is an integer greater than 1; Converting the d segments of partial speech data into text data in parallel according to the first dialect type to obtain d segments of partial text data; Combining the d segments of partial text data in chronological order to obtain reference text data; Perform text calibration on the reference text data to obtain text data corresponding to the first voice data.
6. The method according to claim 5, wherein The segmenting of the second speech data according to the first dialect type and the target speech length to obtain d segments of partial speech data includes: Determine the target professional field dictionary corresponding to the target application scenario; Determine the time intervals corresponding to all professional keywords in the second voice data according to the target professional field dictionary, and obtain e time intervals; e is a natural number; determining a target time window according to the first dialect type and the target speech length; Determine the segmented time points in the second speech data according to the target time window, and obtain f segmented time points; f is a positive integer; Determine whether there is a segmented time point in the f segmented time points that is within the e time intervals; If so, determine the segmented time points and their corresponding time intervals among the f segmented time points, and obtain h segmented time points and h time intervals; h is a natural number less than e; Adjust the h segmentation time points according to the h time intervals so that each professional keyword is completely retained in the same segment of partial voice data, thereby obtaining h target segmentation time points; Segmenting the second voice data according to the h target segmentation time points and fh segmentation time points other than the h segmentation time points among the f segmentation time points to obtain the d segments of partial voice data; If not, the second voice data is segmented according to the f segmentation time points to obtain the d segments of partial voice data.
7. The method according to claim 6, wherein The determining of the target time window according to the first dialect type and the target speech length includes: Obtaining pronunciation feature information corresponding to the first dialect type; determining a reference time window according to the pronunciation feature information; The reference time window is adjusted according to the target speech length to obtain the target time window.
8. A device for updating a speech database, characterized in that: Applied to an electronic device, wherein a preset AI model is provided in the electronic device, the device comprises: an acquisition unit and a database update unit, wherein: The acquisition unit is configured to acquire a basic voice database corresponding to the preset AI model; during the use of the preset AI model, voice data of different users are collected to obtain k pieces of voice data; k is a positive integer; The database updating unit is configured to filter out speech data of non-official language types from the k pieces of speech data to obtain m pieces of speech data; m is a positive integer less than or equal to k; determine the m dialect types corresponding to the m pieces of speech data; convert each piece of speech data from the m pieces of speech data into text data based on the m dialect types to obtain m pieces of text data; determine the speech data from the m pieces of speech data that is not included in the basic speech database based on the m pieces of text data to obtain n pieces of speech data; n is a natural number less than m; add the n pieces of speech data to the basic speech database to obtain a target speech database; and update and train the preset AI model using the target speech database.
9. An electronic device, characterized in that: include: a processor, a memory, a communication interface, and one or more programs; The one or more programs are stored in the memory and configured to be executed by the processor, wherein the programs include instructions for executing the steps of the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that A computer program for electronic data exchange is stored, wherein the computer program enables a computer to execute the method according to any one of claims 1 to 7.