A method and device for generating pictures based on dialect speech

By employing multimodal feature fusion and semantic network mapping technologies, the accuracy problem of generating images from dialect speech was solved, achieving fully automated processing from dialect speech to images, thus improving the accuracy of generated images and user experience.

CN120612921BActive Publication Date: 2026-02-06ANRUI DIGITAL INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510807611.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2026-02-06
Estimated Expiration
2045-06-17

AI Technical Summary

Technical Problem

Existing technologies often fail to accurately recognize dialect speech when generating images, resulting in images that do not match the intended meaning and a poor user experience.

Method used

The acoustic features of dialect speech are extracted by a multimodal feature fusion model, mapped using a pre-built dialect speech dictionary and semantic network, and standardized Mandarin text is generated. Image elements are then combined using a semantic-driven generation model to achieve fully automated processing from dialect speech to images.

Benefits of technology

It improves the accuracy of generated images, reduces human intervention, and expands the acceptance among users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612921B_ABST
    Figure CN120612921B_ABST
Patent Text Reader

Abstract

The application discloses a method and device for generating pictures based on dialect speech. The method comprises the following steps: extracting acoustic features of to-be-processed dialect speech through a multi-modal feature fusion model; finding a dialect corresponding to the acoustic features based on a pre-constructed dialect speech dictionary to generate dialect text of the to-be-processed dialect speech; mapping the dialect text according to a pre-constructed dialect semantic network to obtain standardized Mandarin text corresponding to the dialect text; extracting keywords of the standardized Mandarin text; and combining image elements of the extracted keywords by using a pre-constructed semantic-driven generation model to generate a picture corresponding to the to-be-processed dialect speech.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of speech recognition, and more particularly, to a method and device for generating pictures based on dialect speech. BACKGROUND

[0002] In modern society, the inheritance and protection of dialects have become particularly important. With the development of technology, more and more tools and software have emerged to help us better identify and record dialects. Among them, the software for converting dialect videos into pictures has become a field of concern. In the process of using, when generating pictures using the speech of the dialect, the system recognition is not accurate, there are differences, which leads to the generated pictures not corresponding to the original intention, thereby affecting the response experience. SUMMARY

[0003] In view of the deficiencies of the prior art, the present application provides a method and device for generating pictures based on dialect speech.

[0004] According to one aspect of the present application, a method for generating pictures based on dialect speech is provided, comprising:

[0005] extracting acoustic features of the dialect speech to be processed by a multi-modal feature fusion model;

[0006] finding the dialect corresponding to the acoustic features based on a pre-constructed dialect speech dictionary, and generating a dialect text of the dialect speech to be processed;

[0007] mapping the dialect text according to a pre-constructed dialect semantic network to obtain a standardized Mandarin text corresponding to the dialect text;

[0008] extracting keywords of the standardized Mandarin text;

[0009] adopting a pre-constructed semantic-driven generation model to combine the extracted keywords with image elements to generate a picture corresponding to the dialect speech to be processed.

[0010] Optionally, it further comprises adjusting the mapping of the dialect dictionary according to the feedback of the generated picture using an interactive learning mechanism.

[0011] Optionally, the adaptive learning expression of the interactive learning mechanism is Model_{t+1}=Model_t+α·ΔModel, where Model_t is the current model, a is the learning rate, and △Model is the model update amount calculated according to the dialect characteristics.

[0012] Optionally, the mapping of the dialect text according to the pre-constructed dialect semantic network to obtain a standardized Mandarin text corresponding to the dialect text comprises:

[0013] finding the basic Mandarin vocabulary of the dialect text using the dialect semantic network;

[0014] The ambiguity in the basic common language vocabulary is identified and corrected by using the context awareness model, and the standardized common text corresponding to the dialect text is obtained.

[0015] According to another aspect of the present application, a device for generating a picture based on a dialect voice is provided, comprising:

[0016] A first extraction module is configured to extract acoustic features of the to-be-processed dialect voice by using a multi-modal feature fusion model;

[0017] A first generation module is configured to find a dialect corresponding to the acoustic features based on a pre-constructed dialect voice dictionary, and generate a dialect text of the to-be-processed dialect voice;

[0018] A mapping module is configured to map the dialect text based on a pre-constructed dialect semantic network, and obtain a standardized common language text corresponding to the dialect text;

[0019] A second extraction module is configured to extract keywords of the standardized common language text;

[0020] A second generation module is configured to combine image elements of the extracted keywords by using a pre-constructed semantic-driven generation model, and generate a picture corresponding to the to-be-processed dialect voice.

[0021] According to still another aspect of the present application, a computer readable storage medium is provided, which stores a computer program for executing the method according to any one of the above aspects of the present application.

[0022] According to still another aspect of the present application, an electronic device is provided, which comprises a processor, a memory for storing executable instructions of the processor, and the processor is configured to read the executable instructions from the memory and execute the instructions to implement the method according to any one of the above aspects of the present application.

[0023] Therefore, the present application realizes the full-process automatic processing from the dialect voice input to the picture output, reduces the manual intervention, improves the accuracy of the generated picture, and improves the promotion of the user group. BRIEF DESCRIPTION OF DRAWINGS

[0024] The exemplary embodiments of the present application can be more completely understood by reference to the following drawings:

[0025] Figure 1 is a flowchart of a method for generating a picture based on a dialect voice according to an exemplary embodiment of the present application;

[0026] Figure 2 is a structural schematic diagram of a device for generating a picture based on a dialect voice according to an exemplary embodiment of the present application;

[0027] Figure 3 is a structure of an electronic device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION

[0028] Hereinafter, exemplary embodiments according to the present application will be described in detail with reference to the accompanying drawings. It should be apparent to those skilled in the art that the following described embodiments are only exemplary of the present application and should not be considered to narrow the scope of the present application.

[0029] It should be noted that the relative arrangement of the components and steps, numerical expressions, and numerical values set forth in these embodiments are not limitations on the scope of the application unless otherwise specifically stated.

[0030] Those skilled in the art will understand that the terms "first", "second", and so on in the embodiments of the present application are only used to distinguish different steps, devices or modules, and do not represent any specific technical meaning, nor do they represent a necessary logical order between them.

[0031] It should also be understood that in the embodiments of the present application, "a plurality of" can mean two or more, and "at least one" can mean one, two or more.

[0032] It should also be understood that for any component, data or structure mentioned in the embodiments of the present application, unless specifically limited or if the context clearly indicates otherwise, it can be understood as one or more in general.

[0033] In addition, the term "and / or" in the present application is only a description of the association relationship between the associated objects, which means that there can be three relationships, for example, A and / or B can represent the existence of A alone, the existence of A and B together, and the existence of B alone. In addition, the character " / " in the present application generally represents an "or" relationship between the front and rear associated objects.

[0034] It should also be understood that the description of each embodiment of the present application emphasizes the differences between each embodiment, and the same or similar parts can be referred to each other, and for the sake of brevity, will not be repeated.

[0035] At the same time, it should be understood that for the sake of brevity, the size of each part shown in the drawings is not drawn in accordance with the actual proportional relationship.

[0036] The following description of at least one exemplary embodiment is merely exemplary in nature and is in no way intended to limit the present application or its application or uses.

[0037] Techniques, methods, and devices known to those of ordinary skill in the relevant art can not be discussed in detail herein, but should be considered as part of the specification in appropriate circumstances.

[0038] It should be noted that like reference numerals and characters refer to like elements throughout the following figures and the detailed description, not upon the drawings alone such that, when certain terms are used in the detailed description and / or in the drawings that are common to the figures, they are intended to be interpreted identically, unless otherwise clear from the context.

[0039] Embodiments of the present application can be applied to terminal devices, computer systems, servers, and other electronic devices, which can operate with many other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with terminal devices, computer systems, servers, and other electronic devices include, but are not limited to, personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network personal computers, minicomputer systems, mainframe computer systems, and distributed cloud computing technology environments that include any of the above systems, and the like.

[0040] Terminal devices, computer systems, servers, and other electronic devices can be described in the general context of computer system-executable instructions, such as program modules, being executed by a computer system. Generally, program modules can include routines, programs, objects, components, logic, data structures, and the like, which perform particular tasks or implement particular abstract data types. Computer systems / servers can be implemented in a distributed cloud computing environment, in which tasks are performed by remote processing devices that are linked through a communications network. In a distributed cloud computing environment, program modules can be located in local or remote computer system storage media including storage devices.

[0041] Exemplary method

[0042] Figure 1 is a flowchart of a method for generating a picture based on a dialect voice provided by an exemplary embodiment of the present application. The present embodiment can be applied to electronic devices such as Figure 1 As shown in FIG. 1, the method 100 for generating a picture based on a dialect voice includes the following steps:

[0043] Step 101, extracting acoustic features of the to-be-processed dialect voice through a multi-modal feature fusion model;

[0044] Step 102, finding the dialect corresponding to the acoustic features based on a pre-constructed dialect voice dictionary, and generating dialect text of the to-be-processed dialect voice;

[0045] Step 103, mapping the dialect text according to a pre-constructed dialect semantic network to obtain standardized Mandarin text corresponding to the dialect text;

[0046] Step 104, extracting keywords of the standardized Mandarin text;

[0047] Step 105, using the pre-constructed semantic-driven generation model to combine the extracted keywords with image elements, generating the corresponding picture for the dialect speech to be processed.

[0048] Specifically, the solution idea of the present application is to first recognize and translate classical Chinese into Mandarin, and then perform the picture generation function. The solution is as follows:

[0049] I. Dialect recognition:

[0050] 1. Dialect dictionary construction: First, a dialect dictionary containing various dialect words and phrases needs to be constructed so that the system can accurately recognize dialects.

[0051] 2. Dialect feature extraction: Design an algorithm or model to identify dialects through speech feature extraction techniques, which can use acoustic models or deep learning models.

[0052] The specific implementation is as follows:

[0053] 1) Input association: Receive the dialect text output by the dialect speech recognition module.

[0054] 2) Core technology:

[0055] Obtain basic Mandarin vocabulary through dialect semantic network lookup;

[0056] Use context-aware models to resolve ambiguities (e.g., "Ama" in Minnan can be "mother" or "grandmother" depending on the subsequent verb).

[0057] 3) Output association: Deliver the structured semantic vector to the image generation module.

[0058] II. Picture generation:

[0059] 1. Image database: Establish an image database containing various dialect-related scenes for generating pictures related to dialects.

[0060] 2. Image generation algorithm: Design an image generation algorithm to generate corresponding pictures based on input dialect content, which can use techniques such as generative adversarial networks (GAN) or variational autoencoder (VAE).

[0061] III. Speech and picture association:

[0062] 1. Speech content analysis: Convert the dialect speech to text and extract key information and context.

[0063] 2. Picture content matching: Match appropriate image elements such as scenes, characters, objects, etc. based on dialect content to generate associated pictures.

[0064] IV. System Optimization:

[0065] 1. Model Training: Establish a training set that includes dialect voices and corresponding pictures to train the system's ability to generate dialect-related pictures.

[0066] 2. User Experience Optimization: Optimize the user experience of the system to ensure that the generated pictures are associated with the dialect content and meet user expectations.

[0067] V. Real-time Testing and Adjustment:

[0068] 1. System Testing: Conduct system testing in actual applications to evaluate the accuracy and effectiveness of the system in generating dialect-related pictures.

[0069] 2. Parameter Adjustment: Adjust and optimize the system parameters according to the test results to improve the system's performance and user satisfaction.

[0070] In an embodiment of the present invention, a detailed technical process description:

[0071] 1. Dialect Speech Recognition

[0072] 1) Input: User's dialect voice (e.g., Cantonese "佢喺度食紧饭").

[0073] 2) Processing:

[0074] Extract acoustic features through a multi-modal feature fusion model (Feature = MFCC⊕LPC⊕DeepFeature);

[0075] Call the dialect speech dictionary (phoneme-level mapping, e.g., Cantonese / jy□13 / → "佢").

[0076] 3) Output: Dialect text.

[0077] 2. Dialect → Mandarin Conversion (Core Bridge)

[0078] 1) Input: Dialect text.

[0079] '2) Processing:

[0080] Map the dialect semantic network (e.g., Cantonese "佢" → Mandarin "他");

[0081] The context-aware model corrects the semantics (e.g., "食紧饭" → progressive tense "正在吃饭").

[0082] 3) Output: Standardized Mandarin text.

[0083] 4. Semantic Understanding and Image Generation

[0084] 1) Input: Mandarin text.

[0085] 2) Process:

[0086] Extract keywords ("he" + "eat" + "in progress");

[0087] Semantic-driven generation of model (GAN / VAE) combines image elements:

[0088] Person library -> Asian male

[0089] Action library -> dining action

[0090] Scene library -> restaurant

[0091] 3) Output: Generate picture (Asian male dining in restaurant scene).

[0092] 4. Adaptive optimization closed loop

[0093] 1) User feedback "picture does not match" -> trigger interactive learning mechanism;

[0094] 2) Adjust dialect dictionary mapping (such as associating "eat tight rice" with more accurate animation frames);

[0095] 3) Adaptive learning formula: Model_{t+1} = Model_t + α·ΔModel, where Modelt is the current model, α is the learning rate, and △Model is the model update amount calculated according to the dialect characteristics.

[0096] Further, the innovation of the present application:

[0097] 1. Innovative formula

[0098] Dialect feature extraction formula: Feature = MFCC ⊕ LPC ⊕ DeepFeature

[0099] Where ⊕ represents feature fusion operation, MFCC and LPC are traditional speech features, and DeepFeature is high-level feature extracted by deep learning model.

[0100] Adaptive learning formula: \[\text{Model}_{t+1}=\text{Model}t+\alpha\cdot\Delta\text{Model}\] where \(\text{Model}{t+1}\) is the updated model, Modelt is the current model, α is the learning rate, and △Model is the model update amount calculated according to the dialect characteristics.

[0101] 2. Automation

[0102] Full process automation: Achieve full process automation from dialect speech input to picture output, reduce manual intervention.

[0103] Intelligent optimization algorithm: adopt intelligent optimization algorithm such as genetic algorithm or particle swarm optimization to automatically adjust model parameters and improve system performance.

[0104] 3. Operation innovation

[0105] Interactive learning: design an interactive learning mechanism to allow users to adjust system behavior through feedback and achieve personalized service.

[0106] Multi-task learning: use multi-task learning framework to optimize multiple tasks such as speech recognition, semantic understanding and image generation at the same time, and improve overall efficiency.

[0107] Therefore, the present application realizes the whole process automation from dialect voice input to picture output, reduces manual intervention, improves the accuracy of generated pictures, and improves the promotion of the user group.

[0108] Exemplary apparatus

[0109] Figure 2 is a structural schematic diagram of the device for generating pictures based on dialect voice provided by an exemplary embodiment of the present application. As shown in Figure 2 , the device 200 includes:

[0110] The first extraction module 210 is configured to extract the acoustic features of the to-be-processed dialect voice through a multi-modal feature fusion model.

[0111] The first generation module 220 is configured to find the dialect corresponding to the acoustic features based on a pre-constructed dialect voice dictionary, and generate the dialect text of the to-be-processed dialect voice.

[0112] The mapping module 230 is configured to map the dialect text according to a pre-constructed dialect semantic network to obtain the standardized Mandarin text corresponding to the dialect text.

[0113] The second extraction module 240 is configured to extract the keywords of the standardized Mandarin text.

[0114] The second generation module 250 is configured to combine the extracted keywords with image elements by using a pre-constructed semantic-driven generation model to generate the picture corresponding to the to-be-processed dialect voice.

[0115] Optionally, the device 200 further includes an adjustment module configured to adjust the mapping of the dialect dictionary by using an interactive learning mechanism according to the feedback of the generated picture.

[0116] Optionally, the adaptive learning expression of the interactive learning mechanism is Model_{t+1}=Model_t+α·ΔModel, where Model_t is the current model, α is the learning rate, and △Model is the model update amount calculated according to the dialect characteristics.

[0117] Optionally, the mapping module 230 comprises:

[0118] a searching sub-module, configured to search for the basic Putonghua words of the dialect text by using the dialect semantic network;

[0119] a correcting sub-module, configured to identify and correct the ambiguities in the basic Putonghua words by using the context-aware model, so as to obtain the standardized Putonghua text corresponding to the dialect text.

[0120] Exemplary electronic device

[0121] Figure 3 is a structure of an electronic device provided by an exemplary embodiment of the present application. As shown in Figure 3 The electronic device 30 comprises one or more processors 31 and a memory 32.

[0122] The processor 31 can be a central processing unit (CPU) or other forms of processing units having data processing and / or instruction execution capabilities, and can control other components in the electronic device to perform desired functions.

[0123] The memory 32 can comprise one or more computer program products, which can comprise various forms of computer readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory can comprise, for example, random access memory (RAM), cache memory and / or the like. The non-volatile memory can comprise, for example, read-only memory (ROM), hard disk, flash memory and / or the like. One or more computer program instructions can be stored on the computer readable storage media, and the processor 31 can run the program instructions to implement the methods of the software programs of the various embodiments of the present application described above and / or other desired functions. In one example, the electronic device can further comprise an input device 33 and an output device 34, which are interconnected by a bus system and / or other forms of connection mechanism (not shown).

[0124] In addition, the input device 33 can further comprise, for example, a keyboard, a mouse and / or the like.

[0125] The output device 34 can output various information to the outside. The output device 34 can comprise, for example, a display, a speaker, a printer, a communication network and a remote output device connected thereto, and / or the like.

[0126] Of course, in order to simplify, Figure 3Only some of the components of the electronic device related to the present application are shown, and components such as buses, input / output interfaces, and the like are omitted. In addition, the electronic device can include any other appropriate components depending on the specific application.

[0127] Exemplary computer program product and computer readable storage medium

[0128] In addition to the methods and devices described above, an embodiment of the present application can also be a computer program product that includes computer program instructions that when run by a processor cause the processor to perform steps of the methods according to various embodiments of the present application described in the above "Exemplary Methods" section of the specification.

[0129] The computer program product can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, C++, etc., and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computing device, partly on the user's device, as a stand-alone software package, partly on the user's computing device and partly on a remote computing device or entirely on the remote computing device or server.

[0130] In addition, an embodiment of the present application can also be a computer readable storage medium having stored thereon computer program instructions that when run by a processor cause the processor to perform steps of the methods according to various embodiments of the present application described in the above "Exemplary Methods" section of the specification.

[0131] The computer readable storage medium can be any combination of one or more computer readable media. The computer readable medium can be a computer readable signal medium or a computer readable storage medium. The computer readable storage medium can include, for example, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the above. More specific examples (a non-exhaustive list) of the computer readable storage medium include an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0132] The above describes the basic principles of the present application in conjunction with specific embodiments, but it should be noted that the advantages, benefits, effects and the like mentioned in the present application are only examples and are not limiting, and these advantages, benefits, effects and the like should not be considered as necessary for each embodiment of the present application. In addition, the above specific details disclosed are only for the purpose of illustration and understanding, and are not limiting, and the above details do not limit the present application to be necessarily implemented with the above specific details.

[0133] Each embodiment in the specification is described in a progressive manner, and each embodiment focuses on the difference from other embodiments, and the same or similar parts between each embodiment can be understood by referring to each other. For the system embodiment, since it basically corresponds to the method embodiment, the description is relatively simple, and the relevant parts can be understood by referring to the part of the method embodiment.

[0134] The block diagrams of the devices, systems, apparatuses, systems involved in the present application are only exemplary examples and are not intended to require or imply the connection, arrangement, configuration shown in the block diagram. As those skilled in the art will recognize, these devices, systems, apparatuses, systems can be connected, arranged, configured in any manner. Words such as "include", "contain", "have" and the like are open-ended words, which mean "include but not limited to", and can be used interchangeably. The words "or" and "and" used herein mean the word "and / or", and can be used interchangeably unless the context clearly indicates otherwise. The word "such as" used herein means the phrase "such as but not limited to", and can be used interchangeably.

[0135] The method and system of the present application can be implemented in many ways. For example, the method and system of the present application can be implemented by software, hardware, firmware or any combination of software, hardware and firmware. The above order of steps for the method is only for illustration, and the steps of the method of the present application are not limited to the above specific description, unless otherwise specifically described. In addition, in some embodiments, the present application can also be implemented as programs recorded in recording media, which include machine-readable instructions for implementing the method according to the present application. Therefore, the present application also covers the recording media storing the programs for executing the method according to the present application.

[0136] It should also be noted that in the systems, apparatus, and methods of the present invention, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered equivalents of the present invention. The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use the invention. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of the invention. Therefore, the invention is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.

[0137] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of the invention to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.

Claims

1. A method of generating a picture based on a dialect voice, characterized by, The method comprises the following steps: extracting acoustic features of the to-be-processed dialect speech through a multi-modal feature fusion model; finding the dialect corresponding to the acoustic features based on a pre-constructed dialect speech dictionary to generate dialect text of the to-be-processed dialect speech; mapping the dialect text according to a pre-constructed dialect semantic network to obtain standardized Mandarin text corresponding to the dialect text; extracting keywords of the standardized Mandarin text; generating a picture corresponding to the to-be-processed dialect speech by combining the extracted keywords using a pre-constructed semantic-driven generation model; adjusting the mapping of the dialect speech dictionary using an interactive learning mechanism according to the feedback of the generated picture; The adaptive learning expression of the interactive learning mechanism is: Model_{t+1} = Model_t + α·ΔModel, where Model_t is the current model, α is the learning rate, and ΔModel is the model update amount calculated according to the dialect characteristics.

2. The method of claim 1, wherein, The method of mapping the dialect text according to a pre-constructed dialect semantic network to obtain standardized Mandarin text corresponding to the dialect text comprises: finding basic Mandarin vocabulary of the dialect text using the dialect semantic network; identifying and correcting ambiguities in the basic Mandarin vocabulary using a context-aware model to obtain the standardized Mandarin text corresponding to the dialect text.

3. An apparatus for generating images based on dialect speech, used to implement the method described in any one of claims 1-2, characterized in that, The method comprises the following steps: a first extraction module for extracting acoustic features of the to-be-processed dialect speech through a multi-modal feature fusion model; a first generation module for finding the dialect corresponding to the acoustic features based on a pre-constructed dialect speech dictionary to generate dialect text of the to-be-processed dialect speech; a mapping module for mapping the dialect text according to a pre-constructed dialect semantic network to obtain standardized Mandarin text corresponding to the dialect text; a second extraction module for extracting keywords of the standardized Mandarin text; a second generation module for generating a picture corresponding to the to-be-processed dialect speech by combining the extracted keywords using a pre-constructed semantic-driven generation model.

4. The apparatus of claim 3, wherein, Further comprising: an adjustment module for adjusting the mapping of the dialect speech dictionary using an interactive learning mechanism according to the feedback of the generated picture.

5. The apparatus of claim 4, wherein, The adaptive learning expression of the interactive learning mechanism is: Model_{t+1} = Model_t + α·ΔModel, where Model_t is the current model, α is the learning rate, and ΔModel is the model update amount calculated according to the dialect characteristics.

6. The apparatus of claim 4, wherein, The mapping module comprises: a finding sub-module for finding basic Mandarin vocabulary of the dialect text using the dialect semantic network; a correction sub-module for identifying and correcting ambiguities in the basic Mandarin vocabulary using a context-aware model to obtain the standardized Mandarin text corresponding to the dialect text.

7. A computer readable storage medium characterized in that, The storage medium stores a computer program for executing the method of any one of claims 1-2.

8. An electronic device, comprising: The electronic device comprises: a processor; a memory for storing instructions executable by the processor; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method of any of claims 1-2.

Citation Information

Patent Citations

  • Speech recognition and conversion system

    CN117351938A

  • Head portrait generation method and device, electronic equipment, medium and product

    CN119169132A