Electric power speech synthesis model optimization method and device, and storage medium
By building a power speech synthesis index system and user feedback knowledge base, determining the optimization index sequence and implementing the optimization steps, the problems of poor optimization effects of the power speech synthesis model and low accuracy of speech synthesis in the prior art are solved, and more efficient and accurate speech synthesis model optimization is achieved.
Patent Information
- Application Number
- CN202510161444.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-07-01
AI Technical Summary
The existing power speech synthesis model optimization methods fail to comprehensively evaluate speech synthesis performance, resulting in poor optimization effects and low accuracy of speech synthesis, and the optimization of multiple indicators leads to overlapping performance coupling.
Build a power speech synthesis index system, including first-level indicators such as knowledge, semantics, and rhythm, and determine the optimization index sequence and implement the optimization steps through user feedback knowledge base to avoid the speech coupling phenomenon caused by unified update.
By partially updating each voice, voice coupling is avoided, and the optimization effect of the power speech synthesis model and the accuracy of speech synthesis are improved.
Smart Images

Figure CN120236560A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of optimizing speech synthesis models, and particularly relates to a method for optimizing a power speech synthesis model, a storage device, and a storage medium. Background Art
[0002] With the continuous progress of artificial intelligence technology, the degree of intelligence in the power industry is also getting higher and higher. As the core link of the power system, power grid dispatching has extremely high requirements for the accurate transmission and rapid response of information. Under such a background, power speech synthesis technology has emerged, which converts the instructions and information of power grid dispatching into clear and fluent speech output by simulating human voices.
[0003] Existing methods for optimizing power speech synthesis models usually evaluate the performance of speech synthesis from a standard perspective. Each time the model is optimized, the speech parameters in the model are usually optimized uniformly to facilitate the update and optimization of the model in a short time.
[0004] However, the existing methods for optimizing power speech synthesis models do not consider multi-dimensional indicators such as knowledge, semantics, and prosody, and cannot comprehensively evaluate the performance of power speech synthesis. At the same time, the existing methods also have the problem of performance coupling and overlap caused by multi-index optimization, that is, in the process of multi-index optimization, due to the coupling relationship between different indicators, optimizing one indicator may affect the performance of other indicators or even cause negative optimization. At the same time, due to the possible overlapping information between indicators, this further increases the complexity and uncertainty of the optimization process, resulting in poor optimization effect of single-dimensional indicators of the power speech synthesis model and low accuracy of speech synthesis. Summary of the Invention
[0005] The present invention aims at the problems in the prior art, and provides a method for optimizing a power speech synthesis model, a storage device, and a storage medium, which solves the problems of poor optimization effect and low accuracy of speech synthesis caused by single-dimensional optimization of the speech synthesis model in the power system field in the prior art.
[0006] The technical solution adopted by the present invention is as follows: In a first aspect, the present application provides a method for optimizing a power speech synthesis model, including the following steps: Step S1: Construct a power speech synthesis index system: construct at least two first-level indicators, and at least two second-level indicators are included under each first-level indicator; Step S2: After the power synthesized speech is released, collect user interaction statements and construct a user feedback knowledge base; Step S3: Classify and score the user interaction statements in the user feedback knowledge base based on the second-level indicators, and determine an optimization index sequence based on the scoring results; Step S4: Perform optimizations in sequence according to the optimization index sequence.
[0007] Further, step S3 includes the following steps: Step S3-1: Classify the feature keywords of the user interaction statements collected in the user feedback knowledge base, and divide the corresponding user feedback statements into different secondary indicators according to the feature keywords to generate the corresponding secondary indicator statement sets; Step S3-2: According to the positivity and negativity of the user evaluations, divide the user interaction statements in each secondary indicator into positive feedback statements, neutral statements, and negative feedback statements, and delete the positive feedback statements and neutral statements from the secondary indicator statement sets; Step S3-3: Score each negative feedback statement to obtain the score set of the secondary indicators , indicating the negative feedback score set of the negative feedback statements related to the nth secondary indicator in the ith primary indicator; Step S3-4: Traverse , find the minimum value in and denote it as , and sort
[0008] from small to large to generate the optimization index sequence. Further, step S4 includes the following steps: ; Step S4-1: Determine the indicator to be optimized according to the optimization index sequence ; Step S4-2: Determine the optimization step size corresponding to the indicator to be optimized; perform optimization on the indicator to be optimized ; Step S4-3: After the model parameter optimization is completed, delete the negative feedback statement with the lowest score from the negative feedback statement set corresponding to the indicator to be optimized and synchronously update the negative feedback score set of the indicator to be optimized ; Step S4-4: Judge whether the lowest negative feedback score in the negative feedback score set of the updated indicator to be optimized in step S4-3 is the maximum value in the negative feedback score set with the largest absolute value corresponding to each secondary indicator in step S3-4. If the lowest negative feedback score in the negative feedback score set of the updated indicator to be optimized is the maximum value in the negative feedback score set with the largest absolute value corresponding to each secondary indicator, then perform the optimization of the next secondary indicator in the optimization index sequence, otherwise jump to step S4-2.
[0009] Further, in step S4-1, the indicator to be optimized To optimize the secondary indicator corresponding to the lowest negative feedback score in the indicator sequence.
[0010] Further, in step S4-2, the indicator to be optimized The set of negative feedback scores corresponding to The negative feedback score with the largest absolute value in is denoted as , The corresponding negative feedback statement is denoted as , according to , calculate the optimization step size of the power voice synthesis model ; Specifically expressed as: ; Where Is the maximum optimization step size, Is the maximum reference score.
[0011] Further, step S4 also includes the following steps: Step S4-5: If the lowest negative feedback score in the set of negative feedback scores of the updated indicator to be optimized in step S4-4 Is the maximum value in the set of negative feedback scores with the largest absolute value corresponding to each secondary indicator, jump to step S4-6; otherwise, after the set of negative feedback scores of the indicator to be optimized Is updated, it is published, an observation window W is established to re-determine the negative feedback situation of the indicator to be optimized , and the negative feedback of the updated and published indicator to be optimized within the observation window W is monitored . If there is no negative feedback situation regarding the indicator to be optimized within the observation window W , then jump to step S4-6; if there is still a negative feedback score for the indicator to be optimized , then jump to step S4-2; Step S4-6: Delete the secondary indicator corresponding to the indicator to be optimized specified in step S4-1 , and at the same time update the set of secondary indicator negative feedback scores in step S3-3 . If there are secondary indicators with negative feedback scores in the set of secondary indicator negative feedback scores , jump to step S3-4; otherwise, the optimization of the power voice synthesis model is completed.
[0012] In the second aspect, the present application provides a power voice synthesis model optimization device, the power voice synthesis model optimization device includes a processor, a memory, and a computer program stored on the memory and executable by the processor, wherein when the computer program is executed by the processor, the steps of the power voice synthesis model optimization method described in the first aspect are implemented.
[0013] In a third aspect, the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the power voice synthesis model optimization method as described in the first aspect are implemented.
[0014] As can be seen from the above technical solutions, the present invention has the following advantages: By constructing a power voice synthesis index system and a user feedback database, it is possible to perform partial updates for each piece of voice, avoiding the voice coupling phenomenon caused by unified updates, and solving the problems of poor optimization effect and low accuracy of voice synthesis in the prior art for single-dimensional optimization of voice synthesis models in the field of power systems. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0016] Figure 1 It is the index system diagram in the embodiment of the specific implementation manner of the present invention.
[0017] Figure 2 It is the flowchart of the update method in the embodiment of the specific implementation manner of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] In the following detailed description, various embodiments of the present disclosure will be described more fully. The present disclosure may have various embodiments, and adjustments and changes can be made therein. However, it should be understood that there is no intention to limit the various embodiments of the present disclosure to the specific embodiments disclosed herein, but rather the present disclosure should be understood to cover all adjustments, equivalents, and / or alternative solutions falling within the spirit and scope of the various embodiments of the present disclosure.
[0019] In order to facilitate a clear description of the technical solutions of the embodiments of the present application, the following briefly introduces some terms and technologies involved in the embodiments of the present application: 1. Keyword detection (Spoken Keyword Spotting or Spoken Term Detection) is a subfield in the field of speech recognition, and its purpose is to detect all occurrence positions of a specified word in a speech signal. This method is particularly important in human-computer interaction. For example, when a smart device detects a specific wake-up word, it will activate the voice interaction function.
[0020] 2. Speech Synthesis Model: Also known as the Text-to-Speech (TTS) model, it is a technical model that can convert text information into spoken language voice output. The following is a detailed introduction to the speech synthesis model: The speech synthesis model simulates the human vocalization process and converts the input text information into natural and fluent voice signals. Its basic principles include text analysis, acoustic feature generation, and vocoder synthesis, etc.: Text Analysis: This step preprocesses the input text, including text error correction, normalization, word segmentation, phoneme conversion, and prosody parsing, etc., and converts the text into linguistic features suitable for speech synthesis; Acoustic Feature Generation: According to the linguistic features obtained from text analysis, corresponding acoustic features are generated, such as spectrum, fundamental frequency (pitch), duration, etc. These acoustic features determine the voice quality, pitch, and speech rate, etc. of the voice; Vocoder Synthesis: The vocoder is used to convert the acoustic features into voice signals. The vocoder is an algorithm or model that can convert acoustic parameters into voice waveforms.
[0021] 3. Performance Coupling: Performance coupling refers to the mutual dependence and influence between systems or components in terms of performance. When there is a direct or indirect association between the performance parameters of two or more systems or components, performance coupling is formed. This coupling relationship may cause the performance change of one system or component to affect the performance of other systems or components.
[0022] Existing optimization methods for power speech synthesis models usually evaluate the speech synthesis performance from a standard perspective. Each time the model is optimized, the speech parameters in the model are often uniformly optimized. However, this optimization method does not consider multi-dimensional indicators such as knowledge, semantics, and prosody, and cannot comprehensively evaluate the power speech synthesis performance. At the same time, there is also the problem of overlapping performance coupling caused by multi-index optimization, which further increases the complexity and uncertainty of the optimization process, resulting in poor optimization effects of single-dimensional indicators of the power speech synthesis model and low accuracy of speech synthesis.
[0023] The present invention provides an optimization method, a storage device, and a storage medium for a power speech synthesis model to solve the problems of poor optimization effects and low accuracy of speech synthesis caused by single-dimensional optimization of the speech synthesis model in the prior art in the field of power systems.
[0024] Hereinafter, the term "comprising" or "may comprise" that may be used in various embodiments of the present disclosure indicates the presence of the disclosed functions, operations, or elements, and does not limit the addition of one or more functions, operations, or elements. In addition, as used in various embodiments of the present disclosure, the terms "comprising", "having", and their cognates are only intended to indicate specific features, numbers, steps, operations, elements, components, or combinations of the foregoing items, and should not be construed as precluding the existence or addition of one or more other features, numbers, steps, operations, elements, components, or combinations of the foregoing items first.
[0025] In various embodiments of the present disclosure, the expression "or" or "at least one of A or / and B" includes any combination or all combinations of the recited words. For example, the expression "A or B" or "at least one of A or / and B" may include A, may include B, or may include both A and B.
[0026] Expressions (such as "first", "second", etc.) used in various embodiments of the present disclosure may modify various constituent elements in various embodiments, but do not limit the corresponding constituent elements. For example, the above expressions do not limit the order and / or importance of the elements. The above expressions are only for the purpose of distinguishing one element from other elements. For example, the first user device and the second user device indicate different user devices, although both are user devices. For example, without departing from the scope of various embodiments of the present disclosure, the first element may be referred to as the second element, and similarly, the second element may also be referred to as the first element.
[0027] It should be noted that: if it is described that one constituent element is "connected" to another constituent element, the first constituent element may be directly connected to the second constituent element, and a third constituent element may be "connected" between the first constituent element and the second constituent element. Conversely, when one constituent element is "directly connected" to another constituent element, it can be understood that there is no third constituent element between the first constituent element and the second constituent element.
[0028] Hereinafter, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0029] As Figure 1 and Figure 2 shown, this embodiment provides a method for optimizing a power voice synthesis model, including the following steps: Step S1: Construct an index system for electric power speech synthesis; This application constructs at least two first-level indicators. In this embodiment, three first-level indicators of knowledge, semantics, and prosody are constructed, and each first-level indicator includes at least two second-level indicators; Among them, in this application, the knowledge indicator in the first-level indicators is defined , the semantic indicator , the prosody indicator . Among them, the knowledge in the first-level indicators involves the richness and accuracy of the information contained in the synthesized speech. The semantics in the first-level indicators involves the clarity and accuracy of the synthesized speech at the semantic level. The prosody in the first-level indicators involves the naturalness and auditory comfort of the synthesized speech; Furthermore, the knowledge in the first-level indicators includes knowledge quantity , knowledge density , ……, knowledge matching degree and other second-level indicators. Among them, the knowledge quantity is used to measure the total amount of information contained in the synthesized speech. The knowledge density is used to measure the density of knowledge information transmitted per unit time. The knowledge matching degree is used to measure the matching degree between the content of the synthesized speech and the user's feedback statement; The semantics in the first-level indicators includes ambiguity degree , service attitude , connection response , ……, semantic affirmation degree and other second-level indicators. Among them, the ambiguity degree is used to measure the proportion of content in the synthesized speech that may cause multiple interpretations. The service attitude is used to measure the user service attitude shown when the synthesized speech transmits information. The connection response is used to measure the fluency and logic of the semantic context connection in the synthesized speech during a conversation. The semantic affirmation degree is used to measure the certainty and confidence of the statements in the synthesized speech; The prosody in the first-level indicators includes speech rate , response speed , ……, intonation and other second-level indicators. Among them, the speech rate Used to measure whether the speech rate of the synthesized speech is appropriate. An appropriate speech rate can ensure that the speech information is neither too fast to be understood nor too slow to affect efficiency; response speed Used to measure the response speed of the synthesized speech to user input; intonation Used to measure whether the changes in pitch, intensity, and rhythm of the synthesized speech adapt to the user's auditory experience.
[0030] In this embodiment, = 3, = 4, = 3.
[0031] Through the constructed power speech synthesis index system, partial updates can be performed for each piece of speech, avoiding the speech coupling phenomenon caused by unified updates, and solving the problems of poor optimization effect and low accuracy of speech synthesis in the prior art due to single-dimensional optimization of the speech synthesis model in the power system field.
[0032] Step S2: After the power synthesized speech is released, collect user interaction statements and construct a user feedback knowledge base; After the power synthesized speech is released, collect user interaction statements. Here, the user interaction statements are in the form of one or several combinations of selecting feedback options, converting user oral feedback into text, free text input, emojis or emotion tags, sliders or progress bars, and questionnaires. When using single data collection, classify multiple secondary indicators in each piece of power speech for feedback to improve the accuracy of feedback.
[0033] Among them, the selection feedback option is to provide a series of preset options for users to choose, such as satisfaction rating (such as five-star evaluation), question type selection, etc.; converting user oral feedback into text is that users input feedback through voice, and the system converts it into text for processing; free text input is that users can freely input feedback content in the text box, and before this process, prompt the user to guide the input content, for example, prompt the user to evaluate the accuracy of the words in a certain piece of power speech, etc.; emojis or emotion tags are that users express their feelings about the service or product by selecting emojis or emotion tags; sliders or progress bars are that users quantify their satisfaction or evaluation through sliders or progress bars; questionnaires are to send questionnaires containing multiple questions to users to collect detailed feedback.
[0034] Step S3: Classify and score the user interaction statements in the user feedback knowledge base based on secondary indicators, and determine the optimization index sequence based on the scoring results; Among them, step S3 includes: Step S3-1: Classify the feature keywords of the user interaction statements collected in the user feedback knowledge base, divide the corresponding user feedback statements into different secondary indicators according to the feature keywords, and generate the corresponding secondary indicator statement sets; Step S3-2: According to the positivity and negativity of the user evaluation, divide the user interaction statements in each secondary indicator into positive feedback statements, neutral statements, and negative feedback statements, and delete the positive feedback statements and neutral statements from the secondary indicator statement sets; When classifying user interaction statements, keyword detection and determination are used. By identifying keywords or phrases in the statements, the user intention or the category to which they belong is determined. In this application, one of string exact matching, string partial matching, regular expression matching, or semantic analysis is used, where represents the set of negative feedback statements related to the nth secondary indicator in the ith primary indicator, including negative feedback statements, where: . For example, represents the set of negative feedback statements related to the knowledge quantity indicator, including negative feedback statements related to the knowledge quantity indicator.
[0035] Step S3 also includes Step S3-3: Score each negative feedback statement to obtain the score set of the secondary indicators , represents the negative feedback score set of the negative feedback statements related to the nth secondary indicator in the ith primary indicator.
[0036] Among them, represents the negative feedback score set of the negative feedback statements related to the nth secondary indicator in the ith primary indicator. For example, represents the negative feedback score set of the negative feedback statements related to the knowledge quantity indicator.
[0037] Step S3 also includes Step S3-4: Traverse , find The minimum value in is denoted as , for Sort from small to large to generate an optimized index sequence, is the set The negative feedback score with the largest absolute value in, the highest score is set to 0, and the lowest score is set to , ; The higher the negative degree of the secondary indicator involved in the negative feedback statement, the lower the corresponding score. Then, the set of negative feedback scores with the largest absolute value corresponding to each secondary indicator is denoted as .
[0038] Step S4: Perform optimizations in sequence according to the optimized index sequence; Step S4 further includes step S4-1: determining the index to be optimized according to the optimized index sequence , the index to be optimized is the secondary index corresponding to the lowest negative feedback score in the optimized index sequence; Step S4 further includes step S4-2: determining the optimization step size corresponding to the index to be optimized; optimizing the index to be optimized ; The maximum negative feedback score in the set of negative feedback scores corresponding to the index to be optimized is denoted as , and the negative feedback statement corresponding to is denoted as . According to , calculate the optimization step size of the power speech synthesis model ; Specifically expressed as: ; ; wherein is the maximum optimization step size, is the maximum reference score, The larger it is, the worse the performance, and the larger the optimization step size of the power speech synthesis model; where the adam model update function uses the Adam optimizer, which is a widely used deep learning optimization algorithm that adjusts the learning rate of each parameter by calculating the first-order moment estimate and second-order moment estimate of the gradient.
[0039] Step S4 further includes step S4-3: after the model parameter optimization is completed, delete the negative feedback statement with the lowest score from the set of negative feedback statements corresponding to the index to be optimized , and synchronously update the set of negative feedback scores of the index to be optimized ; Step S4 further includes step S4-4: Step S4-4: Determine whether the lowest negative feedback score in the set of negative feedback scores of the updated index to be optimized in step S4-3 is the maximum value in the set of negative feedback scores with the largest absolute value corresponding to each secondary index in step S3-4. If the lowest negative feedback score in the set of negative feedback scores of the updated index to be optimized is the maximum value in the set of negative feedback scores with the largest absolute value corresponding to each secondary index, then optimize the next secondary index in the optimized index sequence, otherwise jump to step S4-2; ; Step S4-5: If the updated index to be optimized in step S4-4 The lowest negative feedback score in the negative feedback score set is the maximum value in the negative feedback score set with the largest absolute value corresponding to each secondary indicator. Jump to step S4-6. Otherwise, the indicator to be optimized After the negative feedback score set of is updated and published, an observation window W is established to re-determine the indicator to be optimized Regarding the negative feedback situation of, within the observation window W for the indicator to be optimized after the update and publication Perform negative feedback monitoring. If there is no negative feedback situation regarding the indicator to be optimized within the observation window W Jump to step S4-6. If there is still a negative feedback score for the indicator to be optimized Jump to step S4-2; Step S4-6: Delete the secondary indicator corresponding to the indicator to be optimized specified in step S4-1 And update the negative feedback score set of the secondary indicators in step S3-3 If there are secondary indicators with negative feedback scores in the negative feedback score set of secondary indicators Jump to step S3-4; Otherwise, complete the optimization of the power voice synthesis model.
[0040] Among them, the observation window W is a sliding window, which can monitor negative feedback within a certain amount of data or within a certain period of time.
[0041] This solution also provides a power voice synthesis model optimization device. The power voice synthesis model optimization device includes a processor, a memory, and a computer program stored on the memory and executable by the processor. When the computer program is executed by the processor, it implements the steps of the power voice synthesis model optimization method in the above embodiments.
[0042] The power voice synthesis model optimization device may include a processor (such as CPU, GPU, FPGA, etc.), which can execute part or all of the processing in the above-described embodiments according to the program stored in the read-only memory (ROM) or the program loaded from the storage part into the random access memory (RAM). In the RAM, various programs and data required for system operation are also stored. The processor, ROM, and RAM are connected to each other through a bus. The input / output (I / O) interface is also connected to the bus.
[0043] The following components are connected to the I / O interface: an input section including a keyboard, a mouse, etc.; an output section including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section including a hard disk, etc.; and a communication section including a network interface card such as a LAN card, a modem, etc. The communication section performs communication processing via a network such as the Internet. A drive is also connected to the I / O interface as required. A removable medium such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc. is installed on the drive as required so that a computer program read therefrom is installed into the storage section as required.
[0044] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the module, the program segment, or the part of code includes one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0045] The units or modules involved in the embodiments described in the present application may be implemented in software or in hardware. The described units or modules may also be provided in a processor, and the names of these units or modules do not constitute a limitation on the units or modules themselves in some cases.
[0046] Those skilled in the art can understand that the optimized device structure involved in the embodiments of the present invention does not constitute a limitation on the optimized device. The optimized device may include more or fewer components than shown in the figure, or combine some components, or have a different component arrangement. In the embodiments of the present invention, the optimized device includes, but is not limited to, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The optimized device may also represent various forms of mobile devices, such as a personal digital processor, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments described and / or claimed in the present application.
[0047] The optimized device may include a processor, an external memory interface, an internal memory, a Universal Serial Bus (USB) interface, a charging management module, a power management module, a battery, a wireless communication module, an audio module, a speaker, a microphone, a sensor module, keys, a camera, a display screen, and a Subscriber Identity Module (SIM) card interface, etc.
[0048] It can be understood that the structure illustrated in the embodiments of the present application does not constitute a specific limitation on the optimized device. In other embodiments of the present application, the optimized device may include more or fewer components than those illustrated, or combine certain components, or split certain components, or have different component arrangements. The illustrated components may be implemented in hardware, software, or a combination of software and hardware.
[0049] The processor may include one or more processing units. For example, the processor may include a Central Processing Unit (CPU), etc., an Application Processor (AP), a modem processor, a Graphics Processing Unit (GPU), an Image Signal Processor (ISP), a controller, a memory, a video codec, a Digital Signal Processor (DSP), a baseband processor, and / or a Neural-Network Processing Unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.
[0050] Among them, the processor may be the nerve center and command center of the optimized device. The controller may generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.
[0051] A memory may also be provided in the processor for storing instructions and data. In some embodiments, the memory in the processor is a cache memory. This memory may save the instructions or data that the processor has just used or recycled. If the processor needs to use the instruction or data again, it can directly call it from this memory. This avoids repeated accesses, reduces the waiting time of the processor, and thus improves the system efficiency.
[0052] The external memory interface may be used to connect an external memory card, such as a MicroSD card, to implement the storage capacity expansion of the optimized device. The external memory card communicates with the processor through the external memory interface to implement the data storage function. For example, files such as music and videos are saved in the external memory card.
[0053] The internal memory can be used to store computer-executable program code, and the computer-executable program code includes instructions. The processor executes various functional applications and data processing of the optimization device by running the instructions stored in the internal memory. The internal memory can include a program storage area and a data storage area. The internal memory can include a high-speed random access memory and can also include non-volatile memory, such as at least one disk storage device, a flash memory device, a universal flash storage (UFS), etc.
[0054] The wireless communication function of the optimization device can be implemented by an antenna, a wireless communication module, a modulation and demodulation processor, a baseband processor, etc.
[0055] The wireless communication module can provide wireless communication solutions applied to the optimization device, including wireless local area networks (WLANs) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite systems (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc.
[0056] The optimization device can implement audio functions through an audio module, a speaker, a receiver, a microphone, a headphone jack, an application processor, etc.
[0057] The optimization device can implement a shooting function through an ISP, a camera, a video codec, a GPU, a display screen, an application processor, etc.
[0058] The optimization device can implement a display function through a GPU, a display screen, an application processor, etc.
[0059] The GPU is a microprocessor for image processing, connecting the display screen and the application processor. The GPU is used to execute mathematical and geometric calculations for graphics rendering. The processor can include one or more GPUs, which execute program instructions to generate or change display information.
[0060] The display screen is used to display images, videos, etc. The display screen includes a display panel.
[0061] This solution also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the power voice synthesis model optimization method in the above embodiment are implemented.
[0062] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for optimizing an electric power speech synthesis model, characterized in that: The following steps are involved: Step S1: constructing an electric power speech synthesis index system: constructing at least two first-level indicators, each of which includes at least two second-level indicators; Step S2: After the power synthesized speech is released, user interaction sentences are collected to build a user feedback knowledge base; Step S3: classifying and scoring the user interaction sentences in the user feedback knowledge base based on the secondary indicators, and determining the optimization indicator sequence based on the scoring results; Step S4: Optimize in sequence according to the optimization index sequence.
2. The power speech synthesis model optimization method according to claim 1, characterized in that: Step S3 includes the following steps: Step S3-1: classify the user interaction sentences collected in the user feedback knowledge base by feature keywords, divide the corresponding user feedback sentences into different secondary indicators according to the feature keywords, and generate a corresponding secondary indicator sentence set; Step S3-2: according to the positivity and negativity of the user evaluation, the user interaction sentences in each secondary indicator are divided into positive feedback sentences, neutral sentences, and negative feedback sentences, and the positive feedback sentences and neutral sentences are deleted from the secondary indicator sentence set; Step S3-3: Score each negative feedback statement to obtain a score set of secondary indicators , Represents the set of negative feedback scores of the negative feedback sentences involving the nth secondary indicator in the i-th primary indicator; Step S3-4: Traversal ,turn up The minimum value in is recorded as ,right Sort from small to large to generate an optimization indicator sequence.
3. The power speech synthesis model optimization method according to claim 2, characterized in that: Step S4 includes the following steps: Step S4-1: Determine the index to be optimized according to the optimization index sequence ; Step S4-2: Determine the indicator to be optimized The corresponding optimization step size; the optimization index Optimize Step S4-3: After the model parameter optimization is completed, the lowest-scoring negative feedback statement is removed from the list of indicators to be optimized. The corresponding negative feedback statement set is deleted, and the indicator to be optimized is updated synchronously The set of negative feedback scores; Step S4-4: Determine the updated indicator to be optimized in step S4-3 Is the lowest negative feedback score in the negative feedback score set the maximum value in the negative feedback score set with the largest absolute value of each secondary indicator in step S3-4? If the updated indicator to be optimized If the lowest negative feedback score in the negative feedback score set is the maximum value in the negative feedback score set with the largest absolute value corresponding to each secondary indicator, the next secondary indicator in the optimization indicator sequence is optimized, otherwise jump to step S4-2.
4. The method for optimizing the electric power speech synthesis model according to claim 3, characterized in that: In step S4-1, the indicator to be optimized It is the secondary indicator corresponding to the lowest negative feedback score in the optimization indicator sequence.
5. The method for optimizing the electric power speech synthesis model according to claim 4, characterized in that: In step S4-2, the indicator to be optimized The corresponding negative feedback score set The negative feedback score with the largest absolute value is recorded as , The corresponding negative feedback statement is recorded as ,according to , calculate the optimization step size of the power speech synthesis model ; Specifically expressed as: ; in is the maximum optimization step size, is the maximum reference score.
6. The method for optimizing the electric power speech synthesis model according to claim 5, characterized in that: In step S4-2: based on the indicator to be optimized and the optimization step size , update the parameters of the power speech synthesis model; It can be expressed as: ; In the formula, are the parameters of the power speech synthesis model, Update function for the model.
7. The method for optimizing the electric power speech synthesis model according to claim 6, characterized in that: Step S4 also includes the following steps: Step S4-5: If the indicator to be optimized after updating in step S4-4 The lowest negative feedback score in the negative feedback score set is the maximum value in the negative feedback score set with the largest absolute value corresponding to each secondary indicator, and jump to step S4-6, otherwise, the indicator to be optimized The negative feedback score set is updated and published, and an observation window W is established to re-determine the indicator to be optimized Negative feedback of the indicators to be optimized after the update is released within the observation window W Monitor negative feedback. If there is no information about the indicator to be optimized in the observation window W, If there is negative feedback, jump to step S4-6. If the indicator to be optimized If there is still a negative feedback score, jump to step S4-2; Step S4-6: The indicator to be optimized specified in step S4-1 The corresponding secondary indicator is deleted, and the secondary indicator negative feedback score set in step S3-3 is updated. , if the secondary indicator negative feedback score set If there is a secondary indicator with a negative feedback score, jump to step S3-4; otherwise, complete the optimization of the power speech synthesis model.
8. An electric power speech synthesis model optimization device, characterized in that: The electric power speech synthesis model optimization device includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, the steps of the electric power speech synthesis model optimization method as described in any one of claims 1 to 7 are implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein when the computer program is executed by a processor, the steps of the power speech synthesis model optimization method according to any one of claims 1 to 7 are implemented.