Voice model compression method, device and equipment and readable storage medium

By performing layered analysis, iterative adjustment and energy consumption evaluation of the speech model, the problems of reduced accuracy and increased energy consumption in speech model compression are solved, and efficient deployment on local devices is achieved.

CN120356473AActive Publication Date: 2025-07-22AISPEECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510848545.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

When compressing voice models in the prior art, there are problems such as decreasing model accuracy and increasing energy consumption, making it difficult to effectively deploy on local devices.

Method used

By performing hierarchical analysis of the compressed speech model, the contribution score of each channel on each layer is calculated, channels with contribution below the threshold are removed, and dynamic sharing coefficients are allocated through iterative adjustment, clustering and energy consumption evaluation, and the target model is finally obtained.

Benefits of technology

The compressed voice model ensures accuracy while significantly reducing energy consumption, making it suitable for deployment on local devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356473A_ABST
    Figure CN120356473A_ABST
Patent Text Reader

Abstract

The invention discloses a voice model compression method, device and equipment and a readable storage medium, and relates to the technical field of artificial intelligence. Comprising the following steps: firstly, carrying out hierarchical analysis on a to-be-compressed voice model, and calculating a contribution degree score of each channel of each layer in the to-be-compressed voice model; on the basis of the contribution degree scores, removing channels of which the contribution degree scores are lower than a preset score threshold in the to-be-compressed voice model to obtain a first compression model; carrying out iterative adjustment on the first compression model, and reducing the number of channels in the first compression model to obtain a second compression model; clustering each channel in the second compression model, and allocating a dynamic sharing coefficient to each channel in each cluster to obtain a third compression model; and finally obtaining the energy consumption of each channel in the third compression model, and compressing the third compression model again according to the energy consumption to obtain a target model. The accuracy of the compressed target model can still be ensured, and the energy consumption of the to-be-compressed voice model is remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and more particularly to a method, apparatus, device, and readable storage medium for compressing a speech model. Background Art

[0002] In recent years, with the rapid popularization of Internet of Things (IoT) intelligent hardware, speech interaction technology has become the core entry point for natural human-computer interaction. To meet the requirements of low power consumption, high real-time performance, and privacy and security for end-side devices, speech perception models need to be migrated from the cloud to the local side. However, the computing power and storage resources of local end-side devices are limited, and traditional large-scale speech AI models are difficult to directly deploy.

[0003] To overcome the above problems, those skilled in the art usually compress neural network models, such as global channel pruning method, fixed weight sharing method, heuristic channel selection method, and single energy consumption optimization strategy, etc. However, these methods still have problems such as causing a decrease in model accuracy and increasing the running cost of the model (such as an increase in technical duration and energy consumption). Therefore, there is an urgent need for a speech model compression method that can overcome the above defects. Summary of the Invention

[0004] The object of the present invention is to provide a method, apparatus, device, and readable storage medium for compressing a speech model. By scoring, iteratively adjusting, clustering, and evaluating the energy consumption of each channel in the speech model to be compressed, channels with relatively low importance and relatively high energy consumption in the speech model to be compressed are removed, so that the compressed target model can still ensure accuracy and significantly reduce the energy consumption of the speech model to be compressed.

[0005] To achieve the above object, the present invention provides the following technical solutions: In a first aspect, the present invention provides a method for compressing a speech model, the method comprising: Performing hierarchical parsing on the speech model to be compressed, and calculating the contribution degree score of each channel in each layer of the speech model to be compressed; Based on the contribution degree score, removing channels in the speech model to be compressed with a contribution degree score lower than a preset score threshold to obtain a first compressed model; Performing iterative adjustment on the first compressed model to reduce the number of channels in the first compressed model to obtain a second compressed model; Clustering each channel in the second compressed model, and assigning a dynamic sharing coefficient to each channel in each cluster to obtain a third compressed model; Obtaining the energy consumption of each channel in the third compressed model, and further compressing the third compressed model according to the energy consumption to obtain a target model.

[0006] In some embodiments, clustering each channel in the second compression model and assigning dynamic sharing coefficients to each channel in each cluster to obtain a third compression model, including: Using a clustering algorithm to divide each channel in the second compression model into multiple clusters; For each cluster, taking the cluster center as the shared weight and assigning dynamic sharing coefficients to each channel in each cluster to obtain a third compression model.

[0007] In some embodiments, assigning dynamic sharing coefficients to each channel in each cluster includes: Calculating the cosine similarity between each channel and the cluster center; Based on the cosine similarity and a preset precision sensitivity factor, calculating the dynamic sharing coefficient of each channel.

[0008] In some embodiments, obtaining the energy consumption of each channel in the third compression model and further compressing the third compression model according to the energy consumption to obtain a target model, including: Real-time monitoring the energy consumption of each channel in the third compression model and calculating the energy consumption score of each channel according to the energy consumption; Performing weighted summation on the contribution degree score and the energy consumption score to obtain a comprehensive score; Further compressing the third compression model according to the comprehensive score to obtain a target model.

[0009] In some embodiments, performing hierarchical analysis on the speech model to be compressed and calculating the contribution degree score of each channel in each layer of the speech model to be compressed, including: Performing hierarchical analysis on the speech model to be compressed to obtain the channel attributes of each channel in each layer of the speech model to be compressed; Based on the channel attributes, calculating the importance of each channel to obtain the contribution degree score of each channel.

[0010] In some embodiments, iteratively adjusting the first compression model to reduce the number of channels in the first compression model to obtain a second compression model, including: Removing channels in the first compression model one by one and calculating the accuracy of the first compression model after each channel is removed; Based on the accuracy, reducing the number of channels in the first compression model to obtain a second compression model.

[0011] In some embodiments, the method further includes: Testing the target model based on sample data to obtain the test result of the target model; Optimizing and adjusting the target model according to the test result.

[0012] In a second aspect, the present invention further provides a voice model compression device, which includes: A score calculation module, configured to perform hierarchical parsing on the voice model to be compressed and calculate the contribution score of each channel in each layer of the voice model to be compressed; A first compression module, configured to remove channels with contribution scores lower than a preset score threshold in the voice model to be compressed based on the contribution scores, so as to obtain a first compressed model; A second compression module, configured to iteratively adjust the first compressed model to reduce the number of channels in the first compressed model, so as to obtain a second compressed model; A third compression module, configured to cluster each channel in the second compressed model and assign dynamic sharing coefficients to each channel in each cluster, so as to obtain a third compressed model; A fourth compression module, configured to obtain the energy consumption of each channel in the third compressed model and further compress the third compressed model according to the energy consumption, so as to obtain a target model.

[0013] In a third aspect, the present invention further provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the voice model compression method provided in the first aspect is implemented.

[0014] In a fourth aspect, the present invention further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the voice model compression method provided in the first aspect is implemented.

[0015] In a fifth aspect, the present invention further provides a computer program product, including a computer program. When the computer program is executed by a processor, the voice model compression method provided in the first aspect is implemented.

[0016] The beneficial effects of the present invention are as follows: The voice model compression method provided in the present invention first performs hierarchical parsing on the voice model to be compressed, and calculates the contribution score of each channel in each layer of the voice model to be compressed; then, based on the contribution score, removes the channels in the voice model to be compressed whose contribution scores are lower than the preset score threshold to obtain a first compressed model; then iteratively adjusts the first compressed model to reduce the number of channels in the first compressed model to obtain a second compressed model; then clusters the channels in the second compressed model and assigns dynamic sharing coefficients to the channels in each cluster to obtain a third compressed model; finally, obtains the energy consumption of each channel in the third compressed model, and further compresses the third compressed model according to the energy consumption to obtain a target model. By scoring, iteratively adjusting, clustering, and evaluating the energy consumption of each channel in the voice model to be compressed, channels with lower importance and higher energy consumption in the voice model to be compressed are removed, so that the compressed target model can still ensure accuracy and significantly reduce the energy consumption of the voice model to be compressed.

[0017] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly and implement it according to the content of the specification, the following describes in detail with reference to the preferred embodiments of the present invention and the accompanying drawings. Brief Description of the Drawings

[0018] Figure 1 It is a schematic flowchart of a voice model compression method shown in an embodiment of the present invention; Figure 2 It is a schematic flowchart of another voice model compression method shown in an embodiment of the present invention; Figure 3 It is a schematic structural diagram of a voice model compression device shown in an embodiment of the present invention; Figure 4 It is a schematic structural diagram of an electronic device provided in an embodiment of the present application. Detailed Description of the Embodiment

[0019] The technical solution of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0020] It should be noted that the references to "an embodiment", "embodiment", "example embodiment", etc. in this specification mean that the described embodiment may include specific features, structures or characteristics. However, not every embodiment must include these specific features, structures or characteristics. In addition, such expressions do not refer to the same embodiment. Further, when combining specific features, structures or characteristics with an embodiment, it has been shown that it is within the knowledge of those skilled in the art to combine such features, structures or characteristics with other embodiments, whether or not explicitly described.

[0021] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0022] In some embodiments, as Figure 1 shown, a specific process of a voice model compression method is provided: S101, perform hierarchical parsing on the voice model to be compressed, and calculate the contribution score of each channel in each layer of the voice model to be compressed.

[0023] Among them, the voice model to be compressed is a neural network model that needs to be compressed. This neural network model is divided into multiple layers, and each layer has multiple channels. The process of removing some channels from these channels is the process of model compression. The contribution score is a score representing the degree of influence of each channel on the model output result. The higher the score, the more important the channel, and vice versa, indicating that the channel is a redundant channel.

[0024] Specifically, the contribution score of each channel can be calculated according to a preset scoring algorithm.

[0025] Optionally, it can also be to perform hierarchical parsing on the voice model to be compressed to obtain the channel attributes of each channel in each layer of the voice model to be compressed; calculate the importance of each channel based on the channel attributes to obtain the contribution score of each channel.

[0026] Among them, the channel attributes include the weight, activation value and calculation amount index of each channel.

[0027] Specifically, for each channel, the higher the weight, activation value and calculation amount index of the channel, the higher the score of the channel, and vice versa. Therefore, based on the preset weight values of the weight, activation value and calculation amount index, the weight, activation value and calculation amount index of each channel can be weighted and summed to obtain the contribution score of each channel.

[0028] S102, based on the contribution score, remove the channels in the voice model to be compressed whose contribution score is lower than the preset score threshold to obtain a first compressed model.

[0029] Specifically, the preset scoring threshold is set manually and can be set according to the requirements of the compression degree. When the compression requirement is large, the preset scoring threshold can be increased. When the compression requirement is small, the preset scoring threshold can be appropriately decreased. Channels with contribution scores lower than the preset scoring threshold are removed from the speech model to be compressed, and channels with contribution scores greater than or equal to the preset scoring threshold are retained, thus obtaining the first compression model.

[0030] S103. Iteratively adjust the first compression model to reduce the number of channels in the first compression model, obtaining the second compression model.

[0031] Specifically, the influence of each channel in the first compression model on the accuracy of the output result of the first compression model can be tested in turn. When a certain channel has no influence or a small influence on the accuracy of the output result of the first compression model, this channel can be removed from the first compression model to obtain the second compression model.

[0032] Optionally, the method for obtaining the second compression model can also be: removing channels in the first compression model one by one, and calculating the accuracy of the first compression model after each channel is removed; based on the accuracy, reducing the number of channels in the first compression model to obtain the second compression model.

[0033] Specifically, without removing channels, test the original accuracy of the first compression model, and then after removing channels in the first compression model one by one, test again to obtain the accuracy of the first compression model after each channel is removed (exemplarily, first remove the first channel, then test the accuracy of the first compression model, then restore the first channel, then remove the second channel, and then test the accuracy of the first compression model, and so on). For each channel, if the accuracy of the first compression model after this channel is removed does not change or the accuracy increases compared to the original accuracy, directly remove this channel. If the accuracy of the first compression model after this channel is removed has a small decrease compared to the original accuracy, but the decrease amplitude is lower than the preset amplitude threshold, directly remove this channel. If the accuracy of the first compression model after this channel is removed has a relatively obvious decrease, and the decrease amplitude is greater than or equal to the preset amplitude threshold, retain this channel, and finally obtain the second compression model.

[0034] S104. Cluster each channel in the second compression model and assign dynamic sharing coefficients to each channel in each cluster to obtain the third compression model.

[0035] Optionally, use a clustering algorithm to divide each channel in the second compression model into multiple clusters; for each cluster, use the cluster center as the shared weight and assign dynamic sharing coefficients to each channel in each cluster to obtain the third compression model.

[0036] Specifically, there are multiple channels in a neural network model. The functions of some channels are the same or similar. These channels with similar functions can be grouped into one category. For each cluster, calculate the cluster center as the shared weight, and assign a dynamic sharing coefficient to each channel, so as to retain the diversity of channel features during the sharing process. Integrate the generation strategy of the dynamic sharing coefficient into the fine-tuning training, modify the backpropagation algorithm, and update the sharing coefficient in real time to ensure global optimization. Finally, output the dynamic sharing coefficient and apply it to the second compressed model to obtain the third compressed model. The obtained third compressed model significantly reduces the number of parameters while ensuring that information transmission is not affected by excessive sharing.

[0037] Optionally, the method of assigning a dynamic sharing coefficient to each channel in each cluster can also be: calculate the cosine similarity between each channel and the cluster center; calculate the dynamic sharing coefficient of each channel based on the cosine similarity and a preset precision-sensitive factor.

[0038] Specifically, the cosine similarity between each channel and the cluster center can be calculated using the following formula (1): ; where is the cosine similarity, represents the channel weight, represents the cluster center.

[0039] Then, calculate the dynamic sharing coefficient of each channel based on the following formula (2): ; where is the cosine similarity, is the preset precision-sensitive factor, n is the number of channels, is the dynamic sharing coefficient.

[0040] S105. Obtain the energy consumption of each channel in the third compressed model, and compress the third compressed model again according to the energy consumption to obtain the target model.

[0041] Specifically, a hardware energy consumption evaluation tool can be used to monitor the energy consumption of the third compressed model in real time, and remove the channels with energy consumption higher than the energy consumption threshold to obtain the target model.

[0042] Optionally, the method of compressing the third compressed model to obtain the target model can also be: monitor the energy consumption of each channel in the third compressed model in real time, and calculate the energy consumption score of each channel according to the energy consumption; perform a weighted sum of the contribution score and the energy consumption score to obtain a comprehensive score; compress the third compressed model again according to the comprehensive score to obtain the target model.

[0043] Specifically, a hardware energy consumption evaluation tool is used to monitor the energy consumption of the third compression model in real time, obtain the energy consumption of each channel, calculate the energy consumption score of each channel according to the energy consumption (the higher the energy consumption, the lower the score), then perform a weighted sum of the contribution score and the energy consumption score (the weights are preset), obtain the comprehensive score, and remove the channels with a comprehensive score lower than the comprehensive score threshold from the third compression model to obtain the target model. The lower the comprehensive score, the higher the energy consumption and the more redundant the channel, which has little impact on the model prediction accuracy. The finally obtained target model can reduce the overall energy consumption while ensuring the accuracy.

[0044] For the speech model compression method in the above embodiment, first perform hierarchical parsing on the speech model to be compressed, and calculate the contribution score of each channel in each layer of the speech model to be compressed; then, based on the contribution score, remove the channels with a contribution score lower than the preset score threshold from the speech model to be compressed to obtain the first compression model; then perform iterative adjustment on the first compression model to reduce the number of channels in the first compression model to obtain the second compression model; then cluster each channel in the second compression model and assign dynamic sharing coefficients to each channel in each cluster to obtain the third compression model; finally, obtain the energy consumption of each channel in the third compression model, and compress the third compression model again according to the energy consumption to obtain the target model. By scoring, iteratively adjusting, clustering, and evaluating the energy consumption of each channel in the speech model to be compressed, channels with lower importance and higher energy consumption in the speech model to be compressed are removed, so that the compressed target model can still ensure accuracy and significantly reduce the energy consumption of the speech model to be compressed.

[0045] In another embodiment, in order to further improve the accuracy of the target model, the target model can also be tested based on sample data to obtain the test result of the target model; and the target model is optimized and adjusted according to the test result.

[0046] Among them, the sample data includes input data and reference data. The input data is input into the target model, and the target model will output corresponding prediction data. Based on the difference between the prediction data and the reference data, the parameters of the target model are adjusted until the difference between the prediction data and the reference data meets the preset difference range, and then the optimized target model is exported in the format of AP level, embedded or extremely low-power chip, that is, the compression of the model is completed. The compressed model can not only ensure accuracy, but also greatly reduce power consumption, computing power and storage space, and reduce the conditions for model deployment.

[0047] To more comprehensively demonstrate this solution, this embodiment gives an optional way of a speech model compression method, such as Figure 2 shown: S201, perform hierarchical parsing on the speech model to be compressed to obtain the channel attributes of each channel in each layer of the speech model to be compressed.

[0048] S202. Calculate the importance of each channel based on the channel attributes to obtain the contribution score of each channel.

[0049] S203. Based on the contribution scores, remove the channels in the speech model to be compressed with contribution scores lower than the preset score threshold to obtain the first compressed model.

[0050] S204. Remove the channels in the first compressed model one by one, and calculate the accuracy of the first compressed model after each channel is removed.

[0051] S205. Based on the accuracy, reduce the number of channels in the first compressed model to obtain the second compressed model.

[0052] S206. Use the clustering algorithm to divide the channels in the second compressed model into multiple clusters.

[0053] S207. For each cluster, use the cluster center as the shared weight, and assign a dynamic sharing coefficient to each channel in the cluster to obtain the third compressed model.

[0054] Among them, assigning a dynamic sharing coefficient to each channel in the cluster includes: calculating the cosine similarity between each channel and the cluster center; calculating the dynamic sharing coefficient of each channel based on the cosine similarity and the preset accuracy sensitivity factor.

[0055] S208. Real-time monitor the energy consumption of each channel in the third compressed model, and calculate the energy consumption score of each channel according to the energy consumption.

[0056] S209. Perform weighted summation on the contribution scores and the energy consumption scores to obtain a comprehensive score.

[0057] S210. Re-compress the third compressed model according to the comprehensive score to obtain the target model.

[0058] S211. Test the target model based on the sample data to obtain the test result of the target model.

[0059] S212. Optimize and adjust the target model according to the test result.

[0060] For the specific processes of the above S201 - S212, reference can be made to the description of the above method embodiments, and their implementation principles and technical effects are similar, which will not be elaborated here.

[0061] Based on the same inventive concept, an embodiment of the present application further provides a voice model compression device for implementing the above-mentioned voice model compression method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the voice model compression device provided below can refer to the limitations on the voice model compression method in the above text, and will not be elaborated here.

[0062] In one embodiment, as Figure 3 shown, a voice model compression device is provided, and the device includes: A scoring calculation module 30, configured to perform hierarchical parsing on the voice model to be compressed, and calculate the contribution score of each channel in each layer of the voice model to be compressed.

[0063] A first compression module 31, configured to remove channels with contribution scores lower than a preset score threshold in the voice model to be compressed based on the contribution scores, to obtain a first compressed model.

[0064] A second compression module 32, configured to perform iterative adjustment on the first compressed model to reduce the number of channels in the first compressed model, to obtain a second compressed model.

[0065] A third compression module 33, configured to cluster each channel in the second compressed model, and assign dynamic sharing coefficients to each channel in each cluster, to obtain a third compressed model.

[0066] A fourth compression module 34, configured to obtain the energy consumption of each channel in the third compressed model, and further compress the third compressed model according to the energy consumption, to obtain a target model.

[0067] A model optimization module 35, configured to test the target model based on sample data, to obtain the test result of the target model; and optimize and adjust the target model according to the test result.

[0068] In another embodiment, the above Figure 3 The third compression module 33 is specifically configured to: use a clustering algorithm to divide each channel in the second compressed model into multiple clusters; for each cluster, use the cluster center as the shared weight, and assign dynamic sharing coefficients to each channel in each cluster, to obtain a third compressed model. Among them, assigning dynamic sharing coefficients to each channel in each cluster includes: calculating the cosine similarity between each channel and the cluster center; calculating the dynamic sharing coefficient of each channel based on the cosine similarity and a preset precision sensitivity factor.

[0069] In another embodiment, the above Figure 3The fourth compression module 34 in it is specifically configured to: monitor the energy consumption of each channel in the third compression model in real time, and calculate the energy consumption score of each channel according to the energy consumption; perform weighted summation on the contribution score and the energy consumption score to obtain a comprehensive score; and re-compress the third compression model according to the comprehensive score to obtain a target model.

[0070] In another embodiment, the above Figure 3 The scoring calculation module 30 in it is specifically configured to: perform hierarchical parsing on the speech model to be compressed to obtain the channel attributes of each channel in each layer of the speech model to be compressed; calculate the importance of each channel based on the channel attributes to obtain the contribution score of each channel.

[0071] In another embodiment, the above Figure 3 The second compression module 32 in it is specifically configured to: remove the channels in the first compression model one by one, and calculate the accuracy of the first compression model after each channel is removed; reduce the number of channels in the first compression model based on the accuracy to obtain a second compression model.

[0072] The embodiments of the present application further provide an electronic device. In some embodiments, as Figure 4 shown, the electronic device 700 includes an input unit 710, a memory 720, a processor 730, and an output unit 740. The memory 720 stores program instructions that can run on the processor 730, and the processor 730 can execute the speech model compression method and / or technical solution based on the foregoing embodiments by calling the program instructions. The electronic device 700 can be a mobile terminal device such as a mobile phone or a computer.

[0073] In addition, the embodiments of the present application further provide a computer-readable storage medium for storing a computer program for executing the speech model compression method. For example, computer program instructions, when executed by a computer, can call or provide the method and / or technical solution according to the present application through the operation of the computer. The program instructions for calling the method of the present application may be stored in a fixed or removable storage medium, and / or transmitted and / or stored in a storage medium running according to the program instructions through a data stream in a broadcast or other signal-bearing medium.

[0074] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present application can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Optionally, they can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present application is not limited to any specific combination of hardware and software.

[0075] The technical features of the above embodiments can be arbitrarily integrated. For the sake of concise description, not all possible integrations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the integration of these technical features, it should be considered as the scope described in this specification.

[0076] The above embodiments only represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the invention patent should be subject to the appended claims.

Claims

1. A method for compressing a speech model, characterized in that, The method includes: Performing hierarchical parsing on the speech model to be compressed, and calculating the contribution score of each channel in each layer of the speech model to be compressed; Based on the contribution score, removing the channels in the speech model to be compressed with contribution scores lower than the preset score threshold to obtain a first compressed model; Performing iterative adjustment on the first compressed model to reduce the number of channels in the first compressed model to obtain a second compressed model; Clustering each channel in the second compressed model, and assigning dynamic sharing coefficients to each channel in each cluster to obtain a third compressed model; Obtaining the energy consumption of each channel in the third compressed model, and further compressing the third compressed model according to the energy consumption to obtain a target model.

2. The voice model compression method according to claim 1, wherein Clustering each channel in the second compressed model, and assigning dynamic sharing coefficients to each channel in each cluster to obtain a third compressed model, including: Using a clustering algorithm to divide each channel in the second compressed model into multiple clusters; For each cluster, taking the cluster center as the shared weight, and assigning dynamic sharing coefficients to each channel in each cluster to obtain a third compressed model.

3. The voice model compression method according to claim 2, wherein Assigning dynamic sharing coefficients to each channel in each cluster, including: Calculating the cosine similarity between each channel and the cluster center; Based on the cosine similarity and a preset precision sensitivity factor, calculating the dynamic sharing coefficient of each channel.

4. The voice model compression method according to claim 1, wherein Obtaining the energy consumption of each channel in the third compressed model, and further compressing the third compressed model according to the energy consumption to obtain a target model, including: Real-time monitoring the energy consumption of each channel in the third compressed model, and calculating the energy consumption score of each channel according to the energy consumption; Performing weighted summation on the contribution score and the energy consumption score to obtain a comprehensive score; Further compressing the third compressed model according to the comprehensive score to obtain a target model.

5. The voice model compression method according to claim 1, wherein, Performing hierarchical parsing on the speech model to be compressed, and calculating the contribution score of each channel in each layer of the speech model to be compressed, including: Performing hierarchical parsing on the speech model to be compressed to obtain the channel attributes of each channel in each layer of the speech model to be compressed; Based on the channel attributes, performing importance calculation on each channel to obtain the contribution score of each channel.

6. The voice model compression method according to claim 1, wherein Performing iterative adjustment on the first compressed model to reduce the number of channels in the first compressed model to obtain a second compressed model, including: Removing the channels in the first compressed model one by one, and calculating the accuracy of the first compressed model after each channel is removed; Based on the accuracy, reducing the number of channels in the first compressed model to obtain a second compressed model.

7. The voice model compression method according to any one of claims 1-6, characterized in that The method further includes: Testing the target model based on sample data to obtain the test result of the target model; Optimally adjusting the target model according to the test result.

8. A voice model compression device, characterized in that, The device includes: A score calculation module, configured to perform hierarchical parsing on the speech model to be compressed, and calculate the contribution score of each channel in each layer of the speech model to be compressed; A first compression module, configured to remove the channels in the speech model to be compressed with contribution scores lower than the preset score threshold based on the contribution score to obtain a first compressed model; A second compression module for iteratively adjusting the first compression model to reduce the number of channels in the first compression model, thereby obtaining a second compression model; A third compression module for clustering each channel in the second compression model and assigning dynamic sharing coefficients to each channel in each cluster, thereby obtaining a third compression model; A fourth compression module for obtaining the energy consumption of each channel in the third compression model and further compressing the third compression model according to the energy consumption, thereby obtaining a target model.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the speech model compression method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by the processor, it implements the speech model compression method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Deep learning model compression method and related equipment

    CN113128660A

  • Adaptive quantization compression method and system of voice model and electronic equipment

    CN116524941A

  • Neural network model quantitative compression method, electronic equipment and storage medium

    CN116644797A

  • Model compression method and device, computer equipment and storage medium

    CN117689000A

  • Lightweight speech recognition method of unstructured pruning compression based on second-order information

    CN119207382A