Voice recognition method and device based on AI large model, equipment and storage medium

Through the speech recognition method based on AI big model, the speech feature matrix is generated and processed, and the problem of inefficient recording recognition is solved, and the speech recognition efficiency is improved and the precise marking of specific types of speech information in the financial business field is achieved.

CN120260543APending Publication Date: 2025-07-04CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510308881.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Traditional recording recognition methods are inefficient in the financial business field, and it is difficult to efficiently identify excellent language for non-auto insurance insurance, resulting in the service personnel's needs being unable to be quickly met.

Method used

Using a speech recognition method based on AI large model, the initial speech feature matrix is generated by preset audio collection models, and the pre-trained target convolutional neural network model is used for feature extraction and classification, the target speech feature matrix is generated, and the labeling process is performed to generate target speech information.

Benefits of technology

It improves the efficiency of voice recognition in the financial business field, realizes accurate marking and recognition of specific types of voice information, and improves the service quality and efficiency of service personnel.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260543A_ABST
    Figure CN120260543A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of semantic analysis, and discloses a voice recognition method, device and equipment based on an AI large model and a storage medium, and the method comprises the steps: generating an initial voice feature matrix corresponding to each piece of initial voice information through an audio collection model and each piece of initial voice information; determining a target voice feature matrix through a target convolutional neural network model and each initial voice feature matrix; and performing marking processing on the target voice feature matrix to generate target voice information. Through the above mode, the features of the initial voice information are captured through the preset audio collection model, the high-precision initial voice feature matrix is provided for the target convolutional neural network model, and the target voice feature matrix is determined from the initial voice feature matrix and marked. The method can achieve the precise marking of the specific type of voice information, generates the target voice information, and improves the voice recognition efficiency in the financial business field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of semantic parsing technology, and in particular, to a speech recognition method, device, equipment and storage medium based on an AI large model. Background Art

[0002] At present, as the growth of domestic auto insurance related premiums gradually stabilizes and the upward growth trend gradually slows down, insurance companies need to find a breakthrough to maintain stable performance growth. Non-auto insurance is one of the particularly good breakthroughs. However, due to the numerous types of non-auto insurance, including health insurance, liability insurance, credit insurance, surety insurance, marine and cargo insurance, agricultural insurance, special insurance, accident insurance and many other insurance types, the requirements for service personnel are very high. In addition to professional insurance-related qualities, for the service terms of different non-auto insurance, they also need to be flexible and adaptable. Therefore, excellent term learning is particularly important.

[0003] For traditional excellent term recording recognition, most of them search for valuable recordings one by one from a large number of recordings, with extremely large workload. The daily generated recording volume reaches millions, and the efficiency of finding voice information that meets the requirements among numerous recordings is extremely low. Therefore, how to improve the speech recognition efficiency in the financial business field has become an urgent technical problem to be solved. Summary of the Invention

[0004] This application provides a speech recognition method, device, equipment and storage medium based on an AI large model to improve the speech recognition efficiency in the financial business field.

[0005] In a first aspect, this application provides a speech recognition method based on an AI large model, and the method includes:

[0006] Generating an initial speech feature matrix corresponding to each initial speech information through a preset audio collection model and each initial speech information;

[0007] Determining a target speech feature matrix through a pre-trained target convolutional neural network model and each of the initial speech feature matrices;

[0008] Performing a marking process on the target speech feature matrix to generate target speech information.

[0009] In a second aspect, this application also provides a speech recognition device based on an AI large model, and the device includes:

[0010] An initial speech feature matrix generation module, configured to generate an initial speech feature matrix corresponding to each initial speech information through a preset audio collection model and each initial speech information;

[0011] A target voice feature matrix determination module, configured to determine a target voice feature matrix by using a pre-trained target convolutional neural network model and each of the initial voice feature matrices;

[0012] A target voice information generation module, configured to perform a marking process on the target voice feature matrix to generate target voice information.

[0013] In a third aspect, the present application further provides a computer device, where the computer device includes a memory and a processor; the memory is used to store a computer program; the processor is configured to execute the computer program and implement the voice recognition method based on an AI large model as described above when executing the computer program.

[0014] In a fourth aspect, the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement the voice recognition method based on an AI large model as described above.

[0015] The present application discloses a voice recognition method, apparatus, device, and storage medium based on an AI large model. The voice recognition method based on an AI large model includes generating an initial voice feature matrix corresponding to each of the initial voice information through a preset audio collection model and each of the initial voice information; determining a target voice feature matrix by using a pre-trained target convolutional neural network model and each of the initial voice feature matrices; and performing a marking process on the target voice feature matrix to generate target voice information. By the above method, the present application captures the features of the initial voice information through the preset audio collection model, provides a highly accurate initial voice feature matrix for the target convolutional neural network model, determines the target voice feature matrix from the initial voice feature matrix and performs a marking process, so as to achieve accurate marking of specific types of voice information, generate target voice information, and further improve the voice recognition efficiency in the financial business field. Description of the Drawings

[0016] To more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for the description of the embodiments will be briefly introduced below. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0017] Figure 1 It is a schematic flowchart of a voice recognition method based on an AI large model provided by the first embodiment of the present application;

[0018] Figure 2 It is a schematic flowchart of a voice recognition method based on an AI large model provided by the second embodiment of the present application;

[0019] Figure 3 A schematic block diagram of a voice recognition device based on an AI large model provided for an embodiment of the present application;

[0020] Figure 4 A schematic structural block diagram of a computer device provided for an embodiment of the present application. Detailed implementation manners

[0021] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0022] The flowcharts shown in the accompanying drawings are only illustrative examples, and do not necessarily include all contents and operations / steps, nor do they necessarily need to be executed in the described order. For example, some operations / steps can also be decomposed, combined, or partially merged, so the actual execution order may be changed according to the actual situation.

[0023] It should be understood that the terms used in this specification of the present application are only for the purpose of describing specific embodiments and are not intended to limit the present application. As used in this specification of the present application and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an", and "the" are intended to include the plural forms.

[0024] It should also be understood that the term "and / or" used in this specification of the present application and the appended claims refers to any combination and all possible combinations of one or more of the related listed items, and includes these combinations.

[0025] The embodiments of the present application provide a voice recognition method, device, equipment, and storage medium based on an AI large model. Among them, the voice recognition method based on the AI large model can be applied to a server. By capturing the features of the initial voice information through a preset audio collection model, providing a highly accurate initial voice feature matrix for the target convolutional neural network model, determining the target voice feature matrix from the initial voice feature matrix and performing marking processing, accurate marking of specific types of voice information can be achieved, generating target voice information, thereby improving the voice recognition efficiency in the financial business field. Among them, the server can be an independent server or a server cluster.

[0026] Next, some implementation manners of the present application will be described in detail in conjunction with the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0027] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a speech recognition method based on an AI large model provided by the first embodiment of the present application. The speech recognition method based on the AI large model can be applied to a server to capture the features of initial speech information through a preset audio collection model, provide a highly accurate initial speech feature matrix for a target convolutional neural network model, determine a target speech feature matrix from the initial speech feature matrix and perform marking processing, so as to achieve accurate marking of specific types of speech information, generate target speech information, and thus improve the speech recognition efficiency in the financial business field.

[0028] As Figure 1 shown, the speech recognition method based on the AI large model specifically includes steps S10 to S30.

[0029] Step S10: Generate an initial speech feature matrix corresponding to each initial speech information through a preset audio collection model and each initial speech information;

[0030] Specifically, various different initial speech information is obtained by using a preset audio collection model (such as a microphone array, a voice recording device, etc.). The initial speech information can come from different speakers, contain different language contents, be recorded under different environmental conditions, etc. The collected initial speech information is preprocessed, including but not limited to operations such as noise filtering, speech enhancement, and speech segmentation, to improve the speech quality and make it more suitable for subsequent feature extraction.

[0031] Appropriate speech feature extraction methods (such as Mel Frequency Cepstral Coefficients (MFCC), spectrogram, etc.) are used to extract features from the preprocessed speech information to generate an initial speech feature matrix corresponding to each initial speech information. Each feature matrix contains eigenvalue of the speech signal at different time frames or frequencies.

[0032] In a specific embodiment, taking the financial business field as an example, the initial speech information can be obtained from channels such as the internal service team and the customer service center of the company. To ensure that the initial speech information is representative and covers different types of customers, insurance products, and service scenarios, the insurance agents in the initial speech information have accurate knowledge of professional knowledge such as the terms, coverage, and claim conditions of insurance products and can clearly explain them to customers.

[0033] Step S20: Determine a target speech feature matrix through a pre-trained target convolutional neural network model and each of the initial speech feature matrices;

[0034] Specifically, each of the initial speech feature matrices generated in step S10 is sequentially input into a pre-trained target convolutional neural network model. In the model, the feature matrices will be processed through network structures such as convolutional layers, pooling layers, and fully connected layers.

[0035] The target convolutional neural network model is used to perform inference calculations on the input initial speech feature matrices. The model will classify each of the initial speech feature matrices according to the learned feature patterns and classification rules, thereby determining the target speech feature matrices that conform to a specific target type.

[0036] Step S30: Perform a marking process on the target speech feature matrices to generate target speech information.

[0037] Specifically, perform a marking process on the target speech feature matrices determined in step S20. The marking method can be to add specific tags, annotate category information, or modify some elements in the feature matrices to reflect their target characteristics and other operations.

[0038] Based on the marked target speech feature matrices, through speech synthesis technology or other related methods, convert them into the final target speech information, which can include operations such as restoring the feature matrices to time-domain speech signals, or generating speech waveforms corresponding to the marked features, thereby obtaining the target speech information that can be used in actual applications.

[0039] This embodiment discloses a speech recognition method based on an AI large model. The speech recognition method based on the AI large model includes generating initial speech feature matrices corresponding to each of the initial speech information through a preset audio collection model; determining target speech feature matrices through a pre-trained target convolutional neural network model and each of the initial speech feature matrices; performing a marking process on the target speech feature matrices to generate target speech information. By the above method, this application captures the features of the initial speech information through the preset audio collection model, provides high-precision initial speech feature matrices for the target convolutional neural network model, determines the target speech feature matrices from the initial speech feature matrices and performs a marking process, which can achieve accurate marking of specific types of speech information, generate target speech information, and thus improve the speech recognition efficiency in the financial business field.

[0040] Please refer to Figure 2 , Figure 2FIG. 0 is a schematic flowchart of a speech recognition method based on an AI large model provided by the second embodiment of the present application. The speech recognition method based on the AI large model can be applied to a server, and is used to generate an initial speech feature matrix through a preset audio collection model, effectively extract key features in the speech, reduce the data dimension while retaining important information. The initial speech feature matrix is processed by a pre-trained target convolutional neural network model to accurately obtain the target speech feature matrix, and the target speech feature matrix is subjected to dimensionality decomposition processing to generate a feature tag set with a hierarchical structure, further improving the speech recognition efficiency in the financial business field.

[0041] Based on Figure 1 the embodiment shown, this embodiment, as Figure 2 shown, step S30 includes steps S301 to S302.

[0042] Step S301: Perform dimensionality decomposition processing on the target speech feature matrix to generate a feature tag set with at least one level of hierarchical structure;

[0043] Specifically, a suitable dimensionality decomposition method (such as non-negative matrix factorization, tensor decomposition, etc.) is used to decompose the target speech feature matrix. For example, non-negative matrix factorization is used to decompose the feature matrix into multiple non-negative sub-matrices, and these sub-matrices can represent different feature components of the speech.

[0044] The decomposed sub-matrices can be organized according to certain rules to form a feature tag set with at least one level of hierarchical structure. For example, the decomposed sub-matrices are stratified according to the importance degree, frequency, etc. of the features to construct a multi-level feature tag hierarchy.

[0045] Step S302: Map and match each feature tag in the feature tag set with a preset speech information template to generate the target speech information.

[0046] Specifically, the preset speech information template can be a known speech feature pattern, a standard speech signal, etc., and is used as a reference for matching. Each feature tag in the feature tag set is mapped and matched with the preset speech information template. Methods such as similarity calculation and dynamic time warping are used to measure the matching degree between the feature tag and the template, and according to the matching result, the target speech information segment or content corresponding to each feature tag is determined.

[0047] According to the target speech information segment or content obtained by matching, through speech synthesis technology, it is combined into the final target speech information, which may include operations such as restoring the feature tag to a time-domain speech signal, or generating a speech waveform corresponding to the matched feature, so as to obtain the target speech information that can be used in actual applications.

[0048] This embodiment discloses a speech recognition method based on an AI large model. The speech recognition method based on the AI large model includes generating an initial speech feature matrix corresponding to each initial speech information through a preset audio collection model; determining a target speech feature matrix through a pre-trained target convolutional neural network model and each of the initial speech feature matrices; performing dimensionality decomposition processing on the target speech feature matrix to generate a feature marker set with at least one hierarchical structure; and mapping and matching each feature marker in the feature marker set with a preset speech information template to generate the target speech information. By the above method, this application generates an initial speech feature matrix through a preset audio collection model, effectively extracts key features in the speech, reduces the data dimension while retaining important information. The initial speech feature matrix is processed by a pre-trained target convolutional neural network model to accurately obtain the target speech feature matrix. The dimensionality decomposition processing is performed on the target speech feature matrix to generate a feature marker set with a hierarchical structure, further improving the speech recognition efficiency in the financial business field.

[0049] Based on Figure 1 In the embodiment shown, before step S20, it includes:

[0050] Obtain the preset speech feature matrices corresponding to each preset speech information, and construct the target convolutional neural network model according to each of the preset speech feature matrices and the initial convolutional neural network model;

[0051] Train the target convolutional neural network model through each of the preset speech feature matrices to generate the pre-trained target convolutional neural network model.

[0052] Specifically, the steps of constructing the target convolutional neural network model include:

[0053] 1. Determine the input layer: The input is the preprocessed audio feature matrix.

[0054] 2. Determine the convolutional layer:

[0055] 2.1. Perform convolution operations on the input features using multiple convolution kernels to extract different local features.

[0056] 2.2. Convolution formula: For two-dimensional convolution, each element O ij in the output feature map is obtained by performing element-wise multiplication and summation on the local region of the input feature map and the convolution kernel, that is, O ij = ∑ m ∑ n I i+m,j+n K mn , where I is the input feature map and K is the convolution kernel.

[0057] 3. Determine the pooling layer:

[0058] 3.1. Usually, max pooling or average pooling is adopted to reduce the feature dimension, extract the main features, and enhance the robustness of the model.

[0059] 3.2. For example, max pooling takes the maximum value within a local area as the output.

[0060] 4. Determine the fully connected layer:

[0061] 4.1. Map the features after multiple convolutions and poolings into a vector of a fixed length.

[0062] 4.2. The calculation formula of the fully connected layer is similar to that of the traditional neural network, that is, y = f(W x + b), where x is the input vector, W is the weight matrix, b is the bias term, and f is the activation function.

[0063] 5. Output layer:

[0064] 5.1. Use the softmax function for classification and output the probabilities that the recording belongs to excellent and non-excellent conversation scripts.

[0065] 5.2. Softmax function formula: where x i is the i-th element of the input vector.

[0066] Specifically, the steps for training the target convolutional neural network model include:

[0067] 1. Divide the training set, validation set, and test set.

[0068] 2. Select a suitable loss function, such as the cross-entropy loss function. Cross-entropy loss function formula:

[0069] where y i is the true label, is the probability predicted by the model.

[0070] 3. Select an optimization algorithm, such as the Stochastic Gradient Descent (SGD) algorithm.

[0071] 4. Conduct model training and continuously adjust the model parameters to minimize the loss function.

[0072] In a specific embodiment, obtaining the preset speech feature matrices corresponding to the preset speech information and constructing the target convolutional neural network model according to each of the preset speech feature matrices and the initial convolutional neural network model includes:

[0073] Determine the input layer of the initial convolutional neural network model according to the dimensions of each of the preset speech feature matrices;

[0074] Obtain the speech feature maps corresponding to each of the preset speech feature matrices, and determine the convolutional layer of the initial convolutional neural network model according to at least one convolutional kernel and the speech feature maps.

[0075] Extract the features of each of the preset speech feature matrices, and determine the pooling layer of the initial convolutional neural network model according to the features of each of the preset speech feature matrices.

[0076] Determine the fully connected layer of the initial convolutional neural network model according to the features of each of the preset speech feature matrices for convolution and pooling.

[0077] Construct the target convolutional neural network model according to the input layer, the convolutional layer, the pooling layer, and the fully connected layer of the initial convolutional neural network model.

[0078] In a specific embodiment, according to the requirements of the speech processing task, an initial convolutional neural network model is designed and initialized. The structure of the initial convolutional neural network model includes a convolutional layer, a pooling layer, a fully connected layer, etc. The specific number of layers and parameter settings can be determined according to actual needs.

[0079] Input each of the preset speech feature matrices into the initial convolutional neural network model in sequence. In the initial convolutional neural network model, the feature matrix will undergo processes such as feature extraction by the convolutional layer, dimensionality reduction by the pooling layer, and classification by the fully connected layer. According to the performance of the preset speech feature matrix in the initial convolutional neural network model, adjust and optimize the architecture and parameters of the model. For example, adjust the size and number of convolutional kernels, change the window size of the pooling layer, etc., so that the model can better adapt to the input feature data.

[0080] The initial convolutional neural network model after adjustment and optimization is the target convolutional neural network model. This model can effectively extract features and classify the input speech feature matrix.

[0081] In a specific embodiment, training the target convolutional neural network model with each of the preset speech feature matrices to generate the pre-trained target convolutional neural network model includes:

[0082] Divide all the preset speech feature matrices into a training set and a validation set;

[0083] Train the target convolutional neural network model with a preset loss function and the training set to determine the undetermined performance parameters of the target convolutional neural network model;

[0084] Determine the pre-trained target convolutional neural network model through the undetermined performance parameters and the validation set.

[0085] In a specific embodiment, each preset voice feature matrix is used as training data, and corresponding labels (such as the category of the voice, the identity of the speaker, etc.) are prepared to ensure that the training data and the labels correspond one by one and the data volume is sufficient to support the training of the model. Determine the parameters used in the training process, such as the learning rate, batch size, number of training epochs, etc. The settings of these parameters will affect the training effect and convergence speed of the model.

[0086] Use the preset voice feature matrix to train the target convolutional neural network model. During the training process, the model calculates the output result through forward propagation, then uses a loss function (such as cross-entropy loss, etc.) to calculate the error between the prediction result and the true label, and then updates the parameters of the model through the backpropagation algorithm to minimize the value of the loss function.

[0087] Based on Figure 1 the embodiment shown, in this embodiment, step S10 includes:

[0088] Extract the business keywords in each preset voice information through the preset audio collection model, and determine the business labels of each preset voice information according to the business keywords;

[0089] According to the business labels, determine the initial voice information from each preset voice information, and respectively extract the spectral features, speech rate features, pitch features, and zero-crossing features of the initial voice information;

[0090] Integrate the spectral features, the speech rate features, the pitch features, and the zero-crossing features to generate the initial voice feature matrix.

[0091] In one embodiment, perform speech recognition on the collected preset voice information to convert the speech signal into a text. Use natural language processing techniques, such as keyword extraction algorithms or named entity recognition models, to extract business-related keywords from the text, and determine the corresponding business labels for each preset voice information according to the extracted business keywords in combination with business rules or semantic analysis.

[0092] According to the determined business labels, screen out the voice information that meets a specific business type from each preset voice information as the initial voice information, perform preprocessing (such as sampling, quantization, etc.) on the initial voice information, and then calculate the spectrum of the voice signal using methods such as the fast Fourier transform. The spectral features can include indicators such as frequency range, spectral amplitude, and spectral density, reflecting the distribution of the voice signal in the frequency domain.

[0093] Determine the speech rate by calculating the occurrence frequency of phonemes or syllables in the voice signal. For example, count the number of syllables that appear within a unit time, or use a voice activity detection algorithm to identify the duration of the speech segment, thereby calculating the speech rate value, indicating the speed of the speaker.

[0094] Analyze the fundamental frequency contour of the speech signal and extract pitch features. The pitch features can include the average value, variance, change rate, etc. of the fundamental frequency, reflecting the pitch and pitch change of the speech.

[0095] Calculate the zero-crossing rate of the speech signal per unit time, that is, the number of times the signal changes from positive to negative or from negative to positive. The zero-crossing feature can reflect the high-frequency components and noise characteristics of the speech signal, and has a certain effect on distinguishing different types of speech phonemes.

[0096] Normalize the extracted spectral features, speech rate features, pitch features, and zero-crossing features respectively, map the feature values with different dimensions and magnitudes to the same numerical range, and arrange and combine the normalized spectral features, speech rate features, pitch features, and zero-crossing features in a certain order to form a feature vector.

[0097] Arrange the feature vectors corresponding to multiple initial speech information in chronological order or other logical order to form a two-dimensional initial speech feature matrix. Each row or column of the matrix represents the comprehensive features of an initial speech information.

[0098] Based on any of the above embodiments, in this embodiment, step S20 includes:

[0099] Calculate the initial speech scores of each of the initial speech feature matrices;

[0100] Determine the initial speech feature matrix with the initial speech score greater than or equal to the preset score threshold as the target speech feature matrix.

[0101] Please refer to Figure 3 , Figure 3 is a schematic block diagram of a speech recognition device based on an AI large model provided by an embodiment of the present application. The speech recognition device based on the AI large model is used to execute the foregoing speech recognition method based on the AI large model. Among them, the speech recognition device based on the AI large model can be configured in a server.

[0102] As Figure 3 shown, the speech recognition device 400 based on the AI large model includes:

[0103] An initial speech feature matrix generation module 410, configured to generate an initial speech feature matrix corresponding to each of the initial speech information through a preset audio collection model and each of the initial speech information;

[0104] A target speech feature matrix determination module 420, configured to determine a target speech feature matrix through a pre-trained target convolutional neural network model and each of the initial speech feature matrices;

[0105] The target voice information generation module 430 is configured to perform marking processing on the target voice feature matrix to generate target voice information.

[0106] Further, the target voice information generation module 430 includes:

[0107] The feature marking set generation unit is configured to perform dimensionality decomposition processing on the target voice feature matrix to generate a feature marking set with at least one-level hierarchical structure;

[0108] The target voice information generation unit is configured to map and match each feature marking in the feature marking set with a preset voice information template to generate the target voice information.

[0109] Further, the voice recognition device 400 based on the AI large model includes:

[0110] The target convolutional neural network model construction module is configured to obtain the preset voice feature matrices corresponding to the respective preset voice information, and construct the target convolutional neural network model according to the respective preset voice feature matrices and the initial convolutional neural network model;

[0111] The target convolutional neural network model training module is configured to train the target convolutional neural network model through the respective preset voice feature matrices to generate the pre-trained target convolutional neural network model.

[0112] Further, the target convolutional neural network model construction module includes:

[0113] The input layer determination unit is configured to determine the input layer of the initial convolutional neural network model according to the dimensions of the respective preset voice feature matrices;

[0114] The convolutional layer determination unit is configured to obtain the voice feature maps corresponding to the respective preset voice feature matrices, and determine the convolutional layer of the initial convolutional neural network model according to at least one convolutional kernel and the voice feature maps;

[0115] The pooling layer determination unit is configured to extract the features of the respective preset voice feature matrices, and determine the pooling layer of the initial convolutional neural network model according to the features of the respective preset voice feature matrices;

[0116] The fully connected layer determination unit is configured to determine the fully connected layer of the initial convolutional neural network model according to the features of the respective preset voice feature matrices obtained by convolution and pooling;

[0117] The target convolutional neural network model construction unit is configured to construct the target convolutional neural network model according to the input layer, the convolutional layer, the pooling layer, and the fully connected layer of the initial convolutional neural network model.

[0118] Further, the target convolutional neural network model training module includes:

[0119] A training set and validation set determination unit, configured to divide all the preset speech feature matrices into a training set and a validation set;

[0120] An undetermined performance parameter determination unit, configured to train the target convolutional neural network model through a preset loss function and the training set, and determine the undetermined performance parameters of the target convolutional neural network model;

[0121] A target convolutional neural network model training unit, configured to determine the pre-trained target convolutional neural network model through the undetermined performance parameters and the validation set.

[0122] Further, the initial speech feature matrix generation module 410 includes:

[0123] A service label determination unit, configured to extract service keywords in each preset speech information through the preset audio collection model, and determine the service label of each preset speech information according to the service keywords;

[0124] A feature extraction unit, configured to determine the initial speech information from each preset speech information according to the service label, and respectively extract the spectral feature, speech rate feature, pitch feature, and zero-crossing feature of the initial speech information;

[0125] An initial speech feature matrix generation unit, configured to perform feature integration on the spectral feature, the speech rate feature, the pitch feature, and the zero-crossing feature to generate the initial speech feature matrix.

[0126] Further, the target speech feature matrix determination module 420 includes:

[0127] An initial speech score calculation unit, configured to calculate the initial speech scores of the initial speech feature matrices;

[0128] A target speech feature matrix determination unit, configured to determine the initial speech feature matrices with the initial speech scores greater than or equal to a preset score threshold as the target speech feature matrices.

[0129] It should be noted that those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described device and each module can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0130] The above device can be implemented in the form of a computer program, and the computer program can run on a computer device as shown in Figure 4 shown.

[0131] Please refer to Figure 4 , Figure 4 which is a schematic block diagram of a computer device provided by an embodiment of the present application. The computer device may be a server.

[0132] Refer to Figure 4 , the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory may include a non-volatile storage medium and an internal memory.

[0133] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, which when executed, can cause the processor to execute any voice recognition method based on an AI large model.

[0134] The processor is used to provide computing and control capabilities to support the operation of the entire computer device.

[0135] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, it can cause the processor to execute any voice recognition method based on an AI large model.

[0136] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art can understand that Figure 4 the structure shown in [[ ]] is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0137] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0138] Among them, in one embodiment, the processor is used to run the computer program stored in the memory to implement the following steps:

[0139] Collect initial voice information through a preset audio collection model, and generate an initial voice feature matrix corresponding to each of the initial voice information;

[0140] Determine a target voice feature matrix through a pre-trained target convolutional neural network model and each of the initial voice feature matrices;

[0141] Perform marking processing on the target voice feature matrix to generate target voice information.

[0142] In one embodiment, performing marking processing on the target voice feature matrix to generate target voice information is used to achieve:

[0143] Perform dimensionality decomposition processing on the target voice feature matrix to generate a feature marking set with at least one-level hierarchical structure;

[0144] Map and match each feature marking in the feature marking set with a preset voice information template to generate the target voice information.

[0145] In one embodiment, before determining the target voice feature matrix through a pre-trained target convolutional neural network model and each of the initial voice feature matrices, it is used to achieve:

[0146] Obtain preset voice feature matrices corresponding to each preset voice information, and construct the target convolutional neural network model according to each of the preset voice feature matrices and an initial convolutional neural network model;

[0147] Train the target convolutional neural network model through each of the preset voice feature matrices to generate the pre-trained target convolutional neural network model.

[0148] In one embodiment, obtaining preset voice feature matrices corresponding to each preset voice information and constructing the target convolutional neural network model according to each of the preset voice feature matrices and an initial convolutional neural network model is used to achieve:

[0149] Determine the input layer of the initial convolutional neural network model according to the dimensions of each of the preset voice feature matrices;

[0150] Obtain voice feature maps corresponding to each of the preset voice feature matrices, and determine the convolutional layer of the initial convolutional neural network model according to at least one convolutional kernel and the voice feature maps;

[0151] Extract the features of each of the preset voice feature matrices, and determine the pooling layer of the initial convolutional neural network model according to the features of each of the preset voice feature matrices;

[0152] Determine the fully connected layer of the initial convolutional neural network model according to the features of each of the preset speech feature matrices in terms of convolution and pooling;

[0153] Construct the target convolutional neural network model according to the input layer, the convolutional layer, the pooling layer and the fully connected layer of the initial convolutional neural network model.

[0154] In one embodiment, train the target convolutional neural network model with each of the preset speech feature matrices to generate the pre-trained target convolutional neural network model for realizing:

[0155] Divide all the preset speech feature matrices into a training set and a validation set;

[0156] Train the target convolutional neural network model with a preset loss function and the training set to determine the undetermined performance parameters of the target convolutional neural network model;

[0157] Determine the pre-trained target convolutional neural network model with the undetermined performance parameters and the validation set.

[0158] In one embodiment, generate an initial speech feature matrix corresponding to each initial speech information through a preset audio collection model and each initial speech information for realizing:

[0159] Extract the business keywords in each preset speech information through the preset audio collection model, and determine the business labels of each preset speech information according to the business keywords;

[0160] Determine the initial speech information from each preset speech information according to the business labels, and respectively extract the spectral feature, the speech rate feature, the pitch feature and the zero-crossing feature of the initial speech information;

[0161] Integrate the spectral feature, the speech rate feature, the pitch feature and the zero-crossing feature to generate the initial speech feature matrix.

[0162] In one embodiment, determine a target speech feature matrix through the pre-trained target convolutional neural network model and each of the initial speech feature matrices for realizing:

[0163] Calculate the initial speech scores of each of the initial speech feature matrices;

[0164] Determine the initial speech feature matrices with the initial speech scores greater than or equal to a preset score threshold as the target speech feature matrix.

[0165] An embodiment of the present application further provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and the computer program includes program instructions. The processor executes the program instructions to implement any one of the speech recognition methods based on the AI large model provided by the embodiments of the present application.

[0166] Among them, the computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiment, such as the hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device.

[0167] As described above, the above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of various equivalent modifications or replacements, and these modifications or replacements should be covered within the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the protection scope of the claims.

Claims

1. A speech recognition method based on an AI large model, characterized in that, Including: Generating an initial speech feature matrix corresponding to each of the initial speech information through a preset audio collection model; Determining a target speech feature matrix through a pre-trained target convolutional neural network model and each of the initial speech feature matrices; Performing a labeling process on the target speech feature matrix to generate target speech information.

2. The speech recognition method based on the AI large model according to claim 1, wherein The performing a labeling process on the target speech feature matrix to generate target speech information includes: Performing a dimensionality decomposition process on the target speech feature matrix to generate a feature label set with at least one-level hierarchical structure; Mapping and matching each feature label in the feature label set with a preset speech information template to generate the target speech information.

3. The speech recognition method based on the AI large model according to claim 1, wherein Before the determining a target speech feature matrix through a pre-trained target convolutional neural network model and each of the initial speech feature matrices, it includes: Obtaining a preset speech feature matrix corresponding to each preset speech information, and constructing the target convolutional neural network model according to each of the preset speech feature matrices and an initial convolutional neural network model; Training the target convolutional neural network model through each of the preset speech feature matrices to generate the pre-trained target convolutional neural network model.

4. The speech recognition method based on the AI large model according to claim 3, wherein The obtaining a preset speech feature matrix corresponding to each preset speech information, and constructing the target convolutional neural network model according to each of the preset speech feature matrices and an initial convolutional neural network model includes: Determining the input layer of the initial convolutional neural network model according to the dimensions of each of the preset speech feature matrices; Obtaining a speech feature map corresponding to each of the preset speech feature matrices, and determining the convolutional layer of the initial convolutional neural network model according to at least one convolutional kernel and the speech feature map; Extracting the features of each of the preset speech feature matrices, and determining the pooling layer of the initial convolutional neural network model according to the features of each of the preset speech feature matrices; Determining the fully connected layer of the initial convolutional neural network model according to the features of each of the preset speech feature matrices for convolution and pooling; Constructing the target convolutional neural network model according to the input layer, the convolutional layer, the pooling layer and the fully connected layer of the initial convolutional neural network model.

5. The speech recognition method based on the AI large model according to claim 3, characterized in that The training the target convolutional neural network model through each of the preset speech feature matrices to generate the pre-trained target convolutional neural network model includes: Dividing all the preset speech feature matrices into a training set and a validation set; Training the target convolutional neural network model through a preset loss function and the training set to determine the undetermined performance parameters of the target convolutional neural network model; Determining the pre-trained target convolutional neural network model through the undetermined performance parameters and the validation set.

6. The speech recognition method based on the AI large model according to claim 1, wherein The generating an initial speech feature matrix corresponding to each of the initial speech information through a preset audio collection model includes: Extracting service keywords in each preset speech information through the preset audio collection model, and determining the service label of each of the preset speech information according to the service keywords; Determine the initial voice information from each preset voice information according to the service label, and respectively extract the spectral feature, speech rate feature, pitch feature, and zero-crossing feature of the initial voice information; Integrate the spectral feature, the speech rate feature, the pitch feature, and the zero-crossing feature to generate the initial voice feature matrix.

7. The speech recognition method based on the AI large model according to any one of claims 1 to 6, characterized in that Determining the target voice feature matrix through the pre-trained target convolutional neural network model and each of the initial voice feature matrices includes: Calculate the initial voice score of each of the initial voice feature matrices; Determine the initial voice feature matrix with the initial voice score greater than or equal to the preset score threshold as the target voice feature matrix.

8. A voice recognition device based on a large AI model, characterized in that, Including: An initial voice feature matrix generation module, configured to generate an initial voice feature matrix corresponding to each initial voice information through a preset audio collection model and each initial voice information; A target voice feature matrix determination module, configured to determine a target voice feature matrix through a pre-trained target convolutional neural network model and each of the initial voice feature matrices; A target voice information generation module, configured to perform a marking process on the target voice feature matrix to generate target voice information.

9. A computer device, characterized in that, The computer device includes a memory and a processor; The memory is used to store a computer program; The processor is configured to execute the computer program and, when executing the computer program, implement the speech recognition method based on the AI large model according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the processor is caused to implement the speech recognition method based on the AI large model according to any one of claims 1 to 7.